Senior Software Engineer, AI/ML System Infrastructure
About the role
MINIMUM QUALIFICATIONS:
- Bachelor’s degree or equivalent practical experience.
- 5 years of experience working with Go.
- 3 years of experience with developing large-scale infrastructure, distributed
systems or networks, or experience with compute technologies, storage or
hardware architecture.
- 3 years of experience in distributed computing.
- 3 years of experience in infrastructure design.
- 3 years of experience in system architecture.
PREFERRED QUALIFICATIONS:
- Master's degree or PhD in Computer Science or related technical field.
- 5 years of experience with data structures and algorithms.
- 5 years of experience with C and C++.
- 1 year of experience in a technical leadership role.
- Experience developing accessible technologies.
ABOUT THE JOB:
Google's software engineers develop the next-generation technologies that change
how billions of users connect, explore, and interact with information and one
another. Our products need to handle information at massive scale, and extend
well beyond web search. We're looking for engineers who bring fresh ideas from
all areas, including information retrieval, distributed computing, large-scale
system design, networking and data storage, security, artificial intelligence,
natural language processing, UI design and mobile; the list goes on and is
growing every day. As a software engineer, you will work on a specific project
critical to Google’s needs with opportunities to switch teams and projects as
you and our fast-paced business grow and evolve. We need our engineers to be
versatile, display leadership qualities and be enthusiastic to take on new
problems across the full-stack as we continue to push technology forward.
As the Senior Software Engineer, you will drive software development for Tensor
Processing Unit (TPU) system control planes. You will design and implement
health management systems that rely on hardware telemetry. You will build
analytics for detecting hardware problems and develop different health rules
specializing in detecting ICI, OCS, EMD, Tross, and out of band entities failure
modes. You will build algorithms to generate correlated failures and suggest
actions for repair workflows, integrating with both TPU cluster and cloud
infrastructure. You will also build the Diagnoser for the specialized health
rules to maintain and operate these TPU clusters.
Google Cloud accelerates every organization’s ability to digitally transform its
business and industry. We deliver enterprise-grade solutions that leverage
Google’s cutting-edge technology, and tools that help developers build more
sustainably. Customers in more than 200 countries and territories turn to Google
Cloud as their trusted partner to enable growth and solve their most critical
business problems.
Individual pay is determined by factors including job-related skills,
experience, and relevant education or training.
US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits
Learn more about benefits at Google
[https://www.google.com/about/careers/applications/benefits/].
RESPONSIBILITIES:
- Build and develop systems to integrate in continuous running of TPU AI
infrastructure.
- Work cross-functionally to define the requirements for Diagnoser, and how
will the customer benefit from it and define the Critical User Journeys
(CUJs).
- Influence and align the cross-functional teams on roadmap of PodCare and
Diagnoser.
- Deliver high quality code and timely project while working with
cross-functional teams.
- Contribute to CI/CD pipeline and integration and regression test for health
monitoring system.
Responsibilities
- Build and develop systems to integrate in continuous running of TPU AI infrastructure
- Work cross-functionally to define the requirements for Diagnoser and how the customer will benefit from it
- Influence and align the cross-functional teams on roadmap of PodCare and Diagnoser
- Deliver high quality code and timely project while working with cross-functional teams
- Contribute to CI/CD pipeline and integration and regression test for health monitoring system
Qualifications
- Bachelor’s degree or equivalent practical experience
- 5 years of experience working with Go
- 3 years of experience with developing large-scale infrastructure, distributed systems or networks
- 3 years of experience in distributed computing
- 3 years of experience in infrastructure design
- 3 years of experience in system architecture
- Master's degree or PhD in Computer Science or related technical field preferred
- 5 years of experience with data structures and algorithms preferred
Benefits
- 15% bonus target
- equity
- health management systems
- access to cutting-edge technology
Skills mentioned
About Google
A problem isn't truly solved until it's solved for all. Googlers build products that help create opportunities for everyone, whether down the street or across the globe. Bring your insight, imagination and a healthy disregard for the impossible. Bring everything that makes you unique. Together, we can build for everyone. Check out our career opportunities at goo.gle/3DLEokh