Senior Software Engineer, AI/ML System Infrastructure

Google
Sunnyvale, California, United States · Kirkland, Washington, United StatesFull-time$174,000–$252,000Posted Sep 4, 2026

About the role

MINIMUM QUALIFICATIONS:

  • Bachelor’s degree or equivalent practical experience.
  • 5 years of experience working with Go.
  • 3 years of experience with developing large-scale infrastructure, distributed

systems or networks, or experience with compute technologies, storage or

hardware architecture.

  • 3 years of experience in distributed computing.
  • 3 years of experience in infrastructure design.
  • 3 years of experience in system architecture.

PREFERRED QUALIFICATIONS:

  • Master's degree or PhD in Computer Science or related technical field.
  • 5 years of experience with data structures and algorithms.
  • 5 years of experience with C and C++.
  • 1 year of experience in a technical leadership role.
  • Experience developing accessible technologies.

ABOUT THE JOB:

Google's software engineers develop the next-generation technologies that change

how billions of users connect, explore, and interact with information and one

another. Our products need to handle information at massive scale, and extend

well beyond web search. We're looking for engineers who bring fresh ideas from

all areas, including information retrieval, distributed computing, large-scale

system design, networking and data storage, security, artificial intelligence,

natural language processing, UI design and mobile; the list goes on and is

growing every day. As a software engineer, you will work on a specific project

critical to Google’s needs with opportunities to switch teams and projects as

you and our fast-paced business grow and evolve. We need our engineers to be

versatile, display leadership qualities and be enthusiastic to take on new

problems across the full-stack as we continue to push technology forward.

As the Senior Software Engineer, you will drive software development for Tensor

Processing Unit (TPU) system control planes. You will design and implement

health management systems that rely on hardware telemetry. You will build

analytics for detecting hardware problems and develop different health rules

specializing in detecting ICI, OCS, EMD, Tross, and out of band entities failure

modes. You will build algorithms to generate correlated failures and suggest

actions for repair workflows, integrating with both TPU cluster and cloud

infrastructure. You will also build the Diagnoser for the specialized health

rules to maintain and operate these TPU clusters.

Google Cloud accelerates every organization’s ability to digitally transform its

business and industry. We deliver enterprise-grade solutions that leverage

Google’s cutting-edge technology, and tools that help developers build more

sustainably. Customers in more than 200 countries and territories turn to Google

Cloud as their trusted partner to enable growth and solve their most critical

business problems.

Individual pay is determined by factors including job-related skills,

experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google

[https://www.google.com/about/careers/applications/benefits/].

RESPONSIBILITIES:

  • Build and develop systems to integrate in continuous running of TPU AI

infrastructure.

  • Work cross-functionally to define the requirements for Diagnoser, and how

will the customer benefit from it and define the Critical User Journeys

(CUJs).

  • Influence and align the cross-functional teams on roadmap of PodCare and

Diagnoser.

  • Deliver high quality code and timely project while working with

cross-functional teams.

  • Contribute to CI/CD pipeline and integration and regression test for health

monitoring system.

Responsibilities

  • Build and develop systems to integrate in continuous running of TPU AI infrastructure.
  • Work cross-functionally to define the requirements for Diagnoser, and how will the customer benefit from it and define the Critical User Journeys (CUJs).
  • Influence and align the cross-functional teams on roadmap of PodCare and Diagnoser.
  • Deliver high quality code and timely project while working with cross-functional teams.
  • Contribute to CI/CD pipeline and integration and regression test for health monitoring system.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 5 years of experience working with Go.
  • 3 years of experience with developing large-scale infrastructure, distributed systems or networks.
  • 3 years of experience in distributed computing.
  • 3 years of experience in infrastructure design.
  • 3 years of experience in system architecture.
  • Master's degree or PhD in Computer Science or related technical field (preferred).
  • 5 years of experience with data structures and algorithms (preferred).

Benefits

  • 15% bonus target
  • equity
  • benefits

Skills mentioned

GoCC++Distributed SystemsSystem DesignGoogle CloudCI/CDIntegration TestingDebuggingSystems Engineering

About Google

A problem isn't truly solved until it's solved for all. Googlers build products that help create opportunities for everyone, whether down the street or across the globe. Bring your insight, imagination and a healthy disregard for the impossible. Bring everything that makes you unique. Together, we can build for everyone. Check out our career opportunities at goo.gle/3DLEokh

Software Development10,001+ employeesMountain View, CA