Software Engineer III, TPU Performance, Hardware and Software Codesign
About the role
MINIMUM QUALIFICATIONS:
- Bachelor’s degree or equivalent practical experience.
- 2 years of experience with software development in one or more programming
languages.
- 2 years of experience with performance, large-scale systems data analysis,
visualization tools, or debugging.
- 2 years of experience with computer architecture, performance analysis, and
performance modeling.
PREFERRED QUALIFICATIONS:
- Master's degree or PhD in Computer Science or related technical fields.
- 2 years of experience with data structures and algorithms.
- Experience developing accessible technologies.
ABOUT THE JOB:
Google's software engineers develop the next-generation technologies that change
how billions of users connect, explore, and interact with information and one
another. Our products need to handle information at massive scale, and extend
well beyond web search. We're looking for engineers who bring fresh ideas from
all areas, including information retrieval, distributed computing, large-scale
system design, networking and data storage, security, artificial intelligence,
natural language processing, UI design and mobile; the list goes on and is
growing every day. As a software engineer, you will work on a specific project
critical to Google’s needs with opportunities to switch teams and projects as
you and our fast-paced business grow and evolve. We need our engineers to be
versatile, display leadership qualities and be enthusiastic to take on new
problems across the full-stack as we continue to push technology forward.
In this role, you will bridge the gap between ML workloads and custom Tensor
Processing Unit (TPU) hardware. You will analyze and optimize how distributed
systems, compiler architectures—such as Accelerated Linear Algebra (XLA)—and
emerging software abstractions, such as Compound AI and multi-step agentic
systems, execute across Google’s AI infrastructure. You will collaborate with
various product area architects within Google, such as Google Cloud and YouTube,
and external customers to systematically onboard novel workloads with engaged
performance. Your optimizations will directly drive TPU adoption, secure
pre-sales engagements, and shape our future ML infrastructure roadmap.
Google Cloud accelerates every organization’s ability to digitally transform its
business and industry. We deliver enterprise-grade solutions that leverage
Google’s cutting-edge technology, and tools that help developers build more
sustainably. Customers in more than 200 countries and territories turn to Google
Cloud as their trusted partner to enable growth and solve their most critical
business problems.Individual pay is determined by factors including job-related
skills, experience, and relevant education or training.
US: $147000 - $210000 (USD) + 15% bonus target + equity + benefits
Learn more about benefits at Google
[https://www.google.com/about/careers/applications/benefits/].
RESPONSIBILITIES:
- Develop and scale benchmarking and workload characterization strategies to
enable fast grounding-to-silicon, root-cause performance analysis, and TPU
mapping optimization.
- Drive full-stack hardware-software co-design to optimize current and future
ML accelerator architectures for business-critical production models (e.g.,
LLMs and embedding models).
- Partner with Product Areas (e.g., YouTube and Ads) to scale key workload
pipelines efficiently (Perf/$/Watts) during TPU Pilot and General
Availability (GA) transitions.
- Build and upgrade compiler-aware simulator tools, hardware cost-models, and
performance-ladder pathways to baseline and project physical silicon
capabilities.
- Distill complex performance analyses and hardware trade-offs into
presentations to guide TPU roadmap decision-making in core leadership forums
(e.g., ArchForums, NPI, BCR reviews, and TdJs).
Responsibilities
- Develop and scale benchmarking and workload characterization strategies to enable fast grounding-to-silicon, root-cause performance analysis, and TPU mapping optimization.
- Drive full-stack hardware-software co-design to optimize current and future ML accelerator architectures for business-critical production models (e.g., LLMs and embedding models).
- Partner with Product Areas (e.g., YouTube and Ads) to scale key workload pipelines efficiently (Perf/$/Watts) during TPU Pilot and General Availability (GA) transitions.
- Build and upgrade compiler-aware simulator tools, hardware cost-models, and performance-ladder pathways to baseline and project physical silicon capabilities.
- Distill complex performance analyses and hardware trade-offs into presentations to guide TPU roadmap decision-making in core leadership forums (e.g., ArchForums, NPI, BCR reviews, and TdJs).
Qualifications
- Bachelor’s degree or equivalent practical experience.
- 2 years of experience with software development in one or more programming languages.
- 2 years of experience with performance, large-scale systems data analysis, visualization tools, or debugging.
- 2 years of experience with computer architecture, performance analysis, and performance modeling.
- Master's degree or PhD in Computer Science or related technical fields (preferred).
- 2 years of experience with data structures and algorithms (preferred).
- Experience developing accessible technologies (preferred).
Benefits
- 15% bonus target
- equity
- benefits
Skills mentioned
About Google
A problem isn't truly solved until it's solved for all. Googlers build products that help create opportunities for everyone, whether down the street or across the globe. Bring your insight, imagination and a healthy disregard for the impossible. Bring everything that makes you unique. Together, we can build for everyone. Check out our career opportunities at goo.gle/3DLEokh