Sr. Software Engineer- AI/ML, AWS Neuron Apps

Amazon Web Services (AWS)
Seattle, Washington, United StatesFull-time$168,100–$227,400Posted Sep 18, 2026

About the role

Description

Shape the Future of AI Accelerators at AWS Neuron

Join the team behind AWS Neuron — the software stack that powers AWS's purpose-built AI accelerators, Inferentia and Trainium. As a Senior Software Engineer on our Machine Learning Applications team, you will optimize the world's most demanding AI models at a scale few engineers ever get to work on.

What You'll Do

Build and scale distributed inference solutions for leading large language models, including GPT, Kimi, and Qwen

Partner directly with silicon architects and compiler engineers to shape the next generation of AI acceleration

Write custom kernels that optimize LLM computation graphs, improving latency and cost for billions of inference requests worldwide

Optimize state-of-the-art language, vision, and multimodal generative AI models for Neuron hardware

Key job responsibilities

You will drive the Evolution of Distributed AI at AWS Neuron

Technical Impact You'll Drive

Spearhead distributed inference architecture for PyTorch

Engineer breakthrough performance optimizations for AWS Trainium and Inferentia

Develop kernels to improve model efficiency on Amazon AI Accelerators

Transform complex tensor operations into highly optimized hardware implementations

What Makes This Role Unique

Direct influence on AWS's AI infrastructure used by thousands of ML applications

Full-stack optimization from high-level frameworks to hardware-specific primitives

Develop tools that define industry standards for ML deployment

Collaboration with both open-source ML communities and hardware architecture teams

Your Technical Arsenal Should Include

Deep expertise in Transformer architecture, Python and Pytorch internals

Strong understanding of distributed systems and ML optimization

Passion for performance tuning and system architecture

A day in the life

Work/Life Balance

Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives.

Mentorship & Career Growth

Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future.

About The Team

At AWS Neuron, we're revolutionizing how the world's most sophisticated AI models run at scale through Amazon's next-generation AI accelerators. Operating at the unique intersection of ML frameworks and custom silicon, our team drives innovation from silicon architecture to production software deployment.

We pioneer distributed inference solutions for PyTorch and JAX using XLA, optimize industry-leading LLMs like GPT and Llama, and collaborate directly with silicon architects to influence the future of AI hardware. Our systems handle millions of inference calls daily, while our optimizations directly impact thousands of AWS customers running critical AI workloads.

We're focused on pushing the boundaries of large language model optimization, distributed inference architecture, and hardware-specific performance tuning. Our deep technical experts transform complex ML challenges into elegant, scalable solutions that define how AI workloads run in production.

Basic Qualifications

5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience

5+ years of programming experience using Python or C++ and PyTorch.

Experience with AI acceleration via quantization, parallelism, model compression, batching, KV caching, vllm serving

Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators

Fundamentals of Machine learning and deep learning models, their architecture, training and inference lifecycles along with work experience on optimizations for improving the model execution.

Preferred Qualifications

Master's degree in computer science or equivalent

Master's degree in machine learning or equivalent

Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators

Experience in developing CUDA kernels, HPC and inference optimization, tensors operations

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, WA, Seattle - 168,100.00 - 227,400.00 USD annually

Company - Annapurna Labs (U.S.) Inc.

Job ID: A3118746

Responsibilities

  • Build and scale distributed inference solutions for large language models
  • Partner with silicon architects and compiler engineers
  • Write custom kernels to optimize LLM computation graphs
  • Optimize generative AI models for Neuron hardware

Qualifications

  • 5+ years of full software development life cycle experience
  • 5+ years of programming experience using Python or C++ and PyTorch
  • Experience with AI acceleration techniques
  • Fundamentals of machine learning and deep learning models

Benefits

  • Health insurance (medical, dental, vision)
  • 401(k) matching
  • Paid time off
  • Parental leave

Skills mentioned

PythonC++PyTorchDistributed SystemsPerformance OptimizationCUDADeep LearningMachine LearningLarge Language ModelsTransformers

About Amazon Web Services (AWS)

Launched in 2006, Amazon Web Services (AWS) began exposing key infrastructure services to businesses in the form of web services -- now widely known as cloud computing. The ultimate benefit of cloud computing, and AWS, is the ability to leverage a new business model and turn capital infrastructure expenses into variable costs. Businesses no longer need to plan and procure servers and other IT resources weeks or months in advance. Using AWS, businesses can take advantage of Amazon's expertise and economies of scale to access resources when their business needs them, delivering results faster and at a lower cost. Today, Amazon Web Services provides a highly reliable, scalable, low-cost infrastructure platform in the cloud that powers hundreds of thousands of businesses in 190 countries around the world. With data center locations in the U.S., Europe, Singapore, and Japan, customers across all industries are taking advantage of our low cost, elastic, open and flexible, secure platform.

IT Services and IT Consulting10,001+ employeesSeattle, WA