Machine Learning Engineer

Goliath Partners
San Francisco, California, United StatesFull-timePosted Sep 16, 2026

About the role

I’m working with a rapidly growing, well-funded AI company building some of the most technically ambitious real-world AI systems today.

They’re looking for an ML Infrastructure Engineer to join a highly technical team responsible for building the infrastructure that powers large-scale model training, experimentation, and deployment.

This is a hands-on engineering role for someone who enjoys solving difficult systems problems at the intersection of machine learning, distributed systems, and infrastructure.

What you’ll work on:

  • Build and scale infrastructure for training large machine learning models
  • Develop distributed training systems and improve training efficiency, reliability, and throughput
  • Build data pipelines and infrastructure supporting large-scale ML workloads
  • Improve GPU utilization, compute orchestration, checkpointing, and experiment management
  • Develop tooling that enables researchers and ML engineers to iterate faster
  • Diagnose performance bottlenecks across training, data, and compute systems
  • Help take ML systems from experimentation through production and real-world deployment

What they’re looking for:

  • Strong Python and software engineering fundamentals
  • Experience building ML infrastructure, training infrastructure, or large-scale distributed systems
  • Experience working with PyTorch and modern ML training stacks
  • Strong understanding of distributed computing and GPU-based workloads
  • Experience building reliable systems that support ML research or production ML
  • Comfortable working in a fast-moving environment with significant technical ownership

Especially interesting backgrounds include:

  • Distributed training and large-scale model training
  • ML platforms / internal training infrastructure
  • GPU infrastructure and compute orchestration
  • Large-scale data infrastructure and pipelines
  • Performance optimization for ML workloads
  • Infrastructure supporting robotics, embodied AI, computer vision, or other real-world ML systems
  • Experience taking ML systems beyond research and into production or physical-world environments

This is a great opportunity for an engineer who wants to work on hard ML systems problems at scale while being much closer to the models and real-world applications than you would be on a traditional infrastructure team.

📍 Bay Area - in person

If you have a strong ML infrastructure or distributed systems background and want to hear more, apply directly or message me.

Responsibilities

  • Build and scale infrastructure for training large machine learning models
  • Develop distributed training systems and improve training efficiency, reliability, and throughput
  • Build data pipelines and infrastructure supporting large-scale ML workloads
  • Improve GPU utilization, compute orchestration, checkpointing, and experiment management
  • Develop tooling that enables researchers and ML engineers to iterate faster
  • Diagnose performance bottlenecks across training, data, and compute systems
  • Help take ML systems from experimentation through production and real-world deployment

Qualifications

  • Strong Python and software engineering fundamentals
  • Experience building ML infrastructure, training infrastructure, or large-scale distributed systems
  • Experience working with PyTorch and modern ML training stacks
  • Strong understanding of distributed computing and GPU-based workloads
  • Experience building reliable systems that support ML research or production ML
  • Comfortable working in a fast-moving environment with significant technical ownership

Skills mentioned

PythonDistributed SystemsPerformance OptimizationMachine LearningDeep LearningPyTorchData PipelinesMLOpsModel DeploymentCUDA

About Goliath Partners

Driving innovative talent solutions within trading & financial technology that delivers specialized C++ professionals.

Staffing and Recruiting2-10 employeesMiami, Florida