Principal Machine Learning Engineer
About the role
We are partnering with a pioneering technology firm that is at the forefront of innovation, developing critical machine learning systems that drive significant impact across various industries. This company is renowned for its commitment to technical excellence and its ability to solve complex, architectural, and performance challenges at scale.
The Role
Design and evolve mission-critical ML systems, setting technical standards for the organization
Operate across training, inference, evaluation, and infrastructure, addressing complex architectural and performance problems
Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment
Design reproducible, high-performance training pipelines across GPU infrastructure
Implement evaluation pipelines covering performance, robustness, safety, and bias
Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies
What You'll Need
Strong background in deep learning and transformer-based architectures
Artificial Intelligence (AI) experience is required
Hands-on experience training, fine-tuning, or deploying large-scale ML models in production
Proficiency with at least one modern ML framework (e.g., PyTorch, JAX)
Experience with distributed training and inference frameworks (e.g., DeepSpeed, FSDP, Megatron, ZeRO, Ray)
Strong software engineering fundamentals, capable of building robust, maintainable, production-grade systems
What's On Offer
An opportunity to be a deep technical authority in a high-impact, hands-on role
Drive the technical direction of ML systems across the organization
Work on cutting-edge problems in ML, including GPU optimization and large-scale data processing
Comprehensive benefits package including medical, dental, vision, savings plans, and PTO
Apply via Haystack today!
Responsibilities
- Design and evolve mission-critical ML systems
- Operate across training, inference, evaluation, and infrastructure
- Architect and build large-scale ML systems
- Design reproducible, high-performance training pipelines
- Implement evaluation pipelines covering performance, robustness, safety, and bias
- Own production deployment including GPU optimization
Qualifications
- Strong background in deep learning and transformer-based architectures
- AI experience is required
- Hands-on experience training, fine-tuning, or deploying large-scale ML models
- Proficiency with at least one modern ML framework
- Experience with distributed training and inference frameworks
- Strong software engineering fundamentals
Benefits
- Comprehensive benefits package including medical, dental, vision
- Savings plans
- PTO
Skills mentioned
About Haystack
Haystack combines AI & expert vetting to deliver world-class tech candidates who are engaged, aligned, and ready to interview. We're trusted by over 400,000+ tech candidates, working in Software Engineering, Data, Design, DevOps, Cloud, Tech Management, Testing, Product & Delivery, Architecture and more. 100s of employers from startups and scale-ups like Atom Bank, DuckDuckGo and Goodlord to established enterprises like American Express, Dunelm and AWS use Haystack to connect with qualified tech talent that they can't find anywhere else.