AI Infrastructure Engineer
About the role
AI Infrastructure Engineer | 190k | Fully Remote (US-based)
If you have strong HPC and LLM experience, apply below!
Both of these areas are must-haves for this role.
What's in It for You?
Join a highly technical environment supporting cutting-edge AI, machine learning, and high-performance computing platforms.
Play a key role in shaping and scaling enterprise AI infrastructure, including GPU clusters and large-scale model workloads.
Work alongside data scientists, researchers, and engineering teams on innovative projects with real business impact.
Work in a highly respected business consulting firm that works with Fortune 500 companies.
Responsibilities
Act as the LLM specialist on the HPC team, working closely with 4 other HPC engineers.
HPC Responsibilities
Administer NVIDIA GPU infrastructure and software components, including CUDA, cuDNN, NCCL, and the NVIDIA GPU Operator
Work with H100/H200 hardware in production environments
Manage job scheduling through SLURM and other HPC schedulers
Containerise HPC workloads using Docker and Kubernetes
LLM Specialism
Design and maintain GPU clusters for high‑demand AI/ML and HPC workloads
Fine‑tune large models on multi‑GPU clusters
Perform advanced performance tuning for AI/ML and LLM training workloads
Must-Have Requirements
Experience within high-performance computing, research, or enterprise environments.
Experience with Nvidia H200 Chips for HPC workloads
Hands-on experience with GPU infrastructure
Experience supporting AI/ML, LLM workloads
Click Below to Apply!
Responsibilities
- Act as the LLM specialist on the HPC team, working closely with 4 other HPC engineers
- Administer NVIDIA GPU infrastructure and software components, including CUDA, cuDNN, NCCL, and the NVIDIA GPU Operator
- Work with H100/H200 hardware in production environments
- Manage job scheduling through SLURM and other HPC schedulers
- Containerise HPC workloads using Docker and Kubernetes
- Design and maintain GPU clusters for high‑demand AI/ML and HPC workloads
- Fine‑tune large models on multi‑GPU clusters
- Perform advanced performance tuning for AI/ML and LLM training workloads
Qualifications
- Experience within high-performance computing, research, or enterprise environments
- Experience with Nvidia H200 Chips for HPC workloads
- Hands-on experience with GPU infrastructure
- Experience supporting AI/ML, LLM workloads
Skills mentioned
About Franklin Fitch
Franklin Fitch is a specialist recruitment consultancy, offering our clients an extensive range of Talent Solutions. Focused on several core technology areas, the team at Franklin Fitch guarantees integrity and professional recruitment in the following markets: • Infrastructure & Cloud • Security • Data & AI • Software & Application Engineering • Go-To-Market (GTM) Technology Roles We represent candidates for permanent and contract positions at all levels of seniority, administration through to CTO. With expertise and vast experience in our chosen field, Franklin Fitch has quickly established itself as a trusted Technology partner within the UK, US and German markets. _________________________________________________________ Franklin Fitch ist eine Personalberatung, die ihren Kund:innen ein umfassendes Angebot and Talent Solutions bietet. Mit unserem Fokus auf Kerntechnologiebereiche garantiert unser Team professionelles Recruitment in den folgenden Bereichen: • Infrastructure & Cloud • Security • Data & AI • Software & Application Engineering • Go-To-Market (GTM) Technology Roles Wir vertreten Kandidaten sowohl für Festanstellungen als auch Freelancepositionen, von Junior- bis Seniorlevel und von der Administration bis hin zum Vorstand (C-level). Mit Fachkompetenz und langjähriger Erfahrung in IT und Technologie hat sich Franklin Fitch schnell als ein vertrauenswürdiges und führendes Dienstleistungsunternehmen innerhalb Deutschlands, des Vereinigten Königreichs und den USA etalbliert.