Machine Learning Engineer - Inference

Fintal Partners
New York, New York, United StatesFull-timePosted Sep 18, 2026

About the role

A leading quantitative trading firm is seeking a Machine Learning Engineer specializing in inference to join a highly technical AI research team developing and deploying large-scale machine learning models for financial markets.

The team builds powerful foundation models for markets, trained on vast quantities of market and alternative data to predict future market behavior. These models are deployed directly into live trading environments, making inference speed, efficiency, and reliability critical to the business.

As a Machine Learning Engineer, you will work at the intersection of machine learning, high-performance computing, and systems engineering, with a broad mandate to improve large-scale model inference.

What You’ll Work On

Design and optimize high-performance inference systems for large-scale deep learning models

Develop and optimize GPU kernels using CUDA, Triton, Pallas, CuTe DSL, and related technologies

Improve lower-level performance across PyTorch, JAX, XLA, and CUDA Graphs

Explore and develop inference solutions across GPUs, ASICs, and FPGAs

Optimize data streaming and model-serving infrastructure for demanding real-time environments

Work closely with ML researchers to co-design efficient inference architectures

Improve latency, throughput, hardware utilization, and overall inference efficiency

Help shape the team’s broader machine learning and systems research agenda

The inference environment spans multiple platforms deployed globally and supports a variety of model architectures and trading strategies. The work is highly performance-sensitive, technically challenging, and has a direct impact on live trading performance.

Qualifications

2+ years of professional experience building deep learning or machine learning systems

Strong software engineering and systems fundamentals

Experience building deep learning systems in computationally intensive domains such as robotics, recommendation systems, biology, chemistry, physics, audio, video, or similar areas

Ability to translate techniques and approaches across different machine learning domains

Plus experience with one or more of the following:

CUDA, Triton, Pallas, or CuTe DSL kernel development

Lower-level PyTorch, JAX, or XLA development

CUDA Graphs

GPU performance optimization

FPGA or ASIC development

High-performance ML inference systems

Nice to Have

Experience with large-scale or low-latency model inference

LLM or foundation model experience

Experience optimizing GPU kernels or distributed ML workloads

Prior finance or trading experience is not required.

Responsibilities

  • Design and optimize high-performance inference systems for large-scale deep learning models
  • Develop and optimize GPU kernels using CUDA, Triton, Pallas, CuTe DSL, and related technologies
  • Improve lower-level performance across PyTorch, JAX, XLA, and CUDA Graphs
  • Explore and develop inference solutions across GPUs, ASICs, and FPGAs
  • Optimize data streaming and model-serving infrastructure for demanding real-time environments
  • Work closely with ML researchers to co-design efficient inference architectures
  • Improve latency, throughput, hardware utilization, and overall inference efficiency
  • Help shape the team’s broader machine learning and systems research agenda

Qualifications

  • 2+ years of professional experience building deep learning or machine learning systems
  • Strong software engineering and systems fundamentals
  • Experience building deep learning systems in computationally intensive domains such as robotics, recommendation systems, biology, chemistry, physics, audio, video, or similar areas
  • Ability to translate techniques and approaches across different machine learning domains
  • Plus experience with one or more of the following: CUDA, Triton, Pallas, or CuTe DSL kernel development, Lower-level PyTorch, JAX, or XLA development, CUDA Graphs, GPU performance optimization, FPGA or ASIC development, High-performance ML inference systems

Skills mentioned

Machine LearningDeep LearningPyTorchJAXCUDAModel ServingModel DeploymentPerformance OptimizationSystems EngineeringDistributed Systems

About Fintal Partners

Fintal Partners is a boutique executive search firm specializing in the placement of elite talent within the U.S. capital markets. Our focus is sharp and niche: connecting the most exceptional candidates with leading trading shops, hedge funds, and private equity firms. What sets us apart is the strength of our relationships. We have built long-standing partnerships with many of the world’s most prominent and respected institutions, allowing us to offer unparalleled access and insight. With a deep understanding of the capital markets landscape, we deliver tailored search solutions that meet the unique demands of this fast-paced, competitive industry. A core areas of expertise include: Trading - Quantitative Research, Portfolio Management, Algorithmic Development, Systematic & Semi-Systematic Trading Software - Core Software Development, Quantitative Development, Machine Learning, Data, Infrastructure Hardware - R&D, FPGA/ASIC/PCB Design & Verification Back/Middle Office - Risk Management, Analytics, Market Data, Operations, Procurement If you're interested in hearing more, feel free to get in contact: contact@fintalpartners.com

Capital Markets11-50 employeesChicago, Illinois