Machine Learning Engineer, Speech LLM Training

Plaud Inc.
San Francisco, California, United StatesFull-time$200,000–$540,000Posted Aug 30, 2026

About the role

This role is part of the Jobright TNT - the private hiring network connecting top talent with top AI startups like Perplexity, Mercor, Cresta, Suno and 150 more.

This is not a mass job posting. Only select, high-signal candidates are invited to Jobright TNT and recommended directly to hiring teams

Hiring Company: Plaud

One-liner: Plaud Inc. is a hardware and software company making AI-powered voice recorders for note-taking, with over 1M users globally.

Salary: $200K/yr - $540K/yr

Why Join Us:

Role Responsibilities

  • Have a proven track record of building and training large-scale audio or speech models from the ground up, whether that involves unified SpeechLLMs, advanced ASR, expressive TTS, or generative audio architectures
  • Love living at the intersection of research and engineering, eager to design novel sequence modeling architectures one day and debug distributed training clusters the next
  • Are highly comfortable traversing the entire stack—from fundamental signal processing and raw acoustic representations to massive foundation model training and edge-device optimization
  • Possess deep expertise in PyTorch or JAX, with battle scars from optimizing large-scale distributed training runs, managing GPU memory utilization, and resolving complex performance bottlenecks
  • Thrive in a fast-paced, high-growth startup environment where you are expected to take extreme ownership of ambiguous problems and drive them directly into production
  • Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity

Qualitications

Required

  • Have a proven track record of building and training large-scale audio or speech models from the ground up, whether that involves unified SpeechLLMs, advanced ASR, expressive TTS, or generative audio architectures
  • Love living at the intersection of research and engineering, eager to design novel sequence modeling architectures one day and debug distributed training clusters the next
  • Are highly comfortable traversing the entire stack—from fundamental signal processing and raw acoustic representations to massive foundation model training and edge-device optimization
  • Possess deep expertise in PyTorch or JAX, with battle scars from optimizing large-scale distributed training runs, managing GPU memory utilization, and resolving complex performance bottlenecks
  • Thrive in a fast-paced, high-growth startup environment where you are expected to take extreme ownership of ambiguous problems and drive them directly into production
  • Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity

Preferred

  • Text-based LLMs: Hands-on experience with core text-based Large Language Model pretraining, instruction tuning, or RLHF
  • Neural Audio Codecs: Hands-on experience designing and training state-of-the-art neural audio codecs for streamable, high-fidelity audio
  • Generative Architectures: Designing and training diffusion models, flow matching, or autoregressive architectures specifically for speech and voice generation
  • Alignment & Steerability: Applying Reinforcement Learning (RL) techniques (like RLHF or GRPO) to improve conversational cadence, steerability, and alignment in foundation models
  • Deep System Optimization: End-to-end inference and performance optimization, leveraging high-throughput serving frameworks (e.g., vLLM, TensorRT-LLM, SGLang) to minimize latency for real-time cloud streaming
  • Large-Scale Infrastructure: Managing massive GPU clusters, utilizing advanced distributed training frameworks (e.g., FSDP, DeepSpeed), and navigating orchestration tools like Kubernetes

How can I join Jobright TNT:

If this is your first time applying to a Jobright TNT role, the process works as follows:

  • Apply to your first Jobright TNT role
  • We review your background to determine if you meet the TNT quality bar
  • If qualified, your application is directly recommended to the employer
  • Once accepted into TNT, you may be:
  • Invited to apply for other exclusive TNT-only roles
  • Invited to private, invite-only hiring events with top startups

You will be notified of your TNT selection result.

PS: All Jobright TNT roles are 100% real, directly hired by top AI startups we partner with, and come with priority review and higher response rates than the normal application queue.

Responsibilities

  • Build and train large-scale audio or speech models from the ground up
  • Design novel sequence modeling architectures and debug distributed training clusters
  • Traverse the entire stack from signal processing to foundation model training
  • Optimize large-scale distributed training runs and manage GPU memory utilization
  • Drive ambiguous problems directly into production

Qualifications

  • Proven track record in building and training large-scale audio or speech models
  • Deep expertise in PyTorch or JAX
  • Experience with text-based LLMs and neural audio codecs preferred
  • Hands-on experience with generative architectures and reinforcement learning techniques
  • End-to-end inference and performance optimization experience

Skills mentioned

Deep LearningPyTorchJAXLarge Language ModelsDistributed SystemsPerformance OptimizationCUDAKubernetesSignal ProcessingReinforcement Learning

About Plaud Inc.

Every hire is two searches - a person looking for the right role, and a company looking for the right person. For two years, Jobright worked only the first one. 2 million professionals now use our AI agent to find roles, sharpen how they show up, and get in front of the people hiring. Earlier this year, we started working the second. Employers can now put Jobright on a role. Our AI recruiter sources from the professionals who already trust us, verifies every candidate is real, and reaches out under a name candidates recognize, which is why they reply. What reaches the hiring team is a list of people worth interviewing, within 24 hours. All they do is interview. Two searches. One platform. Agents on both sides that keep working until both sides say yes.

Software Development11-50 employeesSanta Clara, California