Senior Machine Learning Engineer
About the role
We build and operate the pipeline that processes, validates, and delivers TB-scale multimodal sensor data — synchronized video, IMU, calibration, and structured metadata. You'll own that pipeline end to end.
Responsibilities
Own the data pipeline end to end: ingest → validate → transcode → package → deliver
Build automated quality gates that turn written specifications into executable checks
Design and version the delivery schema; maintain backward compatibility as specs change
Own storage architecture and cost at PB scale
Build observability and lineage — for any delivered file: which device, which calibration, which pipeline version, why it passed
Interface directly with customer engineering on format, validation, and defect resolution
Scale the system 5× without a rewrite
Requirements
5+ years in production data infrastructure, 2+ at senior/staff scope, still hands-on
Multimodal sensor data pipelines at scale — timestamp synchronization, clock drift, sensor calibration
Robotics/AV data formats: MCAP, ROS bags, protobuf, Foxglove or equivalents
Large-scale video processing: transcoding, codec behavior (H.265/AV1, keyframes, CRF vs CBR), distributed orchestration (Ray/Beam/Argo/Airflow)
Cloud storage and cost engineering at TB–PB scale
Built validation systems, not just run them
Direct technical communication with external stakeholders
Nice to have
Camera calibration models (pinhole/radtan, fisheye/Kannala-Brandt), stereo geometry
Dataset formats for robot policy training (e.g. LeRobot-style conversions)
Dashboards over dataset composition and coverage
Experience where upstream data arrived dirty from external vendors
Not a fit if
Your data experience is primarily text, tabular, or event-stream
You've only worked inside a large company's existing platform
You prefer fully specified requirements
You'd rather not talk to customers
Infrastructure cost isn't something you think about
Why join
Join a fast-growing team and work with world-leading labs and researchers
Full ownership and autonomy — greenfield architecture, no legacy platform to work around
Rare scope for the level: architecture, cost, and customer-facing technical decisions all yours
Small team, high leverage, direct access to leadership
Early in an emerging domain — the standards here aren't set yet
Responsibilities
- Own the data pipeline end to end: ingest → validate → transcode → package → deliver
- Build automated quality gates that turn written specifications into executable checks
- Design and version the delivery schema; maintain backward compatibility as specs change
- Own storage architecture and cost at PB scale
- Build observability and lineage for any delivered file
- Interface directly with customer engineering on format, validation, and defect resolution
- Scale the system 5× without a rewrite
Qualifications
- 5+ years in production data infrastructure, 2+ at senior/staff scope, still hands-on
- Experience with multimodal sensor data pipelines at scale
- Knowledge of robotics/AV data formats
- Experience in large-scale video processing
- Cloud storage and cost engineering at TB–PB scale
- Built validation systems, not just run them
- Direct technical communication with external stakeholders
Benefits
- Join a fast-growing team and work with world-leading labs and researchers
- Full ownership and autonomy
- Rare scope for the level: architecture, cost, and customer-facing technical decisions
- Small team, high leverage, direct access to leadership
- Early in an emerging domain
Skills mentioned
About DeepReach AI
DeepReach is the platform where entrepreneurs build local data businesses serving Physical AI. We enable entrepreneurs to launch and grow independent data businesses in their own communities. DeepReach provides the hardware, software, quality infrastructure, operational support, customer demand, and payments. Entrepreneurs build the businesses. They recruit local experts, develop trusted relationships with local organizations, and continuously expand into new industries, professions, and real-world environments. Every new business unlocks access to new communities, skills, and workplaces. Together, they create the diverse real-world data needed by robot foundation models, world models, and other Physical AI systems. Our vision is to preserve and scale humanity’s physical knowledge by building the world’s largest platform for independent data businesses, enabling every AI system to learn from the collective experience of human work.