AI Engineer

Noösphere
Seattle, Washington, United StatesFull-timePosted Sep 15, 2026

About the role

About the Role

Human-centered physical intelligence starts with models that learn from how people actually do things in the real world. Our systems run in homes and workplaces with people doing real tasks, and every deployment returns hours of continuous, time-synchronized multimodal data. Nothing comparable exists publicly, and the models and benchmarks for this kind of data have yet to be built.

We're looking for an experienced research engineer to be a driving force behind taking this work into production: the training, eval, and data infrastructure our research runs on, the path from experiments to systems operating in the real world, and the inference latency, reliability, and telemetry that make them useful. You'll also be a primary owner of the privacy and governance layer that sits between real-world data and our training pipelines: every deployment involves real people, and the systems that protect them are as much a part of this role as the models. You'll work on the same stack as our research scientists, from data engine to deployed models, with your center of gravity on what ships. Research moves the company forward, and engineering is what turns it into systems people can rely on.

What You'll Do

Take models to production

Drive the path from a trained checkpoint to a deployed system — export, quantization, distillation, and inference optimization across edge and cloud targets under real-time constraints.

Build and operate the serving and streaming infrastructure that runs models over continuous multimodal input, in real time and offline over recorded data.

Build the telemetry that measures how models behave in deployment and feeds failures and edge cases back into the data engine.

Set the engineering bar for the ML stack: reproducible training runs, versioned datasets and models, tested pipelines, and CI for model changes.

Build the tools and abstractions that let the team experiment quickly without hitting infrastructure walls.

Build and run the data engine

Design and operate the full data lifecycle for large volumes of multimodal, time-synchronized data captured from real hardware on real-world tasks.

Build the tooling that tightens the loop between deployment, data, and training.

Grow the engine along two axes at once: raw volume and domain coverage, so our models generalize across the environments, tasks, embodiments, and people they'll encounter.

Turn signal from real-world use into supervision that existing datasets don't offer.

Build privacy and data governance

Build and operate the automated privacy layer that every recording passes through before it reaches training: PII detection and redaction across video, audio, and text (faces, voices, bystanders, license plates, screens, documents), de-identification and pseudonymization, and quality checks that verify redaction actually held.

Enforce consent, purpose limitation, and data minimization in code: consent-scoped access controls, retention and deletion pipelines that honor withdrawal, and audit trails that show what data was used to train what.

Apply privacy-preserving ML where it fits, from on-device and edge redaction before upload to differential privacy and secure aggregation for sensitive signals.

Partner with counsel and research on the controls behind our privacy-by-design commitments, including GDPR, CCPA/CPRA, biometric-privacy laws such as BIPA, and Washington's My Health My Data Act.

Stand up and run the eval suite

Build the eval suite as a repeatable, versioned system the whole team runs against, covering the capabilities we care about over real-world input.

Build novel evals for tasks without an established benchmark, with metrics tied to real-world usefulness.

Extend evals beyond offline benchmarks to online evaluation of how models actually behave in deployment, and turn results into concrete priorities for data and training.

Train and post-train models

Post-train multimodal foundation models on our data — SFT, RL, and preference optimization — and run the experiments that isolate what improves performance.

Run scaling-law studies on how capability moves with data volume, diversity, model size, and compute, and use them to shape the data and infrastructure roadmap.

What We're Looking For

5+ years building and shipping ML systems in production, with a record of models you trained or deployed reaching real users.

Strong software engineering fundamentals: you write production-grade Python and systems code, own services end to end, and are comfortable building infrastructure when the work calls for it.

Deep experience taking models into production — on-device or real-time inference on compute-constrained hardware, quantization, serving, monitoring, and the tradeoffs that come with tight latency and compute budgets.

Hands-on experience building data pipelines for model training at scale: you've owned ingestion, curation, annotation, and quality on real datasets, and you know where data problems hide.

Experience building privacy-preserving data pipelines: PII detection and redaction across video, audio, and text, de-identification, consent and retention enforcement, and the controls that keep sensitive data out of training. You treat privacy as an engineering requirement with tests, not a policy document.

Hands-on experience training and post-training multimodal models (vision, video, audio, language) — SFT, RLHF, DPO or GRPO.

Experience building eval harnesses and benchmarks that a team can run repeatably, and clear opinions about what makes an eval trustworthy.

Practical fluency with AI coding tools and agents as part of your daily workflow, with good judgment about when to lean on them.

Startup DNA: high ownership, comfort with ambiguity, and the judgment to make pragmatic calls on a small, flat, collaborative team.

Stack & Skills

Fluent

PyTorch and the modern training stack — distributed training, mixed precision, experiment tracking, and reproducible pipelines

Model deployment and optimization — ONNX, TensorRT, Triton, vLLM, or similar inference stacks; quantization, distillation, and profiling on edge and real-time targets

Data engineering for ML — large-scale multimodal data processing, dataset versioning, annotation tooling, and quality/coverage measurement

Privacy engineering for ML data — PII detection and redaction (face, voice, text, and document), de-identification and pseudonymization, consent-scoped access control, retention and deletion, and audit logging

Python, plus systems code (C++, Rust, or Go) where a pipeline or serving path needs it

Cloud infrastructure (AWS or GCP), containers, Kubernetes, and orchestration for training, data, and inference jobs

Streaming and real-time inference over continuous sensor input

Working knowledge

Multimodal and vision-language models (VLMs) — architectures, post-training recipes (SFT, RLHF, DPO/GRPO), and RL environments

Eval methodology — benchmark design, held-out set construction, statistical rigor, and human preference data

Video understanding — temporal segmentation, action recognition, and long-horizon reasoning

Speech and audio models — streaming ASR and diarization in real-world conditions

Retrieval — multimodal embeddings and vector search

MLOps tooling — model registries, feature and dataset stores, experiment tracking at scale, and CI/CD for models

Privacy-preserving ML — differential privacy, secure aggregation, federated or on-device learning, and edge-side redaction before upload

Privacy and data-protection regimes as they apply to ML — GDPR, CCPA/CPRA, BIPA and other biometric-privacy laws, and Washington's My Health My Data Act

Bonus

Experience building physical AI data engines — the pipelines, annotation systems, and quality loops behind large real-world datasets

Experience building or running human data-collection programs, including consent and privacy handling

Embedded or edge deployment on constrained hardware, including firmware-adjacent debugging

Experience with active learning, data selection, or scaling-law studies

Published or open-source work grounded in real-world deployment

Cross-embodiment learning, world models, or sim-to-real transfer

Learning from demonstration or policy learning over multimodal sensor streams — vision-language-action models, diffusion policies, or behavior cloning from teleoperation data

Simulation environments — Isaac, MuJoCo, or similar — for training, evaluation, or data generation

Robot data infrastructure and open ecosystems — ingesting and time-aligning multi-camera, proprioceptive, and joint-state streams (ROS bags, MCAP, LeRobot datasets, Open X-Embodiment)

Policy deployment and evaluation on real robots — real-time inference loops, action chunking, safety monitors, task suites, and rollout infrastructure

Who You'll Work With

You'll be working with a tightly knit team that has pioneered frontier AI research, shipped tens of millions of devices, scaled infrastructure used by millions every day, and defined how people and machines interact in the physical world. They've done it at Meta, Amazon, Microsoft, and Valve.

Why Join

Early role at a funded startup founded by the team that built one of the most advanced physical AI data engines in the industry.

Drive the path from research to production for models no one has trained before, on real-world multimodal data that no lab has.

Your work is how research reaches people in the real world.

Set the privacy bar for physical AI. Real-world data is only usable if the people in it are protected, and you'll build the systems that make that true.

Work across the whole loop: data, evals, training, and inference.

Small team, exceptional peers, no bureaucracy.

Compensation: Competitive salary plus meaningful early-stage equity and benefits.

All applicants must be legally authorized to work in the United States. Noösphere will consider sponsorship of qualified candidates for employment visa status where required. Noösphere is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Responsibilities

  • Take models to production
  • Drive the path from a trained checkpoint to a deployed system
  • Build and operate the serving and streaming infrastructure
  • Build the telemetry that measures how models behave in deployment
  • Set the engineering bar for the ML stack
  • Design and operate the full data lifecycle for large volumes of multimodal data
  • Build and operate the automated privacy layer
  • Build the eval suite as a repeatable, versioned system

Qualifications

  • 5+ years building and shipping ML systems in production
  • Strong software engineering fundamentals
  • Experience with Python and systems code
  • Familiarity with cloud infrastructure and containers

Benefits

  • Competitive salary
  • Meaningful early-stage equity
  • Inclusive environment

Skills mentioned

PythonPyTorchModel DeploymentMLOpsModel EvaluationData EngineeringPerformance OptimizationKubernetesDockerAWS

About Noösphere

Noösphere is building human-centered physical intelligence, and is backed by institutional investors including Trilogy and Madrona.

Artificial Intelligence11-50 employeesSeattle, WA