Software Engineer
About the role
Role Overview
Build and operate the autolabeling pipeline that accelerates human annotation throughput for vehicle attribute classification tasks.
Develop a pipeline using foundation models such as Gemini, SigLIP, and CLIP to pre-label tasks for human reviewers to verify and correct.
Own pipeline engineering, including ingesting queued tasks from the annotator service, calling foundation-model APIs at fleet scale, parsing structured predictions, and writing pre-labels back into the labeling workflow.
Partner closely with the team lead, ML engineers, and data infrastructure team to integrate the pipeline with existing Zoox systems.
Responsibilities
Build the autolabeling pipeline to ingest queued annotation tasks, dispatch them to foundation-model APIs, parse structured outputs, and write pre-labels back to the labeling workflow.
Build the observability layer covering per-task latency, per-model cost, per-attribute coverage, and error-mode dashboards.
Set up and execute experiments designed by the team lead and collect outputs in formats suitable for ML engineers to analyze.
Integrate the pipeline with existing Zoox systems in partnership with the data infrastructure team.
Document the system, write runbooks, and ensure a clean handoff at the end of the engagement.
Required Qualifications
3+ years of backend or data pipeline engineering experience.
Strong Python skills with comfort in C++.
Large-dataset experience using PySpark or equivalent.
Understanding of ML fundamentals including model inference, embeddings, structured output, precision, recall, and calibration.
Ability to reason about ML data shapes and integration patterns.
Experience integrating foundation models such as Gemini, OpenAI, or Anthropic at production scale.
Excellent written communication skills for design documents and runbooks.
Bonus Qualifications
Databricks experience.
End-to-end ML pipeline stewardship from data ingest through inference and monitoring.
Experience with annotation tooling or human-in-the-loop ML workflows.
Experience with autonomous-systems data pipelines.
AWS experience, especially S3, ECS/EKS, and Lambda.
Experience working in a shared codebase with ML engineers, including proto schemas and joint deployments.
EEO: All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Cell phones are not allowed at the client site, shooting of photos and/or videos are strictly prohibited while inside the client facility.
Responsibilities
- Build the autolabeling pipeline to ingest queued annotation tasks, dispatch them to foundation-model APIs, parse structured outputs, and write pre-labels back to the labeling workflow.
- Build the observability layer covering per-task latency, per-model cost, per-attribute coverage, and error-mode dashboards.
- Set up and execute experiments designed by the team lead and collect outputs in formats suitable for ML engineers to analyze.
- Integrate the pipeline with existing Zoox systems in partnership with the data infrastructure team.
- Document the system, write runbooks, and ensure a clean handoff at the end of the engagement.
Qualifications
- 3+ years of backend or data pipeline engineering experience.
- Strong Python skills with comfort in C++.
- Large-dataset experience using PySpark or equivalent.
- Understanding of ML fundamentals including model inference, embeddings, structured output, precision, recall, and calibration.
- Ability to reason about ML data shapes and integration patterns.
- Experience integrating foundation models such as Gemini, OpenAI, or Anthropic at production scale.
- Excellent written communication skills for design documents and runbooks.
Skills mentioned
About LanceSoft, Inc.
Established in 2000, LanceSoft is a pioneer in delivering top-notch Global Workforce Solutions and IT Services to a diverse clientele. We pride ourselves on fostering global cross-cultural connections that advance both the careers of our employees and the success of our clients' businesses. At LanceSoft, our mission is clear: to leverage our global network to seamlessly connect businesses with the right talent and individuals with the right opportunities, all without bias. We believe in providing Global Workforce Solutions with a personalized, human touch. Our comprehensive range of services spans various domains, encompassing temporary and permanent staffing, Statement of Work (SOW) arrangements, payrolling, Recruitment Process Outsourcing (RPO), application design and development, program/project management, and engineering solutions. Currently, our team of over 5,000 professionals caters to 110+ enterprise clients worldwide, including Fortune companies. Our client base represents a diverse spectrum of industries, including Banking & Financial Services, Semiconductor/VLSI, Technology, Healthcare & Life Sciences, Government, Telecom & Media, Retail & Distribution, Oil & Gas, and Energy & Utilities. Headquartered in Herndon, VA, LanceSoft operates 32+ regional offices across the North America, Europe, Asia, and Australia. We also have nine delivery centers strategically located in India in Bangalore, Indore, Noida, Baroda, Hyderabad, Bhubaneshwar, Dehradun, Goa, and Aligarh to further enhance our client service capabilities.