Staff AI Engineer
About the role
Title: Staff AI Engineer
Location: San Francisco, CA (Hybrid)
Compensation: Up to $280,000 + Equity
We're partnered with a high-growth, mission-driven SaaS company transforming how businesses build and maintain trust, with AI at the core of their next phase of innovation. The platform is redefining how critical enterprise workflows are automated, particularly in areas where reliability, auditability, and security are essential.
This is a high-impact role where you will help define how AI is architected across the company. You won't just be building features. You'll make foundational decisions around systems, evaluation, and long-term technical direction, working across LLMs, retrieval systems, and agent-based workflows in production environments.
What You'll Do
Design and own production AI systems end-to-end, including LLM pipelines, retrieval systems, and orchestration layers
Build and scale RAG systems, reranking pipelines, and vector-based search infrastructure
Define evaluation frameworks to measure retrieval quality, reasoning accuracy, and system performance
Analyze production behavior, identify failure modes, and drive improvements based on data
Make key architectural decisions across model infrastructure, tooling, and workflows
Partner closely with product, platform, and domain teams to translate complex requirements into scalable systems
Lead best practices for building reliable, observable, and cost-efficient AI systems
Requirements
10+ years of software engineering experience, including 3+ years working on ML or AI systems
Proven experience owning and deploying production LLM systems
Strong background in RAG, embeddings, reranking, and vector databases (e.g., Pinecone, FAISS, Chroma)
Experience designing evaluation systems and improving models through quantitative analysis
Strong Python skills, with solid software engineering fundamentals
Experience making architectural decisions that influence team or org direction
Strong understanding of production systems, including reliability, observability, and cost tradeoffs
Ability to break down ambiguous problems and operate with a high degree of ownership
Clear communication skills and experience working cross-functionally
Nice to Have
Experience in regulated domains such as compliance or security
Familiarity with data platforms or analytics tooling
Experience with orchestration frameworks (e.g., Temporal, Airflow)
Exposure to LLM evaluation platforms or tooling
Contributions to open source, research, or technical communities
If you're interested in shaping how AI systems are built, evaluated, and deployed in high-trust environments, this is an opportunity to have direct influence on both technical direction and real-world impact at a fast-growing company.
Responsibilities
- Design and own production AI systems end-to-end, including LLM pipelines, retrieval systems, and orchestration layers
- Build and scale RAG systems, reranking pipelines, and vector-based search infrastructure
- Define evaluation frameworks to measure retrieval quality, reasoning accuracy, and system performance
- Analyze production behavior, identify failure modes, and drive improvements based on data
- Make key architectural decisions across model infrastructure, tooling, and workflows
- Partner closely with product, platform, and domain teams to translate complex requirements into scalable systems
- Lead best practices for building reliable, observable, and cost-efficient AI systems
Qualifications
- 10+ years of software engineering experience, including 3+ years working on ML or AI systems
- Proven experience owning and deploying production LLM systems
- Strong background in RAG, embeddings, reranking, and vector databases (e.g., Pinecone, FAISS, Chroma)
- Experience designing evaluation systems and improving models through quantitative analysis
- Strong Python skills, with solid software engineering fundamentals
- Experience making architectural decisions that influence team or org direction
- Strong understanding of production systems, including reliability, observability, and cost tradeoffs
- Ability to break down ambiguous problems and operate with a high degree of ownership
Benefits
- Equity
Skills mentioned
About Harnham
Harnham provides specialist Data and AI recruitment and staffing services, along with bespoke training solutions, across multiple industry verticals, operating in the UK, the USA and EU - contact us today to discuss your requirements: info@harnham.com Our recruitment and talent teams cover all aspects of the data and AI pipeline, from collection to consumption, across multiple data roles and functions. Whether you need full-time staff, contract talent, specialized training, Data-qualified graduates, or C-suite executives, Harnham Group is equipped to fulfill all your data talent requirements. Our five core services: * ATD - Rockborne – our graduate development arm – deploys expertly trained data consultants, who have gone through an intensive 12-week data training programme. After two years in the scheme, your consultant could become a * Contract / C2C / Freelance: Whether you're addressing a talent shortage, augmenting a project team, or encountering resource gaps, our specialist consultants offer bespoke interim talent solutions to address your unique challenges and drive success. * Full-Time / Direct Hiring: From early career professionals to senior management, our comprehensive services enable you to unearth outstanding data talent across various specializations, ensuring your organization thrives in the data-driven era. * Executive Search: With our dedicated executive search team and an extensive network encompassing the director to C-suite level, we assist both global corporations and ambitious startups in securing top-tier leadership talent to accomplish their goals. * GenAI, Prompt, LLM Training: Elevate your team's data expertise with our all-encompassing training programs covering essential tools and technologies such as Python, SQL, Machine Learning, LLM’s and AI. Highly bespoke, customised to your team learning requirements and budgets.