Staff AI Engineer

Harnham
San Francisco, CAFull-time$280,000–$280,000Posted Aug 26, 2026

About the role

Title: Staff AI Engineer

Location: San Francisco, CA (Hybrid)

Compensation: Up to $280,000 + Equity

We're partnered with a high-growth, mission-driven SaaS company transforming how businesses build and maintain trust, with AI at the core of their next phase of innovation. The platform is redefining how critical enterprise workflows are automated, particularly in areas where reliability, auditability, and security are essential.

This is a high-impact role where you will help define how AI is architected across the company. You won't just be building features. You'll make foundational decisions around systems, evaluation, and long-term technical direction, working across LLMs, retrieval systems, and agent-based workflows in production environments.

What You'll Do

Design and own production AI systems end-to-end, including LLM pipelines, retrieval systems, and orchestration layers

Build and scale RAG systems, reranking pipelines, and vector-based search infrastructure

Define evaluation frameworks to measure retrieval quality, reasoning accuracy, and system performance

Analyze production behavior, identify failure modes, and drive improvements based on data

Make key architectural decisions across model infrastructure, tooling, and workflows

Partner closely with product, platform, and domain teams to translate complex requirements into scalable systems

Lead best practices for building reliable, observable, and cost-efficient AI systems

Requirements

10+ years of software engineering experience, including 3+ years working on ML or AI systems

Proven experience owning and deploying production LLM systems

Strong background in RAG, embeddings, reranking, and vector databases (e.g., Pinecone, FAISS, Chroma)

Experience designing evaluation systems and improving models through quantitative analysis

Strong Python skills, with solid software engineering fundamentals

Experience making architectural decisions that influence team or org direction

Strong understanding of production systems, including reliability, observability, and cost tradeoffs

Ability to break down ambiguous problems and operate with a high degree of ownership

Clear communication skills and experience working cross-functionally

Nice to Have

Experience in regulated domains such as compliance or security

Familiarity with data platforms or analytics tooling

Experience with orchestration frameworks (e.g., Temporal, Airflow)

Exposure to LLM evaluation platforms or tooling

Contributions to open source, research, or technical communities

If you're interested in shaping how AI systems are built, evaluated, and deployed in high-trust environments, this is an opportunity to have direct influence on both technical direction and real-world impact at a fast-growing company.

Responsibilities

  • Design and own production AI systems end-to-end, including LLM pipelines, retrieval systems, and orchestration layers
  • Build and scale RAG systems, reranking pipelines, and vector-based search infrastructure
  • Define evaluation frameworks to measure retrieval quality, reasoning accuracy, and system performance
  • Analyze production behavior, identify failure modes, and drive improvements based on data
  • Make key architectural decisions across model infrastructure, tooling, and workflows
  • Partner closely with product, platform, and domain teams to translate complex requirements into scalable systems
  • Lead best practices for building reliable, observable, and cost-efficient AI systems

Qualifications

  • 10+ years of software engineering experience, including 3+ years working on ML or AI systems
  • Proven experience owning and deploying production LLM systems
  • Strong background in RAG, embeddings, reranking, and vector databases (e.g., Pinecone, FAISS, Chroma)
  • Experience designing evaluation systems and improving models through quantitative analysis
  • Strong Python skills, with solid software engineering fundamentals
  • Experience making architectural decisions that influence team or org direction
  • Strong understanding of production systems, including reliability, observability, and cost tradeoffs
  • Ability to break down ambiguous problems and operate with a high degree of ownership

Benefits

  • Equity

Skills mentioned

PythonSystem DesignRetrieval-Augmented GenerationLarge Language ModelsEmbeddingsAI AgentsModel EvaluationModel DeploymentMLOpsData Analysis

About Harnham

Harnham provides specialist Data and AI recruitment and staffing services, along with bespoke training solutions, across multiple industry verticals, operating in the UK, the USA and EU - contact us today to discuss your requirements: info@harnham.com Our recruitment and talent teams cover all aspects of the data and AI pipeline, from collection to consumption, across multiple data roles and functions. Whether you need full-time staff, contract talent, specialized training, Data-qualified graduates, or C-suite executives, Harnham Group is equipped to fulfill all your data talent requirements. Our five core services: * ATD - Rockborne – our graduate development arm – deploys expertly trained data consultants, who have gone through an intensive 12-week data training programme. After two years in the scheme, your consultant could become a * Contract / C2C / Freelance: Whether you're addressing a talent shortage, augmenting a project team, or encountering resource gaps, our specialist consultants offer bespoke interim talent solutions to address your unique challenges and drive success. * Full-Time / Direct Hiring: From early career professionals to senior management, our comprehensive services enable you to unearth outstanding data talent across various specializations, ensuring your organization thrives in the data-driven era. * Executive Search: With our dedicated executive search team and an extensive network encompassing the director to C-suite level, we assist both global corporations and ambitious startups in securing top-tier leadership talent to accomplish their goals. * GenAI, Prompt, LLM Training: Elevate your team's data expertise with our all-encompassing training programs covering essential tools and technologies such as Python, SQL, Machine Learning, LLM’s and AI. Highly bespoke, customised to your team learning requirements and budgets.

Staffing and Recruiting201-500 employeesWimbledon, London