Staff AI Engineer - LLM Systems & Evaluation

Altera AI, Inc.
Remote (US)Full-time$200,000–$250,000Posted Aug 26, 2026

About the role

Staff AI Engineer — LLM Systems & Evaluation

Location: Remote (US) · Core working hours aligned with 9am–5pm Central

Type: Full-time

Compensation: $200,000–$250,000 + meaningful early equity

Benefits: Health, dental and vision; PTO; 401(k)

Experience: 7+ years, or equivalent demonstrated depth

About Altera

Most engineers want to join a rocket ship. We want engineers who can build one.

Altera is building an AI litigation associate—not another chatbot or copilot bolted onto a document editor. We’re building a system that takes a real case file, works through it the way a great litigation associate would, and produces work that can actually be filed in court: the right arguments, the right record, the right local rules, and every factual assertion and citation verified before an attorney puts their name on it.

We’re already live, with real lawyers using Altera to produce real court filings. Now we’re scaling. Over the next year, we’re going from two jurisdictions to national coverage of the civil litigation market. Small and mid-size firms are our entry point, not our ceiling.

We’re building in a market with large, well-funded competitors. They have more people and more capital. We have a small team, a product already being used in real litigation, and the conviction that we can build something better. If that sounds like the kind of fight you want to be part of, keep reading.

The Role

What would you build differently if you knew thousands of law firms could eventually depend on the decisions your AI system makes today?

You’d be one of our first five engineers, and the first hired specifically to own Altera’s AI systems end to end. You’re not stepping into a mature AI stack where the important decisions have already been made. You’ll have a disproportionate hand in deciding how Altera reasons, retrieves evidence, drafts legal work, verifies itself, and proves that its output is good enough to file.

Today we run two orchestration approaches side by side: a deterministic pipeline with attorney checkpoints and an agentic lane with constrained tools and hard policy gates. You’ll build both out. You’ll own the prompts, retrieval, grounding, model strategy, agent behavior, evaluation systems, and quality gates behind them. And we’ve committed to letting measured quality—not architectural preference—decide which approach survives. You can’t retire a lane you can’t measure, so you’ll build the instrument that makes that question answerable, and your evidence will drive the decision.

The hardest problem isn’t getting a model to produce persuasive legal prose. It’s knowing when the prose is actually right. A fabricated fact, unsupported assertion, or bad citation that survives into a court filing is a product failure. We need to detect those failures automatically, measure how often we miss them, and be able to state the quality of the system plainly to the attorney putting their name on the filing. Doing that reliably at production quality is still an open problem across the industry, and it’s at the center of this role.

You’ll establish these foundations while they’re still yours to shape: how agents use tools, how claims are grounded against the record, how authorities are retrieved and verified, how models and prompts are evaluated, how attorney feedback becomes ground truth, and how we know a new jurisdiction is ready to ship. As Altera grows, you’ll help decide how those responsibilities evolve into dedicated AI, evaluation, retrieval, and model-quality functions—and help shape the teams that inherit what you built.

This is the ground floor. The AI architecture isn’t settled, the winning orchestration strategy hasn’t been decided, and there aren’t layers of people between you and the outcome. If we build what we believe we can build, you’ll be able to point to the intelligence at the center of Altera and say: I built that.

What you'll own

The agentic lane. The agent loop and a deliberately constrained tool surface — read the record, search case law, verify citations, write sections, validate, request attorney input — plus the policy gates that make our guarantees non-negotiable: verification is mandatory, citations must verify, an attorney confirms before anything is finishable, and the privacy boundary holds. The agent chooses the path. It doesn't choose the guarantees.

The drafting pipeline. The LLM-facing stages of the deterministic lane — extraction, legal theory, strategy, document planning, drafting, verification — and the validator suite that catches fabricated exhibits, phantom authority, and unverified citations before an attorney ever sees them.

Retrieval and grounding. Case-law search over a large public corpus and grounding against the matter record. Whether the right authority surfaces, and whether the draft is actually anchored in the file rather than in the model's priors.

Model and prompt strategy. Migrations across model generations, routing between model tiers, caching and token economics, and the experiment discipline to know whether a change actually helped.

The quality gate. A blocking gate rather than an advisory one: thresholds calibrated against measured noise, meaningful baselines, a no-regression guard, and evidence strong enough to stop a release.

The safety scorers. Faithfulness — does every factual claim in a draft trace back to the record? — and citation integrity: fictitious, mismatched, misattributed holding, conflated authority, wrong statute.

The evaluation corpus. Build the versioned ground-truth datasets behind our quality claims — real legal tasks, attorney adjudication, provenance, disagreement handling, jurisdiction coverage, and hard cases accumulated from production.

Production evaluation. Measurement over appropriately controlled real production drafts, not just an offline fixture set, with drift detection as the jurisdiction count grows.

Proving new jurisdictions. As we expand nationally, a new forum ships only when there's evidence its output is filing-grade. Expansion that outruns measurement just means shipping wrong documents faster — you own the gate that prevents it.

Privacy as a measured system. De-identification scored on false negatives — leaked PII — and false positives, where over-redaction quietly degrades the draft. Both gating. Plus red-teaming for hallucination, leakage, and prompt injection through attorney-supplied documents.

What we're looking for

7+ years in software/ML engineering, or equivalent demonstrated depth, including substantial production AI/LLM systems work — not only notebooks, and not only fine-tuning.

You've shipped agentic or tool-using LLM systems in production. Tool design, loop control, failure and retry handling, cost and latency under real load — and the judgment to know when an agent is the wrong instrument for a problem.

You've built an evaluation system someone actually trusted. Datasets, human and model-based evaluation, LLM-as-judge design and calibration, inter-rater agreement, regression detection. You know why a judge that agrees with itself isn't the same as a judge that's right.

Retrieval systems you've both built and evaluated — and the ability to tell a retrieval failure from a generation failure.

Statistical literacy on messy problems. Variance across seeds, sample size, what a three-point move on a 40-item set does and doesn't mean.

Strong Python engineering. Your code ships, runs in CI, and gets maintained by other people. This is a building role as much as a measuring one.

Fluent with AI coding tools, and accountable for what they produce. We use them heavily and expect you to. You already apply hard-nosed judgment to model output professionally — we want the same judgment applied to your own tooling: verify before you trust, and know the difference between output that looks right and output that is right.

The temperament to say the numbers don't support shipping, with evidence, to a room that wants to ship — including when it's your own work on the line.

Nice to have: legal, medical or financial domains where output errors are consequential; PII detection and NER; self-hosted inference; structured generation and schema-constrained output; open-source or published work on agents or evaluation.

Core stack is Python, Postgres, and frontier LLM APIs, with a self-hosted tracing and evaluation platform. Tooling in this space turns over every few months and we expect ours to.

You don't need a legal background — you need to be willing to sit with attorneys and learn what "wrong" means here, because it isn't obvious from the outside.

How we work

We move fast and we use AI coding tools heavily to do it — and every merge is gated on formatting, static typing, tests, migration checks, and substantive human review. Every pull request ships with a written self-review; CI rejects it otherwise. The speed comes from the tooling, not from a lower bar.

Written-first, evidence-first: claims come with data attached. Small, senior, low-ego. You'll work directly with the CTO and with practicing litigators. We protect focused work and keep meetings scarce.

Altera AI is an equal opportunity employer. We evaluate on demonstrated ability, and we will consider qualified applicants regardless of background.

Responsibilities

  • Own Altera’s AI systems end to end
  • Build and evaluate AI systems for legal work
  • Establish quality gates and evaluation systems
  • Develop retrieval and verification processes
  • Shape the AI architecture and orchestration strategy

Qualifications

  • 7+ years in software/ML engineering
  • Experience with production AI/LLM systems
  • Strong Python engineering skills
  • Experience building evaluation systems
  • Statistical literacy on messy problems

Benefits

  • Health, dental and vision insurance
  • Paid time off (PTO)
  • 401(k) plan

Skills mentioned

PythonLarge Language ModelsAI AgentsRetrieval-Augmented GenerationPrompt EngineeringTool CallingModel EvaluationStatistical AnalysisPostgreSQLModel Monitoring

About Altera AI, Inc.

Altera is building the world’s first AI Litigation Associate. Purpose-built for litigators, Altera works alongside attorneys to perform the substantive legal work behind every case. It reviews thousands of pages of records, analyzes facts and legal issues, identifies strengths and weaknesses, develops case strategy, and drafts attorney-quality work product in a fraction of the time. From pre-suit case evaluations and demand letters to complaints, motions, written discovery, summary judgment briefing, and trial preparation, Altera helps litigation teams move faster while maintaining exceptional quality and consistency. Unlike generic AI tools, Altera is designed specifically for litigation. It reasons through complex legal issues, connects facts to the applicable law, and produces work product that attorneys can review, refine, and confidently file. Founded by the team behind one of America’s fastest-growing law firms, Altera was developed inside a high-volume litigation practice where speed, precision, and results matter. Every feature has been built around the real workflows of trial lawyers, not generic productivity software. Our mission is simple: equip every law firm with an AI Associates that eliminate hours of repetitive legal work, allowing attorneys to focus on strategy, advocacy, and achieving better outcomes for their clients. The future of litigation is not more billable hours. It’s better legal work, delivered faster with an AI Associate by every attorney’s side.

Artificial Intelligence2-10 employees