AI Engineer
About the role
About the role
We're looking for a founding AI Engineer to own the intelligence layer of Helm: the analyst founders talk to, the agents that work in the background, and the evaluation harness that keeps all of it honest.
Helm's AI is built on one rule: deterministic math, narrated by AI. Every metric is computed in SQL or a typed tool, the model only explains, and a guardrail refuses any figure it cannot trace to a tool result. You'll own that architecture end to end, from how the analyst decides which of roughly ninety tools to call, to how a citation is rendered, to how we know a prompt change made things better rather than worse.
You'll set the bar for how Helm builds with language models: the tool contracts, the prompt and model registries, the eval suites, the tracing, the cost discipline. You'll work directly with the founding team and be trusted to make independent technical decisions from day one.
The person who'll succeed in this role has shipped LLM features to real users, has been burned by a model that sounded confident and was wrong, and has strong opinions about how to keep that from happening again.
What you'll do
Own the AI analyst: the agentic chat loop, tool selection, grounding and citation guardrails, refusal behavior, and the streaming experience in the product, in Slack, and over MCP.
Design and extend Helm's typed tool surface. Every tool is a bounded, tenant-scoped, read-only contract with citations attached. Add new tools, retire bad ones, and keep the surface small enough that the model uses it well.
Build and run the evaluation program: golden sets, chat evals, drift and regression checks in CI, and the judgment calls about what "good" means for a founder asking about their runway. Nothing ships to the analyst without evidence it works.
Build agents. Helm runs bounded, preview-and-approve background agents that plan, gather, reduce deterministically, and hand a founder something to approve. Ship the next ones and harden the framework underneath them.
Own model and prompt operations: the model registry and adapter seam, prompt versioning, prompt caching, token budgets, per-tenant limits, and cost per conversation.
Instrument everything. Every AI call is traced with its prompt, tools, tokens, cache hit, and cost. Keep that observability sharp and use it to find the next thing to fix.
Push the frontier where it earns its keep: document and spreadsheet inference for ingestion, meeting intelligence, retrieval over connected context, and whatever the next model release makes possible.
Protect the customer. PII redaction before anything reaches a model, jailbreak resistance, and the habit of writing the cross-tenant test before the feature.
Join customer conversations, hear how founders actually ask about their numbers, and turn it into analyst improvements that ship the same week.
Who we're looking for
An evaluation mindset. You measure before and after, you build golden sets, and you treat "the demo looked good" as the start of the work, not the end.
Judgment about scope. You know when a bigger model, a longer prompt, or a new agent is the answer, and when a deterministic function is.
Security instincts. You understand why a chat product over financial data has a different threat model than a chat product over docs, and you design for it.
Comfort working from ambiguity. You can take a problem like "founders don't trust the churn answer," find out why, and ship the fix.
A bias toward shipping. You can move quickly without lowering the bar.
Qualifications
Production experience with LLM applications: tool calling, agent loops, structured outputs, streaming, and the failure modes that only show up with real users.
Strong TypeScript. Helm is all-in on TypeScript (Next.js, Vercel AI SDK, Anthropic API, Postgres). Python is welcome for evals and analysis, but the product ships in TypeScript.
Experience building evaluation suites for LLM features, including golden datasets and regression checks in CI.
Comfort with data. You can write the SQL a tool needs, reason about a metric's definition, and notice when a number doesn't reconcile.
Experience with LLM observability and cost control: tracing, prompt caching, token budgets, and rate limits.
Fluency with AI coding tools and agents (Claude Code, Codex, and Cursor). Your judgment stays in charge of what ships.
Life at Helm
Ground floor, for real. You'll be Helm's first AI hire, working directly with the founders. The analyst you build is the product our customers talk to every day.
A role with a future. The role begins as a paid contract, with the opportunity to convert to a full-time Founding AI Engineer position.
Room to grow. As Helm scales, so will your role, responsibilities, and compensation.
Flexibility that works. Remote friendly.
How to apply
Send us something you built with language models that real people used, with live links or a short write-up. Include a few sentences on the worst thing a model did in production on your watch and how you fixed it. Tell us why building an AI CFO interests you.
Responsibilities
- Own the AI analyst and the agentic chat loop
- Design and extend Helm's typed tool surface
- Build and run the evaluation program
- Build agents for background tasks
- Own model and prompt operations
- Instrument AI calls for observability
- Document and spreadsheet inference for ingestion
- Protect customer data with security measures
Qualifications
- Production experience with LLM applications
- Strong TypeScript skills
- Experience building evaluation suites for LLM features
- Comfort with data and SQL
- Experience with LLM observability and cost control
- Fluency with AI coding tools and agents
Benefits
- Ground floor opportunity as the first AI hire
- Potential to convert to a full-time position
- Room for growth as the company scales
- Remote friendly work environment
Skills mentioned
About Helm CFO
Helm is a platform that helps founders secure funding and run their financial operations as they grow. We do this with an AI-native solution that organizes their financial data, preps them for investor meetings, and delivers strategic recommendations with our AI CFO.