Senior Software Engineer - Technical Advisor

RJI Search
United StatesContractPosted Aug 27, 2026

About the role

Our well-know AI client is seeking exceptional Software Engineers - Technical Advisors who want to spend a focused period doing something most engineers never get to do: find out exactly where frontier models break on real engineering work.

You will take on hard problems across production-grade codebases, then work out precisely where and why model-generated solutions fall short. You will also build the hard problems those models are tested on. Your judgment about what separates good engineering from plausible engineering is the entire point of the role.

What you'll do

Evaluate Agent Sessions: Analyze how AI coding agents behave over entire coding sessions—tracking what the agent investigates, verifies, assumes, and leaves undone.

Review Model-Written Pull Requests: Audit model-generated pull requests against real production repositories, documenting every identified issue alongside its severity and detailed technical rationale.

Build Evaluation Benchmarks: Design and construct hard, container-based problems that serve as rigorous test benchmarks for frontier models.

Collaborate with AI Researchers: Work directly alongside AI researchers on frontier problems, producing clear written analyses that explain root-cause failures and boundary-condition gaps.

Maintain High Written Rationale Standards: Author original, clear technical rationale for all evaluations (all written deliverables must be independently authored without AI text generation, though AI tools are welcomed for codebase exploration and running test suites).

This isn't a typical "write code and ship features" role. The client is focused on evaluating how frontier AI coding models perform on real, production-grade engineering work, and they need experienced engineers who can act as a high-bar technical authority: reviewing model-generated code, figuring out precisely where and why it breaks, and explaining that reasoning clearly in writing. Your judgment about what separates genuinely good engineering from plausible-but-superficial engineering is the entire point of the role.

What you'll actually do:

  • Review model-generated pull requests against real production repositories, documenting every issue you find along with its severity and your technical rationale.
  • Evaluate full AI coding-agent sessions, tracking what the model investigated, verified, assumed, or skipped.
  • Design and build container-based (Docker) test problems that serve as benchmarks for evaluating frontier models.
  • Write clear, original technical rationale explaining why code fails. All written deliverables must be self-authored, no AI text generators for your written analysis, though AI tools are welcome for exploring codebases or running test suites.
  • Collaborate asynchronously with Anthropic's AI researchers to share findings and refine evaluation criteria.

Who this is a strong fit for:

  • 8+ years of hands-on production engineering experience (Senior, Staff, Principal, Tech Lead, or Open-Source Maintainer backgrounds all fit well)
  • Real experience in codebases with a genuine code review culture, whether at a startup, big tech, or open source
  • Comfortable dropping into unfamiliar languages and codebases on short notice
  • Genuinely comfortable with Docker, git, and the command line to reproduce and debug issues locally
  • Can work across backend, frontend, APIs, data, testing, or dev tooling
  • Write clearly enough to explain why something is broken, not just how to fix it fast

Not the right fit if: you're an Engineering Manager/Director who's stepped away from hands-on code review, your experience is limited to no-code/low-code or single-framework MVP work, you're a QA tester or non-technical prompt engineer, or you don't have real Docker and terminal/CLI experience.

Contract type: W2 Contractor. This role cannot support C2C arrangements or visa sponsorship, so you'll need to confirm US remote work eligibility.

Who you are

Are comfortable dropping into unfamiliar codebases and languages you don't use every day?

the repos change weekly!

Are comfortable with containers and reproducing results locally?

Can explain your reasoning as clearly as you can write the solution?

Are more interested in why a solution is right than in shipping it quickly?

Want to spend a focused period on this rather than committing to a permanent role?

Daily tasks

Solve difficult engineering problems across real, production-grade codebases

Identify where model-generated code fails, and articulate precisely why

Work directly with researchers on problems they are actively investigating

Hold a technical bar that others build on

Required skills

Have deep, demonstrated expertise as a software engineer - we care about the depth of your judgment, not your years!

Write code that other strong engineers learn from

Have reviewed a lot of other people's code, and are known for catching what CI and the author both missed

Candidates must do a CodeSignal Industry Coding Assessment and score 500+ to be considered

Responsibilities

  • Evaluate Agent Sessions: Analyze AI coding agents' behavior over coding sessions.
  • Review Model-Written Pull Requests: Audit model-generated pull requests against production repositories.
  • Build Evaluation Benchmarks: Design and construct container-based problems for testing models.
  • Collaborate with AI Researchers: Work with researchers on frontier problems and produce analyses.
  • Maintain High Written Rationale Standards: Author clear technical rationale for evaluations.

Qualifications

  • 8+ years of hands-on production engineering experience.
  • Experience in codebases with a genuine code review culture.
  • Comfortable with Docker, git, and command line.
  • Ability to work across backend, frontend, APIs, data, testing, or dev tooling.
  • Strong written communication skills.

Skills mentioned

DebuggingGitLinuxDockerContainerizationSoftware TestingIntegration TestingAPI TestingBackend DevelopmentSystem Design

About RJI Search

RJI delivers exceptional opportunities to job seekers and employers by offering a personal touch, regardless of scale. We don’t match résumés to jobs. We take the time to learn about candidates, both professionally and personally, so we can make the best match based on skills, goals, and company culture. The same goes with our Clients. We learn about your organization or business, where you are, and where you are going, and then we find the ideal candidate! RJI is committed to equal employment opportunity and is dedicated to Diversity, Equity, and Inclusion exceeding compliance requirements. Recruiting and retaining a diverse workforce of all backgrounds and perspectives contributes to the overall success of our clients . By helping to foster an equitable culture of inclusivity and belonging, we assist our clients in maintaining an environment in which all staff feel welcomed, valued, and engaged in their work.

Staffing and Recruiting2-10 employees