Senior Software Engineer, Technical Code Evaluation and Root Cause Analysis
About the role
๐ง๐ต๐ถ๐ ๐ถ๐ ๐ป๐ผ๐ ๐ฎ ๐ฐ๐ผ๐ฑ๐ถ๐ป๐ด ๐ท๐ผ๐ฏ. ๐๐ ๐ถ๐ ๐ฎ ๐ท๐๐ฑ๐ด๐บ๐ฒ๐ป๐ ๐ท๐ผ๐ฏ. ๐ฌ๐ผ๐ ๐๐ถ๐น๐น ๐ป๐ผ๐ ๐๐ต๐ถ๐ฝ ๐ณ๐ฒ๐ฎ๐๐๐ฟ๐ฒ๐ ๐ต๐ฒ๐ฟ๐ฒ.
You will read what a coding agent did across an entire session on a real production codebase, work out exactly where its solution falls short, and write up why. To do that you need years of hands-on engineering behind you. The deliverable is your analysis, not your commits.
๐ง๐๐ ๐๐๐๐๐ก๐ง
Pave Talent is hiring on behalf of a frontier artificial intelligence (AI) research lab. This is a contract engagement through their staffing partner.
๐ง๐๐ ๐ช๐ข๐ฅ๐
- Review a coding agent's full session, not the final diff. What it investigated, what it verified, what it assumed, what it left undone
- Locate where model-generated code breaks and explain precisely why, in writing a researcher can act on
- Build the evaluation problems these models get tested against
- Reproduce builds and failures locally in containers
- Work directly with researchers on questions they are investigating right now
๐ง๐๐ ๐ฆ๐๐๐๐๐จ๐๐, ๐๐ก๐ ๐ช๐๐ฌ ๐ฃ๐๐ข๐ฃ๐๐ ๐ง๐๐๐ ๐ง๐๐๐ฆ
30 to 40 hours a week, and you choose when. Weekday evenings. Weekends. A standard Monday to Friday week if you prefer. Some people on this project keep their full-time job and run this alongside it. Others do it as a straight 40-hour week. Both work.
๐ช๐๐๐ง ๐ช๐ ๐๐ฅ๐ ๐๐ข๐ข๐๐๐ก๐ ๐๐ข๐ฅ
- 6+ years writing production code, hands-on. Not managing people who write it
- Heavy code review history. You are the person on your team who catches what continuous integration (CI) and the author both missed
- Comfort in unfamiliar repos and languages outside your daily stack. The codebases rotate weekly
- Working fluency with Docker or equivalent, reproducing builds and failures locally
- You write clearly. If you cannot explain why a solution is wrong, the analysis has no value
- United States work authorization, no sponsorship
๐๐ข ๐ก๐ข๐ง ๐๐ฃ๐ฃ๐๐ฌ ๐๐ ๐๐ก๐ฌ ๐ข๐ ๐ง๐๐๐ฆ๐ ๐๐ฅ๐ ๐ง๐ฅ๐จ๐
We would rather say this plainly than waste your time or ours.
- You have not written production code in the last two years
- You are not willing to sit a proctored coding assessment with one attempt
- You are not willing to verify your identity on video before submission
- You need someone else to interview or work on your behalf
- You are outside the United States
๐๐ข๐ช ๐ช๐ ๐ฆ๐๐ฅ๐๐๐ก
Three steps, in this order, and no exceptions for anyone.
- Technical video screening with us
- Identity verification. You record a brief video holding a government photo ID. You may mask everything except your photo and name
- A CodeSignal Industry Coding Assessment. One attempt, roughly 90 minutes to 2 hours, minimum score 500. We send prep material in advance
After onboarding there are two further assessments the lab built, over about two weeks, that determine whether you continue on the project.
๐ง๐๐ฅ๐ ๐ฆ
๐ฃ๐ฎ๐: $150 per hour, W-2. No corp-to-corp, no 1099
๐๐ผ๐๐ฟ๐: 30 to 40 per week, scheduled by you
๐๐ฒ๐ป๐ด๐๐ต: 6-month contract. Conversion to a permanent role is unlikely, and you should hear that now
๐๐ผ๐ฐ๐ฎ๐๐ถ๐ผ๐ป: Fully remote, United States only
Apply through LinkedIn. Answer the screening questions honestly. Every answer gets verified.
๐ฃ๐ฎ๐๐ฒ ๐ง๐ฎ๐น๐ฒ๐ป๐ | ๐๐ถ๐ฟ๐ถ๐ป๐ด ๐ฅ๐ฒ๐ถ๐บ๐ฎ๐ด๐ถ๐ป๐ฒ๐ฑ
Responsibilities
- Review a coding agent's full session and analyze its performance
- Locate and explain where model-generated code breaks
- Build evaluation problems for model testing
- Reproduce builds and failures locally in containers
- Collaborate with researchers on current investigations
Qualifications
- 6+ years of hands-on experience writing production code
- Strong code review history
- Comfort with unfamiliar repositories and languages
- Fluency with Docker or equivalent
- Ability to write clear explanations of technical issues
Benefits
- Flexible scheduling of 30 to 40 hours per week
- Opportunity to work alongside a full-time job
Skills mentioned
About Pave Talent
Pave Talent helps high-growth companies hire the people who set them apart. The best people are rarely looking. They're already doing great work somewhere else, often for your competitors. Finding them, earning their trust, and bringing them onto your team takes more than a job post and a database. It takes a real process, genuine outreach, and human judgment. That's all we do. We're a boutique, founder-led firm that runs modern AI tooling at every step, sourcing, matching, outreach, and screening, so we move faster than firms many times our size. But the technology is the floor, not the finish. When everyone has the same tools, the difference is still the people: the ones we find for you, and the ones who decide. Your edge isn't the AI. It's the people who run it. Since 2018 we've made 1,000+ placements across autonomous vehicles, technology and software, manufacturing, life sciences, insurance, and healthcare, for everyone from 50-person startups to the Fortune 500. Over 90% of the people we place are still there and thriving, because we hire for fit, not just a keyword match. We also believe in transparency. We tell you exactly where a search stands, what it will cost, and what to expect, the same straight talk our candidates get. Most firms hide their process. We're happy to show you ours. If you're building a team that has to be exceptional, let's talk.