Software Engineer — AI Evaluation & Automation

NLB Services
San Diego, California, United StatesContract$141,440–$162,240Posted Sep 14, 2026

About the role

Job Title: Software Engineer — AI Evaluation & Automation

Work Location: San Diego, CA

Work Arrangement: Remote

Employment Type: 06+ Month Contract

Salary/Wage Range: $68 - $78 / Hrs.

Application Deadline: Applications are accepted on an ongoing basis until the position is filled.

Job Description:

Role Overview

Help build and scale the tooling we use to measure how well AI-powered software development tools actually perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.

Key Responsibilities

  • Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
  • Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
  • Support execution-based benchmarking across quality, productivity, and eJiciency measures, including cost and latency.
  • Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
  • Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.

Required Skills & Experience

  • Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
  • Proficient in at least one general-purpose language such as Python, Java, or JavaScript — the specific language background is flexible.
  • Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
  • Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
  • Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.
  • Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
  • Hands-on experience using AI coding tools and agentic harnesses such as Claude Code, Devin, or Cursor, and command of the best practices for working with them effectively.

Preferred Experience

  • Experience designing benchmarks or evaluations for software systems, especially execution-based grading that verifies against tests.
  • Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them against human raters.
  • Experience with build-system-aware test selection, such as Bazel or mapping changed files to the tests that cover them.
  • Experience building reproducible test environments and managing versioned evaluation datasets.
  • Comfortable writing up methodology and results for engineering leadership.

How to Apply - email: sahil.shaikh@recruiter.nlbtech.com or call 904-425-1558

=====================

Equal Employment Opportunity

NLB Services is an Equal Opportunity Employer. All qualified applicants will receive consideration without regard to race, color, religion, sex (including pregnancy, sexual orientation and gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law.

Reasonable Accommodation

NLB Services provides reasonable accommodations during the recruitment process as required by law. Applicants requiring an accommodation due to a disability, pregnancy-related limitation, or sincerely held religious belief, practice or observance may contact their recruiter or email notifications@nlbservices.com.

Benefits

Based on employment status and assignment, eligible employees may receive benefits, such as Medical, Dental and Vision Insurance, Health Insurance, Disability Insurance, etc. Eligibility and benefit terms vary by assignment.

Other Important Terms

  • Applicants must be legally authorized to work in the United States.
  • Any employment offer may be contingent upon job-related screening, employment eligibility verification and other requirements permitted by applicable law.
  • Following a conditional offer of employment, candidates may be required to complete job-related background screening, which may include criminal history checks and/or drug testing, as applicable and permitted by federal, state, and local law.

Responsibilities

  • Build and integrate evaluation harnesses and automation for software development use cases.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses.
  • Validate and calibrate evaluation approaches against human judgment.
  • Support execution-based benchmarking across quality, productivity, and efficiency measures.
  • Analyze results across repeated runs and find ways to improve workflows.
  • Work with engineering and data teams to improve tooling and document evaluations.

Qualifications

  • Strong software engineering background with experience in automation and developer tooling.
  • Proficient in at least one general-purpose language such as Python, Java, or JavaScript.
  • Solid working knowledge of Git and containerization with Docker.
  • Experience with APIs, CI/CD pipelines, and typical engineering workflows.
  • Understanding of AI evaluation and troubleshooting technical problems.

Benefits

  • Medical, Dental and Vision Insurance
  • Health Insurance
  • Disability Insurance

Skills mentioned

PythonAutomationSoftware TestingGitDockerCI/CDAPI IntegrationPerformance TestingData AnalysisDebugging

About NLB Services

Founded in 2007, NLB Services is a global AI-led technology and business transformation organization that helps enterprises accelerate growth, enhance operational resilience, and unlock digital value at scale. Headquartered in Alpharetta, NLB Services delivers integrated solutions across AI and data engineering, intelligent automation, digital operations, customer experience, talent and workforce transformation, employer branding, skilling, and end-to-end Global Capability Center (GCC) services. Combining deep industry expertise, a robust global delivery network, and a customer-first approach, NLB empowers Fortune 500 and high-growth organizations to modernize operations, build future-ready capabilities, and navigate complex transformation journeys with confidence in an AI-driven world. Awards and Recognition : * Financial Times - The Americas' Fastest Growing Companies 2023 * SIA Top 100 Largest US Staffing Firms for 2023 * SIA Fastest-Growing US Staffing Firms 2019-2023 * Economic Times - India’s Best Organizations for Women in 2023 * Major Contender in Everest Group's US IT Contingent Talent and Strategic Solutions PEAK Matrix® Assessment 2023 * BW Diversity & Inclusion Award 2022 - Winner * BW People & HappyPlus - The Happy Workplaces Award 2024 * Financial Express FuTech Award 2022 * INC. 5000 - Fastest-Growing Private Companies in America 2021 * Happiest Workplaces Award 2023, 2024, 2025

IT Services and IT Consulting5,001-10,000 employeesAlpharetta, Georgia