AI Test & Evaluation Lead

JR Software Solutions Inc.
Alexandria, Virginia, United StatesFull-time$140,000–$180,000Posted Sep 16, 2026

About the role

AI Test & Evaluation Lead: prove, with evidence, whether an AI system is ready for real-world use.

JR Software Solutions Inc. (JRSS) is hiring an AI Test & Evaluation Lead in the Washington, DC area (Alexandria, VA, hybrid) to lead independent testing of generative AI, LLM, retrieval-augmented generation (RAG), and machine learning systems for a federal civilian client.

This is an evaluation role, not a development role. We don't build the systems we test. Our job is to answer one question with evidence: does this system perform to a measurable standard, and is it ready to deploy? You will set the test strategy, measures, datasets, and pass criteria, and your results must be repeatable by anyone who re-runs them.

If you have evaluated LLMs for groundedness and hallucination, know why an LLM-as-judge score needs checking, and can write a test plan that holds up to review, this is your role.

About JRSS

JRSS is a certified Women-Owned Small Business (WOSB), Small Disadvantaged Business (SDB), and HUBZone company headquartered in Tampa, FL. We deliver Cloud, AI/ML, Big Data and BI Analytics, ERP, and IT program management to Federal Civilian agencies and Fortune 500 clients. Our AI Assurance practice gives agencies independent evidence that their AI systems perform, behave safely, and are ready to deploy.

What you'll do

Write the AI Test & Evaluation Plan, test strategies, evaluation methods, test plans, and test cases.

Choose measures that fit each system: accuracy, groundedness, citation fidelity, instruction following, refusal behavior, bias, consistency, latency, and stability.

Evaluate RAG retrieval and generation separately; score ML models on precision, recall, F1, AUC, and calibration.

Build mission-specific test sets from the client's domain terminology and realistic scenarios.

Make every result repeatable: versioned datasets and prompts; recorded model versions, settings, and seeds; immutable run logs.

Check AI-assisted scoring against human review, and cross-check vendor scores (e.g., Azure AI Foundry) with independent methods such as RAGAS.

Work with red team, IV&V, and cloud engineers on a FedRAMP-authorized Azure Government testing platform.

Map test evidence to the NIST AI Risk Management Framework to support agency AI governance and ATO decisions.

Write results, findings, and deployment readiness reports, and brief technical staff and leadership.

What you bring

Bachelor's degree in computer science, data science, engineering, statistics, or a related field.

8+ years in software or model testing and evaluation, including hands-on evaluation of generative AI, LLM, or RAG systems.

Working knowledge of the NIST AI RMF and its Generative AI Profile (NIST AI 600-1).

Strong Python for evaluation harnesses, scoring scripts, and analysis notebooks.

Experience with LLM evaluation metrics, including LLM-as-judge scoring and how to validate it.

A track record of test strategies, plans, and reports that stand up to review.

You live in the DC / Northern Virginia / Maryland area and can be on site as directed.

Nice to have

Azure AI Foundry evaluations, RAGAS, promptfoo, or DeepEval.

AI red teaming with PyRIT, Garak, or NIST Dioptra; MITRE ATLAS and the OWASP Top 10 for LLM Applications.

Federal IV&V, ATO, or NIST SP 800-53 experience.

Azure Government or other FedRAMP environments.

Master's or PhD; Azure AI Engineer, ISTQB, CSTE, or CISSP certification.

An active Public Trust.

Citizenship and work authorization — please read before applying

This is a full-time, W-2 position with JR Software Solutions. It is not available on a corp-to-corp or 1099 basis, and we do not sponsor visas for this role.

Because of federal client requirements, U.S. citizenship is required, and you must have lived in the United States for at least 3 of the last 5 years. You must be able to obtain and maintain a federal Public Trust.

What we offer

$140,000 – $180,000 per year, based on experience, certifications, and clearance · medical, dental, and vision coverage · 401(k) · paid time off and holidays · hybrid schedule · professional development and certification support.

We are interviewing immediately.

JR Software Solutions Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by law.

Responsibilities

  • Write the AI Test & Evaluation Plan, test strategies, evaluation methods, test plans, and test cases.
  • Choose measures that fit each system: accuracy, groundedness, citation fidelity, instruction following, refusal behavior, bias, consistency, latency, and stability.
  • Evaluate RAG retrieval and generation separately; score ML models on precision, recall, F1, AUC, and calibration.
  • Build mission-specific test sets from the client's domain terminology and realistic scenarios.
  • Make every result repeatable: versioned datasets and prompts; recorded model versions, settings, and seeds; immutable run logs.
  • Check AI-assisted scoring against human review, and cross-check vendor scores with independent methods.
  • Work with red team, IV&V, and cloud engineers on a FedRAMP-authorized Azure Government testing platform.
  • Map test evidence to the NIST AI Risk Management Framework to support agency AI governance and ATO decisions.

Qualifications

  • Bachelor's degree in computer science, data science, engineering, statistics, or a related field.
  • 8+ years in software or model testing and evaluation, including hands-on evaluation of generative AI, LLM, or RAG systems.
  • Working knowledge of the NIST AI RMF and its Generative AI Profile.
  • Strong Python for evaluation harnesses, scoring scripts, and analysis notebooks.
  • Experience with LLM evaluation metrics, including LLM-as-judge scoring and how to validate it.
  • A track record of test strategies, plans, and reports that stand up to review.

Benefits

  • Medical, dental, and vision coverage
  • 401(k)
  • Paid time off and holidays
  • Hybrid schedule
  • Professional development and certification support

Skills mentioned

PythonSoftware TestingPerformance TestingMachine LearningModel EvaluationGenerative AILarge Language ModelsRetrieval-Augmented GenerationData AnalysisStatistical Analysis

About JR Software Solutions Inc.

Interested in solving public health and federal missions? You're in the right place. For two decades, JR Software Solutions (JRSS) has helped CDC, DHS, and other federal, state, and local agencies turn technology into mission outcomes — building enterprise data analytics platforms, secure GenAI assistants, and AI agents that run inside federal networks, and modernizing the applications and infrastructure underneath them. Our six core capabilities: AI & Machine Learning · Intelligent Automation & AI Agents · Data & Analytics · Application Development · Cybersecurity & Zero Trust · DevSecOps & Secure Cloud — delivered on Azure, Azure Government, AWS, Databricks, and Snowflake. Why agencies and primes work with us: → Certified small business: WOSB · HUBZone · SDB · MBE (NMSDC) → Contract vehicles: GSA MAS #47QTCA24D0005 · NASA SEWP VI Category C (80TECH26D0769) → Triple ISO certified: 9001 · 27001 · 20000-1, plus CMMC Level 1 → Past performance rated Exceptional · three consecutive years on the Inc. 5000 Founded in 2005 · Headquartered in Tampa, FL, with offices in Atlanta · Learn more: www.jrssinc.com We're hiring across data, cloud, security, and AI — nine open roles at www.jrssinc.com/careers

IT Services and IT Consulting11-50 employeesTampa, Florida