Data Scientist — AI Evaluation Analytics

Invovia
United StatesContractPosted Aug 29, 2026

About the role

Position -1

Data Scientist — AI Evaluation Analytics

Role Overview

Provide data science and analytical expertise to improve the measurement, reliability, and interpretation of AI system performance through data-driven evaluation approaches.

Key Responsibilities

Support development and refinement of evaluation metrics, scoring approaches, and benchmark methodologies.

Validate datasets, data quality criteria, and evaluation inputs to improve reliability of results.

Analyze evaluation data, performance trends, and measurement signals to generate actionable insights.

Apply statistical analysis and experimentation techniques to assess AI system effectiveness.

Develop analytical approaches to identify trends, patterns, and opportunities for improvement.

Required Skills & Experience

Strong background in data science, analytics, statistics, or related fields.

Experience designing metrics, analytical frameworks, or measurement approaches.

Proficiency in Python, SQL, and data analysis techniques.

Experience working with large datasets and extracting meaningful insights.

Strong communication skills with the ability to explain analytical results clearly.

Preferred Experience

Experience evaluating AI/ML systems and agentic AI applications.

Experience with experimentation, benchmarking, or performance analysis.

Experience working with software engineering, productivity, or quality metrics.

Position 2

Software Engineer — AI Evaluation & Automation

Role Overview

Help build and scale the tooling we use to measure how well AI-powered software development tools actually perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.

Key Responsibilities

Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.

Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.

Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.

Support execution-based benchmarking across quality, productivity, and efficiency measures, including cost and latency.

Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.

Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.

Required Skills & Experience

Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.

Proficient in Python, and comfortable in at least one of Java, JavaScript, or a similar language.

Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.

Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.

Some familiarity with how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.

Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.

Responsibilities

  • Support development and refinement of evaluation metrics, scoring approaches, and benchmark methodologies.
  • Validate datasets, data quality criteria, and evaluation inputs to improve reliability of results.
  • Analyze evaluation data, performance trends, and measurement signals to generate actionable insights.
  • Apply statistical analysis and experimentation techniques to assess AI system effectiveness.
  • Develop analytical approaches to identify trends, patterns, and opportunities for improvement.

Qualifications

  • Strong background in data science, analytics, statistics, or related fields.
  • Experience designing metrics, analytical frameworks, or measurement approaches.
  • Proficiency in Python, SQL, and data analysis techniques.
  • Experience working with large datasets and extracting meaningful insights.
  • Strong communication skills with the ability to explain analytical results clearly.

Skills mentioned

PythonSQLData AnalysisStatistical AnalysisData ScienceAutomationGitDockerCI/CDSoftware Testing

About Invovia

Innovation Powers Business Excellence Your Involved and Viable Partner in Innovation and Excellence Invovia is more than just a company; we are your collaborative partners in achieving innovation and excellence across various industries using our service offerings. With a steadfast commitment to pushing boundaries, we provide a range of cutting-edge solutions and services that empower businesses to thrive in a rapidly changing world. Invovia offers a diverse portfolio of services, spanning healthcare, technology, talent development, and beyond. From state-of-the-art healthcare solutions to ground-breaking technology advancements. We are dedicated to drive progress and to ensure our clients and partners are ahead of the curve in their respective fields. ☑ Our Mission Our mission is to drive positive change empower individuals and organizations; and deliver innovative solutions that transform industries. We are dedicated to fostering collaborative partnerships, promoting diversity and inclusivity, and upholding the highest standards of integrity. Our mission is to make a meaningful impact in the world by continuously pushing the boundaries of what’s possible, enabling our clients and partners to achieve their full potential, and contributing to a brighter and a more inclusive future. ☑ Our Vision Our vision is to be a global leader in Inc500, in driving innovation, fostering collaboration, and making a meaningful impact across diverse industries. We envision a world that fosters limitless creativity, stimulates advancement, and generates new standards. We are committed to transgress boundaries, embrace diversity, and shape a futuristic environment that is reachable to all stakeholders.

Technology11-50 employeesPleasanton, California