Data Scientist – Conversational AI & GenAI Evaluation

ClifyX
Johnston, Rhode Island, United States · Westwood, Massachusetts, United StatesContractPosted Sep 18, 2026

About the role

Job Title: Data Scientist – Conversational AI & GenAI Evaluation

Work Location : Johnston, RI or Westwood, MA (Hybrid Onsite)

Contract duration: 12+ months Contract

Detailed job description - Skill Set:

We are seeking an experienced Data Scientist to support the development and evaluation of AI-powered fraud self-service voice agents and conversational AI systems. The primary responsibility is not model deployment or engineering implementation, but designing evaluation frameworks, measuring system performance, identifying failure patterns, conducting root-cause analysis, and optimizing model behavior through data-driven experimentation.

Key Responsibilities

Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems.

Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes.

Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation.

Perform prompt tuning and experimentation to improve model accuracy and reliability.

Conduct root-cause analysis of conversational failures and recommend remediation strategies.

Measure performance across different model configurations, prompts, and guardrail implementations.

Partner with AI Engineering and Product teams to validate solutions before production deployment.

Build dashboards and reports that communicate model effectiveness and operational impact.

Support fraud-related customer service use cases, including intent detection and multi-turn conversation flows.

Success Criteria

Develop reliable evaluation methodologies for conversational AI systems.

Quantify the effectiveness of fraud self-service voice agents.

Optimize prompts, retrieval strategies, and guardrails using empirical evidence.

Deliver actionable insights that improve customer experience and model performance.

Establish measurable KPIs for intent detection and multi-turn conversation success.

Must have:

Strong background in Data Science, Machine Learning, Generative AI, or a related quantitative field.

Hands-on experience evaluating LLM, RAG, Agentic AI, or Conversational AI solutions.

Deep understanding of model evaluation techniques and metrics, including:

Precision@K

Recall@K

Mean Reciprocal Rank (MRR)

F1 Score

Retrieval and generation quality assessment

Experience performing experimentation, statistical analysis, and performance benchmarking.

Strong Python programming skills.

Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy, and related tools.

Ability to communicate technical findings succinctly to highly technical stakeholders.

Desired Skills:

Experience with:

Generative AI and LLM ecosystems

Multi-agent systems

RAG/Agentic RAG architectures

Amazon Bedrock

AWS SageMaker

Databricks

MLflow

LangSmith

Weights & Biases

Knowledge of conversational AI, IVR systems, digital assistants, and voice agents.

Experience in financial services, fraud detection, or customer service automation.

Important Note

This role is primarily a Data Science and AI Evaluation position, not an AI Engineering or deployment-focused role. The emphasis is on measuring, analyzing, validating, and improving AI system performance rather than building production deployment pipelines.

Responsibilities

  • Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems
  • Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes
  • Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation
  • Perform prompt tuning and experimentation to improve model accuracy and reliability
  • Conduct root-cause analysis of conversational failures and recommend remediation strategies
  • Measure performance across different model configurations, prompts, and guardrail implementations
  • Partner with AI Engineering and Product teams to validate solutions before production deployment
  • Build dashboards and reports that communicate model effectiveness and operational impact

Qualifications

  • Strong background in Data Science, Machine Learning, Generative AI, or a related quantitative field
  • Hands-on experience evaluating LLM, RAG, Agentic AI, or Conversational AI solutions
  • Deep understanding of model evaluation techniques and metrics
  • Experience performing experimentation, statistical analysis, and performance benchmarking
  • Strong Python programming skills
  • Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy

Skills mentioned

PythonData ScienceMachine LearningGenerative AILarge Language ModelsRetrieval-Augmented GenerationAI AgentsModel EvaluationStatistical AnalysisScikit-learn

About ClifyX

ClifyX has immense experience of over 20 years in Digitalization, Service Management, Cloud Transformation, and now in enabling implementation and deployments for its end customer through Salesforce, ServiceNow, Selenium, and Automation. Our rich experience, combined with our unyielding care for our employees, is the driving force behind all we do. And we deliver!

Staffing and Recruiting501-1,000 employeesSouth Plainfield, NJ

H-1B sponsorship history

Historical employer filing data was found for ClifyX. The employer record includes 9 historical certified applications. This is employer-level history, not a guarantee that this role currently offers sponsorship.