Senior Data Scientist – Generative AI / Conversational AI Evaluation
About the role
Role: Senior Data Scientist – Generative AI / Conversational AI Evaluation
Location: Johnston, RI or Westwood, MA
Hybrid/Onsite preferred, though location flexibility may be considered
Domain: Cards, Banking, FSI
Description:
We are seeking an experienced Data Scientist to support the development and evaluation of AI-powered fraud self-service voice agents and conversational AI systems. The primary responsibility is not model deployment or engineering implementation, but designing evaluation frameworks, measuring system performance, identifying failure patterns, conducting root-cause analysis, and optimizing model behavior through data-driven experimentation.
Key Responsibilities
- Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems.
- Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes.
- Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation.
- Perform prompt tuning and experimentation to improve model accuracy and reliability.
- Conduct root-cause analysis of conversational failures and recommend remediation strategies.
- Measure performance across different model configurations, prompts, and guardrail implementations.
- Partner with AI Engineering and Product teams to validate solutions before production deployment.
- Build dashboards and reports that communicate model effectiveness and operational impact.
- Support fraud-related customer service use cases, including intent detection and multi-turn conversation flows.
Success Criteria
- Develop reliable evaluation methodologies for conversational AI systems.
- Quantify the effectiveness of fraud self-service voice agents.
- Optimize prompts, retrieval strategies, and guardrails using empirical evidence.
- Deliver actionable insights that improve customer experience and model performance.
- Establish measurable KPIs for intent detection and multi-turn conversation success.
Mandatory Skills:
- Strong background in Data Science, Machine Learning, Generative AI, or a related quantitative field.
- Hands-on experience evaluating LLM, RAG, Agentic AI, or Conversational AI solutions.
- Deep understanding of model evaluation techniques and metrics, including:
o Precision@K
o Recall@K
o Mean Reciprocal Rank (MRR)
o F1 Score
o Retrieval and generation quality assessment
- Experience performing experimentation, statistical analysis, and performance benchmarking.
- Strong Python programming skills.
- Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy, and related tools.
- Ability to communicate technical findings succinctly to highly technical stakeholders.
Desired Skills:
- Experience with:
o Generative AI and LLM ecosystems
o Multi-agent systems
o RAG/Agentic RAG architectures
o Amazon Bedrock
o AWS SageMaker
o Databricks
o MLflow
o LangSmith
o Weights & Biases
- Knowledge of conversational AI, IVR systems, digital assistants, and voice agents.
- Experience in financial services, fraud detection, or customer service automation
Responsibilities
- Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems.
- Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes.
- Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation.
- Perform prompt tuning and experimentation to improve model accuracy and reliability.
- Conduct root-cause analysis of conversational failures and recommend remediation strategies.
- Measure performance across different model configurations, prompts, and guardrail implementations.
- Partner with AI Engineering and Product teams to validate solutions before production deployment.
- Build dashboards and reports that communicate model effectiveness and operational impact.
Qualifications
- Strong background in Data Science, Machine Learning, Generative AI, or a related quantitative field.
- Hands-on experience evaluating LLM, RAG, Agentic AI, or Conversational AI solutions.
- Deep understanding of model evaluation techniques and metrics.
- Experience performing experimentation, statistical analysis, and performance benchmarking.
- Strong Python programming skills.
- Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy.
Skills mentioned
About BURGEON IT SERVICES
Burgeon is a provider of technical and IT staff augmentation solutions. We are in business to help you maintain your competitive advantage by cost-effectively delivering highly skilled consultants when and how you need them most. Burgeon is located in USA, AUSTRALIA & INDIA. Burgeon offers Contract, Contract-to-Hire, Direct Hire and Payroll Services to companies of all sizes. Burgeon understands the reaction time required to consistently provide companies with the technical resources necessary for growth. Burgeon places heavy emphasis on internal hiring, database technology integration, ongoing internal training and customer service. Our goal is to build a working relationship with you that is mutually beneficial. As a result, we are highly motivated to consistently deliver the best consultants and the best customer service available. Burgeon has earned a reputation amongst companies and consultants as being fair dealing and highly customer service oriented. We invite you to experience the difference working with Burgeon can make. Our Culture Burgeon has a unique culture and a very distinct, progressive, and professional attitude. We are results driven and passionate about matching the right people with the right companies. We take pride in providing quality consultants to reputable clients in order to ensure a mutually beneficial relationship. We fully understand the competitiveness of our business and rely on each other to maintain an adaptable, professional atmosphere. Our employees work in a fast-paced environment focused on goal setting and accomplishing those goals. Vision To provide the highest quality IT resources for small, medium, and Fortune 500 companies. Mission At Burgeon we are passionate about matching great people with great companies. Values Achieving our mission requires intelligent, energetic and committed people who possess:
H-1B sponsorship history
Historical employer filing data was found for BURGEON IT SERVICES. The employer record includes 2 historical certified applications. This is employer-level history, not a guarantee that this role currently offers sponsorship.