Staff Data Scientist

Vi
Boston, Massachusetts, United StatesFull-timePosted Sep 15, 2026

About the role

Vi Engage puts predictive models into the daily operations of the largest health systems and health plans in the country — driving care navigation, specialty capture, and the workflows that follow from them. This role builds the modeling engine those deployments run on.

You will own the pipelines that turn longitudinal claims, EHR, lab, and online behavioral intent data into predictions about which patients face upcoming healthcare utilization and which are high-propensity to enroll in preventative care programs. Not one model for one customer — the config-driven machinery that trains, selects, scores, and delivers models for any customer, with automatic feature engineering and model selection doing the work that bespoke engineering does today.

The measure of this work is generalization. A model that lifts enrollment for one health plan is a good result; a pipeline that reproduces that lift for the next twenty without per-customer engineering is the product. You set the modeling standard that Forward Deployed Data Scientists build their customer deployments against, and you are accountable for the improvements that hold across all of them.

What You'll Own

ML pipelines; Built for scale. Config-driven pipelines that train, score, and deliver predictions from large longitudinal datasets.

Mass customization. Automatic feature engineering and model selection that produces a customer-specific model without customer-specific work.

Cross-customer generalization. Find the modeling improvements that hold up across every deployment.

Building blocks for the field. Implementing and maintaining the DS components that Forward Deployed Data Scientists compose into running customer deployments, so the next deployment is faster than the last.

Productization. Turn pilots and one-off proofs into product capabilities that survive their fourth customer.

What We're Looking For

Modeling longitudinal data in production. You have personally shipped uplift, survival, or propensity models whose output changed how an organization spent money or who it reached. This is the capability we screen hardest on.

POC to product. You have taken a pilot or proof of concept and turned it into something repeatable that survived a second, third, and fourth customer. You know which parts of a one-off are the product and which parts are the customer.

Real data science depth. Segmentation, campaign optimization, and the identification strategy behind an uplift estimate. You can defend a modeling choice to a skeptical internal analytics team, evaluate a model honestly, and say when a simpler approach is the right answer.

Python and ML engineering. Fluent in Python and the working stack — pandas, sklearn, PySpark, airflow. You build the pipeline, not just the model inside it.

MLOps. Model tracking and deployment tooling — mlflow, SageMaker, or comparable

AWS and Cloud. You don’t need to hand-off to a dedicated engineer. You can get your models running at scale in the cloud using our AWS stack: S3, Glue, EMR, MWAA, SageMaker.

Nice To Have

Healthcare or life sciences domain knowledge — claims, EHR, HL7/FHIR, lab data, or population health analytics

Familiarity with HIPAA and healthcare compliance and data governance frameworks

Experience designing pilots and efficacy studies that tie model performance to a business outcome

Experience building internal platforms or frameworks that other engineers build on top of

What This Role Is Not

This is an applied, in-production role. It is not a research position — the work is measured by pipelines that run and models that hold up across customers, not by novelty. It is primarily not a customer-facing role (Applied Data Scientists own the customer accounts) but you may interface with design partner clients on occasion. This is not a management role — you will be hands-on architecting and building this product with the team.

Responsibilities

  • Own ML pipelines built for scale
  • Implement mass customization for customer-specific models
  • Ensure cross-customer generalization of modeling improvements
  • Build components for faster customer deployments
  • Productize pilots into repeatable product capabilities

Qualifications

  • Experience modeling longitudinal data in production
  • Ability to turn POCs into repeatable products
  • Depth in segmentation and campaign optimization
  • Fluency in Python and ML engineering tools
  • Experience with MLOps and cloud deployment

Skills mentioned

PythonPandasScikit-learnMachine LearningData PipelinesApache SparkApache AirflowAWSMLflowAmazon SageMaker

About Vi

The AI execution layer for healthcare, life sciences, and wellness Vi turns fragmented health data into precise, measurable action across the systems enterprises already run. We've helped support 190M+ patients and members, generated $2B+ in measurable value, and helped bring 50+ life-changing drugs to market. At the foundation is the Vi Data Web — a privacy-safe intelligence layer covering 190M+ de-identified patient and member records and licensed signals across 96% of U.S. households. It powers our AI Applications: Vi Activate — Precise, predictive targeting to activate new patients, members, and HCPs Vi Engage — Predictive engagement that reaches the right patient, member, or care team at the right moment Vi Operate — An agentic suite built to drive operational excellence Every action and outcome flows through Vi Pulse, the real-time interface for insights, ROI, governance, and agentic deployment. Vi deploys into existing systems from day one — neutral, interoperable, and model-agnostic. No rip and replace. Our 4X return model, with 1X downside protection, means we win when our partners win. Our vision is health abundance in our lifetime — a future where every person has access to precise, predictive, and affordable care. Because health abundance won't come from more data or dashboards. It comes from execution.

Software Development51-200 employeesNew York, NY