Applied Data Scientist
About the role
Vi Engage puts predictive models into the daily operations of the largest health systems and health plans in the country — driving care navigation, specialty capture, and the workflows that follow from them. Applied Data Scientists are the people who turn a new customer into a running deployment on Vi's platform.
You will own 3–5 enterprise healthcare accounts at a time, end to end. For each one you are the data expert in the room: you design the pilot that proves value, integrate the customer's data (tokens, claims, EHR, marketing) into Vi's platform and build the automated pipelines that train, score, and deliver insights on a schedule. You stay the technical owner through operations and drive accounts to “autopilot”.
This is a hands-on-keyboard role with direct exposure to stakeholders. You are expected to sit with customer teams with expertise in analytics, marketing, and clinical operations to understand what KPIs they need to move, and how to leverage Vi’s capabilities to make it a reality. What you learn in the field becomes the product: you will spot the patterns across your accounts and turn them into the requirements that shape Vi's roadmap.
What You'll Own
Data integration. Own the pipeline between customer data systems and Vi's data platform — ingestion, mapping, quality, and the judgment calls about what the data can and cannot support.
Pilot design. Design studies that tie model performance to a business outcome the customer will underwrite. Power, controls, confounders, and a metric that survives scrutiny from a sophisticated internal analytics team.
Automated production pipelines. Build per-customer pipelines that use Vi’s ML capabilities to train, score, and deliver insights without manual intervention — and keep them running.
Customer ownership. Be the technical authority across onboarding and ongoing operations, including the ad-hoc data and reporting questions that come with a live account.
Product leverage. Identify what's common across customers and synthesize it into high-leverage requirements for the platform, so the next deployment is faster than the last.
What We're Looking For
Experimental design you can defend. You have designed and run studies where the result mattered commercially — and you can explain, to a skeptical stakeholder, why the design supports the conclusion. This is the capability we screen hardest on.
Real data science depth. You can explain how a model works, evaluate it honestly, and say what it is not good for. You know when a simpler approach is the right answer.
Python and ML engineering. Fluent in Python and the working stack — pandas, sklearn, airflow. You ship code to production that automates client deliverables.
PySpark at scale. You have built distributed data and ML pipelines on Spark against large, messy datasets.
Client-facing command. You can run a working session with clinical, IT, and analytics leaders at a major health system or health plan on your own — translating between their business problem and what the data will actually support, and holding the line when the answer isn't what they hoped. You will occasionally travel onsite with clients.
Ownership instinct. You treat your accounts as yours: you find the problem before the customer does and you fix it.
Nice To Have
Healthcare or life sciences domain knowledge — claims, EHR, HL7/FHIR, lab data, or population health analytics
Experience designing and optimizing the targeting of marketing and engagement campaigns
AWS data infrastructure: S3, Glue, EMR, MWAA, SageMaker
Familiarity with HIPAA and healthcare compliance and data governance frameworks
Experience taking a product from its first customer deployment to a repeatable one
What This Role Is Not
This is an applied, in-production role. It is not a research position — the work is measured by deployments that run and outcomes customers can point to, not by novelty. And it is not a delivery or engagement management role: you will be architecting and writing the pipelines yourself, not coordinating someone else who does.
Responsibilities
- Own the pipeline between customer data systems and Vi's data platform
- Design studies that tie model performance to business outcomes
- Build per-customer pipelines that use Vi’s ML capabilities
- Be the technical authority across onboarding and ongoing operations
- Identify common requirements across customers for faster deployments
Qualifications
- Experience in experimental design and data science
- Fluent in Python and ML engineering
- Experience with PySpark at scale
- Strong client-facing skills
- Ownership instinct for account management
Skills mentioned
About Vi
The AI execution layer for healthcare, life sciences, and wellness Vi turns fragmented health data into precise, measurable action across the systems enterprises already run. We've helped support 190M+ patients and members, generated $2B+ in measurable value, and helped bring 50+ life-changing drugs to market. At the foundation is the Vi Data Web — a privacy-safe intelligence layer covering 190M+ de-identified patient and member records and licensed signals across 96% of U.S. households. It powers our AI Applications: Vi Activate — Precise, predictive targeting to activate new patients, members, and HCPs Vi Engage — Predictive engagement that reaches the right patient, member, or care team at the right moment Vi Operate — An agentic suite built to drive operational excellence Every action and outcome flows through Vi Pulse, the real-time interface for insights, ROI, governance, and agentic deployment. Vi deploys into existing systems from day one — neutral, interoperable, and model-agnostic. No rip and replace. Our 4X return model, with 1X downside protection, means we win when our partners win. Our vision is health abundance in our lifetime — a future where every person has access to precise, predictive, and affordable care. Because health abundance won't come from more data or dashboards. It comes from execution.