Data Scientist - Onshore (B)
About the role
V4C is seeking a Data Scientist with strong experience in Databricks, Python, machine learning, and Master Data Management (MDM) to help build data-driven solutions that support healthcare and member engagement initiatives. The ideal candidate will have experience working with large, complex healthcare datasets and transforming disparate data sources into reliable, analytics-ready data.
Key Responsibilities
Develop and productionize machine learning models, statistical analyses, and predictive analytics using Python and Databricks
Build and maintain scalable data science workflows using Databricks, PySpark, SQL, Delta Lake, and related cloud data technologies
Work with MDM processes and frameworks to establish consistent, accurate, and trusted master data across multiple source systems
Analyze and resolve data quality, duplication, matching, and entity-resolution issues across member, provider, patient, and other healthcare-related datasets
Partner with Data Engineering, Product, Analytics, and business stakeholders to translate healthcare business problems into data science solutions
Develop data validation, profiling, and quality-monitoring approaches to improve reliability of analytical datasets
Perform exploratory data analysis and identify trends, patterns, and insights that can support member engagement and healthcare outcomes
Contribute to feature engineering, model evaluation, experimentation, and deployment of data science solutions into production
Ensure data solutions follow applicable healthcare data privacy, security, and governance requirements, including HIPAA where applicable
Document models, datasets, assumptions, methodologies, and data lineage to support reproducibility and governance
Required Qualifications
8+ years of experience in Data Science, Machine Learning, Advanced Analytics, or a related field
Strong hands-on experience with Databricks and PySpark
Advanced Python and SQL skills
Experience developing and deploying machine learning or predictive models
Strong understanding of Master Data Management (MDM) concepts, including:
Data matching and deduplication
Entity resolution
Golden/master records
Data standardization
Data quality
Reference/master data
Experience working with large-scale structured and semi-structured datasets
Experience with Delta Lake / Lakehouse architecture
Strong understanding of data governance, data quality, and data lineage
Experience working with healthcare, payer, provider, patient, or other regulated data is preferred
Experience working in a HIPAA-regulated environment is highly desirable
Preferred Qualifications
Experience with healthcare member/patient data and healthcare data models
Experience with MDM platforms such as Informatica MDM, Reltio, IBM MDM, or similar technologies
Experience with cloud platforms such as Azure or AWS
Experience with MLflow or similar model lifecycle management tools
Experience with Power BI, Tableau, or other analytics/visualization platforms
Experience building production-grade ML/data science pipelines
Familiarity with healthcare interoperability standards such as FHIR, HL7, or claims data is a plus
Core Skills
Data Science: Python, Machine Learning, Statistics, Predictive Analytics
Databricks: Databricks, PySpark, Delta Lake, MLflow
Data: SQL, Data Quality, Data Governance, Data Lineage, Data Modeling
MDM: Master Data Management, Entity Resolution, Matching, Deduplication, Golden Records
Healthcare: Healthcare Data, HIPAA, Patient/Member Data, FHIR/HL7
Cloud: Azure/AWS
Responsibilities
- Develop and productionize machine learning models, statistical analyses, and predictive analytics using Python and Databricks
- Build and maintain scalable data science workflows using Databricks, PySpark, SQL, Delta Lake, and related cloud data technologies
- Work with MDM processes and frameworks to establish consistent, accurate, and trusted master data across multiple source systems
- Analyze and resolve data quality, duplication, matching, and entity-resolution issues across healthcare-related datasets
- Partner with Data Engineering, Product, Analytics, and business stakeholders to translate healthcare business problems into data science solutions
- Develop data validation, profiling, and quality-monitoring approaches to improve reliability of analytical datasets
- Perform exploratory data analysis and identify trends, patterns, and insights that can support member engagement and healthcare outcomes
- Contribute to feature engineering, model evaluation, experimentation, and deployment of data science solutions into production
Qualifications
- 8+ years of experience in Data Science, Machine Learning, Advanced Analytics, or a related field
- Strong hands-on experience with Databricks and PySpark
- Advanced Python and SQL skills
- Experience developing and deploying machine learning or predictive models
- Strong understanding of Master Data Management (MDM) concepts
- Experience working with large-scale structured and semi-structured datasets
- Experience with Delta Lake / Lakehouse architecture
- Strong understanding of data governance, data quality, and data lineage
Skills mentioned
About v4c.ai
v4c.ai is a premier IT services consultancy specializing in Databricks to help organizations unlock the full potential of their data. We partner with enterprises to accelerate their journey to becoming data-driven by delivering end-to-end Databricks services across Lakehouse implementation, data engineering, AI/ML, and governance. Our expertise in integration, optimization, and enablement empowers clients to unify disparate data sources, modernize analytics, and build AI-ready platforms. By aligning Databricks capabilities with strategic business goals, we help organizations achieve faster insights, stronger competitive advantage, and scalable innovation.