Junior Data Scientist
About the role
Junior Data Scientist
Phase IV | Pharmaceutical Intelligence Fully remote, United States $100,000 to $120,000 base, depending on experience Full-time | Start date: as soon as we find the right person
About Phase IV
Patients write about their medicines in public, at length, and in more detail than they give anyone in a clinic. They describe the side effect that made them stop taking something, the twelve weeks it took before anything changed, the switch an insurer forced on them, the symptom they never mentioned to their prescriber because it sounded trivial. Most of it is written in forums and on drug review platforms, in a dozen languages, and almost none of it reaches the company that made the drug.
Phase IV turns that material into structured evidence. We collect first-hand patient accounts of medicines at scale, extract what is actually being described, normalize it against the standard vocabularies the industry runs on, and make it queryable for commercial, medical affairs, HEOR and safety teams.
Two things define the work. The first is the insistence on first-hand accounts. A great deal of what is written about a drug online is promotional or secondhand: someone relaying something they read elsewhere. Separating the two is a hard NLP problem and it is the core of what we sell, because an account of what someone experienced is worth something and an account of what they heard is not.
The second is that we sit upstream of pharmacovigilance case-management systems such as Oracle Argus and Veeva Vault Safety. We are not a replacement for them and we do not do regulatory reporting. We are the layer that tells a company what patients are saying before it shows up anywhere else.
We are early and deliberately small. The corpus already runs into millions of posts and comments across the major condition forums and drug review platforms, with named-entity and relation models trained on a purpose-built annotation schema.
The role
This is a first or second job in data science, and the point of it is range. You will work across the collection and cleaning of patient-authored text, the annotation that produces our training data, the normalization of drugs and events against RxNorm and MedDRA, and the evaluation that tells us whether a model is good enough to put in front of a client.
Producing the code is no longer the difficult part of this job, and we are not hiring on typing speed. A model will draft much of what you write and we expect you to let it. The difficult part is deciding what should be built, recognizing when what came back is subtly wrong, and being able to say why one approach was right and the alternatives were not. In this domain that matters more than most: a label that is nearly right produces a finding that is confidently false about a real medicine.
You will be working directly with the founding team, including a technical founder who holds a PhD in machine learning and has published peer-reviewed research on this method. There is no layer between you and the decisions. That is the main reason to take a job like this one.
What you will actually do
Work out how a piece of data should be collected, cleaned or linked before anything is written, and be able to justify the choice
Build and maintain pipelines in Python that gather patient-authored text from forums and drug review platforms, including sources that do not want to be gathered from
Annotate text against our entity and relation schema, contribute to the gold set, and measure agreement honestly rather than flatteringly
Review machine-generated annotation, find where it fails systematically, and feed that back into the guidelines
Normalize drug mentions against RxNorm and event mentions against MedDRA, and investigate the cases that will not map
Handle the awkward parts of real patient language: negation, hypotheticals, dosage written six different ways, and accounts that turn out to be about someone else
Query and manage data in PostgreSQL, and reason about what a query is actually doing at scale
Build and run the checks that tell us whether a pipeline is quietly failing
Analyze model outputs, look for patterns worth reporting, and say so when a pattern does not hold up
Document what you built, and why you built it that way, well enough that someone else can pick it up
What we need from you
Essential:
A sound grasp of the fundamentals: statistics, sampling and bias, evaluation design, how a relational database executes what you asked it for, and what a model is doing when it returns a number
The ability to reason from a problem to an approach, weigh it against the alternatives, and defend the choice when someone questions it
A critical reading eye. You should be able to look at code, a label set or an analysis and find what is wrong with it
Working Python and SQL. You need to be fluent enough to work unaided and to judge whether output is correct, which is a lower bar than it used to be and a different one
Patience with detailed, careful work. Annotation quality decides model quality, and there is no way around that
The ability to work through an ambiguous problem without waiting to be told what to do next
Clear written and spoken English, because you will be explaining your work to people who are not data scientists
A degree in a quantitative, computational, scientific or life-sciences subject, or equivalent demonstrable ability
Useful, and worth mentioning if you have it:
Any exposure to NLP, text classification, named-entity recognition, or transformer models
Web scraping, APIs, or reverse-engineering how a site actually serves its data
Reading knowledge of a language other than English, particularly German, French, Spanish, Portuguese or Japanese
Any background in pharmacy, pharmacology, life sciences or clinical research
Familiarity with medical vocabularies such as RxNorm, MedDRA, ICD-10 or SNOMED
AWS
We are not expecting all of the above. Sound judgment, curiosity and a habit of checking things will get you further with us than a long list.
On AI tools
We use them constantly and we expect you to. There is no exercise here where you write code unaided, and no credit for doing by hand what a model does in seconds. A significant part of our annotation pipeline is model-driven.
The condition is that you understand and can defend everything you put your name to. If you cannot explain why an approach was taken over the alternatives, or walk through what a piece of code does and why it is correct, it is not finished, whoever or whatever produced it. Plausible output and correct output look identical until somebody checks, and checking is most of the job now.
How we work
Remote, asynchronous, and light on meetings. We care what you ship and how well it holds up, not when you were at your desk. We review work seriously and give direct feedback. You will be expected to do the same in return.
What you get
$100,000 to $120,000 base salary, set by experience rather than by negotiation
Health, dental and vision coverage, and a 401(k)
Paid time off with a floor of 20 days plus federal holidays, and an expectation that you take them
An annual budget for training, conferences and equipment, chosen by you
Direct mentorship from the founding team, and a realistic path into owning a modeling area within 18 months
Deep, genuinely transferable expertise in biomedical NLP, which is a narrow field and a well-paid one
The chance to join a company at the point where what you build still shapes what it becomes
Applying
Please apply directly via LinkedIn.
Responsibilities
- Work out how a piece of data should be collected, cleaned or linked before anything is written, and justify the choice
- Build and maintain pipelines in Python that gather patient-authored text from forums and drug review platforms
- Annotate text against entity and relation schema, contribute to the gold set, and measure agreement honestly
- Review machine-generated annotation, find systematic failures, and feed that back into the guidelines
- Normalize drug mentions against RxNorm and event mentions against MedDRA
- Handle real patient language complexities
- Query and manage data in PostgreSQL
- Build and run checks to monitor pipeline performance
Qualifications
- A sound grasp of statistics, sampling and bias, evaluation design, and relational databases
- Ability to reason from a problem to an approach and defend choices
- Critical reading eye for code and analysis
- Fluency in Python and SQL
- Patience with detailed, careful work
- Ability to work through ambiguous problems independently
- Clear written and spoken English
- Degree in a quantitative, computational, scientific or life-sciences subject, or equivalent ability
Benefits
- Health, dental and vision coverage
- 401(k)
- Paid time off with a minimum of 20 days plus federal holidays
- Annual budget for training, conferences and equipment
- Direct mentorship from the founding team
- Path into owning a modeling area within 18 months
- Transferable expertise in biomedical NLP
Skills mentioned
About Phase IV
Patients describe their medicines in public — continuously, unprompted, and in far more detail than any survey captures. Phase IV reads that record and turns it into evidence. We collect first-hand patient accounts from drug review platforms, regulators, condition forums, and international patient communities. We take this complex, unstructured, and multilingual feedback and extract all of the insight: every medicine resolved to its ingredient through RxNorm, every reported effect mapped to MedDRA, every claim annotated at the relation level — drug, effect, dose, duration, sentiment — rather than scored as a block of text. The result is queryable in plain English and traceable back to coded terminology. Brand and commercial insights teams use it for launch tracking, adherence barriers, and switching drivers. Medical affairs teams use it for unmet need in the language patients actually use. HEOR teams use it for patient-reported burden at a scale interviews cannot reach. Safety teams use it as an early-warning layer upstream of their case-management system — feeding it, not replacing it. Phase IV is a system of intelligence, not a system of record. Public sources only. PII scrubbed at ingestion. GDPR-compliant processing, with operations in the European Union and the United States. Interested? Get a free brief on your product within 72 hours. No procurement process, no implementation project, no sales call – just email us at hello@phase-iv.ai for your insights.