Data Scientist
About the role
Our client, a leading global litigation law firm, is seeking a Data Scientist to join its AI & Data Analytics team and help deliver sophisticated AI solutions across active legal matters and firmwide initiatives. This hands-on, highly consultative role will partner directly with case teams to manage AI projects, deploy proprietary technology, design scalable AI capabilities, and evaluate emerging tools to determine the right build-versus-buy approach. The ideal candidate is a technically strong, flexible generalist with hands-on experience across Python, SQL, LLMs, RAG, NLP, data pipelines, and AI evaluation who can move comfortably between technical development, consulting, and practical deployment. This is a hybrid position based near one of the firm's U.S. or London offices.
About The Role
We are looking for a Data Scientist to join the AI & Data Analytics Group. This is a hands-on, matter-facing role at the intersection of data science, artificial intelligence, and trial practice.
You will build and deploy tools and models that change how our litigators find, analyze, and act on information - from large-scale document review and information extraction to case intelligence and knowledge retrieval. Roughly half your time will be spent embedded directly with case teams on live matters, at times including client-facing work; the rest will go to firm-wide tooling, evaluation, and capability building.
This is a role for someone who wants to solve hard problems with substantial real-world impact, on a compressed litigation clock.
What You'll Do
Matter-embedded data science
Partner with trial teams to scope data problems on active matters and turn ambiguous asks into structured, executable solutions, often under deadline pressure set by the court.
Build custom workflows for document-heavy matters: LLM-based information extraction, semantic chunking, entity recognition, and classification.
Support early case assessment and dispute intelligence — turning raw case materials into structured knowledge that lets attorneys identify patterns and test case theories.
Design and execute data-modeling projects across discovery datasets, court filings, deposition transcripts, and third-party legal data sources.
Support trial preparation and witness preparation workflows, and maintain case intelligence that evolves as a matter progresses.
Produce clear visualizations and analyses that attorneys can act on directly, and that hold up under scrutiny from opposing counsel.
AI solution development
Prototype, build, and deploy retrieval-augmented generation (RAG) pipelines against firm and matter document corpora.
Develop and maintain the data pipelines that feed them, working with both structured data (timekeeping, matter metadata, docket data) and unstructured document sets.
Move promising prototypes into production in partnership with IT and Information Security.
Evaluation and verification
Define and run evaluation frameworks for AI outputs — precision, recall, extraction accuracy, hallucination rates, and robustness.
Design and maintain the human verification protocols that sit between model output and attorney work product.
Monitor deployed model performance over time and flag drift or degradation.
Support the firm's AI governance program on responsible use, data privacy, confidentiality, privilege, and compliance with firm policies and client requirements.
Capability building
Serve as a technical resource on AI for the group and the broader firm.
Contribute to firm-wide enablement, including AI Office Hours and attorney training sessions.
Evaluate emerging tools, models, and vendor offerings, and give clear recommendations on what is worth adopting.
Document your work so others can maintain and extend it.
What You'll Bring
Required
Experience: Demonstrated experience in data science, machine learning, or advanced analytics. Experience in a law firm, legal services, professional services, or other confidentiality-sensitive environment is strongly preferred.
Education: Bachelor's degree in Data Science, Computer Science, Statistics, Mathematics, or a related quantitative field. Master's preferred.
Programming: Strong Python for data analysis, modeling, and prototyping (pandas, scikit-learn, and at least one deep learning framework). Advanced SQL. Claude Skills.
NLP: Practical experience with text classification, named entity recognition, semantic search, summarization, and document parsing.
LLMs: Hands-on work with modern LLM platforms and tooling (e.g., Anthropic Claude, OpenAI, Hugging Face, LangChain), including RAG architectures and vector databases.
Pipelines: Demonstrated experience building data pipelines and preparing both structured and unstructured datasets for production use.
Visualization: Proficiency with Power BI, Tableau, or equivalent, plus strong Excel.
Evaluation: Experience measuring model and LLM output quality with defined metrics.
Judgment under pressure: Comfort working to litigation deadlines, and the discipline to state plainly what a model does and does not establish.
Communication: Ability to explain technical concepts to attorneys and clients in plain language, and to push back constructively when a proposed use of AI is a poor fit.
Version control: Git or equivalent.
Preferred
Familiarity with eDiscovery platforms and workflows, particularly Relativity, including structured analytics and custom development.
Working knowledge of legal data types — matter metadata, timekeeping data, docket and court filing data, billing data.
Exposure to the litigation lifecycle and the practical constraints of privilege, confidentiality, work product, and protective orders.
Experience supporting expert work, damages analysis, or other quantitative analysis offered in a contested proceeding.
Cloud experience (Azure, AWS, or GCP).
Front-end skills (JavaScript, HTML, CSS) sufficient to build lightweight internal tools and dashboards.
Familiarity with legal technology platforms: document management, docketing, and case management systems.
Expected salary for this role is $165,000 - $225,000, commensurate with experience, training, skills, qualifications, and other market factors.
Job ID: 7629
Responsibilities
- Partner with trial teams to scope data problems on active matters
- Build custom workflows for document-heavy matters
- Support early case assessment and dispute intelligence
- Design and execute data-modeling projects across discovery datasets
- Produce clear visualizations and analyses for attorneys
- Prototype, build, and deploy retrieval-augmented generation pipelines
- Define and run evaluation frameworks for AI outputs
- Serve as a technical resource on AI for the group
Qualifications
- Demonstrated experience in data science, machine learning, or advanced analytics
- Bachelor's degree in Data Science, Computer Science, Statistics, Mathematics, or a related field
- Strong Python for data analysis and modeling
- Practical experience with text classification and named entity recognition
- Hands-on work with modern LLM platforms
- Experience building data pipelines
- Proficiency with Power BI, Tableau, or equivalent
Skills mentioned
About TruLegal (formerly TRU Staffing)
TruLegal (formerly TRU Staffing Partners) is an award-winning global leader in delivering bespoke AI-enabled talent solutions for modern legal teams. With a network of 100,000+ legal professionals across 75+ countries, we place legal pros in legal operations, litigation & eDiscovery, data privacy, governance, product counsel, and cybersecurity. For 15+ years, we’ve delivered contract, direct hire, attorney secondee and executive search services to the Fortune 1000 and AmLaw 200. To learn more about us, visit www.trulegal.ai or get in touch at info@trulegal.ai.