Software Engineer- Benchmarking
About the role
For our client, we are seeking a Software Engineer- Benchmarking to join the team of a leader in the Enterprise Technology space.
This role will focus on building the infrastructure that supports rigorous model benchmarking, evaluation, and analysis. You will work closely with researchers and technical stakeholders to turn evaluation methodologies into reliable software systems that produce accurate, reproducible results. The position combines hands-on Python engineering with data preparation, containerized environments, and tools that help teams understand model performance and failure modes.
Location: Remote - US based candidates only, no visa sponsorship available
Compensation: $160,000 – $210,000 annually
Responsibilities
Prepare and maintain benchmark datasets through cleaning and validation
Build and maintain evaluation pipelines for consistent model assessments
Create containerized environments for evaluating model performance
Support fine-tuning of small open-source LLMs and analyze performance
Develop scoreboards and leaderboards for published evaluation results
Create analysis tools to identify model failure modes and performance tracking
Collaborate with teams to ensure data accuracy and integration in publications
Qualifications
4+ years of professional software engineering experience with a focus on Python
Strong capability in preparing and maintaining datasets for accurate comparisons
Experience with containerization using Docker to build reproducible environments
Ability to collaborate with researchers to translate methodologies into functional systems
Familiarity with AI evaluation frameworks is a plus
Benefits
Medical, dental, and vision coverage
Monthly wellness and fitness stipend
Paid time off along with company holidays
Annual company off-sites in various locations
Parent-friendly policies, including remote flexibility and paid family leave
Our client is an equal opportunity employer. We encourage you to apply even if you don’t meet every qualification—your background could be exactly what this team needs.
Responsibilities
- Prepare and maintain benchmark datasets through cleaning and validation
- Build and maintain evaluation pipelines for consistent model assessments
- Create containerized environments for evaluating model performance
- Support fine-tuning of small open-source LLMs and analyze performance
- Develop scoreboards and leaderboards for published evaluation results
- Create analysis tools to identify model failure modes and performance tracking
- Collaborate with teams to ensure data accuracy and integration in publications
Qualifications
- 4+ years of professional software engineering experience with a focus on Python
- Strong capability in preparing and maintaining datasets for accurate comparisons
- Experience with containerization using Docker to build reproducible environments
- Ability to collaborate with researchers to translate methodologies into functional systems
- Familiarity with AI evaluation frameworks is a plus
Benefits
- Medical, dental, and vision coverage
- Monthly wellness and fitness stipend
- Paid time off along with company holidays
- Annual company off-sites in various locations
- Parent-friendly policies, including remote flexibility and paid family leave
Skills mentioned
About Ladders
We lead the leaders. Ladder’s community of 10 mm leading professionals come to us to be inspired, grow in their careers, and get ahead in their professional lives. We focus on the $100K+ job market, representing 25% of all jobs and about 50% of income in the US & Canada. We build data-driven tools, marketplaces, news and entertainment products for our growing audience, and the companies, recruiters, and advertisers that want to reach them. We work best with people who enjoy using their talents, commitment and hard work to achieve great successes for millions of real life users through products that impact their careers. We are deeply technology-driven -- about half the company are engineers, technical product people or product designers -- but we never build technology for technology’s sake. We’re always trying to connect our efforts to the good it can do in people’s lives.
H-1B sponsorship history
Historical employer filing data was found for Ladders. The employer record includes 8 historical certified applications. This is employer-level history, not a guarantee that this role currently offers sponsorship.