Software Engineer, RL Data
About the role
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
Software Engineer, Reinforcement learning
Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.
About The Role
As a Software Engineer on the RL Data team at Cursor, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.
What You’ll Do
Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.
Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.
Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
Partnering with research on whether a dataset is actually teaching the thing we think it is.
You may be a fit if
You write careful, fast code and have strong software engineering fundamentals.
You like setting tasks: breaking a fuzzy capability into something concrete you can measure.
You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.
You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
Software Engineer, Reinforcement learning
Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.
What You’ll Do
Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.
Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.
Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
Partnering with research on whether a dataset is actually teaching the thing we think it is.
You may be a fit if
You write careful, fast code and have strong software engineering fundamentals.
You like setting tasks: breaking a fuzzy capability into something concrete you can measure.
You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.
You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.
Responsibilities
- Designing a task set that teaches a specific agent capability
- Reading agent traces to find failure modes or surprising behaviors
- Building systems that surface more agent behaviors
- Partnering with research to evaluate dataset effectiveness
Qualifications
- Strong software engineering fundamentals
- Experience with infra, data, or distributed systems
- Ability to break down fuzzy capabilities into measurable tasks
- Experience with reinforcement learning is a plus
Skills mentioned
About Cursor
Cursor is a coding agent for building ambitious software. Our goal is to help you engineer anything. Our work includes training the world’s most widely used coding models, creating infrastructure that supports billions of requests per day, and building better ways for humans and AIs to work together.