logo

AI Job Summary

You'll design and build full-stack tools that researchers and vendors use to create, review, and monitor reinforcement-learning training data at scale. The role requires shipped full-stack experience with TypeScript, React, and backend languages like Python or Go, plus strong skills building data-facing interfaces such as review queues, transcript viewers, or observability dashboards. You should have experience designing workflows where users have incentives to circumvent checks, and ideally familiarity with design systems. This is embedded in a research team in a flat organization.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

About the role

As a Software Engineer on the RL Data team, you’ll design and build the tools that researchers and external contributors use to create, review, submit, and monitor the environments and tasks behind Cursor’s reinforcement-learning runs. This is a full-stack product-engineering role embedded in a research team.

You’ll own the review and acceptance experience end to end: from rollout and transcript inspection, task-quality signals grader and reward-hacking analysis, to the workflows that move a submission into training. From there, you’ll build authoring interfaces that let researchers, vendors, and domain experts create and improve environments and tasks quickly and confidently.

Your work will significantly shorten the loop from a task idea or data sources, to candidate task, to trusted training data.

What you’ll work on

  • Create fast, trustworthy workflows for vendors and research team to interact effectively with each other — vendor task creation and iteration, vendor submissions, and task acceptance into training.

  • Build review tools for inspecting and comparing rollouts, transcripts, grader outputs, and other signals of task quality.

  • Develop environment-health, failure-search, versioning, and catalog experiences that make training data easy to understand, manage, and extend.

  • Establish a shared component kit, then use it to build self-serve interfaces for creating and improving tasks with quality checks inline.

You may be a fit if

  • You’ve shipped full-stack products and owned systems from user interface through storage or services, using technologies such as TypeScript and React alongside Node, Python, or Go.

  • You’ve built dense, data-facing tools such as transcript viewers, diffing systems, review queues, observability products, or operational dashboards—and you have strong opinions about how structured data should be rendered.

  • You’ve built or maintained a design system or component library and can establish durable product and engineering conventions for a fast-moving team.

  • You’ve designed review, QA, moderation, fraud, or acceptance workflows where users had an incentive to get past the checks, and you know how to keep those systems honest.

  • You care about data quality, and are willing to inspect raw data. Experience with evaluations, graders, reinforcement learning, or data-quality systems is helpful but not required.

  • You move quickly under ambiguity, collaborate closely with researchers and domain experts, and take open-ended problems from rough need to reliable product.

Applying

If there appears to be a fit, we’ll schedule two or three short technical interviews focused on frontend craft for dense data and system design for a review-and-acceptance workflow. After that, we’ll invite you onsite to work on a small project using real rollouts, discuss ideas, and meet the team.

#LI-DNI

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Software Engineer, Research Tools
  • Experience 3-4 years
  • Education Any
  • Work type On-site
  • Location United States
Check my match (free)
Cursor
Developer Tools & Platforms · 100-200 Members · San Francisco, CA, United States

Cursor, developed by Anysphere, is an AI-native code editor that integrates large language models directly into the development workflow, from autocomplete to multi-file edits and agentic coding.

Founded in 2022, it has become one of the most widely adopted AI coding tools.

All jobs at Cursor

718 more Software Engineer roles in San Francisco

Job Overview

Approx. salary range

263K – 393K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 93 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
23 hours ago
Job Type
Full Time
Education
Any
Experience
3-4 years

Share This Job: