(USA) Principal, Data Scientist
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Cursor
NewYou'll design and build full-stack tools that researchers and vendors use to create, review, and monitor reinforcement-learning training data at scale. The role requires shipped full-stack experience with TypeScript, React, and backend languages like Python or Go, plus strong skills building data-facing interfaces such as review queues, transcript viewers, or observability dashboards. You should have experience designing workflows where users have incentives to circumvent checks, and ideally familiarity with design systems. This is embedded in a research team in a flat organization.
Written from this posting by Neural Jobs AI. The full description is below.
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
As a Software Engineer on the RL Data team, you’ll design and build the tools that researchers and external contributors use to create, review, submit, and monitor the environments and tasks behind Cursor’s reinforcement-learning runs. This is a full-stack product-engineering role embedded in a research team.
You’ll own the review and acceptance experience end to end: from rollout and transcript inspection, task-quality signals grader and reward-hacking analysis, to the workflows that move a submission into training. From there, you’ll build authoring interfaces that let researchers, vendors, and domain experts create and improve environments and tasks quickly and confidently.
Your work will significantly shorten the loop from a task idea or data sources, to candidate task, to trusted training data.
Create fast, trustworthy workflows for vendors and research team to interact effectively with each other — vendor task creation and iteration, vendor submissions, and task acceptance into training.
Build review tools for inspecting and comparing rollouts, transcripts, grader outputs, and other signals of task quality.
Develop environment-health, failure-search, versioning, and catalog experiences that make training data easy to understand, manage, and extend.
Establish a shared component kit, then use it to build self-serve interfaces for creating and improving tasks with quality checks inline.
You’ve shipped full-stack products and owned systems from user interface through storage or services, using technologies such as TypeScript and React alongside Node, Python, or Go.
You’ve built dense, data-facing tools such as transcript viewers, diffing systems, review queues, observability products, or operational dashboards—and you have strong opinions about how structured data should be rendered.
You’ve built or maintained a design system or component library and can establish durable product and engineering conventions for a fast-moving team.
You’ve designed review, QA, moderation, fraud, or acceptance workflows where users had an incentive to get past the checks, and you know how to keep those systems honest.
You care about data quality, and are willing to inspect raw data. Experience with evaluations, graders, reinforcement learning, or data-quality systems is helpful but not required.
You move quickly under ambiguity, collaborate closely with researchers and domain experts, and take open-ended problems from rough need to reliable product.
If there appears to be a fit, we’ll schedule two or three short technical interviews focused on frontend craft for dense data and system design for a review-and-acceptance workflow. After that, we’ll invite you onsite to work on a small project using real rollouts, discuss ideas, and meet the team.
#LI-DNI
Here is what this employer asked for. Sign in and we will fill in your half.
Cursor, developed by Anysphere, is an AI-native code editor that integrates large language models directly into the development workflow, from autocomplete to multi-file edits and agentic coding.
Founded in 2022, it has become one of the most widely adopted AI coding tools.
718 more Software Engineer roles in San Francisco
263K – 393K
Our estimate — this employer did not publish a salaryOur estimate, not the employer’s. Worked out from the middle half of 93 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
Search by role, company, or anything a posting mentions.