Research Operations, Reinforcement Learning
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Design and build agent environments, develop AutoRaters calibrated against human evaluation, and create diagnostic tools to identify agent failure modes. Research how agents can automatically generate verifiers and iteratively improve. Build datasets and reward signals from multi-turn trajectories to train frontier models. Requires two years with agentic AI workflows, one year with machine learning and deep learning, software engineering experience, and proficiency in Python, TensorFlow, or PyTorch.
Written from this posting by Neural Jobs AI. The full description is below.
Here is what this employer asked for. Sign in and we will fill in your half.
Google builds internet, software, cloud, and AI products used by consumers, developers, and organizations. Its portfolio includes Search, YouTube, Android, Chrome, Maps, Gmail, Workspace, Google Cloud, advertising platforms, devices, and Gemini AI products. The company develops large-scale computing infrastructure and research that power information retrieval, communication, productivity, media, navigation, and machine learning. Google is the largest operating business within Alphabet and earns a substantial share of its revenue from digital advertising.
251K – 367K
Our estimate — this employer did not publish a salaryOur estimate, not the employer’s. Worked out from the middle half of 88 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
Search by role, company, or anything a posting mentions.