logo

AI Job Summary

Design and build agent environments, develop AutoRaters calibrated against human evaluation, and create diagnostic tools to identify agent failure modes. Research how agents can automatically generate verifiers and iteratively improve. Build datasets and reward signals from multi-turn trajectories to train frontier models. Requires two years with agentic AI workflows, one year with machine learning and deep learning, software engineering experience, and proficiency in Python, TensorFlow, or PyTorch.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

  • Design, build, and scale realistic agent environments and task suites. Research and develop Agentic AutoRaters (AR), calibrate them against human evaluation, and benchmark agent capabilities against industry-leading frontier models.
  • Enable agents to automatically generate and iterate on verifiers (such as automated test cases), facilitating effective exploration and iterative problem-solving.
  • Develop trajectory analysis frameworks and diagnostic tooling to identify root-cause agent failure modes (e.g., passivity, hallucination, or brittle tool execution), and automatically optimize agent harnesses to drive continuous self-improvement.
  • Harvest complex multi-turn interaction trajectories into high-quality datasets and reward signals to power SFT and RL flywheels for frontier Gemini models.

Minimum qualifications:

  • 2 years of experience with agentic AI workflows, frameworks, and approaches.
  • 1 year of experience with machine learning and deep learning.
  • Experienece in software engineering and cloud-based development.
  • Experience with Python, TensorFlow, PyTorch, or similar ML frameworks.

Preferred qualifications:

  • In-depth knowledge of machine learning algorithms, including supervised learning and reinforcement learning.
  • 2 years of experience with agentic AI workflows, frameworks, and approaches.
  • 1 year of experience with machine learning and deep learning.
  • Experienece in software engineering and cloud-based development.
  • Experience with Python, TensorFlow, PyTorch, or similar ML frameworks.
Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Research Engineer, Advancing Agent Quality, DeepMind
  • Experience 3-4 years
  • Work type On-site
  • Location United States
Check my match (free)
Google
AI Research Lab · 500+ Members · Mountain View, CA, United States

Google builds internet, software, cloud, and AI products used by consumers, developers, and organizations. Its portfolio includes Search, YouTube, Android, Chrome, Maps, Gmail, Workspace, Google Cloud, advertising platforms, devices, and Gemini AI products. The company develops large-scale computing infrastructure and research that power information retrieval, communication, productivity, media, navigation, and machine learning. Google is the largest operating business within Alphabet and earns a substantial share of its revenue from digital advertising.

All jobs at Google
Job Overview

Approx. salary range

251K – 367K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 88 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
6 days ago
Job Type
Full Time
Experience
3-4 years

Share This Job: