logo

AI Job Summary

This role focuses on building evaluation systems and quality measurement frameworks for autonomous agent capabilities. You'll architect automated pipelines to measure agent performance across email, calendar, browser, and business software tasks, define success criteria for complex multi-step workflows, and design grading systems including LLM-as-a-judge approaches. The position requires 5+ years in data science or ML roles with proven experience in evaluation frameworks and production metrics, strong Python and SQL skills, solid statistical knowledge, and familiarity with LLM behavior and failure modes. Based on-site in Palo Alto; visa sponsorship unavailable.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

About the Role

This role sits at the intersection of applied data science and AI product quality for a small, fast-moving AI productivity startup building autonomous agents that handle email, calendar, browser, and business software tasks. You will own the measurement of agent quality end-to-end: turning ambiguous product behavior into rigorous, actionable evaluation systems that directly guide engineering and product decisions.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering common workflows, edge cases, ambiguous requests, and adversarial scenarios.

  • Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure grader agreement, false positives, and false negatives.

  • Analyze traces, tool calls, model outputs, and production outcomes to identify root causes and build a useful failure taxonomy.

  • Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.

  • Build dashboards and release-quality signals that make evaluation results understandable and actionable for engineering, product, and leadership.

  • Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.

What We're Looking For

  • 5+ years in data science, machine learning, or analytics roles, with a focus on evaluation systems, metrics frameworks, or quality measurement for production systems.

  • Demonstrated experience designing and implementing evaluation frameworks, grading systems, and success criteria for ML or AI systems in production.

  • Strong Python and SQL proficiency with the ability to build automated data pipelines and production-quality analysis code at scale.

  • Solid statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems.

  • Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance.

  • Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes of language model systems.

  • Ability to connect quantitative patterns to individual system traces and identify failure origins across model, prompt, context, tools, data, and application logic.

  • Experience communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders.

  • Comfort operating with high ownership in ambiguous, fast-moving environments, independently turning open-ended quality questions into evaluation systems.

  • Experience with LLM-as-a-judge systems, agentic or multi-step task evaluation, or benchmarking platforms for AI systems is a strong plus.

Location

On-site in Palo Alto, California, United States. Visa sponsorship is not available for this role.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Data Scientist, Agent Evaluations & Quality
  • Experience 3-4 years
  • Work type Remote
  • Location United States
Check my match (free)
Clera
AI & Machine Learning · 500+ Members · United States

Clera is an AI recruiting platform that introduces candidates directly to hiring managers at the companies they want to work for.

All jobs at Clera
Job Overview

Approx. salary range

209K – 362K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 38 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Hires remotely in the United States.
Workplace
Remote
Job Posted:
5 days ago
Job Type
Full Time
Experience
3-4 years

Share This Job: