logo

AI Job Summary

You'll lead the agent intelligence layer for an early-stage AI company, owning core architecture decisions and team leadership while writing production code. The role requires 7+ years of software engineering with at least 2 years building LLM-based agents taking real-world actions, deep expertise in agentic systems architecture, and demonstrated experience shipping AI tooling on top of engineering software like CAD or PLM systems. Strong Python skills, rigorous evaluation framework experience, and hands-on technical leadership are essential. On-site in San Francisco; $160,000–$250,000 annually plus equity.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

About the Role

This is a senior technical leadership role at an early-stage AI software company building intelligent agents for hardware engineers, sitting at the intersection of applied agentic AI, user research, and product delivery. You will own the core agent intelligence layer that turns engineers' intent into reliable, cost-efficient multi-step workflows across desktop engineering tools, reporting directly to the CTO. The work you do here will determine the real-world value this product delivers to enterprise customers.

What You'll Do

  • Lead development of the agent intelligence layer that executes multi-step workflows across complex desktop engineering software.

  • Serve as technical lead for a small team of AI engineers, a user researcher, and domain expert contractors.

  • Own the full product loop: define agent capabilities from user stories, build implementations, and benchmark against real workflows.

  • Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating on completion metrics.

  • Set and enforce per-task token budgets and track cost per completed workflow to ensure commercial viability.

  • Build rigorous, reproducible evaluation infrastructure grounded in validated user stories.

  • Lead user story mapping and validation through interviews and close collaboration with domain experts.

  • Translate validated user stories into testable evals, closing the loop between user research and benchmarking.

  • Own agent architecture decisions: tool-calling, state management, error recovery, model routing, and context management.

  • Act as a player-coach: write production code, review designs, unblock the team, and raise engineering standards.

  • Collaborate cross-functionally with integrations, product, and customers during POCs to align agent behavior with real-world usage.

What We're Looking For

  • 7+ years of software engineering experience, including at least 2 years building LLM-based agents that take real-world actions.

  • Exceptional technical depth in complex agentic systems, not just familiarity.

  • Strong hands-on individual-contributor capability as a player-coach rather than a primarily managerial profile.

  • Deep experience with LLM application architecture: model selection, context management, retrieval, tool calling, and orchestration.

  • Experience building rigorous evaluation and benchmarking frameworks for agent task completion, cost efficiency, and failure modes.

  • Strong Python proficiency and familiarity with LLM function calling, tool APIs, observability/tracing, and evaluation tooling.

  • Experience setting technical direction and reviewing code for a small engineering team while continuing to write production code.

  • Demonstrated experience shipping AI or LLM tooling on top of proprietary engineering data or desktop engineering software, such as agents over CAD/PLM APIs, MCP servers over BOM and revision data, or RAG over firmware, schematics, or manufacturing telemetry. General-purpose chatbot or web-app RAG work alone is not sufficient.

  • Experience with desktop automation or programmatic control of applications (COM or similar).

  • Background in mechanical engineering, CAD/CAE, PLM, or an adjacent engineering-software domain is a strong plus.

  • Understanding of enterprise deployment constraints on locked-down corporate workstations.

  • Track record contributing to public benchmarks, publications, or open-source agentic AI projects is a plus.

Compensation & Benefits

Base salary: $160,000 to $250,000 USD annually, plus equity. Visa sponsorship is not available for this role.

Location

On-site in San Francisco, California, United States.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Staff Engineer, Agentic AI
  • Experience 8-9 years
  • Work type On-site
  • Location United States
Check my match (free)
Clera
AI & Machine Learning · 500+ Members · United States

Clera is an AI recruiting platform that introduces candidates directly to hiring managers at the companies they want to work for.

All jobs at Clera

713 more Software Engineer roles in San Francisco

Job Overview
Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
3 days ago
Job Type
Full Time
Experience
8-9 years

Share This Job: