Senior Software Engineer, AI Infrastructure
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Rebuild how the world works, to make institutions work better for the people they serve.
Brain Co. builds AI-native operating systems for large, regulated institutions. Each system is built for a specific industry, powered by agents that push real workflows forward. Underneath it all is Atlas, our proprietary platform that keeps customers in control, secure by design, and never locked into one model.
Brain Co. is entering its next phase of production deployments on a national scale with an elite team built from Palantir, Google, Meta, and Nvidia, and a growing footprint across government, insurance, health, and financial services.
Joining now means shaping both the company and a new category of applied AI. Every project here ships to production and is expected to create measurable customer value and impact.
You'll work alongside exceptional peers on some of the hardest problems in applied AI. It’s the kind of work you'll still be proud of in ten years from now.
You'll join the team that builds and enables agentic workflows across Brain Co. For every engineer, operator, and business team internally, and for the production AI systems we deploy to governments, healthcare systems, and critical industries. This is a platform role at the center of the company's agent-first strategy: you'll build foundational systems used by every engineering team, and the bar is product-grade because the entire company depends on them.
Own the foundations of how LLMs are used across the company: cost visibility and controls, data privacy, identity and access, routing, and the security posture around all provider traffic.
Design the sandboxing, orchestration, audit, and guardrail layers that product teams build their agents on, so verticals don't need to invent their own abstraction.
Solve the hard problems: prompt-injection defenses, scoped credentials, kill switches, multi-tenant isolation (including VM-level pod isolation), and runaway-cost controls.
Design the orchestration, isolation, and resource models that make this viable: cold-start vs. always-on tradeoffs, credential and token lifecycle, fan-out and fan-in patterns, fairness and quota enforcement across tenants, and the observability needed to debug at that volume.
Make AI-assisted development a first-class platform layer: coding agents that review and ship code, automate CI, refactor at scale, and run as background workers across the codebase, together with the canonical scaffolding and guardrails that govern them.
Build the systems that let every team; engineering, operations, and the business, run their own agents reliably and safely against the tools they already use, with the right credentials, scheduling, memory, and audit underneath.
End-to-end ownership: architecture, implementation, rollout, observability, on-call, and iteration based on internal user feedback.
Partner closely with security, infrastructure, and product teams to make agent deployments safe by default.
Have 5+ years building backend systems in production, with deep proficiency in at least one of Python, TypeScript, Go, or Rust.
Bring strong fundamentals in distributed systems: consistency, idempotency, retries, failure modes, queueing, scheduling.
Have designed and operated APIs and services that other engineers depend on.
Have a proven track record building shared infrastructure, internal platforms, or developer-facing services that real users adopted.
Have strong intuition for developer experience, long-term maintainability, and where to draw abstraction boundaries.
Are comfortable owning the full lifecycle: writing the design doc, shipping the MVP, hardening it, and driving adoption across the company.
Have owned services with real uptime and operational responsibility, and are comfortable with observability stacks, incident response, and SLOs.
Bring cloud-native experience: Kubernetes, infrastructure-as-code, OAuth/OIDC, secrets management.
Experience building or operating LLM infrastructure: gateways, inference systems, prompt routing, cost attribution, evaluation harnesses.
Experience with agent frameworks, tool-use systems, or sandboxed code execution.
Security instincts around prompt injection, supply-chain risk in agent ecosystems, and credential scoping for autonomous systems.
Background in multi-tenant, regulated, or government deployments (HIPAA, SOC2).
Open-source contributions to AI infrastructure, agent tooling, or developer platforms.
Collaborate with industry veterans from Tesla, DeepMind, Databricks, and more
Accelerate your career with ownership based on impact, not tenure
Earn competitive compensation + meaningful equity in a high-growth company
Thrive in a culture built on speed, curiosity, and impact
Competitive salary plus equity
Daily lunches
Commuter benefits
401(k)
Medical, Dental and Vision
Unlimited PTO
Here is what this employer asked for. Sign in and we will fill in your half.
Brain Co. builds agent-native operating systems for modern institutions.
167 more Infrastructure & Platform Engineer roles in San Francisco
203K – 310K
Our estimate — this employer did not publish a salaryOur estimate, not the employer’s. Worked out from the middle half of 97 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
Already have an account? Sign in
Continue without an account and apply on the Brain Co. website
Search by role, company, or anything a posting mentions.