Site Reliability Engineer
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Atlassians can choose where they work – whether in an office, from home, or a combination of the two. That way, Atlassians have more control over supporting their family, personal goals, and other priorities. We can hire people in any country where we have a legal entity.
Our organization is dedicated to driving AI innovation across all Atlassian products and platforms. We aim to deliver seamless AI experiences while establishing a robust Atlassian AI infrastructure for the future. Our purpose is to:
Develop horizontal AI capabilities and infrastructure that can be leveraged across all products.
Establish a centralized Search, Q&A, and Conversational AI system that integrates seamlessly with all Atlassian products.
Explore integrating Atlassian products with AI solutions beyond the Atlassian ecosystem.
Our team’s goal is to build the foundations to democratize AI and Machine Learning for Atlassian’s teams, customers, and ecosystem. We aim to build productive, reliable tools that empower Atlassian teams to harness AI. These tools will facilitate the development, deployment, measurement, and operation of AI & ML models and experiences.
As a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding).In this role, you are expected to:
Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
Optimize latency and throughput of model inference under real production workloads.
Build reliable, high-concurrency serving systems that serve billions of requests reliably
Benchmark, fine-tune, and accelerate inference engines.
Create robust CI/CD infrastructure for seamless model deployment and inference engine updates.
Partner with senior ML engineers to fine‑tune open-source LLMs and deploy
On your first day, we’ll expect you to have
5+ years of software engineering experience, 2+ years of system performance optimization experience
Deep low-level systems programming (C/C++ or Rust)
Experience with large-scale, high-concurrent production serving.
Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
Strong background in system optimizations: batching, caching, load balancing, parallelism.
It would be great, but not required if you have
Low-level inference optimizations: GPU kernels
Algorithmic inference optimizations: quantization, speculative decoding, distillation
Experience with testing, benchmarking, and reliability of inference services.
Experience designing and implementing CI/CD infrastructure for inference.
Benefits & Perks
Atlassian offers a wide range of perks and benefits designed to support you, your family and to help you engage with your local community. Our offerings include health and wellbeing resources, paid volunteer days, and so much more. To learn more, visit
go.atlassian.com/perksandbenefits.
About Atlassian
At Atlassian, we're motivated by a common goal: to unleash the potential of every team. Our software products help teams all over the planet and our solutions are designed for all types of work. Team collaboration through our tools makes what may be impossible alone, possible together.
We believe that the unique contributions of all Atlassians create our success. To ensure that our products and culture continue to incorporate everyone's perspectives and experience, we never discriminate based on race, religion, national origin, gender identity or expression, sexual orientation, age, or marital, veteran, or disability status. All your information will be kept confidential according to EEO guidelines.To provide you the best experience, we can support with accommodations or adjustments at any stage of the recruitment process. Simply inform our Recruitment team during your conversation with them.
To learn more about our culture and hiring process, visit go.atlassian.com/crh. In line with local law, identity verification (which may include use of biometric data) is a condition of employment with Atlassian for employment fraud purposes.
Here is what this employer asked for. Sign in and we will fill in your half.
Atlassian is a software company founded in Sydney in 2002.
Teams around the world plan, build and support their work with Jira, Confluence, Trello and Loom.
109 more Infrastructure & Platform Engineer roles in San Francisco
202K – 294K
Our estimate — this employer did not publish a salaryOur estimate, not the employer’s. Worked out from the middle half of 64 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
Already have an account? Sign in
Continue without an account and apply on the Atlassian website
Search by role, company, or anything a posting mentions.