Embed with a small number of EMEA accounts as the sole technical owner of their AI deployments on Fireworks, from scoping through production. You'll optimize model serving, run fine-tuning work, and drive infrastructure performance. This requires depth in software engineering (Python plus systems languages), machine learning (SFT, LoRA, RLHF, evaluation design), and GPU infrastructure (quantization, speculative decoding, multi-GPU serving). You need formal ML training, post-training LLM experience, and founder or startup background. Based in EMEA with customer embedding required.
Written from this posting by Neural Jobs AI. The full description is below.
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.
As an Applied Machine Learning Engineer, you are the technical owner of a customer engagement. This is what is globally known as the Forward Deployed Engineer role; at Fireworks in EMEA it carries the AMLE title. You embed with a small number of accounts, work in their environment, and own whatever stands between them and production on Fireworks: the model, the serving path, the integration, the performance. You scope the work before signature and deliver what you scoped. Depth over breadth: few accounts, owned end to end.
What we value most is first principles thinking: reasoning about unfamiliar problems from the ground up rather than from playbooks. You will be evaluated on this directly in the interview process, and strength here can outweigh gaps elsewhere in your background.
Software engineering. You ship production systems and are comfortable dropped into an unfamiliar codebase or a customer environment with a deadline. Backend and systems depth, strong Python plus at least one systems language, and fluency with AI-assisted and agentic engineering as a core part of how you build.
Machine learning. You take open models and make them better on a customer's task: SFT, LoRA and other parameter-efficient methods, reinforcement learning (RLHF, RLVR), distillation. Just as important, you design evaluations that reflect the real task and let them drive training decisions, and you know when fine-tuning is not the answer.
Infrastructure and performance. You know your GPUs: memory bandwidth, kernels, parallelism, what utilization numbers actually mean. Serving and inference optimization: quantization, speculative decoding, batching, KV cache behavior, latency vs throughput tradeoffs, multi-GPU and multi-node serving, capacity planning. You can take a deployment that works and make it fast and economical, and reason about where the bottleneck is before touching a profiler.
Customer engagement. You want to be in the room. The point of this role is creating specialized intelligence for each customer: the best model for their task, their data, their constraints. That takes real interest in the customer's problem, not just the technical one. Exceptional here looks like founder energy with customers; the floor is genuine appetite to embed, explain, and own outcomes with them.
Own Engagements: Take customer deployments from scoping through production, as the single accountable technical owner.
Embed: Work directly inside customer teams; understand their domain, data, and constraints well enough to make decisions they trust.
Train and Tune: Run fine-tuning and post-training work where the engagement needs it, from data to evaluation to a model serving in production.
Optimize: Drive serving performance and cost on Fireworks: model choice, deployment shape, and inference optimization.
Scope Honestly: Assess what is achievable before we commit, and inherit what you scope.
Product Signal: Feed what customers need back to product and engineering, with enough precision to be actionable.
Spikes in two or all three areas.
For the ML spike: formal training (PhD, PhD in progress, or a strong MSc) and experience post-training LLMs specifically.
For the infrastructure spike: experience with GPU infrastructure, distributed serving, or inference engines.
Founder or founding engineer experience.
Experience in a startup or fast-paced environment.
Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.
Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.
Check my match
Fireworks AI is a generative-AI inference platform focused on fast, low-cost serving and fine-tuning of open models for production workloads.
Founded in 2022 by former PyTorch engineers, it optimises model serving for latency, throughput and cost.
United Kingdom
Search by role, company, or anything a posting mentions.