Senior Manager, Forward Deployed Engineering (UAE)
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Own the real-time voice layer for a conversational AI team, building streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony systems while optimizing end-to-end latency under 800 milliseconds. You need at least five years shipping production software, including two-plus years with voice or real-time audio systems, plus strong Python or TypeScript skills and hands-on experience with audio stacks like LiveKit, Pipecat, or Twilio. Fully remote with core overlap 13:00–17:00 UTC. $96,000 annually; visa sponsorship unavailable.
Written from this posting by Neural Jobs AI. The full description is below.
As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency.
Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.
Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.
Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.
Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.
Compare voice providers and models through evidence-based testing, and make changes based on results.
Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.
Work directly with founders and make technical decisions in a fast-moving team.
At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems.
Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.
Strong Python or TypeScript skills and comfort working in both.
Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.
Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.
Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions.
Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus.
Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available.
Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.
Here is what this employer asked for. Sign in and we will fill in your half.
Clera is an AI recruiting platform that introduces candidates directly to hiring managers at the companies they want to work for.
Search by role, company, or anything a posting mentions.