Staff Research Data Scientist, Google Labs
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
We are on the lookout for a Senior Data Scientist to join our Content tribe.
We are building the ratings and reviews systems — social proof — that shape how millions of people decide what to order, across dozens of markets and languages. It is a greenfield space: the current system is early, and the interesting decisions have not been made yet.
The raw material is the hard kind. Millions of short, noisy, multilingual, contradictory pieces of user-generated text, which have to become something a person can act on in two seconds on a restaurant page. Getting that right is an LLM systems problem: aspect extraction, sentiment, summarisation, model selection across providers, systematic prompt optimisation, and the evaluation infrastructure that tells you whether any change made things better.
We are honest about the situation. The team is rebuilding ownership of these systems during a transition, and a lot is undefined. That is the offer: you will not inherit a technical direction; you will set it — and you will set the LLM bar for a team with the appetite and the room to clear it.
Why is this one different?
You define the direction. Architecture, model strategy, evaluation approach — these are open questions, and they become yours.
You are the LLM authority for the squad, not a contributor to someone else's. Part of the job is raising everyone else's ceiling.
Real scale, real consequence. Your models move conversion across Delivery Hero's global platforms.
Genuinely unsolved problems. Multilingual UGC at scale, empirical multi-provider model selection, prompt optimisation, and evaluation infrastructure for generative output — none of these have a settled answer here or anywhere.
Agentic development is a first-class part of how we work, used pragmatically to move faster without losing system understanding.
Your Mission
You'll own the LLM systems behind social proof end-to-end — quality, reliability, and coverage across languages, platforms, and use cases — including the unglamorous parts: drift, miscalibration, silent quality degradation, and data issues in production.
You'll set the standard for how this team builds with LLMs. Model and provider selection based on empirical cost, latency, and quality evidence. Prompt strategy that is systematic rather than folkloric. The judgment on when an LLM is the wrong tool.
You'll drive the roadmap through problem discovery, finding high-impact gaps, quantifying the business value, and turning them into scoped initiatives — treating cost of inference as a product decision, not only an engineering constraint.
You'll define and operationalise meaningful metrics, and build the evaluation infrastructure behind them — offline eval suites, LLM-as-judge frameworks, and annotation processes — so that evaluation reflects true business value rather than misleading proxies, and a prompt change or model swap becomes a one-day decision.
You'll take prototypes to production, shaping architecture and data flows with backend and data engineering, and building the feedback loops that let the system keep improving without constant manual intervention.
You'll raise the bar beyond your own work, through best practices, mentoring, and a culture of ownership and pragmatism.
NLP & LLM systems depth. You have built LLM-based systems that reached production, preferably on noisy, multilingual, user-generated text. You can explain the architecture, the failure modes you hit, and why you made the calls you made. You have worked across more than one model family and more than one architecture, and you can reason empirically about cost, latency, and quality rather than defaulting to a familiar vendor or pattern. You know when to prompt, when to add context or retrieval, when to fine-tune, when a small task-specific model wins, and when not to use an LLM at all.
Evaluation rigour for generative systems. You know an online A/B test is too slow and too blunt to iterate an LLM system on. You have designed offline evaluation suites, used LLM-as-judge patterns while being honest about their biases, structured annotation for generative outputs, and measured faithfulness, hallucination, and quality at scale.
Observability for LLM pipelines. You instrument what you build: tracing chained calls, watching token usage and cost, and capturing intermediate outputs so a multi-step pipeline can actually be debugged. You have a view on where this is non-negotiable and where it is overhead.
Product thinking and problem framing. You turn an ambiguous business need into a well-scoped DS problem, spot the high-impact opportunity nobody put on the roadmap, and connect your technical work to outcomes people outside data science care about.
Ambiguity and an MVP instinct. You move from zero to one on imperfect information, ship in increments, and resist over-engineering.
Leverage beyond yourself. You improve how the people around you work — practices, standards, rigour — and you are the reason a team gets better at something, not just faster at it.
Also valuable
Production-grade Python and analytical SQL, with familiarity with ML lifecycle practices, orchestration, monitoring, and modern engineering standards such as version control and CI/CD.
A strong foundation in statistics and causal inference, and the instinct to challenge a result that looks too good.
Experience with data collection and labelling pipelines, including working with annotation teams.
Agentic development tools used pragmatically — including applying them to model improvement itself, such as automated prompt optimisation or LLM-assisted annotation — with the judgment to know when to verify the output.
Ensuring you and all our Heroes are looked after, happy, and healthy is always on the menu. Because if you’re in good shape, then we’re in good shape.
Ready to join our team? If you’re excited to grow, collaborate and be part of the world’s leading delivery platform, we’d love to hear from you. Apply today!
We believe diversity and inclusion are key to creating not only an exciting product, but also an amazing customer and employee experience. Fostering this starts with hiring - therefore we do not discriminate on the basis of racial identities, religious beliefs, color, national origin, gender identities or expressions, sexual orientations, age, marital or disability statuses, or any other aspect that makes you, you.
We encourage you to let us know if you need any accommodations or specific accessibility support to ensure a smooth interview experience—just let us know with an email to our Inclusion Officer at [email protected].
Severely disabled applicants with equal qualifications will be given preferential consideration.
You're welcome to share your pronouns (he/she/they) right from the start so we can address you respectfully from our first contact.
Here is what this employer asked for. Sign in and we will fill in your half.
Delivery Hero operates local delivery platforms that connect consumers with restaurants, grocery stores, and other merchants. Its marketplaces, logistics network, and quick-commerce operations support ordering and delivery across markets in Asia, the Middle East, Europe, and Latin America.
Already have an account? Sign in
Continue without an account and apply on the Delivery Hero website
Search by role, company, or anything a posting mentions.