We are seeking a highly motivated and experienced AI Quality Engineer to join our Retail and Wealth Risk Engineering team under the Enterprise Risk Technology platform. This role spans the full spectrum of modern AI quality engineering — from Agentic AI flow testing and RAG pipeline validation to AI safety, test automation, and performance & reliability engineering.
You will be the quality pillar for complex autonomous AI systems, ensuring they are safe, accurate, explainable, resilient, and production-ready at scale. This is a high-impact, highly technical role that requires both depth in AI/ML and breadth across testing disciplines.
Responsibilities
Agentic AI Testing
- Design and execute end-to-end test strategies for Agentic AI pipelines, including single-agent and multi-agent workflows.
- Validate agent reasoning, planning, and decision-making chains (e.g., ReAct, Chain-of-Thought, Plan-and-Execute, Reflexion).
- Test tool-use correctness — ensuring agents invoke the right tools, with correct parameters, at the right time.
- Evaluate agent memory systems (short-term, long-term, episodic) for accuracy and context retention across sessions.
- Validate agent handoff and delegation logic in multi-agent orchestration frameworks (e.g., AutoGen, CrewAI, LangGraph).
- Test termination conditions, loop detection, and infinite loop prevention in autonomous agent loops.
RAG (Retrieval-Augmented Generation) Testing
Test Automation
- Build and maintain automated test harnesses for Agentic and RAG systems, including agent trajectory replay, tool mock injection, and prompt simulation.
- Develop automated evaluation pipelines integrated into CI/CD workflows for continuous model and agent validation.
Create data validation and data quality frameworks (using Great Expectations, Deequ, or custom tooling) for training, retrieval, and inference data.
- Build prompt regression suites to detect behavioral drift across LLM versions or prompt changes.
- Implement determinism and reproducibility tests for stochastic LLM-based decisions.
- Automate vector database validation — index integrity, embedding drift, and retrieval consistency checks.
AI Safety & Security Testing
- Conduct red-teaming and adversarial testing to uncover jailbreaks, prompt injection vulnerabilities, and goal misalignment in LLM-based systems.
- Test output guardrails and content filters for unsafe, biased, toxic, or out-of-scope model behavior.
- Validate privilege escalation controls — ensuring agents do not exceed permitted actions or access unauthorized resources.
- Perform data poisoning and backdoor attack simulations to assess model robustness.
- Evaluate models for bias, fairness, and discrimination using frameworks such as AI Fairness 360 and Aequitas.
- Test PII leakage and data privacy controls in RAG and agent pipelines in accordance with GDPR, CCPA, and internal data governance policies.
- Conduct security testing aligned with the OWASP Top 10 for LLM Applications, including:
- Prompt Injection (Direct & Indirect)
- Insecure Output Handling
- Training Data Poisoning
- Insecure Plugin / Tool Design
- Sensitive Information Disclosure
- Validate constitutional AI constraints, RLHF-aligned behavior boundaries, and system prompt integrity.
- Collaborate with cybersecurity teams on AI-specific threat modeling and vulnerability management.
- Maintain safety testing playbooks and document red-team findings with severity ratings and remediation recommendations.
Performance & Reliability Testing
Domain Knowledge
Preferred Qualifications
- Experience with MCP (Model Context Protocol) or similar agentic communication standards.
- Exposure to multi-modal agent testing (agents handling text, images, code, documents).
- Experience in regulated industries (banking, finance, healthcare) with strict compliance requirements.
- Familiarity with chaos engineering tools (Chaos Monkey, Gremlin, LitmusChaos).
Education
Bachelor’s degree in Computer Science, Engineering, or a related field.
Master’s degree is a plus.
Experience
------------------------------------------------------
Job Family Group:
Technology
------------------------------------------------------
Job Family:
Technology Quality
------------------------------------------------------
Time Type:
Full time
------------------------------------------------------
Most Relevant Skills
Please see the requirements listed above.
------------------------------------------------------
Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.
------------------------------------------------------
Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.
View Citi’s EEO Policy Statement and the Know Your Rights poster.