Senior Product Manager - Tech, GenAI


Location
London
Hours
Full Time
Salary
Competitive, based on experience
About the Role
Amazon's Rufus AI team is building the future of conversational shopping. Rufus helps hundreds of millions of customers find and discover products through natural language, powered by an automated quality measurement system using LLM-as-a-Judge (LLMAJ) technology. As a Senior Product Manager - Tech, you will own the quality governance, global scaling, and operational excellence of this judge portfolio. You will collaborate closely with Language Engineers, Product Managers, Data Scientists, and Engineering teams to drive end-to-end ownership of your domain, making key decisions and delivering impactful results.
This role sits at the intersection of AI evaluation, product management, and applied tooling. You will lead the governance framework for dozens of LLM judges that power critical evaluation metrics used for release decisions, competitive benchmarking, and leadership reporting. Your responsibilities include driving localization of judges from en-US to 5+ international marketplaces, facilitating model evaluation and debugging workflows, and building purpose-built tools and agents to automate governance operations at scale.
Key responsibilities include:
- Owning the LLMAJ governance framework: judge registry, versioning standards, quality validation gates, deprecation policies, and agreement rate monitoring across the judge portfolio
- Leading international LLMAJ expansion: driving judge localization, identifying coverage gaps, defining remediation plans, and validating judge quality per locale
- Facilitating model evaluation and debugging by collaborating with Language Engineers and Scientists to trace response quality issues and root-cause judge disagreements or regressions
- Building automation tools and agents to streamline governance workflows, judge monitoring, data extraction, and reporting
- Defining and owning partner-facing quality metrics powered by LLMAJ, including defect rates, agreement rates, and evaluation dimension reporting
- Driving human-in-the-loop validation workflows, coordinating between evaluation platforms and annotation teams to maintain judge calibration
- Enforcing discipline on evaluation requests by requiring data-driven problem statements, clear scoping, and definition of done before work begins
- Writing business requirements documents, contributing to leadership updates, and representing LLMAJ governance in cross-functional forums
A typical day includes monitoring agreement rate dashboards for drift, triaging alerts, debugging judge regressions with engineers, presenting international judge coverage in reviews, shipping governance automation updates, and managing evaluation request scopes.
About the team
We are responsible for measuring whether Amazon's AI shopping assistant delivers quality experiences. Our team builds LLM judges, defines quality standards, and runs evaluations that directly influence what ships to hundreds of millions of customers. We work fast, value measurement rigor, and believe that automatic quality measurement is essential to scalable improvement.
Experience
- Bachelor's degree
- Experience in technical product management, program management, or engineering
- Proven track record owning and driving roadmap strategy and product definition
- Experience delivering end-to-end product features and managing tradeoffs
- Ability to contribute to engineering discussions on technology decisions and product strategy
- Experience representing and advocating for diverse customers and stakeholders during executive-level prioritization and planning
About you
- Strong analytical and problem-solving skills
- Comfortable working autonomously with high ownership
- Collaborative mindset to work across engineering, science, and product teams
- Passionate about AI, product quality, and operational excellence
- Detail-oriented with a focus on data-driven decision making
Qualifications
- Preferred experience with analytical tools such as Tableau, Qlikview, or QuickSight
- Experience building and driving adoption of new tools and automation frameworks

