- Mountain View, CA
- Full-Time
- 29 days ago
GDM Staff Software Engineer, Agent Data Quality, DeepMind.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
GDM Staff Software Engineer, Agent Data Quality, DeepMind: our view in 3 lines...
- The Role:This role is for a staff software engineer focused on agent data quality and evaluation systems for Google DeepMind's AI research.
- The Person:The person will build agent testing systems, create quantitative benchmarks and automated evaluation frameworks, develop agentic data pipelines, and productionize tools for metrics, judges, and automated insights.
- Requirements:The ideal candidate has a bachelor's degree in Computer Science or Engineering, or equivalent practical experience, and experience with Large Language Models, NLP, or Generative AI.
About the role
Minimum qualifications:
- Bachelor's degree in Computer Science or Engineering, or equivalent practical experience.
- Experience with Large Language Models (LLMs), NLP, or Generative AI.
Preferred qualifications:
- MBA or Master's degree in Computer Science.
About the job:
At Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.
Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Construct rigorous quantitative benchmarks and automated evaluation frameworks (including LLM-as-a-judge) to measure agent capabilities in reasoning, planning, and tool use.
- Develop end-to-end data flywheel on agent evalset curation, and automate and productionize the flywheel with agentic data pipelines.
- Develop a dedicated platform for "Self-Service Judges," allowing researchers to rapidly prototype and test new LLM-based classifiers and annotations against historical trajectory data.
- Establish a formal ecosystem for metric lifecycle management, including a sandbox for validating new metrics against a library of high-confidence calibration experiments.
- Productionize automated insights tools to identify the top movers of metric deltas across complex slices such as task type and user cohort.

