- Sunnyvale, CA
- Full-Time
- <48 Hours
Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI: our view in 3 lines...
- The Role:This role is for a senior site reliability engineer working on Workspace AI infrastructure for Google Cloud.
- The Person:The person will own AI infrastructure architecture, lead production incident response, drive blameless postmortems, and work with development teams on reliable, scalable, cost effective systems.
- Requirements:The ideal candidate has a bachelor’s degree in Computer Science or a related field, 8 years of software development experience, 4 years of technical leadership, and experience with machine learning/AI in a software development environment.
About the role
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development in one or more programming languages.
- 4 years of experience leading projects, and providing technical leadership.
- 3 years of experience in designing, analyzing, and troubleshooting distributed systems.
- Experience with machine learning/AI in a software development environment.
Preferred qualifications:
- Master's degree in Computer Science or Engineering.
About the job:
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
As a part of the Workspace AI SRE team, your mission is to safely and responsibly enable rapid iteration of high-quality AI (e.g. Gemini) services and features using common infrastructure within Workspace. You will partner with development teams to ensure new AI-powered features for products like Gmail, Docs, Meet, and more are reliable, scalable, and efficient. As Workspace AI products move fast to stay competitive, SRE plays a critical role in building the necessary automation and safety guardrails.
Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next-generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $262000 - $364000 (USD) + 25% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Own the architecture and design of the inference and training AI infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while working closely with senior technical leads in the development teams.
- Drive AI-first development to help the team leapfrog in its current efforts, serving as the AI advocate who introduce new and novel ways of working.
- Lead the on-call response to production incidents, driving blameless postmortems and ensuring preventative measures are implemented.
- Engage broadly with SRE leaders across other product teams to ensure the best scalable solutions for all of Google, acting as the conduit to bring solutions in and take them across Google.

