- New York, NY
- Full-Time
- <48 Hours
Staff Software Engineer, Colossus SRE.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Staff Software Engineer, Colossus SRE: our view in 3 lines...
- The Role:This role is for a staff-level software engineer focused on SRE work for Google Cloud storage systems.
- The Person:The person will improve the lifecycle of services, support system design and launch reviews, scale systems with automation, and drive scalability, reliability, efficiency, and security for Colossus SRE.
- Requirements:The ideal candidate has 8 years of software development experience, 3 years leading projects, and 3 years designing, analyzing, and troubleshooting distributed systems, with experience in distributed storage.
About the role
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development in one or more programming languages.
- 3 years of experience leading projects.
- 3 years of experience designing, analyzing, and troubleshooting distributed systems.
Preferred qualifications:
- Master's degree in Computer Science or Engineering.
- Experience with technical direction of team members, performance, system architecture, systems data analysis and debugging
- Experience with distributed storage.
About the job:
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
Colossus Site Reliability Engineering (SRE) is responsible for the low-level filesystems that form the basis of Google's storage systems: one of the largest data storage systems in the world. Our goal is to provide efficient and scalable storage for Google's data. We work closely with our dev partners to provide system analysis and design, automation, tooling, and customer consultation to meet these goals.
Our team works closely with developers across storage systems up and down the stack including D, Colossus, Chronicle, Blobstore/Google Cloud Storage (GCS) and others.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Engage in and improve the whole lifecycle of services—from inception and design, through to deployment, operation and refinement.
- Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning and launch reviews.
- Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity.
- Partner with Development and SRE leadership to design and drive efforts to increase the scalability, reliability, efficiency and security of cluster local filesystem services.
- Drive technical direction and planning of integrating new services and flows into the Colossus SRE support.
