- FactSet Europe Limited
- London, Greater London
- Full-Time
- <72 Hours
Lead Site Reliability Engineer (Kubernetes Required) - Hybrid.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Lead Site Reliability Engineer (Kubernetes Required) - Hybrid: our view in 3 lines...
- The Role:This role is for a Lead Site Reliability Engineer focused on keeping production systems reliable, scalable and performant in a financial data and analytics environment.
- The Person:The person will monitor production systems, respond to incidents, run post-mortems, define SLOs and SLIs, build automation, support on-call, and contribute to capacity planning, performance optimisation and runbooks.
- Requirements:The role requires Kubernetes, Helm, cloud platforms such as AWS, GCP or Azure, CI/CD tooling, Prometheus, Grafana, OpenTelemetry, Terraform, Pulumi, Ansible, Puppet, Chef, Python, Go and Bash.
About the role
FactSet creates flexible, open data and software solutions for over 200,000 investment professionals worldwide, providing instant access to financial data and analytics that investors use to make crucial decisions. Â
At FactSet, our values are the foundation of everything we do. They express how we act and operate, serve as a compass in our decision-making, and play a big role in how we treat each other, our clients, and our communities. We believe that the best ideas can come from anyone, anywhere, at any time, and that curiosity is the key to anticipating our clients’ needs and exceeding their expectations. Â
About the RoleÂ
We are looking for a skilled and motivated Lead Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our systems and services. You will work closely with development and operations teams to build and maintain robust infrastructure, automate processes, and drive engineering best practices.Â
Â
Key ResponsibilitiesÂ
- Monitor, maintain, and improve the reliability and availability of production systemsÂ
- Respond to and resolve incidents, conducting thorough post-mortems to prevent recurrenceÂ
- Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)Â
- Collaborate with development teams to build reliability into services from the ground upÂ
- Design and implement automation to reduce toil and improve operational efficiencyÂ
- Participate in an on-call rotation to support critical systemsÂ
- Contribute to capacity planning and performance optimization effortsÂ
- Document systems, processes, and runbooks to support the wider teamÂ
Â
Â
Required Technical Skills:
Kubernetes (Required)Â
- Hands-on experience deploying, managing, and troubleshooting workloads in KubernetesÂ
- Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and IngressÂ
- Experience with Kubernetes cluster management and administrationÂ
- Familiarity with Helm for application packaging and deploymentÂ
- Understanding of Kubernetes networking, storage, and security best practicesÂ
- Bachelors degree in computer science or relevant degree.
- Willing to work a hybrid model
- Must be fluent in English both verbal and written
- undefined
Â
Additional Technical SkillsÂ
- Cloud Platforms: (e.g. AWS, GCP, Azure)Â
- CI/CD Tooling: (e.g. GitHub Actions, ArgoCD, Harness)Â
- Monitoring & Observability: (e.g. Prometheus, Grafana, Coralogix, OpenTelemetry)Â
- Infrastructure as Code: (e.g. Terraform, Pulumi)Â
- Config Management: (e.g. Ansible, Puppet, Chef)Â
- Programming/Scripting: (e.g. Python, Go, Bash)Â
Â
Â
Soft Skills & General RequirementsÂ
- Strong problem-solving and analytical skills with a methodical approach to troubleshootingÂ
- Excellent communication skills with the ability to collaborate across technical and non-technical teamsÂ
- A proactive mindset with a focus on automation and continuous improvementÂ
- Ability to work effectively under pressure, particularly during incident responseÂ
- Commitment to a blameless culture and continuous learningÂ
Â
Â
Nice to HaveÂ
- Experience contributing to open-source projectsÂ
- Familiarity with SRE principles as defined by the Google SRE handbookÂ
- Previous experience in a DevOps or Platform Engineering roleÂ
Company Overview:Â
FactSet (NYSE:FDS | NASDAQ:FDS) helps the financial community to see more, think bigger, and work better. Our digital platform and enterprise solutions deliver financial data, analytics, and open technology to more than 8,200 global clients, including over 200,000 individual users. Clients across the buy-side and sell-side, as well as wealth managers, private equity firms, and corporations, achieve more every day with our comprehensive and connected content, flexible next-generation workflow solutions, and client-centric specialized support. As a member of the S&P 500, we are committed to sustainable growth and have been recognized among the Best Places to Work in 2023 by Glassdoor as a Glassdoor Employees’ Choice Award winner. Learn more at www.factset.com and follow us on X and LinkedIn.Â
At FactSet, we celebrate difference of thought, experience, and perspective. Qualified applicants will be considered for employment without regard to characteristics protected by law.Â

