- IPolarity LLC
- Whippany, NJ
- Full-Time
- 34 days ago
Technical Program Manager SRE Kubernetes Cloud AI.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Technical Program Manager SRE Kubernetes Cloud AI: our view in 3 lines...
- The Role:This role is for an experienced Technical Program Manager focused on SRE, Kubernetes, cloud, and AI initiatives.
- The Person:The person will lead SRE, cloud, Kubernetes, and platform reliability initiatives, manage Kubernetes platforms, build automation tools, improve observability, and drive incident management, high availability, and disaster recovery.
- Requirements:The ideal candidate has 10–14 years of experience and strong hands-on understanding of Kubernetes, GKE, SRE, cloud automation, Python, Java, Node.js, Datadog, Splunk, Grafana, AppDynamics, Apigee, and GenAI/LLM/AIOps.
About the role
Hiring: Technical Program Manager – SRE / Kubernetes / Cloud / AI
We are hiring an experienced Program Manager to lead complex technology initiatives focused on platform reliability, observability, automation, and application performance across large-scale systems.
Location: Hartford, CT
Experience: 10–14 Years
Experience: 10–14 Years
Key Responsibilities
• Lead complex SRE, cloud, Kubernetes, and platform reliability initiatives
• Manage and improve Kubernetes platforms, including GKE and Rancher RKE2
• Build automation and operational tools using Python, Java, and Node.js
• Leverage Generative AI/LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident response, and operational automation
• Implement reliable API and microservices solutions using Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies
• Drive high availability, active-active deployments, disaster recovery, and multi-datacenter Kubernetes environments
• Develop and enhance observability using Splunk, Datadog, Grafana, and AppDynamics
• Drive SRE best practices, incident management, reliability engineering, and continuous improvement
• Collaborate with cross-functional engineering, operations, and business teams to deliver measurable outcomes
• Lead complex SRE, cloud, Kubernetes, and platform reliability initiatives
• Manage and improve Kubernetes platforms, including GKE and Rancher RKE2
• Build automation and operational tools using Python, Java, and Node.js
• Leverage Generative AI/LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident response, and operational automation
• Implement reliable API and microservices solutions using Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies
• Drive high availability, active-active deployments, disaster recovery, and multi-datacenter Kubernetes environments
• Develop and enhance observability using Splunk, Datadog, Grafana, and AppDynamics
• Drive SRE best practices, incident management, reliability engineering, and continuous improvement
• Collaborate with cross-functional engineering, operations, and business teams to deliver measurable outcomes
Ideal Candidate
We are looking for a Technical Program Manager / SRE leader with strong hands-on understanding of:
Kubernetes & GKE
SRE & Cloud Automation
Python / Java / Node.js
Datadog / Splunk / Grafana / AppDynamics
Apigee & API/Microservices Reliability
High Availability & Disaster Recovery
GenAI / LLM / AIOps
Technical Program & Cross-Functional Leadership
We are looking for a Technical Program Manager / SRE leader with strong hands-on understanding of:
Kubernetes & GKE
SRE & Cloud Automation
Python / Java / Node.js
Datadog / Splunk / Grafana / AppDynamics
Apigee & API/Microservices Reliability
High Availability & Disaster Recovery
GenAI / LLM / AIOps
Technical Program & Cross-Functional Leadership
Flexible work from home options available.

