- Apple
- Sunnyvale, CA
- Full-Time
- 8 days ago
Software Engineer, Reliability Engineering, AiDP.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Software Engineer, Reliability Engineering, AiDP: our view in 3 lines...
- The Role:This role is for a software engineer focused on reliability engineering for AI and data platform services.
- The Person:The person will build and operate large-scale platform and distributed systems, troubleshoot complex production issues, improve uptime and availability, and work on GenAI, ML, inference, and big data platform projects.
- Requirements:The ideal candidate has BS/MS in computer science or equivalent experience, 2+ years in Python, Java, or Go, 2+ years in Kubernetes or Docker, and experience with CI/CD pipelines, Linux, Spark, Flink, Iceberg, Ray, or MLflow.
About the role
The Applied Machine Learning team in AI and Data Platform org has been at the forefront of accelerating digital transformation through machine learning across Apple's enterprise ecosystem. We build and operate ML, GenAI, Inference and Data Platforms and Services to provide a comprehensive suite of capabilities—serving business-critical needs across Apple's enterprise. We work on interesting and hard challenges related to scale and performance across diverse set of open-source and cutting edge technologies.
Description
We are looking for a talented engineer to join our team and bring passion for building and operating large scale platform and distributed systems leveraging cutting edge open source technologies across hybrid cloud environments.As a software engineer in AiDP reliability engineering you will work on one or many projects related to GenAI, ML, Inference and Big data platform.
Minimum Qualifications
BS/MS in computer science or equivalent experience.
2+ years experience programming skills in one of the following areas: Python, Java, or Go.
2+ years experience in Kubernetes, Docker or other container orchestration framework.
Preferred Qualifications
Ability to read and explain open source codebase.
Experience deploying and managing CI/CD pipelines.
Strong expertise in troubleshooting complex production issues.
Should be able to understand complex architectures and be comfortable working with multiple teams.
Ability to conduct performance analysis and troubleshoot large scale distributed systems.
Should be highly proactive with a keen focus on improving uptime/availability of our mission-critical services.
Experience with big data technologies - Spark, Flink, Iceberg or emerging GenAI/ML like Ray/MLflow/model serving) technologies.
Experience of Linux, database and security concepts.

