- Mindrobotics
- Palo Alto, CA
- Full-Time
- 44 days ago
DevOps Engineer.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
DevOps Engineer: our view in 3 lines...
- The Role:This role is for a DevOps Engineer supporting the infrastructure behind physical AI and robotic systems for industrial use.
- The Person:The person will build and maintain cloud infrastructure, CI/CD pipelines, Kubernetes clusters, monitoring systems, and deployment automation for robotics software, embedded applications, and machine learning workflows.
- Requirements:The ideal candidate has 4+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering, plus strong experience administering Linux systems, Kubernetes, Docker, Terraform, and CI/CD pipelines.
About the role
About Mind:
Mind Robotics is building Physical AI for real-world industrial deployment, starting with the factory floor. We believe the hardest problems in AI are solved when researchers and engineers are hands-on with the physical world every day - and we're looking for people who are passionate about robotics, value ownership, and are excited to tackle difficult problems. Join us if you want to move beyond digital intelligence and put intelligence into motion.
About the team and the role:
At Mind Robotics, we're building generalized physical AI—robotic systems capable of dexterous, adaptive, and reasoning-intensive work in real-world industrial environments.
Our robots rely on a sophisticated software platform that spans cloud infrastructure, robotics middleware, simulation, machine learning, and deployment to physical hardware. We're looking for a DevOps Engineer to build and operate the infrastructure that enables engineers to develop, test, deploy, and monitor robotic systems reliably at scale.
You'll work across software, ML, and robotics teams to automate infrastructure, streamline deployments, improve developer productivity, and ensure our robots can be continuously updated and monitored both in the lab and in production.
Responsibilities:
-
Design, deploy, and maintain scalable cloud infrastructure using AWS, GCP, or Azure.
-
Build and maintain Infrastructure as Code using Terraform or similar tools.
-
Develop CI/CD pipelines for robotics software, embedded applications, and machine learning workflows.
-
Automate software deployment across development robots, test environments, and production fleets.
-
Improve build systems, artifact management, release processes, and version control workflows.
-
Manage Kubernetes clusters and containerized services supporting robotics applications.
-
Monitor infrastructure health, application performance, and production systems using modern observability tools.
-
Improve system reliability, uptime, security, and disaster recovery processes.
-
Build internal developer tooling that increases engineering productivity.
-
Partner with robotics, firmware, and ML teams to create reproducible development environments.
-
Support simulation infrastructure and large-scale testing pipelines.
-
Optimize cloud resource utilization and infrastructure costs.
-
Implement security best practices for infrastructure, secrets management, networking, and access controls.
-
Participate in incident response, root cause analysis, and continuous operational improvements.
Requirements:
-
4+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering.
-
Strong experience administering Linux systems.
-
Experience with cloud platforms such as AWS, GCP, or Azure.
-
Experience with Kubernetes and Docker in production environments.
-
Strong Infrastructure as Code experience (Terraform preferred).
-
Experience building CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, Buildkite, or similar platforms.
-
Experience with configuration management and automation tools.
-
Proficiency with Python, Go, or Bash for infrastructure automation.
-
Experience implementing monitoring and logging solutions (Prometheus, Grafana, Datadog, ELK, OpenTelemetry, etc.).
-
Strong understanding of networking, security, authentication, and distributed systems.
-
Experience troubleshooting production systems under operational constraints.
Nice to haves:
-
Experience supporting robotics, autonomous systems, or embedded software teams.
-
Experience deploying software to fleets of physical devices.
-
Experience supporting GPU infrastructure and ML training clusters.
-
Experience with NVIDIA GPUs, CUDA environments, or distributed training systems.
-
Experience managing artifact repositories and large binary assets.
-
Experience supporting simulation environments.
-
Knowledge of software supply chain security and SBOMs.
-
Experience with edge computing or IoT deployments.

