HOME Cloud & Infrastructure Engineering Staff Engineer - Customer Facing
  • Ddn
  • Santa Clara Colocation,
  • Full-Time
  • 18 days ago
Ddn VERIFIED EMPLOYER

Staff Engineer - Customer Facing.

Cloud & Infrastructure Engineering Full-Time

Staff Engineer - Customer Facing: our view in 3 lines...

  • The Role:This role is for a staff engineer focused on customer-facing work for a software-defined storage platform built for AI and accelerated computing.
  • The Person:The person will handle complex customer escalations, lead incident response and root-cause analysis, debug distributed storage issues, reproduce customer problems, and feed findings into product and reliability improvements.
  • Requirements:The ideal candidate has significant experience in enterprise storage, distributed systems or cloud infrastructure, deep understanding of S3, POSIX, NFS and storage performance, strong Linux systems knowledge, and coding ability in Python or C++.

About the role

DDN is seeking a Staff Engineer to join our Infinia Core team. This is a hands-on technical role combining deep distributed-systems engineering with direct engagement with customers running Infinia in production.

You'll own complex technical escalations end-to-end, from root-cause analysis and incident response through to mitigation, customer communication and product improvements. You'll also help shape engineering best practice, mentor other engineers and drive the use of AI and automation to improve reliability and diagnostics.

If you love the technical depth but want to stay behind the curtain, this probably isn't the right fit but if you want to combine serious engineering with real customer impact, read on.

About Infinia

Infinia is DDN's next-generation, software-defined storage platform, built from the ground up for AI and accelerated computing. It combines separate control and data planes, all-flash performance, sub-millisecond latency and multi-tenancy for demanding enterprise and hyperscale AI and GPU workloads.

What You'll Do

  • Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.

  • Own complex customer escalations from diagnosis through to resolution, mitigation and RCA.

  • Lead live incident response, war rooms and cross-functional investigations with Engineering, QA and Field teams.

  • Debug complex distributed-systems, storage and performance issues across the system, protocol and application layers.

  • Reproduce customer issues and feed findings into product and reliability improvements.

  • Develop runbooks, troubleshooting guidance and performance-tuning practices.

  • Act as a technical authority on Infinia internals, mentoring engineers and influencing architectural best practice.

  • Partner with Field CTOs, Solutions Architects and Sales Engineers on strategic customer issues.

  • Use AI, automation and observability to improve diagnostics, reliability and MTTR.

  • Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.

  • This position requires participation in an on-call rotation to provide after-hours support as needed.

What You'll Bring

Must-Haves

  • Significant experience in enterprise storage, distributed systems or cloud infrastructure, with technical leadership at Senior or Staff level.

  • Deep understanding of file systems and storage technologies, including S3, POSIX, NFS and storage performance.

  • Strong Linux systems knowledge, including kernel-level troubleshooting and debugging.

  • Strong coding ability in Python or C++.

  • Proven ability to diagnose complex issues using tools such as strace, tcpdump and perf.

  • Genuine interest in working directly with customers and taking ownership of complex problems through to resolution.

Nice-to-Haves

  • Experience with DDN, VAST, Weka or similar scale-out storage/file systems.

  • Familiarity with observability platforms such as Prometheus, Grafana, ELK or OpenTelemetry.

  • Knowledge of replication, consistency models and data integrity mechanisms.

  • Experience supporting AI/ML, LLM training or other high-performance computing environments.

  • Experience using AI tools for log analysis, troubleshooting, automated RCA or reducing MTTR.

Published September 10, 2026
Location Santa Clara Colocation, United Kingdom
Job Type Full-Time