HOME DevOps & Engineering Member of Technical Staff
  • Morph
  • San Francisco, CA
  • Full-Time
  • 58 days ago
  • $175,000 – $350,000
Morph VERIFIED EMPLOYER

Member of Technical Staff.

DevOps & Engineering Full-Time

Member of Technical Staff: our view in 3 lines...

  • The Role:This role is for a technical staff engineer focused on inference infrastructure and performance for frontier-scale open models.
  • The Person:The person will improve latency and throughput, optimize batching, scheduling, routing, quantization, and distributed execution, and build benchmarks and observability for the inference stack.
  • Requirements:The ideal candidate is strong in Python, CuTEdsl, GPU performance, memory bandwidth, collectives, and inference serving.

About the role

Morph was a 1 person company from 0 → 10M of revenue. We will be the first 10 person $10b company.

Every employee should contribute >30M of revenue/yr to the company.

The best candidates would be top 1% at multiple parts of the inference stack.

work on PD disaggregation research

Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.

What you’ll do

  • Find the gap between theoretical hardware performance and production performance
  • Trace latency and throughput regressions from the API layer down to individual kernels
  • Optimize batching, scheduling, routing, quantization, and distributed execution
  • Build benchmarks and observability that make bottlenecks obvious
  • Work with NVLink and RoCE
  • Validate that every optimization preserves model quality and correctness

You might be a fit if you

  • Have optimized complex production systems
  • Can juggle 8+ Codex/Claude/other coding agents concurrently
  • Understand GPU performance, memory bandwidth, collectives, and inference serving
  • Are strong in Python, CuTEdsl, and comfortable navigating unfamiliar codebases
  • Can turn profiling data into clear engineering decisions
  • Care about tokens per second, tokens per dollar, and correctness equally

You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.

2 day work trial

Morph builds specialized code-generation models and serves them on a custom inference stack.

Technical work involves autoresearch for kernels and custom speculative-decoding models.

Published August 1, 2026
Location San Francisco, CA
Category DevOps & Engineering  
Job Type Full-Time