- Beam
- San Francisco, CA
- Full-Time
- 75 days ago
- $140,000 – $200,000
Applied AI Research Engineer.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Applied AI Research Engineer: our view in 3 lines...
- The Role:This role is for an engineer focused on LLM inference research and optimisation for an AI inference platform.
- The Person:The person will own inference research, reduce cost per token and latency, optimize production workloads, and turn customer learnings into platform improvements.
- Requirements:The ideal candidate has a systems or research background in LLM inference, deep understanding of LLM serving, and experience shipping products or research in production-like scenarios.
About the role
Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.
About the Role
We’re looking to hire someone to own inference research hands-on and find ways to lower cost per token and latency on our customer workloads.
- Low-level inference optimization, from speculative decoding, quantization, KV-cache and memory management
- Work directly with customers to optimize their production workloads, and apply your learnings to our platform as product improvements
- High-level of autonomy to find the highest upside bets and guide the future of our inference platform based on your work
Skills & Experience
- Systems or research background in LLM inference
- Deep understanding of LLM serving, from the kernel to the scheduler
- History of shipping products or research that people use in production-like scenarios, whether academic or industry
- Excited to collaborate closely with customers
- Enthusiasm for developer tools, cloud native technologies, and open source software
Benefits
- Competitive salary and meaningful equity
- Join a fast-growing pre-series A company at the ground floor
- Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
- Opportunities to participate in events across the cloud native community
- Fitness stipend, learning budget, and much, much more
Cloud computing is broken.
AI has introduced a new generation of workloads, like GPU inference, sandboxes, and agents. These aren't ordinary applications that can be run as Lambdas, or Dockerized apps on VMs: they're massive, stateless containers that need to spin up in <1s, often across multiple clouds and regions.
Today, engineers are hacking together infra that breaks under real-world loads. That's where we come in.
Our mission is to build the world's best compute platform for AI. Our first product is a serverless inference platform, used by companies like Coca Cola, Geospy and hundreds more. We've built our own container runtime, called beta9, which is designed for launching GPU-backed containers in under 1s.
We're a small, highly-technical team, with backgrounds in distributed systems and robotics. We've raised $7M from YC, Tiger, Guy Podjarny (Founder of Snyk), and Jason Warner (former CTO of Github).
We're searching for intensely curious, passionate, and hard-working engineers to join our mission in rebuilding the cloud for the age of AI.
