- Meta
- Sunnyvale, CA
- Full-Time
- 82 days ago
- $154,003–$217,000 / year
Software Engineer, ML Compiler and Performance.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Software Engineer, ML Compiler and Performance: our view in 3 lines...
- The Role:This role is for a software engineer focused on performance work for AI training and inference workloads on MTIA chips.
- The Person:The person will identify performance bottlenecks, analyze and report results, develop optimizations, and work with compiler and client teams on AI model performance.
- Requirements:The ideal candidate has experience developing and deploying optimizations at the level of PyTorch/Aten, optimizing multi-node distributed compute, and optimizing runtimes and/or kernels for accelerator platforms.
About the role
Meta's Training and Inference Accelerators (MTIA) team is developing novel HW to enable efficient execution of AI training and inference workloads. In this role, you will have end-to-end responsibility for the performance of in-production AI models in their transition from stock HW to MTIA chips, with a focus on models that require multi-node compute.
To learn more about MTIA, explore the links below:
- https://ai.meta.com/blog/meta-training-inference-accelerator-AI-MTIA/
- https://ai.meta.com/blog/next-generation-meta-training-inference-accelerator-AI-MTIA/
- https://dl.acm.org/doi/full/10.1145/3695053.3731409
- Identifying bottlenecks and quantifying opportunities for improving performance
- In-depth, end-to-end performance analysis and reporting
- Developing optimizations to address identified bottlenecks
- Optimizing compute/communication overlap
- Work closely with other compiler teams as well as client teams (Recommendation Systems, Generative AI, etc)
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- Experience developing and deploying optimizations at the level of PyTorch/Aten or comparable stacks
- A PhD degree and 2+ years in-domain experience
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Experience optimizing multi-node distributed compute
- Experience optimizing runtimes and/or kernels for accelerator platforms
- A Master's degree and 4+ years of in-domain experience
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
$154,003/year to $217,000/year + bonus + equity + benefits

