- Periodic Labs
- Menlo Park, CA
- Full-Time
- 47 days ago
- $250,000 – $350,000
Research Engineer - Midtraining.
Before you go
Before you leave us, sign up for our email alerts
We don't do job spam, just the best digital jobs delivered straight to your inbox.
Research Engineer - Midtraining: our view in 3 lines...
- The Role:This role is for a research engineer working on frontier models for scientific discovery in materials, energy, and related areas.
- The Person:The person will curate scientific data, generate synthetic data, build evaluations, run large-scale training experiments, and develop tools to study how data choices affect model intelligence.
- Requirements:The ideal candidate has experience training LLMs on curated mixes of trillions of tokens, working on a dedicated evals team, using self-distillation or on-policy distillation, and applying scaling laws and compute-optimal hyperparameters.
About the role
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.
About the Role
We're training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Your work will also lay the groundwork for our pre-training efforts down the line.
What You'll Do
-
Identify, process, and curate novel sources of scientific data for large-scale model training.
-
Generate high-quality synthetic data to fill gaps in scientific knowledge and reasoning.
-
Build evaluations that correlate with downstream scientific task performance, working closely with RL researchers, physicists, and chemists.
-
Develop and apply techniques such as self-distillation and on-policy distillation to improve model capability.
-
Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.
-
Build tools for yourself and the team to investigate how data choices shape model intelligence.
You Will Thrive in This Role If You Have
-
Experience training LLMs on curated mixes of trillions of tokens.
-
Experience on a dedicated evals team supporting a large production training run.
-
Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline.
-
Experience with scaling laws and compute-optimal hyperparameters.
-
Comfort working across data, evals, and training infrastructure.
Especially Strong Candidates May Also Have
-
Experience optimizing throughput and reliability for large-scale distributed training runs.
-
A background in AI for science or training on specialized domain data (e.g., protein, materials, or other scientific datasets).
-
Experience creating evals or synthetic data for non verifiable tasks and tracking performance over live runs.
Mechanics
-
Minimum education: Bachelor's degree or similar experience
-
Location: Menlo Park, CA (Soon: San Francisco, too)
-
Compensation: $250,000–$350,000 + equity
-
Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

