Skip to main content

Research Scientist/Research Engineer, Midtraining

Periodic LabsResearch
Pay
$250K–$350K
per year
Work mode
On-site
Full Time
Education
Bachelor's

Menlo Park, CA at a glance

Rent
#2 of 51
$3,490/mo+46% vs US avg
Weather
#17 of 51
292 mild days0 hot · 0 cold
Income tax
#1 of 51
13.3% top rateCalifornia

What you need

  • Experience training LLMs on trillions of tokens
  • Experience on dedicated evals team for large training run
  • Hands-on self-distillation/on-policy distillation in training pipeline
  • Experience with scaling laws and compute-optimal hyperparameters
  • Comfort across data, evals, and training infrastructure

What you'll do

  • Curate novel scientific data sources for large-scale model training
  • Generate high-quality synthetic data for scientific reasoning gaps
  • Build evaluations correlating with downstream scientific task performance
  • Develop and apply self-distillation/on-policy distillation techniques
  • Design and run large-scale training experiments across thousands of GPUs

We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.

 

About the Role

We're training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Your work will also lay the groundwork for our pre-training efforts down the line.

 

What You'll Do

  • Identify, process, and curate novel sources of scientific data for large-scale model training.

  • Generate high-quality synthetic data to fill gaps in scientific knowledge and reasoning.

  • Build evaluations that correlate with downstream scientific task performance, working closely with RL researchers, physicists, and chemists.

  • Develop and apply techniques such as self-distillation and on-policy distillation to improve model capability.

  • Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.

  • Build tools for yourself and the team to investigate how data choices shape model intelligence.

     

You Will Thrive in This Role If You Have

  • Experience training LLMs on curated mixes of trillions of tokens.

  • Experience on a dedicated evals team supporting a large production training run.

  • Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline.

  • Experience with scaling laws and compute-optimal hyperparameters.

  • Comfort working across data, evals, and training infrastructure.

     

Especially Strong Candidates May Also Have

  • Experience optimizing throughput and reliability for large-scale distributed training runs.

  • A background in AI for science or training on specialized domain data (e.g., protein, materials, or other scientific datasets).

  • Experience creating evals or synthetic data for non verifiable tasks and tracking performance over live runs.

     

Mechanics

  • Minimum education: Bachelor's degree or similar experience

  • Location: Menlo Park, CA

  • Compensation: $250,000–$350,000 + equity

  • Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

Optimize your resume for this job

Get a match score and the keywords you're missing

Optimize resume

About Periodic Labs

We're building AI scientists and the autonomous laboratories for them to operate. Join us: https://jobs.ashbyhq.com/periodic-labs

Similar Research roles

Research Scientist/Research Engineer, Midtraining
$250K–$350K · Periodic Labs
Apply