Skip to main content

Research Scientist, Scaling RL

Periodic LabsResearch
Pay
$250K–$350K
per year
Work mode
On-site
Full Time
Experience
5+ yrs
Bachelor's

Menlo Park, CA at a glance

Rent
#2 of 51
$3,490/mo+46% vs US avg
Weather
#17 of 51
292 mild days0 hot · 0 cold
Income tax
#1 of 51
13.3% top rateCalifornia

What you need

  • 5+ yrs experience in RL/LLM training
  • Hands-on training LLMs with reinforcement learning
  • Design small-scale RL setups transferring to large scale
  • Implement, debug, test ideas in complex training stack
  • Rigorous scientific approach to problem-solving

What you'll do

  • Design experiments to understand RL scaling with compute, data, reward quality
  • Develop RL algorithms: policy optimization, advantage estimation, exploration, credit assignment
  • Build adaptive sampling and curriculum methods for task difficulty and rollouts
  • Study bias and stability in RL training, including importance-sampling corrections
  • Improve compute efficiency via hyperparameter experiments (length penalties, rollout counts, batch sizes)

About Periodic Labs

We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.


About the Role

We're training frontier models to develop deep scientific knowledge and reasoning for scientific tasks. You’ll study how RL scales with training compute, develop better algorithms, and take ideas from controlled experiments to our largest runs like Periodic Neon.


What You'll Do

  • Design experiments to understand how RL performance scales with compute, model size, data, and reward quality, building on work such as ScaleRL

  • Develop better RL algorithms, spanning policy optimization, advantage estimation, exploration, and credit assignment for long-horizon RL tasks

  • Build adaptive sampling and curriculum methods that adjust task difficulty, problem selection, and the number of rollouts as models improve

  • Study bias and stability during RL training, including importance-sampling corrections and methods to tackle policy staleness and training–inference mismatch, as discussed here.

  • Improve compute efficiency across training and inference through experiments with hyperparameters, such as length penalties, rollout counts, batch sizes, and update schedules.


You Will Thrive in This Role If You Have

  • Hands-on experience training LLMs with reinforcement learning

  • Strong attention to detail and rigorous approach to answer questions scientifically.

  • Coming up with small-scale RL setups that transfers to large-scale training runs.

  • Comfort working across a complex training stack to implement, debug, and test new research ideas.


Mechanics

Minimum experience: 5+ years
Minimum education: Bachelor’s degree or similar experience

Location: Menlo Park, CA

Compensation: $250,000-$350,000 base + equity

Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

 

Optimize your resume for this job

Get a match score and the keywords you're missing

Optimize resume

About Periodic Labs

We're building AI scientists and the autonomous laboratories for them to operate. Join us: https://jobs.ashbyhq.com/periodic-labs

Similar Research roles

Research Scientist, Scaling RL
$250K–$350K · Periodic Labs
Apply