Skip to main content

Head of Simulation Engineering

Aaru
New York, NY
Full Time
Compensation
$425,000–$525,000/year

Job Description

Location: New York City · In person, five days a week
Compensation: $425,000–$525,000 base

About Aaru

Aaru builds simulations of human behavior. Each simulation contains a population of agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and policy changes to critical communications. Because the agents are simulated rather than recruited, they can reason through complex hypotheticals without fatigue or the response effects common in human studies.

The role

Simulation engineering builds, tests, and evaluates methods of simulation and deploys them into production. For deployment, methods must be accurate, calibrated, fast, measurable, and reliable. Day-to-day work resembles building a great AI-native product, albeit with much higher stakes—rather than informing an email draft or a Python file, these simulations determine new market entries, product decisions, and acquisitions.

As the Head of Simulation Engineering, you will build and lead the interface that the research, platform, and infrastructure teams use. This is a hands-on leadership role reporting directly to the founders. Early on, you will design systems, write and review code, debug model behavior, and expand the team of simulation engineers. As it expands, you will move to building the organization out, hiring managers and formalizing functions while still staying grounded in the day-to-day technical work.

What you will do

  • Own simulation quality in production across behavioral fidelity, calibration, latency, cost, reproducibility, and reliability.

  • Define the architecture and operating model that carries a method from research prototype through evaluation, rollout, observation, and improvement.

  • Build evaluation harnesses, benchmarks, ablations, graders, and regression tests that distinguish a faithful simulation from a merely plausible answer.

  • Make large population runs observable and debuggable: version models, prompts, data, agent definitions, environments, and experiment configuration so results can be reproduced and explained.

  • Partner with research to decide when a new method is ready to ship and what evidence is required before it becomes a customer-facing capability.

  • Build tight feedback loops with deployment. Turn field failures and surprising outcomes into concrete hypotheses, experiments, fixes, and new research questions.

  • Set clear interfaces and ownership across Simulation Research, Simulation Engineering, Infrastructure, and Platform.

  • Hire, coach, and retain an exceptional team while continuing to unblock the hardest technical problems yourself.

Representative problems

  • A method improves an offline benchmark but changes customer conclusions unpredictably. Determine why and establish the evidence required for rollout.

  • Two population runs with the same inputs diverge. Find the source of nondeterminism and make every material dependency inspectable.

  • A simulation is directionally accurate overall but miscalibrated for an important subgroup. Build the diagnostics and correction loop.

  • Reduce the cost and latency of a hundred-thousand agent run without eroding behavioral fidelity or hiding uncertainty.

  • Turn a fragile research workflow into a self-serve system with automated guardrails, launch criteria, monitoring, and rollback.

You might thrive in this role if

  • You have built and operated an AI-native product where model behavior was part of the product—not a feature hidden behind an API.

  • You can take an ambiguous behavioral problem and turn it into a hypothesis, an evaluation, a system, and a shipped improvement.

  • You have led engineers in a fast-moving environment and still enjoy doing the hardest technical work yourself.

  • You design evaluations before you trust a result, and you treat unexplained model regressions as production incidents.

  • You can turn research-grade code into a reproducible, testable, observable, and efficient system.

  • You make clear tradeoffs among quality, latency, cost, reliability, and iteration speed.

  • You communicate credibly with researchers, product engineers, deployment teams, and customers.

  • You want to build in person, in New York, at high speed.

Strong candidates may also have

  • Experience with coding agents, copilots, autonomous workflows, multi-agent systems, or other agentic products.

  • Experience with post-training, model evaluation, inference systems, experimentation platforms, or LLM orchestration.

  • Experience building simulation, synthetic-data, or distributed-compute systems at meaningful scale.

  • Time as a founder or early engineer at a fast-growth AI company.

Success in this role looks like

  • Aaru has one trusted, legible quality bar from research experiments through customer deployment.

  • New simulation methods move into production faster because evaluation, rollout, and observability are built into the system.

  • Large runs are reproducible; regressions are caught early; and failures can be traced to concrete causes.

  • Research, Platform, and deployments have clear interfaces and fast feedback loops.

  • A small, exceptional Simulation Engineering team owns the system end to end.

Compensation and benefits

Base salary of $425,000–$525,000, equity, and full benefits. Final compensation depends on experience and sits within our internal bands.

Optimize Your Resume for This Job

Get a match score and see exactly which keywords you're missing

Optimize Resume

Job Details

Category
Software
Employment Type
Full Time
Location
New York, NY
Posted
Compensation
$425,000 - $525,000 per year

About Aaru

Aaru is a Rethinking the science of prediction.

Found this role interesting?

Head of Simulation Engineering
Aaru
Apply