Skip to main content

Evaluation Researcher

Aaru
New York, NY
Full Time
Compensation
$200,000–$600,000/year

Job Description

About Aaru

Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes.

Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions.

We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.

About Evaluation Research

Evaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities.

The function builds both rails and carts. Rails are reusable evaluation infrastructure: datasets, harnesses, libraries, experiment standards, leaderboards, reporting systems, and ways to translate technical evidence into decisions. Carts are the specific evaluations that run on those rails: a historical backtest, a prospective forecast study, a population-coherence test, a reproduction of an observed behavioral pattern, or an end-to-end comparison with a resolved outcome.

Evaluation Research is not conventional QA and it is not benchmark administration. It is an independent research function. The work requires understanding the systems deeply, collaborating closely with their builders, and remaining willing to conclude that an attractive method did not improve what matters.

The role

As an Evaluation Researcher, you will own difficult measurement problems at the boundary of machine learning, statistics, behavioral science, and product decision-making. You will define constructs, assemble or create evaluation data, design studies, write analysis and evaluation code, inspect individual failures, quantify uncertainty, and communicate what the evidence does and does not support.

Some projects will build reusable rails used across the research organization. Others will be focused studies intended to resolve one important uncertainty. In both cases, the goal is the same: create an evaluation that is valid enough to trust, diagnostic enough to guide improvement, and clear enough to inform a real decision.

You will work closely with Population Research, Prediction Research, Simulation Engineering, Product Engineering, Research Product, and Deployment while protecting the independence and integrity of final measurements.

What you will do

  • Own an important evaluation or measurement area across population construction, predictive systems, individual agent behavior, group dynamics, or end-to-end simulations.

  • Turn broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs and explicit decision criteria.

  • Design studies using historical backtests, temporal holdouts, prospective outcomes, observational records, controlled experiments, expert judgment, or mixed methods as the problem requires.

  • Build tests of individual-profile quality, including internal coherence, contradictions, impossible combinations, unsupported specificity, stability, and whether a profile induces behavior consistent with the represented person.

  • Build tests of population quality, including marginal and joint distributions, conditional relationships, coverage of rare but plausible profiles, subgroup fidelity, and sensitivity to sampling choices.

  • Evaluate predictions using calibration, proper scoring rules, ranking quality, selective prediction, temporal validity, subgroup performance, and the decision cost of different errors.

  • Compare simulations with transactions, product usage, behavioral traces, operational outcomes, market movements, resolved events, surveys, and longitudinal decisions.

  • Design longitudinal and interaction-based evaluations that test how agents change over time, respond to new information, and influence one another.

  • Build end-to-end studies that determine whether a component improvement actually changes the quality of the conclusion a customer receives.

  • Establish strong baselines and compare agent-based simulation with direct forecasting, conventional statistical models, simpler segment-level methods, and human or market benchmarks where appropriate.

  • Find failures hidden by aggregate metrics, especially failures concentrated in important subgroups, rare cases, changing environments, or ambiguous labels.

  • Build diagnostic evaluations that localize why a system failed and whether a proposed fix generalizes beyond the development set.

  • Develop reusable evaluation datasets, harnesses, libraries, graders, experiment schemas, leaderboards, and reporting tools where shared infrastructure will accelerate future research.

  • Work with Simulation Engineering to version and automate evaluations while keeping protected holdouts and final claims insulated from development leakage.

  • Convert customer surprises, production incidents, and resolved real-world outcomes into durable test cases.

  • Write clear technical reports that distinguish exploratory evidence from claim-supporting evidence and communicate uncertainty, limitations, and alternative interpretations.

  • Report negative, null, and inconclusive findings with the same care as positive results.

Representative research directions

You might investigate questions such as:

  • Construct a population using data available at one point in time, then test whether its future transactions, choices, or behavioral outcomes match what later occurred.

  • Recreate a historical decision environment using only information available before the outcome and compare the simulation with the observed result.

  • Run prospective evaluations in which Aaru records predictions before outcomes are known and tracks performance as those outcomes resolve.

  • Determine whether profile-coherence scores, population-distribution metrics, or forecast calibration predict end-to-end simulation quality.

  • Measure whether simulated groups reproduce observed patterns in information diffusion, coordination, influence, polarization, or collective choice.

  • Study how population size, heterogeneity, interaction structure, model capability, context length, and computation affect simulation fidelity.

  • Build tests for memorization, leakage, prompt sensitivity, unsupported certainty, judge bias, and benchmark-specific overfitting.

  • Develop evaluation methods for settings where outcomes are delayed, noisy, partially observed, selected, or open to more than one reasonable interpretation.

  • Compare automated graders, expert review, behavioral outcomes, and human-subject measurements to determine when each is a valid proxy.

  • Estimate whether the measured gain is large enough to matter for the decision, not merely statistically distinguishable from zero.

How we work

A useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held-out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. When ground truth is imperfect, the quality and limits of the outcome data are part of the research problem rather than a footnote.

Evaluation should make research faster by giving teams clear signals about what improved, what did not, and why. It should also make Aaru more trustworthy by exposing failures early and keeping product and external claims aligned with the evidence.

Researchers own the full arc of the work: construct definition, data provenance, implementation, analysis, failure inspection, interpretation, communication, and the decision that follows. A polished metric without a valid construct is not success.

You might thrive in this role if

  • You have a record of rigorous work in ML evaluation, statistics, behavioral science, computational social science, economics, psychometrics, experimental design, or a related field.

  • You have built an evaluation that changed a research direction, model capability, product decision, or scientific conclusion.

  • You can define a difficult construct precisely enough to measure it without losing the underlying question.

  • You are comfortable with observational data, sampling, statistical power, uncertainty, causal threats, selection, leakage, and condition shift.

  • You can write code, analyze large datasets, build evaluation systems, and inspect individual examples rather than relying only on aggregate dashboards.

  • You naturally ask what a benchmark actually measures, what behavior it incentivizes, and where it can be gamed or contaminated.

  • You can work closely with a system's builders while maintaining independent judgment about its quality.

  • You care more about an accurate conclusion than a favorable one and are willing to revise your own evaluation when it proves inadequate.

  • You communicate technical evidence clearly to researchers, engineers, product teams, customers, and non-specialists.

  • You want to work in person in New York with a team that moves quickly and treats inconvenient evidence as valuable.

Strong candidates may also have

  • Work in forecasting evaluation, econometrics, psychometrics, causal inference, survey methodology, experimental economics, measurement theory, or model behavior.

  • Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, simulations, or decision-support tools.

  • Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, or operational validation.

  • Experience building evaluation platforms, regression suites, experiment-tracking systems, shared datasets, model scorecards, or scientific reporting tools.

  • Experience with automated graders, human evaluation, rubric design, inter-rater reliability, benchmark contamination, or adversarial evaluation.

  • A record of finding an important failure that standard metrics missed and developing a better way to measure it.

  • Experience communicating scientific results in customer-facing, public, policy, legal, or regulatory settings.

Candidates need not have

  • Prior experience in population simulation or a job title containing the word “evaluation.”

  • A PhD, provided you can demonstrate equivalent research depth and empirical rigor.

  • Expertise in every statistical or machine-learning method listed above. We care most about measurement judgment, technical execution, and intellectual honesty.

What success looks like

  • You create evaluations that resolve important uncertainty rather than merely producing additional metrics.

  • Your measurements are valid enough to support decisions, diagnostic enough to guide improvement, and reproducible enough for others to challenge.

  • Important failures are found early, explained clearly, and converted into durable datasets, tests, or research questions.

  • Protected evidence remains trustworthy while development teams still receive fast feedback.

  • Reusable rails reduce duplicated work and make results comparable across projects, systems, and versions.

  • Product and research claims become more precise because their supporting evidence and limitations are explicit.

  • Your work helps Aaru distinguish a genuine general improvement from overfitting, leakage, a proxy failure, or a favorable anecdote.

  • Colleagues trust your conclusions because you combine scientific rigor with a practical understanding of how systems improve.

Location and benefits

This role is based in New York City. Aaru is an in-person company, working five days a week in the office. Candidates should be located in the New York metropolitan area or open to relocation.

Aaru offers a competitive base salary, equity participation, comprehensive medical, vision, and dental coverage, visa sponsorship and relocation support, and other benefits and perks. Final compensation depends on level and experience and is set within Aaru's internal bands.

Optimize Your Resume for This Job

Get a match score and see exactly which keywords you're missing

Optimize Resume

Job Details

Category
Research
Employment Type
Full Time
Location
New York, NY
Posted
Compensation
$200,000 - $600,000 per year

About Aaru

Aaru is a Rethinking the science of prediction.

Found this role interesting?

Evaluation Researcher
Aaru
Apply