Skip to main content

AI/ML Engineer

LemmaSoftware
Pay
$120K–$200K
per year
Work mode
On-site
Full Time
Level
Entry

San Francisco, CA at a glance

Rent
#2 of 51
$2,680/mo+46% vs US avg
Weather
#17 of 51
295 mild days0 hot · 0 cold
Income tax
#1 of 51
13.3% top rateCalifornia

What you need

  • High growth trajectory; new grads welcome
  • Shipped products people actually use
  • Production LLM experience: evals, LLM-as-judge, embeddings
  • Research taste: detect real improvements without ground truth
  • Bonus: ran agents in production, experienced failures

What you'll do

  • Own detection quality: find silent failures across traces
  • Turn implicit user signals into failure evidence
  • Build evals to measure precision without ground truth
  • Make patch generation trustworthy: reproduce, verify, stay quiet
  • Optimize LLM-as-judge cost per event at scale

Lemma is production monitoring for AI agents. We catch the silent failures your observability tools and evals miss (think bad tool calls, lost context, and infinite loops) before your users find them.

Why this role exists

Agents break silently. They call the wrong tool, forget what the user said three turns ago, and loop until someone pulls the plug. Soon they'll be responsible for the majority of the world’s economic work, and most teams won't even know when they fail.

Making agents reliable is the problem Lemma exists to solve. That means catching the unknown unknowns, the failures nobody thought to write an eval for, and closing the loop end to end so they get fixed, not just flagged. It's the foundation for building agents people can actually trust and the future of self-improving systems.

The hardest part of our product is deciding what counts as a failure.

There is no ground truth here, and no benchmark to climb. Every customer's agent is different, what "wrong" means changes from one to the next, and we have to get it right across production without anyone telling us what to look for. Being confidently wrong often costs us more trust than being right fifty times earns.

This role owns the intelligence in the loop: what we flag, how sure we are, and whether the fix we propose actually fixes it.

What you’ll do

  • Own detection quality. Find the failures that don't look like failures: compliant but wrong, omissions, and patterns that only show up across thousands of traces
  • Turn implicit signals into evidence. Rephrasing, abandonment, retries, and the other ways users tell you something broke without saying so
  • Build the evals for our own system. If we can't measure precision on a problem with no labels, we can't improve it
  • Make patch generation trustworthy. Reproduce the failure, verify the fix, and know when to stay quiet instead of opening a bad PR
  • Keep it affordable. LLM-as-judge on every event is easy. Doing it at a cost per event that doesn't eat the business is the actual job
  • Read real customer traces every week. The best ideas here come from staring at production, not papers

What we’re looking for

  • High slope over years of experience. New grads and dropouts welcome
  • A track record of shipping things people actually use
  • Hands-on with LLMs in production: evals, LLM-as-judge, embeddings, and knowing when a smaller model or no model is the right call
  • Real research taste. You can tell a real improvement from noise, even when there's no ground truth to check against
  • Bonus: you were the customer once. You ran agents in production and got burned

Who you’ll work with

You'll join a team of dropout founders and engineers from Amazon, Together, and Zoom. We've been founding operators at unicorns and at startups that went on to be acquired.

Onsite in San Francisco. We don’t sponsor visas.


Interview Process

We keep this short on purpose. Target is an offer within two weeks of first contact.

Intro call with a founder (30 minutes). What we're building, what you've built, whether the shape of the role actually fits what you want next.

Take-home project and deep dive. We give you a real problem from Lemma. You work on it on your own time, then we walk through it together. The conversation matters more to us than the artifact.

Paid work trial (onsite in-person). Real problem, real codebase, sitting with the team. You find out what working here actually feels like before you commit, which tells you more than anything we could say about it.

May ask for additional references.

Optimize your resume for this job

Get a match score and the keywords you're missing

Optimize resume

About Lemma

Lemma catches the silent, semantic failures your observability tools miss, where your AI agent looks like it worked but didn’t. We scan every trace to surface issues before users complain, identify root causes, and help you fix them without manual digging, so your agents improve over time.

Similar Software roles

AI/ML Engineer
$120K–$200K · Lemma
Apply