
Harness Engineer
CA, CA at a glance
- Rent
- #2 of 511-bedroom+46% vs US avg
- Weather
- #17 of 51179 mild daysstatewide median
- Income tax
- #1 of 5113.3% top rateCalifornia
What you need
- 2+ yrs building software with agent systems/harnesses
- Strong conversational agent experience (primary focus)
- Strong database and system design skills
- Shipped agents to real users, handled production failures
- Comfortable with ambiguity, no established playbook
What you'll do
- Own both AI agents end-to-end including failure handling
- Build harness coordinating LLMs, tools, memory, async workflows
- Develop robust tooling: schema validation, retries, permissions, error handling
- Create eval frameworks for task completion, accuracy, safety, latency
- Implement observability: traces, logs, failure analysis across workflows
Most of our users never really touch our app. They text us. Restaurant owners and creators handle almost everything through conversation, and on our side that's a set of AI agents doing the booking, the follow-ups, the scheduling, and the problem solving.
You own those agents. Not the prompts alone, the whole system underneath them: how they understand what someone actually wants, how they call tools without breaking, how they hold context across a conversation that's been going for 6 months, and how they take real actions in the real world without us having to check their work.
Any text that comes in, your systems handle it.
What you'll own
The agents. Full ownership of both of our agents end to end. How they're built, what they can do, and what happens when they get something wrong.
The harness. The layer coordinating LLMs, tools, memory, and async workflows. This is the actual engineering problem here and it's most of your week.
Tooling that doesn't fall over. Schema validation, retries, permissions, error handling. An agent that calls a tool correctly 95% of the time is not good enough when it's booking real visits at real restaurants.
Evals. Frameworks that measure task completion, accuracy, safety, latency, and failure modes. If we change a prompt on Tuesday, we should know by Tuesday whether it made things worse.
Observability. Traces, logs, and failure analysis across agent workflows. When something goes wrong in a conversation, you should be able to see exactly where.
Turning vague into dependable. A restaurant owner texts something ambiguous at 11pm. Getting from that to a reliable agent behavior is the hard part, and it's the part we care most about.
What we're looking for
Required
- 2+ years building software, with real experience building agent systems or harnesses. Not just calling an API in a side project
- Strong conversational agent experience. This is the thing we weigh most heavily. Our product is a conversation, and someone who has only built single-turn or task-runner agents will struggle here
- Strong database and system design
- You've shipped agents that real people used and dealt with the fallout when they broke
- Comfortable with ambiguity. There's no established playbook for most of this
- Based in Orange County / willing to relocate and able to work onsite
- This isn't a 9-to-5
Nice to have
- iMessage agent experience
- You've built eval frameworks, not just run them
- Observability and tracing work on LLM systems
- Restaurant industry or creator economy experience
Optimize your resume for this job
Get a match score and the keywords you're missing
About TryNearby
TryNearby connects local businesses with creators who live nearby. Restaurants subscribe monthly and automatically get booked visits from local creators who film, edit, and post a review to the local algorithm. This creates a steady stream of content that compounds into discovery on TikTok, Instagram, and increasingly AI search. Over 100 restaurants and thousands of creators are on the platform across Southern California, with AI agents handling matching, booking, and every creator conversation.
Similar Software roles


