
Job Description
The Opportunity
Sixtyfour turns a single name, email, or domain into a full, verified picture of a person or company — by sending AI agents out to research the open web the way a sharp analyst would, then checking and scoring what they find. You'll build real parts of that: the agents that reason and gather evidence, the systems that run them at scale, and the product people use to see the results.
How We Work
We spend most of our time — call it 80% — understanding the problem deeply, planning, and designing the system before a line gets written. Getting the design right is the hard part and the best part. We hold a high bar, we stay on the edge of what's possible, and everything we ship has to hold up at scale. If you love the part of engineering that happens on the whiteboard — arguing the right design, the failure modes, the tradeoffs — you'll fit here.
What You'll Do
- Build AI agents for OSINT and deep web research — design agents that investigate people and companies across the open web, public records, social platforms, and other sources, then cross-reference and structure what they find.
- Own the thinking, not just the code — dig into the problem, weigh the designs, and write the plan before you build, because that's where the real leverage is.
- Design systems that hold up at scale — reason through data volume, concurrency, latency, and cost up front, so what you build survives real load.
- Ship features end to end — design, build, test, deploy — so customers get something new in weeks, not quarters.
- Build and sharpen the AI agents that research people and companies, so enrichment returns more accurate, better-sourced answers.
- Add new data sources and tools to the enrichment engine, so agents can reach information they couldn't before.
- Write evals and tests that prove whether a model or agent change actually made results better, so the team improves on evidence instead of hope.
- Make long-running jobs fast and reliable — batching, caching, retries, orchestration — so millions of records enrich without falling over.
What We're Looking For
Must-have — this is a high bar, and we mean it:
- Strong engineering fundamentals. You understand how real systems work underneath — concurrency, APIs, databases, how the web fits together — and why they're built that way. Syntax is the easy part; you get the concepts beneath it.
- System-design and architecture instinct. Hand you a fuzzy problem and you can break it into pieces, find the failure modes, weigh the tradeoffs, and design something that holds. You think before you build.
- You think at scale by default. You reason about data volume, concurrency, latency, and cost without being told to — and you can point to real examples where you built, scaled, or seriously worked through large systems.
- You've shipped something real and can defend every decision — a project, open source, research, a hackathon — and go deep on why you designed it the way you did.
- You write solid code in at least one language and learn new ones fast. Our stack is mostly Python (backend and AI) and TypeScript/React (product) — you need one and the ability to pick up the other.
- Genuine curiosity about LLMs and agents. You want to build with them, not just use them.
- You move fast, own your work, and can work in person in San Francisco.
Nice-to-have — bonus, not required:
- Strong OSINT experience is a major plus — you’ve done deep online investigations, identity resolution, reverse username research, entity mapping, or similar open-source intelligence work.
- You've built something with LLMs or agents — a research tool, a RAG app, a scraper, an agent loop.
- Experience with distributed systems, queues, workflow engines, or high-throughput pipelines.
- Some React or Next.js, or experience building data-heavy UIs.
- You've run systems in production — databases (SQL/Postgres), Redis, search, observability.
- A sharp eye for data quality — you notice when an answer looks right but is subtly wrong.
If you came up a non-traditional path or haven't touched a specific tool, apply anyway. We'll teach you our stack. What we won't compromise on is how you think about problems.
What You'll Get
- Real ownership. You'll own features that reach production and paying customers, with your work clearly yours.
- Direct mentorship from engineers building genuinely hard applied AI — research agents, evals, and the large-scale systems that run them — who will review your designs and push you to be great.
- A team that thinks before it builds, so you'll leave a much stronger engineer: better at architecture, systems, and judgment, not just faster at typing.
- A generous budget for the best LLMs and dev tools. We want you building with the strongest tools available, not rationing tokens.
- Founder habits: scope your own work, make the call, and watch it land.
The Details
- Location: In person, San Francisco. This role is not remote.
- Level: Junior and up — strong students, new grads, and early-career engineers all welcome.
- Pay: $6,000–$10,000 / month, based on level and experience.
- Perks: Lunch in the office, team offsites, and a real budget for LLM usage and tooling.
How to Apply
Show us something you built and be ready to go deep on why — the design you chose, the ones you rejected, and how it would hold up at scale. If you see a hard problem and your instinct is to understand it fully and own it end to end, we want to meet you. We're a small team, we move fast, and we hire people who want to be exceptional. Come build with us.
Interview Process
-
Screening (15 min)
-
Take-home assessment
-
Virtual onsite (1 hr)
-
Call with the founders
Optimize Your Resume for This Job
Get a match score and see exactly which keywords you're missing
Job Details
- Category
- Software
- Employment Type
- Internship
- Location
- San Francisco, CA
- Posted
- Compensation
- $6,000 - $10,000 per month
About Sixtyfour
Sixtyfour is a AI based reasech service company that automates to enrich specialized professionals, company data, and insights.
More Roles at Sixtyfour
Similar Software Roles



Found this role interesting?