Case 01 · Calibr
Building an AI product that refuses to lie
Calibr rewrites your resume in a target company's language. Its hardest requirement was not quality. The output had to be something you could defend out loud in an interview, so the model is not allowed to invent anything.
- Role
- Founder. The only engineer and designer.
- Timeframe
- March 2026 to now.
- Team
- 8 people I hired and lead.
- Live at
- calibr.co
In 30 seconds
- A hard rule: the model may not invent anything. The output reads less flashy than a chatbot's, and every line survives an interview.
- An eval, not my taste. An LLM-as-judge rubric, gated in CI, took output quality from 67 to 93 out of 100.
- Run like a product. Pricing, onboarding and feature tests took it to 250+ users and grew paid users 52%.
01The problem
Every company reads a resume differently, so tailoring one by hand took me about two hours. Keyword tools were safe but useless. General AI chatbots wrote better, and invented metrics users then had to defend in interviews.
- 25+Scripted user interviews.
- 8+Competitors torn down, including Teal, Rezi and Jobscan.
- 100+Public complaints coded into themes.
The gap was not quality. It was a product that could be both good and true.
02What I considered
| Approach | For | Verdict |
|---|---|---|
| Keyword matching | Cheap, no model risk. | Rejected. Matches strings, not language. |
| One big prompt to a general model | Best raw writing, ships in a weekend. | Rejected. It fabricates, invisibly. |
| Live research, then a guarded per-bullet rewrite | Company-specific and traceable to real work. | Taken. Slower and costlier per run. |
The trade-off: output that looks less impressive next to a chatbot, and a review step for every change. The guardrails have only ever been tightened.
03How it works
Built alone in TypeScript and React on Supabase, Vercel and the Anthropic API: 230 commits, 10+ deploys a week behind feature flags. I also set up agent skills, prompts and evals for 3 teammates.
One calibration, end to end
Every rewrite passes a grounding check before the user sees it.
- 1 · Research Tool-calling research agent Web search with parallel tool calls, prompt caching and a cheaper model for research. 3.5 min → 70 s
- 2 · Rewrite Per-bullet rewrite A frontier model, structured JSON validated against a schema, with retries and model fallbacks.
- 3 · Guard Grounding check Any number, tool or outcome not in the input is blocked and sent to human review. Human-in-the-loop
- 4 · Approve User reviews each change In-app ratings feed a golden regression set.
Before any prompt change ships: the golden set is re-scored by the LLM-as-judge eval in CI. If quality drops, the change does not merge.
04How I knew it worked
For six weeks I judged quality by reading it. Then I wrote a 10-criteria rubric and had a model score 100+ real outputs, checked against 7 human graders. It caught quantification failing on technical roles while the average looked fine. Quality went from 67 to 93.
| BCGBusiness Analyst | McKinseyStrategy | AccentureMgmt Consulting | MercerStrategy & Ops | PwCFinancial Advisory | GoogleData Eng. | UberAnalytics Eng. | AmazonSoftware Eng. | AppleMobile Eng. | MetaAI Research | |
|---|---|---|---|---|---|---|---|---|---|---|
| Company fit | 9 | 9.5 | 9 | 9 | 8.5 | 9 | 8.5 | 9 | 9 | 9 |
| Role relevance | 9 | 10 | 9.5 | 9.5 | 9 | 9.5 | 9 | 9.5 | 9.5 | 9.5 |
| Preserves original | 9 | 9 | 9 | 8.5 | 7.5 | 9 | 8.5 | 8.5 | 8 | 8.5 |
| Accuracy | 9 | 9 | 9 | 8.5 | 8 | 9 | 8.5 | 8.5 | 8.5 | 8.5 |
| Formatting | 10 | 10 | 10 | 10 | 8.5 | 8.5 | 8.5 | 8.5 | 8.5 | 8.5 |
| Word choice | 9 | 9.5 | 9 | 9 | 9 | 9.5 | 9 | 9.5 | 9.5 | 9.5 |
| Impact focus | 9 | 10 | 9 | 9 | 9 | 8.5 | 8.5 | 9 | 9 | 8.5 |
05What moved the numbers
| Problem, and what I did | Result |
|---|---|
| New users quit during a slow first run. Interviews and funnels found it; I redesigned onboarding. | Drop-off 80% → 52% |
| Monthly plans did not fit how students job-hunt. I A/B tested a recruiting-season pass. | Paid users +52% |
| Session replay showed a useful feature was buried. I unbundled it. | Optimizations +27% |
| Users finished a run and stalled. I built a profile-strengthening feature from 50+ usability tests. | 2-week retention 15% → 21% |
| Research was slow. Prompt caching, model routing and parallel tool calls. | Research 3.5 min → 70 s |
| A usage spike in the LLM logs. I capped a duplicate research call. | Cost per run −60% |
Where it got to.
250+
Users, from the first 12.
2 hrs → 20 min
To tailor a resume to one company.
67 → 93
Output quality, gated in CI.
06What I would do differently
Build the eval in week one, not week six. Every change before it is unmeasured.
Track cost per run from day one. The duplicate call ran for weeks before I saw it.
Pick the north-star metric earlier. Activation and retention took turns until I did.