Joshua Lee

Case 01 · Calibr

Building an AI product that refuses to lie

Calibr rewrites your resume in a target company's language. Its hardest requirement was not quality. The output had to be something you could defend out loud in an interview, so the model is not allowed to invent anything.

Role
Founder. The only engineer and designer.
Timeframe
March 2026 to now.
Team
8 people I hired and lead.
Live at
calibr.co

In 30 seconds

  1. A hard rule: the model may not invent anything. The output reads less flashy than a chatbot's, and every line survives an interview.
  2. An eval, not my taste. An LLM-as-judge rubric, gated in CI, took output quality from 67 to 93 out of 100.
  3. Run like a product. Pricing, onboarding and feature tests took it to 250+ users and grew paid users 52%.

01The problem

Every company reads a resume differently, so tailoring one by hand took me about two hours. Keyword tools were safe but useless. General AI chatbots wrote better, and invented metrics users then had to defend in interviews.

  • 25+Scripted user interviews.
  • 8+Competitors torn down, including Teal, Rezi and Jobscan.
  • 100+Public complaints coded into themes.
The gap was not quality. It was a product that could be both good and true.

02What I considered

Three ways to build the rewrite engine.
ApproachForVerdict
Keyword matchingCheap, no model risk.Rejected. Matches strings, not language.
One big prompt to a general modelBest raw writing, ships in a weekend.Rejected. It fabricates, invisibly.
Live research, then a guarded per-bullet rewriteCompany-specific and traceable to real work.Taken. Slower and costlier per run.

The trade-off: output that looks less impressive next to a chatbot, and a review step for every change. The guardrails have only ever been tightened.

Calibr's research panel for Oliver Wyman: the firm's values, what to emphasize, the language it uses and what to avoid, with cited sources.
Live research on the target company, every claim sourced. Research shapes wording, never facts.

03How it works

Built alone in TypeScript and React on Supabase, Vercel and the Anthropic API: 230 commits, 10+ deploys a week behind feature flags. I also set up agent skills, prompts and evals for 3 teammates.

One calibration, end to end

Every rewrite passes a grounding check before the user sees it.

  1. 1 · Research Tool-calling research agent Web search with parallel tool calls, prompt caching and a cheaper model for research. 3.5 min → 70 s
  2. 2 · Rewrite Per-bullet rewrite A frontier model, structured JSON validated against a schema, with retries and model fallbacks.
  3. 3 · Guard Grounding check Any number, tool or outcome not in the input is blocked and sent to human review. Human-in-the-loop
  4. 4 · Approve User reviews each change In-app ratings feed a golden regression set.

Before any prompt change ships: the golden set is re-scored by the LLM-as-judge eval in CI. If quality drops, the change does not merge.

04How I knew it worked

For six weeks I judged quality by reading it. Then I wrote a 10-criteria rubric and had a model score 100+ real outputs, checked against 7 human graders. It caught quantification failing on technical roles while the average looked fine. Quality went from 67 to 93.

Eval scores out of 10 for one scoring run, by criterion and target role.
BCGBusiness Analyst McKinseyStrategy AccentureMgmt Consulting MercerStrategy & Ops PwCFinancial Advisory GoogleData Eng. UberAnalytics Eng. AmazonSoftware Eng. AppleMobile Eng. MetaAI Research
Company fit99.5998.598.5999
Role relevance9109.59.599.599.59.59.5
Preserves original9998.57.598.58.588.5
Accuracy9998.5898.58.58.58.5
Formatting101010108.58.58.58.58.58.5
Word choice99.59999.599.59.59.5
Impact focus9109998.58.5998.5
One scoring run, 10 roles by 7 of the 10 criteria. Read the rows, not the average: formatting scored 10 on consulting roles and 8.5 on every technical one.

05What moved the numbers

North-star metric: free-trial-to-paid conversion, tracked in PostHog.
Problem, and what I didResult
New users quit during a slow first run. Interviews and funnels found it; I redesigned onboarding.Drop-off 80% → 52%
Monthly plans did not fit how students job-hunt. I A/B tested a recruiting-season pass.Paid users +52%
Session replay showed a useful feature was buried. I unbundled it.Optimizations +27%
Users finished a run and stalled. I built a profile-strengthening feature from 50+ usability tests.2-week retention 15% → 21%
Research was slow. Prompt caching, model routing and parallel tool calls.Research 3.5 min → 70 s
A usage spike in the LLM logs. I capped a duplicate research call.Cost per run −60%

Where it got to.

250+

Users, from the first 12.

2 hrs → 20 min

To tailor a resume to one company.

67 → 93

Output quality, gated in CI.

06What I would do differently

Build the eval in week one, not week six. Every change before it is unmeasured.

Track cost per run from day one. The duplicate call ran for weeks before I saw it.

Pick the north-star metric earlier. Activation and retention took turns until I did.