Joshua Lee

Anyone can ship with AI now. Knowing if it’s good is the job.

Two questions: was it worth building, and is what shipped any good. Both are really about waste: where the effort went, and how much of it counted. That is what I cannot leave alone.

Calibr Capital One Bain & Company Relate QANDA United Nations Command
200+ users on a product I built alone 80,000 employees in a model I shipped at Capital One $20M acquisition my analysis argued against at Bain 300+ cross-border missions led in the JSA

01

Product

Two things I built, and the decisions behind them.

Calibr

Founder · live at calibr.co

I built an AI product that is not allowed to lie, and then proved it.

25+

I started from the complaint, not the idea.

Interviews on a fixed script, eight tools torn down, a hundred public complaints coded into themes. Keyword tools were safe and useless. ChatGPT was useful and unsafe: it invented the metrics people then had to defend in a room.

Calibr's research panel for a target firm: its stated values, what to emphasise, and the language it uses, each drawn from cited sources.
What a rewrite is built from, every line traced to a cited source. The rule is printed at the bottom: never to invent facts.

9 / 13

So I made it worse on first impression, on purpose.

If a rewritten line contains a number the user never wrote, it does not ship. Next to ChatGPT, Calibr often looks less impressive. I took that trade, because a resume you cannot defend in an interview is a liability, not a feature.

Calibr's result view: a match score of 74 out of 100 against a specific firm, a note that the rewrite was built from live research on that firm, and a counter showing 9 of 13 lines changed.
Nine of thirteen lines changed, each one clickable against the original. The user approves the document, not the model.

67 → 93

Then I stopped trusting my own taste.

Judging quality by reading it is a mood, not a method. I wrote ten criteria and had a model score every run against them. It caught what I had been missing for weeks: quantification failing on technical roles while the average looked fine.

Read the full case
The evaluation matrix: target roles across the top, scored qualities down the side, each cell holding a score out of ten and the written reason behind it.
Four of the eleven target roles. Down a column is one rewrite. Across a row is the criterion failing everywhere, which is what an average hides.

200+

Users, from the first 12. 500+ resumes processed.

2 hrs → 20 min

To tailor a resume to one company.

8

People on the team I lead across product, strategy and marketing.

Relate

Product Manager · Relate, YC S22

The plan I was handed described a product reps would not have bought.

18 → 3

I cut the roadmap before writing any code.

A teardown of twelve competing tools, plus research on what reps complained about. Eighteen screens became the three features reps named unprompted. Waitlist enrollment rose 45%, and the product reached its first 200 reps.

Then I respecified the build as acceptance criteria per screen, not descriptions. A description can be satisfied several ways; a criterion cannot. Three weeks of estimate shipped in five days.

Read the full case
The sitemap for the deck product, branching into dozens of controls.
The product as originally mapped. Every green box is a control somebody has to build, document and support.

02

Analysis

Twice, the deliverable was a decision.

Neither of these ended in a feature. Both ended in a number someone had to bet on.

Bain & Company

Associate Consultant Intern · Seoul · 2025

The client walked away from a $20M acquisition.

I rebuilt fifteen years of cohort economics and stress-tested three downside cases. The analysis exposed an estimated 40% erosion in projected returns. They did not proceed.

$20M

Acquisition the analysis argued against.

$2.5M

Ad budget reallocated on a separate case, after decomposing return on spend across 100+ SKUs.

Chart: the seller's base-case return projection indexed to 100 at the hurdle, against three stress-tested downside cases finishing at 88, 81 and 74. Every downside case lands below the hurdle.
The seller's case cleared the hurdle. Not one of my three downside cases did, and that gap is the whole recommendation.
Capital One

Product Management & Analytics Intern · Richmond, VA · 2026

Four workforce trackers, four different answers.

Fifteen leaders, and no two counted a head the same way. Getting them to agree on one definition was the actual work. The SQL that unified four systems came after, and the self-serve forecasting after that.

80,000

Employees covered by the forecasting model.

6 hrs → minutes

A manual process, automated end to end, across 10+ lines of business.

Table: four workforce systems reporting four different headcounts because each counted contractors, open requisitions and pending starts differently. One agreed definition resolves them to 80,000.
What fifteen interviews actually surfaced. The systems did not disagree about the data, they disagreed about who counted as an employee.

03

Leadership

Before any of it, two years standing between four countries.

Joshua Lee with his Joint Security Area unit, standing beneath the flags of the United Nations Command member nations.

I learned what a bad translation costs before I learned what a PRD was.

United Nations Command, Joint Security Area. Squad leader and interpreter, 2023 to 2024.

I led a 10-person operations team across 300+ cross-border missions, coordinating live communication between the US Army, UN Command and four other nations. When a line had to stay open, I was the person keeping it open. The unit named me its best squad leader, and my proposal won its development competition.

+40%

Operating log accuracy, after I overhauled the workflows behind it.

5+

Companies that adopted the real-time speaker-identifying transcriber I built.

300+

Cross-border missions run by the team I led.