Anyone can ship with AI now. Knowing if it’s good is the job.
So I built both. Calibr is live with 250+ users, and its eval took output quality from 67 to 93. I also shipped a 0 to 1 product at Relate (YC S22), automated an 80,000-employee forecast at Capital One, and exposed 40% return erosion on a $20M acquisition at Bain.
Selected work
- 01 CalibrFounder · 2026 to now An AI resume product that is not allowed to invent anything, and the eval that proves it. 250+users, 500+ optimizations Read case
- 02 Relate, YC S22Product Manager · 2026 Refocused the founders' roadmap on what reps needed, then shipped the product in 5 days instead of 3 weeks. +45%waitlist enrollment Read case
- 03 QANDAProduct Manager Intern · 2024 Killed a feature US students did not want, and launched the one they did. +24%completion, A/B tested Read case
- 04 Capital One and BainAnalytics and strategy · 2025 to 2026 Two analyses that ended in a decision: an 80,000-employee forecast and a deal the client walked away from. $20Macquisition not pursued Below
01
Product
Three products I built, and the decisions behind them.
An AI product that is not allowed to lie, and the eval that proves it.
67 → 93
I built the eval before I trusted my own taste.
A model grades every output on 10 criteria, checked against human graders and gated in CI. Then pricing, onboarding and feature tests grew it, while I hired and led an 8-person team.
Read the case250+
Users. I am the only engineer and designer.
+52%
Paid users from a pricing A/B test.
2 hrs → 20 min
To tailor a resume to one company.
−60%
Cost per run, after tracing a spike in the LLM logs.
I refocused the founders' plan on what reps needed, then shipped it in five days.
3 wks → 5 days
A 12-tool teardown showed the plan leaned on generic CRM fields. Reps needed to know who opened their deck. I put view-tracking at the center, then wrote acceptance criteria as the Claude Code spec and built all 18 screens.
Read the case+45%
Waitlist enrollment after the roadmap change.
+25%
Reply rates in the rep pilot.
200
Reps on the product.
14 → 3
Days between releases.
I killed a feature US students did not want, and launched the one they did.
+24%
Mock Exam completion, A/B tested.
200+ US interviews showed students wanted to test themselves, not read the summary notes Korean students loved. I killed the planned feature and wrote the PRD for Mock Exam instead.
Read the case6 wks
Engineering time saved by the kill (est.).
+30%
Retention after AI Tutor shipped first.
+36%
Task completion after cutting MVP scope.
02
Analysis
Twice, the deliverable was a decision.
There was no standard for counting a head, so we wrote one.
I defined the KPIs with 15+ leaders, reconciled 3 workforce systems in SQL, and built a self-serve AI forecast I presented to the executive committee.
80,000
Employees in the forecast.
6 hrs → min
Forecast time, once built by hand.
4
Planning teams adopted it after UAT.
1,000+
Roles flagged over- or under-staffed.
From four trackers to one forecast leaders use
- Before4 trackers, no shared definitionEach line of business counted headcount its own way.
- DefineOne set of KPIs15+ leader interviews and design-thinking sessions.
- ReconcileSQL across 3 systemsData-quality checks, then one trusted source.
- ShipAI forecast + Tableau dashboardExecutive committee readout, UAT, 4 teams adopted.
The client walked away from a $20M acquisition.
I rebuilt 15+ years of cohort economics and stress-tested three downside cases. None cleared the hurdle.
40%
Projected return erosion exposed (est.).
$2.5M
Ad spend reallocated on a second case, across 100+ SKUs.
03
Leadership
Before any of it, two years at the Joint Security Area.
I learned what a bad translation costs before I learned what a PRD was.
I led a 10-person, 24/7 operations team coordinating the US Army, UN Command and 4+ countries, and interpreted 100+ briefings. Named Best Squad Leader.
300+
Cross-border missions run.
+42%
Log accuracy, after postmortems on 50+ incidents.
5+
Frontline units adopted the transcriber I built.
Joshua Lee
easyhoon75@gmail.comGraduating from UNC Kenan-Flagler in May 2027. Looking for an APM role.