Make the comparison concrete
I built an arena that runs up to four unedited outputs from the same brief. Coding tasks include dashboards, games and simulations; writing tasks use the same side-by-side approach.
Let people judge what AI models actually make, rather than asking them to trust a launch demo.
A comparison is only useful if the inputs and judging conditions are understandable. The product needed consistent briefs, reproducible artifacts and a voting experience that did not reveal the model first.
Solo product, interface and generation-pipeline work. Community preference is the measured outcome; the arena does not claim a universal ranking of model ability.
I can turn a technical comparison into a usable product while keeping its assumptions and measurement limits visible.
testingmodels.com is a coding and writing arena. Visitors compare working outputs side by side, vote with model names hidden and inspect cost, generation time and the published method.

English AI narration · Subtitles included · Play with sound, or read the transcript.
0:00 I built a way to compare AI models using the same prompt and their actual outputs.
0:07 Challenges bring the results, vote counts and current leaders into one place.
0:13 Each pane runs the page that a model produced, so visitors can inspect and use it.
0:19 Blind mode hides model names and shuffles the results before people vote.
0:25 The same comparison works for games, with playable versions of one shared brief.
0:31 Writing challenges compare stories, essays and other formats side by side.
0:37 The leaderboard puts preference rankings next to cost, token use and generation time.
0:44 The methodology explains how those numbers work, and what they do not tell us.
I built an arena that runs up to four unedited outputs from the same brief. Coding tasks include dashboards, games and simulations; writing tasks use the same side-by-side approach.
New visitors see shuffled Model A to D panes. Voting reveals the names, and the leaderboard calculates rankings from the complete match log.
Prompts, effort variants, generation times and estimated costs are visible. The method page explains both the calculation and the limits of a preference-based arena.
The public vote API recorded 3,691 votes, including ties, on 6 October 2026.
The 6 September 2026 artifact manifest contained 4,530 outputs across 54 challenges and 148 model variants. Those artifacts can be opened and judged, rather than existing only as a benchmark score.
Domain property: testingmodels.com (all subdomains)
Open Google Search Console report · Checked 8 Oct 2026
Domain property: testingmodels.com (all subdomains). Provider timezone: America/Los_Angeles.
| Month | Google search clicks |
|---|---|
| 1 Apr - 30 Apr 2026 | 0 |
| 1 May - 31 May 2026 | 0 |
| 1 Jun - 30 Jun 2026 | 0 |
| 1 Jul - 31 Jul 2026 | 26 |
| 1 Aug - 31 Aug 2026 | 3 |
| 1 Sept - 30 Sept 2026 | 12 |
testingmodels.com and www.testingmodels.com only · UTC
Recorded activity is affected by consent, blockers, tracking gaps and internal traffic.
| Metric | 7 Sept - 4 Oct 2026 | 10 Aug - 6 Sept 2026 |
|---|---|---|
| Recorded pageviews | 224 | 477 |
| Recorded sessions | 180 | 274 |