< cd ~/raba.pl

Public products & communities

testingmodels.com

Let people judge what AI models actually make, rather than asking them to trust a launch demo.

  • live
  • June 2026 - present
  • Solo: product, code, generation pipeline

The problem

A comparison is only useful if the inputs and judging conditions are understandable. The product needed consistent briefs, reproducible artifacts and a voting experience that did not reveal the model first.

My part

Solo product, interface and generation-pipeline work. Community preference is the measured outcome; the arena does not claim a universal ranking of model ability.

What this shows

I can turn a technical comparison into a usable product while keeping its assumptions and measurement limits visible.

Where it started

testingmodels.com is a coding and writing arena. Visitors compare working outputs side by side, vote with model names hidden and inspect cost, generation time and the published method.

See it in action

Product walkthrough · 1:00 · 60 FPS · English narration
testingmodels.com interface

English AI narration · Subtitles included · Play with sound, or read the transcript.

Read the video transcript

0:00 I built a way to compare AI models using the same prompt and their actual outputs.

0:07 Challenges bring the results, vote counts and current leaders into one place.

0:13 Each pane runs the page that a model produced, so visitors can inspect and use it.

0:19 Blind mode hides model names and shuffles the results before people vote.

0:25 The same comparison works for games, with playable versions of one shared brief.

0:31 Writing challenges compare stories, essays and other formats side by side.

0:37 The leaderboard puts preference rankings next to cost, token use and generation time.

0:44 The methodology explains how those numbers work, and what they do not tell us.

The decisions behind the product

Make the comparison concrete

I built an arena that runs up to four unedited outputs from the same brief. Coding tasks include dashboards, games and simulations; writing tasks use the same side-by-side approach.

Separate the work from its label

New visitors see shuffled Model A to D panes. Voting reveals the names, and the leaderboard calculates rankings from the complete match log.

Publish the method alongside the result

Prompts, effort variants, generation times and estimated costs are visible. The method page explains both the calculation and the limits of a preference-based arena.

What came out of the work

A comparison people participated in

The public vote API recorded 3,691 votes, including ties, on 6 October 2026.

A broad set of inspectable outputs

The 6 September 2026 artifact manifest contained 4,530 outputs across 54 challenges and 148 model variants. Those artifacts can be opened and judged, rather than existing only as a benchmark score.

  • 3,691community votes, ties includedPublic vote API · 6 Oct 2026
  • 4,530live artifacts to open and judgeArtifact manifest · build 6 Sep 2026
  • 148model variants across 24 familiesArtifact manifest · build 6 Sep 2026
  • 54challenges: 43 coding, 11 writingArtifact manifest · build 6 Sep 2026

Traffic over time

Google Search Console

Google search clicks

Domain property: testingmodels.com (all subdomains)

Latest month
121 Sept - 30 Sept 2026
First month
01 Apr - 30 Apr 2026
Google search clicks per month, 1 Apr to 30 Sept 2026Google Search Console. Domain property: testingmodels.com (all subdomains). Monthly clicks. 1 Apr: 0; 1 May: 0; 1 Jun: 0; 1 Jul: 26; 1 Aug: 3; 1 Sept: 12. Vertical axis starts at zero.015301 Apr - 30 Apr 2026: 0 clicks1 May - 31 May 2026: 0 clicks1 Jun - 30 Jun 2026: 0 clicks1 Jul - 31 Jul 2026: 26 clicks1 Aug - 31 Aug 2026: 3 clicks1 Sept - 30 Sept 2026: 12 clicksAprMayJunJulAugSept
6 complete months · 1 Apr - 30 Sept 2026 · Download chartFirst month Latest month
Source, methodology and monthly data

Open Google Search Console report · Checked 8 Oct 2026

Domain property: testingmodels.com (all subdomains). Provider timezone: America/Los_Angeles.

  • Six complete calendar months: 1 April to 30 September 2026. Current is September; comparison baseline is April. Both comparison months have 30 days. This is not a comparison against the preceding six months.
  • Final web-search data, all countries and devices, America/Los_Angeles reporting timezone. Search clicks are not website visits or unique visitors.
  • Domain properties include all subdomains with property aggregation. Documentation pages retain exact page filters from the preceding evidence collection and measure documentation interest, not adoption.
  • Monthly chart values and the six-month total are reconciled with independently queried totals.
  • The percentage, where shown, compares September 2026 with April 2026, the last and first complete months of this window.
  • April has fewer than 100 events, so absolute counts are shown without a percentage growth claim.
  • For fewer than 100 events in the baseline period, the comparison shows absolute counts instead of a percentage.
Complete monthly totals
MonthGoogle search clicks
1 Apr - 30 Apr 20260
1 May - 31 May 20260
1 Jun - 30 Jun 20260
1 Jul - 31 Jul 202626
1 Aug - 31 Aug 20263
1 Sept - 30 Sept 202612
PostHog recorded activity224 pageviews in 28 days

testingmodels.com and www.testingmodels.com only · UTC

Recorded activity is affected by consent, blockers, tracking gaps and internal traffic.

Recorded totals, separate from search data
Metric7 Sept - 4 Oct 202610 Aug - 6 Sept 2026
Recorded pageviews224477
Recorded sessions180274

Inside the implementation

Explore the features, architecture and quality checks

What I built

  • Coding and writing arenas: up to four models answer the same brief in side-by-side panes, each running the model's unedited output
  • 43 coding briefs, from landing pages and dashboards to playable games, simulations and three Godot builds, plus 11 writing briefs from essays to satire
  • Blind-first judging: new visitors see shuffled panes labelled Model A to D, and a vote reveals who wrote what
  • A per-arena Elo ledger built from pairwise votes and recomputed from the full match log on every read
  • Every thinking-effort level published as its own variant, with estimated tokens, cost and generation time per output
  • A leaderboard, a method page with a changelog, a renders gallery and a blog, in English and Polish; cookieless Umami, PostHog only after consent

How it works

  1. BriefOne prompt per challenge, written once and handed verbatim to every model. The full prompt is public.
  2. GenerateSingle-file tasks are one-shot HTML with no follow-ups. Godot tasks iterate until the build passes or 40 turns run out.
  3. BakeBuild scripts index every artifact, estimate tokens as characters / 4 and price them at published output rates.
  4. JudgePanes load side by side, blind for new visitors. Heavy tasks start as posters, so nothing wins by loading first.
  5. RankA vote scores a win over every shown opponent. Elo starts at 1500, with a K-factor of 32 for a variant's first 30 matches, then 16.

Quality and reliability

  • No per-model prompt tuning or quiet retries; an entry re-run after a crash with zero output is marked as regenerated
  • Votes validated with Zod at the API; one vote per task per browser, changeable at any time
  • Home page numbers are computed from the artifact manifest at build, so they cannot drift from what shipped
  • Strict TypeScript with a tsc --noEmit typecheck script
  • The method page states what the arena is not: no pass@k, no held-out test sets, no significance testing

Technology

  • Next.js 16
  • React 19
  • TypeScript
  • Tailwind CSS 4
  • Neon Postgres
  • Zod
  • Godot
  • Playwright
  • Umami
  • PostHog
  • Vercel

Explore another problem I worked on

>_ Your next project

Liked what you saw?
Let’s get in touch!

Have an idea, a challenge, or a role in mind? I’d love to hear about it.

Let’s talk

>_ Start a conversation

Liked what you saw?
Let’s get in touch!

Tell me what you’re working on. Let’s see how I can help.

Email mepatryk@raba.pl
Send a message