Engraved medieval helmet and gorget in a museum — photo by Blackcurrant Great on Pexels
Suit of medieval armour holding a shield in a museum — photo by Alisa Skripina on Pexels

Creative intelligence.

Where models compete on real creative work. Commissioned battles and long-form research, drawn from the network for creative judgement.

The Human Craft Benchmark The first eval that scores models the way working creatives do.
Latest 06/01/2026BattleLumen v4 won 46.2% of typography matchups. 06/01/2026ProfilePrism reliably edits, but can it keep the rest of the image still? 05/29/2026BattleQuill took 60% of head-to-heads. Atlas Code kept the rest.

Best performing models


Image Video Web design

Bradley–Terry leaderboard pooled across 9 image-generation studies and 240 completed tournaments.
Higher Elo means stronger aggregate head-to-head performance across the studies in the pool.

BT Elo · 9 image studies
1142
1Lumen Image 21142
2Lumen Image 1.51061
3Orrery 5.0 Lite1039
4Prism 3 Pro Image Preview1024
5Prism 3.1 Flash Image Preview998
6Halo 2 Large981
7Orrery 4.5964
8Vantage.2947

Methods & standards

Latest research

28 studies

Every battle, model profile, and field note, sorted by newest first, tagged by domain.


Renaissance fresco on a grand painted ceiling — photo by Magda Ehlers on Pexels
06/01/2026ImageBattle

Lumen v4 won 46.2% of typography matchups.

10 designers, 4 models, 240 images. Spelling is solved. Typographic craft and client-readiness are where Lumen v4 pulls away.

Read

Typography win-rate, pooled

Lumen v4 46.2%
Prism 3.1 Flash 34.8%
Vantage.2 (max) 13.9%
Orrery Imagine 1.0 5.1%

Share of head-to-head typography prompts · 4 image models · 20 prompts × 12 iterations · pooled 1 – 5

Classical marble bust of a woman in a garden — photo by Tamula Aura on Pexels
06/01/2026ImageProfile

Prism reliably edits, but can it keep the rest of the image still?

11 production-style sessions. Prism made the edit 75% of the time, kept the rest of the frame intact 58% of the time.

Read
9 / 12Sessions that passed the edit-isolation test
7 / 12Sessions that passed the pose-lock test
5 / 12Sessions that passed both controllability tests

Prism controllability checks · 11 sessions, four tasks per session, both graders in agreement

Ancient statues in a museum hall, black and white — photo by Mesut Yalcin on Pexels
05/21/2026Web designField note

In Atlas Design, your opening prompt decides the ceiling.

5 designers, 5 openings, 1 luxury brief. The first prompt set what each session could reach.

Read

Specificity score (0–1)

1.00.80.60.40.20.0 00.350.480.580.78 P1P2P3P4P5 ConfirmAsset swapFrameworkFull briefHybrid

First-prompt specificity by participant, blind-scored · 5 designers, 1 brief

Cast of the Discobolus statue among greenery — photo by Talha Usman on Pexels
05/20/2026Web designProfile

Atlas Design gets you 40%, Rivet gets the rest.

5 sessions, 5 designers, 1 real-world client brief. Strong as a starting structure, breaks under precision edits.

Read
55 → 100%Designers flagging layout & spacing, Edit 1 → Edits 4–5
5 / 5Sessions where layout was the recurring failure mode
≤38%Designer verdict: use it to here, then hand off

Atlas Design designer evaluation · 5 designers, 1 real-world client brief

Baroque ceiling fresco of angels — photo by Regan Dsouza on Pexels
05/19/2026Cross-cuttingField note

Creatives keep telling us the same thing: every output looks the same.

12 models, 5 creative domains. One repeated complaint from working evaluators: the work all looks the same.

Read
High sameness Low sameness Good practice Rare practice “Creative pattern” “Full-spectrum revolt” “Trainable” “Opinionated engine”

Convergence themes coded across 5 domains and clustered by two independent graders

Vibrant Renaissance fresco of a religious scene — photo by Magda Ehlers on Pexels
05/18/2026Cross-cuttingField note

AI isn’t replacing creative professionals. It’s making the best ones better.

Survey of 340 high-earning independent creatives. What they actually do with AI on real client work.

Read
<22%

of AI output makes it to final deliverables

The rest is stripped, redrawn, or replaced

Self-reported survey · 340 independent creatives · fielded over six weeks

Renaissance-style portrait of a young woman in a ruff collar — photo by Freek Wolsink on Pexels

Connecting with the missing signal: taste

Vellum connects working creative minds with the teams training models to understand taste. This is expert input, not crowd labour — the creative layer under the next generation of tools.

Designers Writers Art directors Engineers Community leads Video editors & animators Sound designers
1.28M+creative experts
380+skills and tools represented
$214M+verified expert earnings