Anthea for Training Models
We work alongside frontier training teams to give their models reference, restraint and a working sense of proportion.
Instead, there can be a room full of odd, specific, well-made things.
That means learning to measure, rank, retrieve and encode subjective domains into data a model can train on, and tools an agent can call.
So that systems return work that is not merely correct, but worth keeping.
We are making the unscoreable scoreable, beginning with design.
Anthea across the stack
We work alongside frontier training teams to give their models reference, restraint and a working sense of proportion.
Design me a quiet, unhurried site for a studio that wants working artists to apply
We work with the app layer, coding agents and creative tooling, supplying the context, evaluation and verification rails so what they ship reads considered and on brand.
Bring your judgment to post-training, retrieval, evaluation, data, design and everything around them, and help us make subjective domains legible.






Product & operations staff
We are half research lab, half infrastructure company. Here is what is currently on our minds.
Explore writing
A conversation with Odile Marchetti, who co-led our $14.6M seed, on why preference is infrastructure.
Explore
Generic output is a measurement problem. Here is how we score work that resists a rubric.
Explore
Production got cheap and discernment got scarce. Notes on grading what will not sit still.
ExploreDrag a tile