The test got easier. The model did not get better.

Looking for a job? Apply now

Horizon

Every public benchmark saturates within a year of release. The scores keep going up. The complaints from your users do not go down. Recast writes the expert data, the evaluations and the environments that still tell you something after the model has seen everything else.

Expert data · practitioners, own field, own name
Evaluations · built from the job, held out
Environments · real work, messy parts left in
Task 4127hepatology · spec v3
differential, 14 findings
AuthorH-0412 · 14 yr in practice
board certified, registry checked
GradersH-0087, H-0203 on conflict
batch κ 0.81
Batch 0931 · delivered 2026-09-12 14:03
Authored1,204 items · 31 authors
median 48 h from spec
Accepted831 · ships with record
agreement per item attached
Rejected373 · reason logged
#4131: rationale copied from a model
A delivery, as it arrives. Drag the sun to turn it.

Most expert data was written by someone with a search bar.

The job posting said domain expert. The rate said otherwise. What arrived was a file with a score in one column and no way to tell who earned it.

Recast pays practitioners at their own rate to write in their own field, under their own name. A second practitioner grades it. A third settles the disagreements. What did not pass stays in the batch, with the reason.

At fifteen dollars an hour you are not buying expertise. You are buying a search bar.
Path of one itemrings · inside out
SPEC author · writesreview · editsadjudication · on conflictdelivery 373 rejected · reason logged every dot keeps its record
One item's path, inside out. Rejected items fall outside the last ring and keep their record.

20,000,000experts

Clinicians, lawyers, engineers and accountants who do the job for a living.

Your task gets the few who have done it before.

20,000,000 on the network

A benchmark stops measuring on the day it is published.

Once the questions are public they are in the next training run. The chart is one public test scored across five model generations, next to a task set the models never saw.

Recast builds the evaluation from the job you want done, scores it against your bar, and keeps it off the internet. The line that matters is the one that moves slowly.

If the score went up and the product did not, you measured the wrong thing.
Score vs. model generationplotter 02
1007550250 gen 1gen 2gen 3gen 4gen 5 public benchmark · 98.1 held-out task set · 45.4 SCORE (%)
model a 72.6
model b 68.2
model c 61.1
model d 52.2
model e 41.9
Same models, two tests. Below, the held-out set, one dot per five points.

Agents fail in the part of the job nobody wrote down.

The spec covers the happy path. The job is the rest: the ticket with two asks in it, the document that is out of date, the tool that returns nothing and does not say why.

Recast builds environments from real work with the messy parts left in, and grades the run the way the person who owns the job would grade a new hire.

Nobody gets promoted for the happy path.
Run 88 · claims adjustmentenv v2
  • 1 open claim C-2291
  • 2 read policy · finds 2018 revision
  • 3 request repair estimate
  • 4 estimate arrives · two line items
  • 5 check coverage · line 1 covered
  • 6 approves line 2 · excluded since 2021
  • 7 writes customer · polite, wrong
  • · grader H-0771: "would have caught it on day three."
One run, graded by the person who used to do the job.

For experts · We're hiring

You know things a model cannot look up.

Fourteen years in a clinic teach what the textbook leaves out. That is the part worth paying for, and the part nobody has been paying for.

Recast hires practising clinicians, lawyers, accountants and engineers at practitioner rates to write and grade in their own field. Your name stays on the work. You see how it was graded and by whom.

Get paid for the part of the job that took a decade to learn.
01 / 04
Hepatology
.88
in practice
14 yr
items
412
agreement
0.88 · vs. 2 peers
02 / 04
Tax & controversy
.84
in practice
9 yr
items
287
agreement
0.84 · vs. 2 peers
03 / 04
Distributed systems
.79
in practice
11 yr
items
536
agreement
0.79 · vs. 3 peers
04 / 04
Structural engineering
.91
in practice
18 yr
items
194
agreement
0.91 · vs. 2 peers
Four authors, as data. The ring is agreement with peers.

Send us the job you want the model to do.

Get started
Earthset from Orion, Artemis II: NASA