Every public benchmark saturates within a year of release. The scores keep going up. The complaints from your users do not go down. Recast writes the expert data, the evaluations and the environments that still tell you something after the model has seen everything else.
The job posting said domain expert. The rate said otherwise. What arrived was a file with a score in one column and no way to tell who earned it.
Recast pays practitioners at their own rate to write in their own field, under their own name. A second practitioner grades it. A third settles the disagreements. What did not pass stays in the batch, with the reason.
Clinicians, lawyers, engineers and accountants who do the job for a living.
Your task gets the few who have done it before.
Once the questions are public they are in the next training run. The chart is one public test scored across five model generations, next to a task set the models never saw.
Recast builds the evaluation from the job you want done, scores it against your bar, and keeps it off the internet. The line that matters is the one that moves slowly.
The spec covers the happy path. The job is the rest: the ticket with two asks in it, the document that is out of date, the tool that returns nothing and does not say why.
Recast builds environments from real work with the messy parts left in, and grades the run the way the person who owns the job would grade a new hire.
For experts · We're hiring
Fourteen years in a clinic teach what the textbook leaves out. That is the part worth paying for, and the part nobody has been paying for.
Recast hires practising clinicians, lawyers, accountants and engineers at practitioner rates to write and grade in their own field. Your name stays on the work. You see how it was graded and by whom.