Create your own eval

Describe your agent. Get five evals.

Two questions, then five evals matched from the Prefactor eval library. They are a starting set: some will need their criteria or wording amended for your agent, and each one links to the library entry it came from.

1. Where is your agent?This decides what the five evals are for.
What changes between the two
Still buildingIn production
Runs onYour cases, on demand, marked purpose EvalEvery live run, as it happens
CatchesThe failures you already know aboutThe failures you did not imagine, by category
CostsTime to write the casesModel tokens on judge evals, nothing on rule checks
Gives youA golden set to gate every releaseA verdict on every run and a monthly cost line
Same in bothThe question each eval asks. Only where it runs and what it costs change.
2. Your agentPlain English. The worked example is filled in.