Two questions, then five evals matched from the Prefactor eval library. They are a starting set: some will need their criteria or wording amended for your agent, and each one links to the library entry it came from.
| Still building | In production | |
|---|---|---|
| Runs on | Your cases, on demand, marked purpose Eval | Every live run, as it happens |
| Catches | The failures you already know about | The failures you did not imagine, by category |
| Costs | Time to write the cases | Model tokens on judge evals, nothing on rule checks |
| Gives you | A golden set to gate every release | A verdict on every run and a monthly cost line |
| Same in both | The question each eval asks. Only where it runs and what it costs change. | |