1.0.0-beta5 - Per-judge models and trace review in error analysis.

Highlights since 1.0.0-beta4:
- Each LLM judge can use its own provider and model; runs, validations
  and exports record the model each judge used (#3594751).
- Reviewed conversations carry failure modes and feed failure rates on
  trace review and the test workflow page (#3594752).
- Promote to dataset writes criteria and an expected fact, so promoted
  examples are scored on the next run (#3594752).
- Error analysis steps on the test workflow page do what they say, and
  the guided test no longer mixes up example rows (#3594752).

Run drush updatedb (updates 10040-10044).