Tags give the ability to mark specific points in history as being important
-
1.0.0-beta6
d9bdd6e4 · ·1.0.0-beta6 - Riverside Library example recipe from the demo videos, with provider and models as recipe inputs, and installer/riverside.sh to build a new DDEV site with it in one command (#3594754). - A target can reference an entity dataset by UUID, so a target can ship in a recipe or config sync together with its dataset (#3594753). No database updates.
-
1.0.0-beta5
1e3da0d1 · ·1.0.0-beta5 - Per-judge models and trace review in error analysis. Highlights since 1.0.0-beta4: - Each LLM judge can use its own provider and model; runs, validations and exports record the model each judge used (#3594751). - Reviewed conversations carry failure modes and feed failure rates on trace review and the test workflow page (#3594752). - Promote to dataset writes criteria and an expected fact, so promoted examples are scored on the next run (#3594752). - Error analysis steps on the test workflow page do what they say, and the guided test no longer mixes up example rows (#3594752). Run drush updatedb (updates 10040-10044).
-
1.0.0-beta4
c72c3277 · ·1.0.0-beta4 - Guided test workflow and honest gate figures. Highlights since 1.0.0-beta3: - A guided three-step test workflow with drafts, shared datasets, run comparison and recovery of interrupted runs (#3594748). - Gate aggregates count only runs that carry gate evidence (#3594741). - Import every_eval_ever envelopes from the results dashboard (#3594742). - markdown_structure rubric checks now run (#3594743). - Result pages no longer crash when the summary links a rubric or dataset (#3594750). Upgrading: run drush updatedb (10038, 10039). File datasets are now refused in full when invalid (#3594732). See CHANGELOG.md.
-
1.0.0-beta3
7fa303e5 · ·1.0.0-beta3 - Portable ground truth and outside-in verdicts. Highlights since 1.0.0-beta2: - Row assertions accept fixture UUIDs, so shipped datasets grade any fresh provision without re-keying (#3594736). - every_eval_ever envelopes import into the results table, so outside-in and in-process verdicts share one screen (#3594739). - Unreadable assertion input fails in place instead of grading what is left (#3594738).
-
1.0.0-beta2
df0b261e · ·1.0.0-beta2 - Trust hardening and a documentation site. Sprint 6 and the post-sprint trust work, plus the first published documentation site. Highlights since 1.0.0-beta1: - Judge discrimination resolution: validation now estimates the smallest quality difference a judge can actually see, and the optimizer refuses to claim a gain below that floor. - Tool assertions can name arguments and order, and the run record carries each call's arguments (scalars only, truncated, credential-looking names redacted). The judges that read the record are shown arguments too. - An agent-mode run says so when nothing it uses can observe behavior rather than text, read from a new observes_actions flag on the grader attribute. - Per-grader score scales, so a binary judge is no longer forced onto a 0-5 gradient. - Documentation site at https://project.pages.drupalcode.org/ai_eval, 23 pages built by MkDocs, plus a first ai_eval.api.php. Full detail in CHANGELOG.md under 1.0.0-beta2.
-
1.0.0-beta1
0d53e0ae · ·1.0.0-beta1 — Trustworthy Beta (Sprint 4). Adds pluggable rubric check kinds (target_match, command, score_delta, chrF) and check-executor-provided schemas (#3594701, #3594704) on top of the errored-judge trust fix (#3594703).
-
1.0.0-alpha10
87451ac3 · ·Release 1.0.0-alpha10 Adds Wilson confidence interval and Cochran sample-size math to scoring/reporting, fixes three optimizer review findings (aggregated baseline gate, propose-margin ordering, bounded LLM judge input), exposes five previously config-only settings in the admin form, and rewrites the Cochran sample-size warning for clarity (now fires on any below-floor dataset, with plain-language framing and concrete actions). Supersedes the unpublished alpha9 tag.
-
1.0.0-alpha8
6b745783 · ·Release 1.0.0-alpha8 Fixes form TypeError when saving EvalTarget (question_pass_threshold / response_char_limit / threshold typed properties rejected raw response char limit feature and judge-validation hardening.
-
1.0.0-alpha7
4cad6317 · · -
1.0.0-alpha6
e78c7cdd · ·1.0.0-alpha6: Config entities, chat mode, CallerContextInterface, PHP 8.3 attributes, OOP hooks
-
1.0.0-alpha4
dcd4c8d1 · ·1.0.0-alpha4: UI polish, select fields, meters, dashboard layout, docs update
-
1.0.0-alpha2
07f21550 · · -
1.0.0-alpha1
1e63237f · ·