Prerequisites
- A project and one task that all compared Runs address.
- At least two existing Runs, or one or more launch recipes for controlled participants.
- A compatible evaluation method and trusted judge agents when AI judging is required.
Create a Study
Open Project → Evaluations and create a Study for one task. Add participants in either mode:- Observed: select an existing Run. Use this for a retrospective comparison.
- Launched: define an immutable recipe and let the Study create new Runs. Use this when the experiment must control the execution axes.
Design useful variants
Change one important axis at a time when you need a causal answer. Examples:
Use replicates when model variance could be larger than the expected difference.
A single Run can reveal a defect, but it rarely establishes a stable ranking.
What the Study freezes
Starting an evaluation seals the participant set and bounded evidence snapshots. A participant added later does not enter an in-flight execution. The method, judge panel, effective policy, and candidate evidence are also snapshotted so a catalog edit cannot rewrite the meaning of an old score. Evidence can include diffs, test and lint reports, judgments, result objects, Run-tree metrics, duration, and token use. Unavailable evidence stays unavailable; MAIster does not convert missing data into a zero.Evaluation methods
An evaluation can combine several tool types:
Judge candidates are blinded and receive only the evidence allowed by the
method. AI judgments are advisory. Only a person can record the conclusive
Study verdict, correct it with a superseding verdict, or standardize a winning
recipe.
Read the result
Evaluate each participant across four dimensions:- Correctness and evidence: required checks, artifact freshness, diff, and promotion readiness.
- Quality judgment: criterion scores, pairwise results, judge agreement, and the reasons behind them.
- Execution behavior: retries, rework, crashes, child Runs, and human wait.
- Economics: input, output, cache-read, and cache-creation tokens; elapsed time; runner/model attribution; resume overhead; and budget events.
Compare recursive work
For orchestrated Runs, objective providers can measure child Run count, invalid result count, result collection and consumption ratios, rework, crashes, tree tokens, tree wall-clock, and promotion readiness. These facts show whether a recursive harness used its children effectively; a judge or person still decides whether the trade-off was worthwhile.Standardize a winner
After a conclusive human verdict, a project member can standardize an eligible launched recipe for a project slot. MAIster reruns preflight at confirmation and appends an audit revision. Rollback appends another revision; it does not rewrite history. Standardization does not promote the participant and does not change any Run status. Promotion remains a separate human-governed action.Failure signals
- Preflight refusal: fix the named compatibility, trust, artifact, or runner condition before creating the batch.
- Partial evaluation: one or more objective checks, pairs, or judge attempts did not produce sufficient valid evidence.
- Review required: judges disagree or quorum was not met.
- Inconclusive verdict: the evidence does not justify a winner; preserve it rather than forcing a ranking.