CI integration & regression gating¶
Wire an eval into continuous integration so a pull request that degrades your skill
fails the build. The mechanism is deliberately small: you declare per-judge
thresholds in eval.yaml, and the scorer
exits non-zero when a run falls below them — which is all a CI job needs to gate on.
This is an evolving area
The harness ships the primitives for CI gating (thresholds + a non-zero exit), and the GitHub Actions job below works today. Turnkey CI recipes, reusable actions, and caching patterns are still being fleshed out — see remaining work. Treat the example as a starting point to adapt, not a drop-in.
How gating works¶
Scoring aggregates each judge by value type — boolean judges become a pass_rate,
numeric judges become a mean, and a --baseline pairwise
comparison produces a win_rate. Regression detection compares those aggregates against
the thresholds block and, on any breach, returns exit code 1.
flowchart TD
A[score.py judges] --> B{thresholds<br/>configured?}
B -- no --> P[exit 0]
B -- yes --> C[detect_regressions<br/>aggregates vs thresholds]
C --> D{any breach?}
D -- no --> E["print REGRESSIONS: 0<br/>exit 0"]
D -- yes --> F["print REGRESSIONS: N detected<br/>exit 1 → CI fails"]
Thresholds¶
thresholds is a top-level map of judge name → gate. Each gate sets one or more of
four keys; a run regresses if the matching aggregate misses its bound.
| Key | Applies to | Aggregate compared | Passes when |
|---|---|---|---|
min_pass_rate |
boolean judges | pass_rate (fraction of cases that returned True) |
pass_rate >= min_pass_rate |
min_mean |
numeric judges (e.g. 1–5 LLM scores) | mean (average value across cases) |
mean >= min_mean |
min_win_rate |
a pairwise judge (needs --baseline) |
win_rate |
win_rate >= min_win_rate |
max_error_rate |
any judge (opt-in) | fraction of cases where the judge errored | error_rate <= max_error_rate |
judges:
- name: has_content # boolean judge → pass_rate
check: |
content = outputs.get("main_content", "")
return (len(content.strip()) >= 100,
f"{len(content.strip())} chars")
- name: output_quality # numeric judge (1–5) → mean
feedback_type: int
score_range: [1, 5] # declare the scale — omitting it warns at config load
prompt: "Score the output 1-5 for completeness, clarity, and accuracy."
thresholds:
has_content: { min_pass_rate: 1.0 } # every case must pass
output_quality: { min_mean: 3.5 } # average score must stay >= 3.5
An unavailable metric counts as a regression
If a threshold is set but its aggregate is None, that is reported as a regression,
not silently skipped. This is almost always a config mistake — the judge was skipped
for every case (its if: condition, or it errored), or the key targets the wrong
judge type (e.g. min_pass_rate on a numeric judge, whose pass_rate is always
None). Match the key to the judge's value type.
Exit-code behavior¶
Two entry points enforce thresholds; both sys.exit(1) on regression so a CI runner
fails the step automatically.
score.py judges runs the judges and, if thresholds is set, checks them at the end
of scoring — so a normal /eval-run already gates. It prints REGRESSIONS: 0 or
REGRESSIONS: N detected (one line per breach) before exiting.
Runs live under $AGENT_EVAL_RUNS_DIR/<eval-name>/<run-id>/ (default eval/runs/), so
pass a stable --run-id (e.g. the commit SHA) to find the same directory later in the job.
Comparing against a baseline¶
Beyond absolute floors, you can gate a run relative to a previous one with
--baseline <run-id>. The baseline must be a prior run under the same eval-name.
/eval-run --baseline <run-id>adds a position-swapped pairwise judge on top of the regular judges. Each case is judged both A/B and B/A, and only a consistent preference counts as a win;summary.yamlgains apairwisesection (wins_a,wins_b,ties). Gate it with amin_win_ratethreshold.score.py regression --baseline <run-id>also does a direct aggregate comparison: formeanandpass_rate, a current value more than0.5below the baseline is flagged asDegraded vs baseline— catching drift even where no absolute floor was crossed.
# Score this run with a pairwise comparison against last week's baseline,
# then gate both the absolute thresholds and the vs-baseline deltas.
python3 skills/eval-run/scripts/score.py pairwise \
--run-id "$RUN_ID" --baseline 2026-07-09-opus --config eval.yaml
python3 skills/eval-run/scripts/score.py regression \
--run-id "$RUN_ID" --baseline 2026-07-09-opus --config eval.yaml
Absolute vs relative gates
Use min_mean / min_pass_rate for a hard quality floor that must always hold, and
--baseline for catch-any-drift protection between a known-good run and the PR under
test. They compose — a run can pass its absolute floors yet still fail because it
dropped sharply versus the baseline.
A GitHub Actions job¶
Illustrative skeleton — not drop-in runnable
The score.py regression step below is the real, copy-pasteable gate
(exit 1 fails the job). The run step is pseudocode: /eval-run is a
Claude Code slash command, not a binary on PATH. Wire
Claude Code headless or Harbor before using this
in a real repo.
name: skill-eval
on: [pull_request]
jobs:
eval:
runs-on: ubuntu-latest
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
AGENT_EVAL_RUNS_DIR: eval/runs
RUN_ID: ci-${{ github.sha }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install the harness
run: pip install -e .
# PSEUDOCODE — replace with a real driver (not a bare /eval-run shell call).
# Options: Claude Code headless (guides/headless.md) or Harbor
# (`/eval-run --runner harbor …`). Must produce:
# eval/runs/<eval-name>/$RUN_ID/summary.yaml
- name: Run the eval
run: |
your-headless-or-harbor-driver \
--run-id "$RUN_ID" \
--model sonnet
# Gate: exits 1 on any threshold breach, failing the job.
- name: Check for regressions
run: |
python3 skills/eval-run/scripts/score.py regression \
--run-id "$RUN_ID" --config eval.yaml
- name: Upload report
if: always()
uses: actions/upload-artifact@v4
with:
name: eval-report
path: eval/runs/**/report.html
Driving the run in CI
Replace your-headless-or-harbor-driver with Claude Code headless
(or an equivalent driver that prepares workspaces, executes cases, and scores).
For heavier or containerized CI, use the same eval.yaml on Harbor
with --runner harbor — the config is unchanged; only the substrate flag
differs.
Where to go next¶
-
Threshold semantics
The full reference for
min_mean,min_pass_rate,min_win_rate, andmax_error_rate. -
Pairwise & sampling
How
--baselinecomparisons and repeated-sample stability work. -
Run headless
Auto-answer questions and gate external services for unattended runs.