Architecture overview¶
The harness turns a single eval.yaml into a scored HTML report (and optional
MLflow run). This page follows the data as it flows through the moving parts, then
tours the agent_eval Python package so you know where each concept lives.
The data flow¶
Everything downstream is driven by one config file. eval.yaml is parsed into an
EvalConfig and strictly validated at load time, a
runner invokes the target on an execution backend,
the produced artifacts are collected, judges score them, and the
results become a report plus optional MLflow tracking.
flowchart TD
Y["eval.yaml"] -->|"EvalConfig.from_yaml"| C["EvalConfig<br/>(validated at load)"]
C --> W["Per-case workspace<br/>(input.yaml + workspace.files)"]
W --> R["Runner<br/>(EvalRunner ABC)"]
R --> B{"Execution backend<br/>(--runner flag)"}
B -->|local| L["Agent CLI subprocess"]
B -->|harbor| H["Harbor task package<br/>(container)"]
B -->|evalhub| E["In-process in Job pod"]
L --> O["Collected outputs<br/>artifacts + tool calls + traces"]
H --> O
E --> O
O --> J["Judges<br/>(check / LLM / agent / builtin / module)"]
J --> T["Thresholds<br/>(regression gate)"]
J --> RW["Reward scalar<br/>(optional, RL)"]
J --> REP["report.html + results"]
REP --> M["MLflow<br/>(optional: traces, metrics, feedback)"]
C -.->|"tags, experiment"| M
Stage by stage¶
| Stage | What happens | Where it lives |
|---|---|---|
| Config | eval.yaml → EvalConfig; mutually-exclusive keys, enums, and reward formulas fail at load |
agent_eval/config.py |
| Prepare | Isolated workspace per case; input.yaml staged, dataset.workspace.files copied in, execution.env injected |
skills/eval-run/scripts/workspace.py |
| Execute | Runner invokes the skill or prompt headlessly, one call per case (mode: case) or one call for all cases (mode: batch) |
agent_eval/agent/, skills/eval-run/scripts/execute.py |
| Collect | Gather outputs[].path artifacts and outputs[].tool calls; map them back to cases (via batch_pattern in batch mode) |
skills/eval-run/scripts/collect.py |
| Score | Run each judge over the per-case outputs record; pairwise/regression as configured |
skills/eval-run/scripts/score.py |
| Report | Per-judge pass rates / means, per-case detail, diffs, cost & token metrics | skills/eval-run/scripts/report.py |
| Track | Optional: sync dataset, log run params/metrics/artifacts, attach trace feedback | agent_eval/mlflow/, skills/eval-mlflow/ |
Describe what, not where
eval.yaml describes what to evaluate. The execution backend is always
the --runner CLI flag — never a config key — so the identical file runs Local,
on Harbor, or on EvalHub unchanged.
What gets executed vs. how¶
The harness keeps two dimensions orthogonal — how many invocations
(execution.mode) and what to execute (execution.skill or execution.prompt,
mutually exclusive). See the execution model for the full grid.
Argument templating
arguments (and prompt) support two auto-detected placeholder styles resolved
against each case's input.yaml: Jinja2 ({{ input.field }}, StrictUndefined)
and braces ({field} required, {field?} optional). See
resolve_arguments in agent_eval/config.py.
The agent_eval package¶
The skills under skills/ are thin orchestration around the reusable agent_eval
Python package. The package is where the harness logic actually lives.
agent_eval/
├── config.py # eval.yaml → EvalConfig (+ strict validation)
├── state.py # shared key-value state persistence
├── agent/ # runners: the EvalRunner abstraction
│ ├── base.py # EvalRunner ABC + RunResult
│ ├── claude_code.py # Claude Code CLI runner (claude --print)
│ ├── cli_runner.py # opaque CLI runner (command templates)
│ └── stream_capture.py # stream-json → events, timestamps, usage, hooks
├── harbor/ # containerized execution (task packages, reward bridge,
│ # Podman + Kubernetes environments, results parsing)
├── evalhub/ # in-process adapter for the EvalHub platform
├── tools/ # interception.py — PreToolUse interception generation
├── judges/ # builtin judges (auto-discovered by category)
├── prompts/ # builtin generation prompts (synthetic datasets)
├── mlflow/ # experiment setup, dataset sync, trace builder, feedback
└── cli/ # claude-trace standalone tracing CLI
Key handoffs between package and skills:
EvalConfigis the contract every layer reads.resolve_skill()andis_prompt_mode()decide skill vs. prompt;eval_name()derives the run/experiment identifier;resolve_path()resolvesdataset.pathrelative to the config file.EvalRunner+RunResult(agent/base.py) is the runtime-agnostic seam. Therunner.typediscriminator (claude-code,cli, …) selects the implementation; new agent runtimes plug in here. See Runners.harbor/andevalhub/are alternative execution substrates behind the same config. Harbor emits self-contained container task packages (with the judge engine bundled asreward.json); EvalHub runs the eval in-process inside a Job pod. See Execution backends.
Schema fields are documentation, not a parser spec
dataset.schema and outputs[].schema are natural-language descriptions read by
LLM agents and judges. The scripts move file paths around — there are no
hardcoded field names, and nothing parses the schema text into a struct.
The outputs record judges see¶
Collection produces one outputs dict per case, and that is the sole input each judge
receives. It carries collected artifacts (keyed per outputs[].schema/types),
captured tool calls, the traces you enabled (stdout, stderr, events,
metrics), and the case's annotations.yaml under outputs["annotations"]. LLM
judges additionally get {{ conversation }} and {{ tool_trace }} template variables.
See Judges & scoring.
Where to go next¶
- The execution model — the case/batch × skill/prompt grid
- Runners — the
EvalRunnerabstraction andrunner.type - Execution backends — Local, Harbor, EvalHub from one config
- Datasets & case provenance — case anatomy and the three strategies
- Judges & scoring — the five judge types and the outputs record
- Regression thresholds — how a run is gated
- The Reward API — collapsing judges into an RL scalar
- Tool interception — the PreToolUse hook
- Lifecycle hooks — before/after all/each pipeline hooks
- MLflow tracing — hierarchical traces from stream-json
- The HTML report — what the generated report shows
- The eval.yaml schema — every config key