runner¶
The runner block selects which agent runtime executes your skill or prompt and
carries runtime-specific knobs. Its type is a discriminator; the remaining fields are
read selectively — a field a runner doesn't understand is simply ignored, so the same
block is safe to share across runs.
Runtime, not backend
runner.type picks the agent (Claude Code, Codex, an opaque CLI, or the OpenAI
Responses API). The execution backend (Local, Harbor, EvalHub) is a separate --runner
CLI flag, never a config key. See backends and the
runners concept for the distinction.
The four runner types¶
flowchart TD
T{runner.type}
T -->|claude-code<br/>default| CC[ClaudeCodeRunner<br/>claude --print]
T -->|codex| CX[CodexRunner<br/>codex exec]
T -->|cli| CLI[CliRunner<br/>arbitrary command template]
T -->|responses-api| RA[ResponsesAPIRunner<br/>OpenAI Responses + Skills API]
type |
Runtime | Use it for |
|---|---|---|
claude-code (default) |
Claude Code CLI in headless mode (claude --print --output-format …) |
The primary path — full tracing, tool interception, permission enforcement, subagent capture |
codex |
Codex CLI in non-interactive mode (codex exec --json) |
Native Codex execution, skill staging, JSONL usage parsing, and sandbox-mode mapping |
cli |
Any command you provide, via a placeholder template | Wrapping OpenCode, a custom agent, or a shell script. See the opaque CLI runner contract |
responses-api |
OpenAI Responses API with the Shell tool + Skills API | Apples-to-apples comparison of the same skill on an OpenAI model |
Field reference¶
Not every runner reads every field. The matrix below shows where each field lands.
| Field | Type | claude-code |
codex |
cli |
responses-api |
|---|---|---|---|---|---|
type |
str |
discriminator | discriminator | discriminator | discriminator |
effort |
str (enum) |
--effort flag |
model_reasoning_effort |
{effort} placeholder |
— |
permission_mode |
str (enum) |
--permission-mode flag |
mapped to Codex sandbox mode | — | — |
settings |
dict |
merged into workspace .claude/settings.json |
-c config overrides for codex exec |
— | connection settings (see below) |
plugin_dirs |
list[str] |
one --plugin-dir per entry (workspace-staged copy for out-of-workspace paths) |
skills copied into .agents/skills |
— | — |
env |
dict |
injected on the safe allowlist | injected on the safe allowlist | — (uses execution.env) |
— |
system_prompt |
str |
--append-system-prompt |
prepended to the prompt | {system_prompt} placeholder |
developer message |
command |
str | list[str] |
— | — | required — command template | — |
workspace_mode |
None | "repo" |
harness-level (all runners) | harness-level | harness-level | harness-level |
Unset ≠ empty behavior
A field ignored by the active runner is harmless — it just does nothing. But two
fields are easy to misplace: runner.env has no effect on the cli runner
(which inherits the full caller environment and reads execution.env for
additions), and runner.settings means completely different things to
claude-code versus responses-api (see below).
type¶
Selects the runner implementation. One of claude-code (default), codex, cli, or
responses-api. Any other value fails to resolve at run time.
effort¶
Reasoning-effort level for the agent. The accepted values are runner-specific:
| Runner | Valid values |
|---|---|
claude-code |
low, medium, high, xhigh, max |
codex |
minimal, low, medium, high, xhigh |
An invalid value raises at construction time for claude-code and codex. The CLI --effort flag
overrides this field. For the cli runner it is exposed as the {effort} placeholder
(empty string if unset); responses-api ignores it.
permission_mode¶
Claude Code permission mode, passed as the --permission-mode CLI flag. Because
it is a CLI flag (not a settings-file key), it applies even in the untrusted,
isolated per-case workspaces where .claude/settings.json permissions.allow /
additionalDirectories are trust-gated and the trust dialog can't appear in
headless mode. One of:
| Value | default |
acceptEdits |
plan |
auto |
dontAsk |
bypassPermissions |
|---|---|---|---|---|---|---|
An invalid value raises at construction time for claude-code and codex; cli and
responses-api ignore the field. For a prompt-free, deny-by-default headless
run, pair dontAsk (allows only what's pre-approved) with a complete
permissions.allow list (fed to --allowed-tools, also
trust-independent). bypassPermissions skips all prompts — isolated
environments (containers/VMs) only.
Codex preserves the same intent using its available sandbox modes: plan maps to
read-only; bypassPermissions maps to Codex's explicit dangerous bypass; and the
remaining modes map to workspace-write. Codex cannot translate Claude Code's
fine-grained tool allow/deny rules exactly, so the runner emits a warning when such
rules are configured.
runner:
type: claude-code
permission_mode: dontAsk # default | acceptEdits | plan | auto | dontAsk | bypassPermissions
settings¶
A dict whose meaning depends on the runner:
Merged into each case workspace's generated .claude/settings.json (after the
harness defaults, so your scalars win and lists are extended). Use it to add
Claude Code settings — model defaults, env, MCP servers — without forking the
harness.
Connection and container settings for the OpenAI Responses API. Recognized keys:
base_url, api_key, default_model, network_policy, memory_limit_mb
(default 512). Missing base_url / api_key / default_model fall back to the
OPENAI_BASE_URL, OPENAI_API_KEY, and OPENAI_MODEL env vars.
Each key is passed to codex exec as a -c key=value config override.
model_reasoning_effort is also accepted here as a fallback when
runner.effort is unset; the top-level effort field takes precedence
and is validated against Codex's effort values.
The cli runner ignores settings.
plugin_dirs¶
claude-code and codex. Both runners copy plugin content into the case workspace
rather than exposing the configured path to the session. Claude Code stages each
entry's discoverable content (manifest, skill roots, commands/, agents/,
hooks/, scripts/) into .staged-plugins/ and passes the staged copy to
--plugin-dir — the configured path would otherwise land verbatim in session
context, where Bash (not path-gated) can follow it out of the workspace. An entry
already inside the workspace is passed through unchanged, and
workspace_mode: repo skips staging entirely — the workspace is the real
project there, so there is nothing to isolate. Codex copies each plugin's skills into the case workspace's
.agents/skills directory for the duration of the run. Relative paths are always
resolved from the project root.
A lexically external path such as ../shared-skills is an explicit opt-in like an
absolute path; a path declared inside the project may not escape through a symlink.
env¶
claude-code and codex. Extra environment variables injected into the runner
subprocess, additive to the runner's built-in safe allowlist (PATH, HOME,
provider credentials, MLFLOW_TRACKING_URI, …). A value starting with $ is resolved
from the caller's environment; missing vars are dropped.
runner:
type: claude-code
env:
ANTHROPIC_AUTH_TOKEN: $ANTHROPIC_AUTH_TOKEN # forward from caller
FEATURE_FLAG: "enabled" # literal
Where env vars belong
runner.env forwards vars into the agent runtime. To make a var available to
the skill and its hooks inside each case workspace, use
execution.env instead. The cli runner reads only
execution.env — runner.env is a no-op there. See
environment variables.
system_prompt¶
Extra system-prompt text prepended to the agent's context.
claude-code— passed via--append-system-prompt, composed with the harness prompt.codex— prepended to the user prompt.cli— exposed as the{system_prompt}placeholder in the command template.responses-api— sent as adeveloperrole message.
command¶
cli only, and required for it. A command template — a string (shell-parsed via
shlex) or a list of arguments (safer; no shell parsing). Placeholders are substituted
before execution and string values are shell-quoted.
Common placeholders (full list in the contract):
| Placeholder | Value |
|---|---|
{agent} |
Skill name (empty in prompt mode) |
{workspace} |
Absolute case workspace path |
{output_dir} |
{workspace}/output (write artifacts here) |
{model} |
Resolved model (--model or models.skill) |
{subagent_model} |
Subagent model (empty if unset) |
{args} |
Resolved execution.arguments |
{effort} / {system_prompt} |
From runner.effort / runner.system_prompt |
{timeout} / {max_budget_usd} |
From execution (budget is advisory only) |
{field} |
Any field from the case input.yaml |
Contract obligations
An opaque command must exit non-zero on failure and write artifacts to
{output_dir} (or the outputs[*].path dirs). To surface token/cost data it must
write {output_dir}/metrics.json — otherwise cost tables are empty. Tool
interception, stream-json tracing, subagent capture, and budget enforcement are
Claude-Code-only and do not work with cli. See the
cross-runner cookbook.
workspace_mode¶
Execution context, honored by all runners. Validated at load time — only None or
"repo" are accepted (a typo raises rather than silently changing behavior).
| Value | Meaning |
|---|---|
unset (None) |
Isolated workspace (default) — each case runs in its own temp workspace with symlinked project resources |
repo |
In-repo — the agent runs against the real repository checkout |
workspace_mode: repo is meaningful for prompt mode evals that need the agent to
navigate the actual repository (e.g. agentic-docs testing). Pair it with
permissions to keep the agent from writing to the repo.
Examples¶
See also¶
- runners concept — how runtimes differ and when to reach for each
- models — model-per-role and the
{model}/{subagent_model}resolution - execution —
execution.env, timeout, budget, parallelism - permissions — allow/deny, essential for
workspace_mode: repo - headless execution — how the Claude Code runner is driven non-interactively