Skip to content

OpenShell Backend

The OpenShell backend runs AI agents inside OpenShell sandboxes with network policy enforcement, filesystem isolation, and Landlock-based access control.

New to OpenShell? Start with Local Development with OpenShell, a step-by-step guide to running agents in a sandbox on your workstation.

How It Works

The backend manages three components: a gateway (control plane), a provider (credentials), and a sandbox (isolated execution environment). On each agentic-ci run --backend openshell, it:

  1. Starts the OpenShell gateway with TLS and mTLS auth
  2. Creates a GCP, Anthropic, or OpenAI credential provider (none for a Claude subscription OAuth token)
  3. Creates a sandbox container from the specified image
  4. Applies a network policy and waits for it to activate
  5. Runs setup steps on the host (if configured in .agentic-ci/config.yml)
  6. Uploads the workdir (including setup step outputs) into the sandbox
  7. Uploads an env script with agent configuration
  8. Executes the agent inside the sandbox
  9. Downloads the workdir back to the host and restores the host's git control files (see Host git after the run)
  10. Tears everything down on completion

OpenShell Commands

Below is the exact sequence of openshell CLI commands that agentic-ci executes. All management commands go through the openshell client CLI, which talks to the running openshell-gateway server over gRPC.

Gateway Setup

# Check if gateway is already running
openshell status

# Generate TLS certificates for sandbox JWT auth
openshell-gateway generate-certs \
  --output-dir ~/.local/state/openshell/tls \
  --server-san host.openshell.internal

# Start the gateway server (background process)
# Reads config from ~/.config/openshell/gateway.toml
openshell-gateway --db-url sqlite::memory: --log-level info

# Register the gateway with the CLI
openshell gateway add https://localhost:17670 --local --name ci

# Wait for the gateway to become healthy (retries)
openshell status

The gateway config (gateway.toml) is generated by agentic-ci:

[openshell]
version = 1

[openshell.gateway]
bind_address = "0.0.0.0:17670"
compute_drivers = ["podman"]

# Only added when OPENSHELL_SUPERVISOR_IMAGE is set
[openshell.drivers.podman]
supervisor_image = "custom-supervisor:latest"

Provider Setup

The provider injects credentials into the sandbox. The setup differs by auth mode.

Vertex AI with User OAuth (local development)

openshell provider get ci-gcp                    # check if exists
openshell provider create \
  --name ci-gcp \
  --type google-cloud \
  --from-gcloud-adc \
  --config project_id=<PROJECT> \
  --config region=global

Requires gcloud auth application-default login to have been run first. The --from-gcloud-adc flag reads the user's OAuth refresh token from ~/.config/gcloud/application_default_credentials.json and mints an initial access token synchronously.

Vertex AI with Service Account (CI)

openshell provider get ci-gcp                    # check if exists
openshell provider create \
  --name ci-gcp \
  --type google-cloud \
  --credential GCP_SA_ACCESS_TOKEN=placeholder \
  --config project_id=<PROJECT> \
  --config region=global \
  --config service_account_email=<EMAIL>

# Configure JWT-based token refresh from the service account key
openshell provider refresh configure \
  --credential-key GCP_SA_ACCESS_TOKEN \
  --strategy google-service-account-jwt \
  --material client_email=<EMAIL> \
  --material private_key=<PRIVATE_KEY> \
  --secret-material-key private_key \
  ci-gcp

# Mint the initial token immediately (refresh worker runs on 60s interval)
openshell provider refresh rotate \
  --credential-key GCP_SA_ACCESS_TOKEN \
  ci-gcp

The three-step flow is needed because --from-gcloud-adc rejects service account keys. The refresh rotate call triggers immediate token minting instead of waiting for the 60-second background sweep.

API Key (direct Anthropic API)

openshell provider profile export agentic-ci-anthropic   # import if missing
openshell provider profile import \
  -f src/agentic_ci/backends/openshell/profiles/agentic-ci-anthropic.yaml
openshell provider get ci-gcp                    # check if exists
openshell provider create \
  --name ci-gcp \
  --type agentic-ci-anthropic \
  --credential ANTHROPIC_API_KEY

For Codex, agentic-ci creates an OpenAI provider from its agentic-ci-openai profile and uses OPENAI_API_KEY. The sandbox env script does not carry the key: the provider sets OPENAI_API_KEY in the sandbox to an OpenShell placeholder, Codex's login --with-api-key stores that placeholder, and the proxy swaps it for the real key only on requests to api.openai.com. The endpoint is L4, but the proxy still terminates TLS and replaces the placeholder in the headers of every request on the tunnel, including the WebSocket upgrade that carries Codex's Authorization header. The websocket-credential-rewrite and request-body-credential-rewrite endpoint options are not needed, and OpenShell accepts them only on L7 (rest or websocket) endpoints.

OpenShell refuses to inject a key that contains a line break or NUL and answers the request with HTTP 500 (credential_unavailable), which Codex reports as "We're currently experiencing high demand". A key copied into a CI secret often ends with a newline, so agentic-ci strips surrounding whitespace from OPENAI_API_KEY before it stores the key in the provider, the same way codex login --with-api-key does, and stops with an error if a line break remains inside the key. When it strips anything, the run log shows Stripped surrounding whitespace (a line break) from OPENAI_API_KEY (or spaces or tabs), never the key itself.

openshell provider create \
  --name ci-gcp \
  --type agentic-ci-openai \
  --credential OPENAI_API_KEY

agentic-ci ships its own provider profiles (src/agentic_ci/backends/openshell/profiles/) because OpenShell's builtin openai and anthropic profiles add a rule that lets /usr/bin/curl and /usr/local/bin/curl reach the API host with the key injected, and a sandbox policy cannot remove that rule. agentic-ci's profiles drop that curl rule, and for Codex the raw OpenAI key no longer enters the sandbox. The key can still be spent through an agent binary, for example codex sandbox -- curl, because agentic-ci's own policy lets the agent binaries reach the API host and the proxy injects the key for any of them. The profiles live in the gateway database, so agentic-ci imports them before it creates the provider, and it replaces a provider created from a builtin profile by an earlier release. The sandbox identity file records a short SHA-256 fingerprint of OPENAI_API_KEY (never the key), so after a key rotation agentic-ci recreates the sandbox and stores the new key in the existing provider. Updating the provider alone would not be enough: a running sandbox picks the update up only after about ten seconds, and a placeholder issued before the update keeps resolving to the old key.

OAuth Token (Claude subscription)

When the Claude Code harness runs with CLAUDE_CODE_OAUTH_TOKEN and no ANTHROPIC_API_KEY, agentic-ci creates no provider. OpenShell's Anthropic provider profiles only carry ANTHROPIC_API_KEY as an x-api-key header, and no profile covers a subscription bearer token. The sandbox is created without --provider, the token is exported by the sandbox env script, and the oauth auth mode is recorded only in the sandbox identity file (~/.config/agentic-ci/openshell-sandbox.json).

API-key exposure

With the Anthropic API key, or the Claude subscription OAuth token, the sandbox env script still exports the real credential, so the agent and its child processes can read it. Only Codex's OpenAI key is kept out of the sandbox so far.

Sandbox Resources

When these are omitted, no --memory, --cpu or --gpu flag is passed and the effective limits come from the compute driver and host configuration rather than from agentic-ci. That limit can sit far below what the host has: a sandbox created without --memory on a 16 GB host was OOM-killed by its memory cgroup at under 4 GiB. Pass memory, cpu and gpu to size the sandbox explicitly:

backend = create_backend(
    "openshell",
    harness=harness,
    memory="8Gi",  # accepts 512Mi, 4Gi, 8G
    cpu="4",  # accepts 500m, 1, 2.5
    gpu=1,  # GPU count
)

All three default to None, which passes no flag and leaves current behavior unchanged. A sandbox profile with resources fills in any of the three the caller did not pass (sandbox_profile= on create_backend, or SkillConfig.sandbox_profile); explicit values win.

Resource limits are fixed when the sandbox is created. setup() reuses an existing sandbox rather than recreating it, so these values do not apply to one that already exists; the backend logs a warning saying so. Delete the sandbox to apply new values.

Two failure modes worth recognizing, because neither says what it is:

  • Exceeding a memory limit does not raise. The cgroup OOM-kills the process, the supervisor log stops mid-line, and the next command reports sandbox is not ready. The kernel is the only witness — journalctl -k on the host shows Memory cgroup out of memory: Killed process.
  • Not requesting a GPU leaves the accelerator invisible to the agent even when the host has one and the container running agentic-ci can see it. nvidia-smi -L inside the sandbox returns nothing.

Sandbox Lifecycle

openshell sandbox get ci                         # check if exists

# Create sandbox with the provider attached (omitted for OAuth token auth).
# --memory / --cpu / --gpu are only passed when the caller sets them.
openshell sandbox create \
  --name ci \
  --no-tty \
  --no-auto-providers \
  --provider ci-gcp \
  --from <SANDBOX_IMAGE> \
  --memory 8Gi \
  --cpu 4 \
  --gpu 1 \
  --detach \
  -- sleep infinity

# Keep the sandbox's main process alive; agent commands run via exec.

# Apply network policy and wait for the supervisor to compile and load it.
# Built-in defaults are always included. If .agentic-ci/openshell-policy.yml
# exists in the workdir, its endpoints are merged in automatically.
openshell policy update --wait \
  --binary /usr/local/bin/claude \
  --binary /usr/bin/opencode \
  --add-endpoint github.com:443:full \
  --add-endpoint *.github.com:443:full \
  --add-endpoint gitlab.com:443:full \
  --add-endpoint pypi.org:443:read-only \
  --add-endpoint files.pythonhosted.org:443:read-only \
  --add-endpoint aiplatform.googleapis.com:443:read-write \
  --add-endpoint *.aiplatform.googleapis.com:443:read-write \
  --add-endpoint oauth2.googleapis.com:443:read-write \
  --add-endpoint api.anthropic.com:443:read-write:::allow-uninspected-credentials \
  ci

# Upload env script with agent configuration
openshell sandbox upload --no-git-ignore ci <env-script-file>
openshell sandbox exec --name ci --no-tty -- \
  bash -c "mv <filename> /tmp/.agentic-ci-env.sh"

# Run the agent
openshell sandbox exec --name ci --no-tty -- \
  bash -c ". /tmp/.agentic-ci-env.sh && exec \"$@\"" -- \
  claude --permission-mode bypassPermissions --model <MODEL> \
  --output-format stream-json --verbose -p "<PROMPT>"

Teardown

openshell sandbox get ci                         # check if exists
openshell sandbox delete ci
openshell gateway remove ci                      # deregister from CLI
# Gateway and podman service processes are killed by PID

Host git after the run

The workdir download copies the sandbox's .git over the host repository, and the agent can write that .git. Host-side git that runs afterwards (git rm --cached, git commit --amend, git push) would otherwise honor whatever the agent configured: core.hooksPath, core.fsmonitor, core.sshCommand, credential helpers, filter and diff drivers, aliases, include.path, url.*.insteadOf rewrites, hook scripts, and so on, with every secret the host job holds.

run() therefore records the host's git control files before the agent starts and puts them back once the download finishes, even if it fails (snapshot_git_control() and restore_git_control() in agentic_ci/git.py). The restored paths under .git are config, config.worktree, commondir (which redirects git to another config and hooks directory), hooks/, and info/ (which holds attributes). Objects, refs and the index come back from the sandbox unchanged, so the agent's commits are kept. Any hardening applied before the run, such as harden_git_config(), is part of the restored config.

Restoring first renames each path into a quarantine directory inside .git, which git never reads, and then rewrites it from the recorded copy, so a symlink the agent planted is never written through. Nothing the agent wrote is read or walked before that rename, so a deeply nested tree or a huge file cannot make the restore fail while the agent's config is still live. The quarantine is compared with the recorded copy only for the log line below and is then deleted. If the host copy cannot be written back, .git is moved aside and deleted and the run raises GitControlTamperError, so host git finds no repository rather than the agent's config. If the agent replaced .git itself with a file or symlink, the run raises GitControlTamperError and removes that entry. A .git the agent created in a workdir that had none is removed, since host git run from the workdir would otherwise use it even when the workdir sits inside another repository. When the restore runs as root, the recorded owner of each path is put back as well. Changes the agent made are logged as Agent changed git control files in ...; restored host copy of: .... Git config the agent sets for its own use does not persist to the host.

Setup Steps

Because the sandbox has no internet access by default, repositories that need dependency installation (e.g. npm ci for Node.js projects) can define setup steps that run on the host before the workdir is uploaded. See Project Configuration for full details.

# .agentic-ci/config.yml
setup:
  - name: Install dependencies
    run: npm ci

Network Policy

Endpoints are applied via openshell policy update --wait after sandbox creation. The --wait flag blocks until the supervisor confirms the policy rules are compiled and active. This prevents a race condition where the agent starts before the policy is ready.

Each endpoint must specify explicit binary paths (--binary /usr/local/bin/claude). Using --binary "*" as a wildcard does not work for CONNECT tunnel requests, which is how HTTPS clients establish connections through the supervisor proxy.

The default endpoints cover:

Endpoint Access Purpose
github.com:443 full GitHub API and git operations
*.github.com:443 full GitHub subdomains (raw, API, etc.)
gitlab.com:443 full GitLab API and git operations
pypi.org:443 read-only Python package index
files.pythonhosted.org:443 read-only Python package downloads
aiplatform.googleapis.com:443 read-write Vertex AI (global endpoint)
*.aiplatform.googleapis.com:443 read-write Vertex AI (regional endpoints)
oauth2.googleapis.com:443 read-write GCP token exchange
api.anthropic.com:443 read-write Anthropic API (API key or OAuth token auth)
api.openai.com:443 read-write OpenAI API (Codex, API key auth)
chatgpt.com:443 read-write Codex ChatGPT backend API

Hosts that the attached provider profile marks as credentialed (api.anthropic.com, api.openai.com) carry the allow-uninspected-credentials endpoint option. OpenShell v0.0.116 and later reject L4-only rules for credentialed hosts without it. OAuth token auth attaches no provider, so its api.anthropic.com rule is plain L4.

Project-specific endpoints

Projects can declare additional endpoints in .agentic-ci/openshell-policy.yml. See Project Configuration for details.

OpenShell Artifacts

OpenShell is consumed from UBI9-based artifacts. All three components are pinned to the same version via OPENSHELL_VERSION / OPENSHELL_IMAGE_TAG in images/ci/Containerfile.openshell:

Component Source How it is consumed
CLI (openshell) wheel from the RHOAI package index uv pip install --no-deps (standalone binary; the Python SDK is not used)
Gateway (openshell-gateway) quay.io/opendatahub/odh-openshell-gateway binary copied into the CI image via a multi-stage COPY --from
Supervisor (openshell-sandbox) quay.io/opendatahub/odh-openshell-supervisor pulled at runtime by the gateway's podman driver

The OpenShell CI image (Containerfile.openshell) uses a UBI9 final stage. The sandbox images build directly on their hardened Hummingbird agentic images and use the matching digest-pinned Hummingbird builder to install python3 and supporting tools through a temporary dnf mount. The final sandbox images retain no package manager. The podman backend images (Containerfile.podman, Containerfile.base) are unaffected and remain on UBI10. scripts/bump-versions.py bumps the CLI wheel version and the image tag together to keep the three OpenShell components in sync.

Supervisor Image

The sandbox supervisor runs inside each sandbox container and enforces policies. It is mounted as a read-only image volume by the gateway's podman driver. The default is quay.io/opendatahub/odh-openshell-supervisor (set via OPENSHELL_SUPERVISOR_IMAGE in the CI image).

To override the supervisor image, set the OPENSHELL_SUPERVISOR_IMAGE environment variable before running agentic-ci. This is written into the gateway's TOML config under [openshell.drivers.podman] supervisor_image.

Known Issues and Workarounds

--binary "*" does not work for CONNECT requests

The wildcard * in openshell policy update --binary "*" fails to match binaries making HTTPS CONNECT tunnel requests. Use explicit paths instead:

--binary /usr/local/bin/claude --binary /usr/bin/opencode

--from-gcloud-adc rejects service account keys

The google-cloud provider's --from-gcloud-adc flag only accepts user OAuth credentials (from gcloud auth application-default login). Service account JSON keys must be configured via the three-step create + refresh configure + rotate flow described above.

Credential refresh worker does not mint initial tokens

After openshell provider refresh configure, the gateway's refresh worker runs on a 60-second interval. Without an explicit openshell provider refresh rotate, the agent may start before the first token is minted. Always call rotate after configure for service accounts.

OPENSHELL_SUPERVISOR_IMAGE is not a gateway env var

The gateway binary does not read OPENSHELL_SUPERVISOR_IMAGE from the environment. It reads supervisor_image from the [openshell.drivers.podman] section of gateway.toml. The env var is a convention used by agentic-ci (and OpenShell's own dev scripts) to pass the image name into config generation.