Tools How-to

Run Codex from a script: codex exec, structured output and CI

The non-interactive mode turns Codex into a pipeline step. The defaults are conservative, and one warning matters a lot.

Codex isn’t only a terminal app you sit in front of. codex exec runs it without the interactive interface, so it can be a step in a script, a scheduled job or a CI pipeline. The non-interactive mode docs lay out the behavior, and the defaults tell you how OpenAI expects you to use it.

What you get by default

codex exec streams progress to stderr and prints only the final message to stdout. That split is what makes it pipeable: the answer goes down the pipe, the chatter doesn’t. It runs in a read-only sandbox, so by default it can read your repo and report, not change anything. It also expects to be inside a Git repository, unless you pass --skip-git-repo-check.

codex exec "summarize the repo structure"

Feed it context by piping, either as extra context for an instruction or as the whole prompt:

cat build.log | codex exec "find the first real error in this log and say which file to look at"
cat prompt.txt | codex exec -

Letting it edit

To allow changes, opt in to a writable sandbox. The docs recommend the explicit flag in new scripts:

codex exec --sandbox workspace-write "add missing unit tests for src/parse.ts"

danger-full-access exists for broader access and the docs say to use it only in controlled environments. In a CI runner you’d normally reach for it never.

Structured output

Free text is fine for a human and poor for a script. Two flags help. --output-schema takes a path to a JSON Schema and constrains the final response to it. -o (or --output-last-message) writes that final message to a file.

A small schema for a triage job, which is ours:

{
  "type": "object",
  "properties": {
    "severity": { "type": "string", "enum": ["low", "medium", "high"] },
    "summary": { "type": "string" }
  },
  "required": ["severity", "summary"],
  "additionalProperties": false
}
codex exec --output-schema triage.schema.json -o triage.json \
  "triage the failing test output in ci.log"
jq -r .severity triage.json

For richer plumbing, --json turns stdout into a JSON Lines stream of every event, including types like thread.started, turn.completed, item.* and error. Use it when you want to log what the agent did, not just what it concluded.

Resume a session

A multi-stage job can continue the same session:

codex exec resume --last "now write the fix"
codex exec resume <SESSION_ID>

Make runs reproducible

CI wants determinism. The docs list flags for that: --ephemeral skips saving session files, --ignore-user-config skips $CODEX_HOME/config.toml, and --ignore-rules skips user and project execpolicy rule files. A run that ignores your laptop’s config behaves the same on a runner.

The warning to read twice

For authentication, the docs show setting CODEX_API_KEY inline for a single run:

CODEX_API_KEY=<api-key> codex exec --json "triage open bug reports"

Then comes the caution. Never set OPENAI_API_KEY or CODEX_API_KEY as job-level environment variables when repository-controlled code runs in the same job. If a test script, install hook or build step from the repo can read the job’s environment, so can anyone who can open a pull request. Scope the key to the single step that needs it.

Next step: wire a read-only codex exec into a non-blocking CI job that summarizes failed builds, and look at a week of its output before you give it write access.