Assistants How-to

GPT-6.1 Sol is in Codex. Test it against Astra on your own task before you switch

OpenAI's pitch is near-Astra quality at lower cost. Here's how to check that on your repo with one script.

Two overlapping circles, one filled with halftone dots and one hatched, with a bright overlap in the middle

OpenAI added GPT-6.1 Sol to Codex on September 29, and Codex CLI 0.161.0 (October 7) makes it the default model in the bundled catalog, according to the Codex changelog . The pitch: near-Astra performance for complex work, at lower cost than Astra. That’s a claim about OpenAI’s tasks. Whether it holds on your codebase is a ten-minute experiment, and OpenAI’s own model selection guide tells you to run it: compare the models on the same task and keep the lightest setting that meets your quality bar.

What the docs say about the lineup

The Codex models page puts three models side by side. Astra is the most capable, for the hardest end-to-end work. Sol sits in the middle and is aimed at repeated, long-running work. Luna is the cheap one, for clear, high-volume, repeatable tasks.

Sol is included with Plus, Pro, Business, Enterprise and Edu plans, in the Codex app and CLI. Enterprise and Edu admins have to switch it on, since it’s off by default. Free and Go plans don’t get it at launch. Standard and Fast modes are available now; Ultrafast is listed as coming later.

You can pick it three ways:

# one-off
codex -m gpt-6.1-sol

# non-interactive
codex exec -m gpt-6.1-sol "Review the current changes"

# permanent, in config.toml
model = "gpt-6.1-sol"

Inside a session, /model switches models, and choosing “More reasoning” afterwards exposes the Max and Ultra levels. Per the models page, Max gives the model more time on a single task, while Ultra uses subagents to parallelise work that splits cleanly. Both depend on your settings.

A same-task comparison

Pick a task you’ve already done by hand or have a known-good diff for: a bug with a failing test, a small refactor, a migration of one module. Then give each model its own worktree so the runs can’t contaminate each other.

#!/usr/bin/env bash
set -euo pipefail

TASK="Fix the failing test in tests/test_billing.py. Do not change the test."
base=$(git rev-parse HEAD)

for model in gpt-6.1-sol "$ASTRA_ID"; do
  dir="../cmp-$model"
  git worktree add -f --detach "$dir" "$base"
  ( cd "$dir" && time codex exec -m "$model" "$TASK" )
done

$ASTRA_ID is whatever identifier /model shows for Astra in your client. I haven’t hard-coded it because the changelog and models page only give the identifier for Sol.

Run your test suite in each worktree afterwards, then read both diffs. Three things to compare, in this order:

  1. Does the suite pass, and did the model touch anything it was told not to?
  2. How big is the diff? A fix that rewrites three modules is a worse fix than one that changes four lines, even if both pass.
  3. How long did it take, and how much of your plan’s usage did it burn?

Codex on a ChatGPT plan draws on plan usage limits, so the “cost” you can see is usage and wall-clock time, not a per-token bill. If you’re on API keys, the cost is on your invoice.

Where it will mislead you

One task is an anecdote. Two models can tie on a bug fix and diverge on a refactor that needs context from twelve files, so repeat the comparison on a second, harder task before you change a team default. Run each model more than once if the task is flaky, because a single lucky run proves little.

OpenAI’s guide pairs each model with a reasoning level: Sol at Medium for complex technical work, Sol at Extra high for polished deliverables and conflicting evidence. Treat those as starting points. Reasoning level is a second variable, so change one thing per run: Sol at Medium against Astra at Medium, then Sol at a higher level if the first result is close.

If Sol matches Astra on your tasks, make it your config.toml default and keep Astra for the work where it visibly pulls ahead. The guide says that if cost and latency aren’t a concern you can default to Astra, which is fair, but most teams running agents all day are concerned about both.