THE WORLD IS NOT STANDING STILL.RSS
BIG CHANGE.

Markdown edition

# Choose a GPT-6 model for a production API workflow

> A task-based guide to GPT-6 Astra, GPT-6.1 Sol and Luna, with current API rates and a worksheet for evaluating quality, latency and cost before deployment.

By BIG CHANGE Editorial

Published: 2026-10-09T20:51:30.540Z
Updated: 2026-10-09T20:51:30.540Z
Canonical: https://bigchange.ai/blog/choose-gpt-6-model-production-api-workflow

![Conceptual charcoal illustration of blank task cards entering three equal, distinct channels that converge at an empty circular checkpoint.](https://bigchange.ai/api/media/file/gpt6-task-routing-hero-v3.png)
AI-generated conceptual editorial illustration by BIG CHANGE.

OpenAI published a [GPT-6 family guide](https://openai.com/index/practical-guide-building-gpt-6/) on October 2, 2026. It brings model choice, instructions, long-running tasks and deployment checks together. A software team still needs to find the combination that passes its own task checks within its cost and response-time limits.

The current family choices in that guide are GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. We checked OpenAI's API documentation and prices on October 3. The selection method and worksheet below are proposals; we did not run the models or measure a production workload.

## The big change

- **What changed:** OpenAI now presents the GPT-6 family as a set of workload choices, with adjustable reasoning effort and tools for work that spans multiple steps. Teams can configure each part of a workflow instead of making one model setting carry every task.
- **Why it matters:** A builder can measure whether a focused extraction step needs Luna, whether a coding or research step warrants Sol, and where Astra's additional capability earns its price. The answer depends on completed tasks, latency and total workflow cost, including tool use and failed attempts.
- **What to watch:** Long runs need explicit handoffs and checks. Steering can update instructions during a run, while asynchronous tool results and delegated work still need to be reconciled before the final answer is accepted.

## Choose by task, then measure the whole workflow

Start with a representative set of real tasks and an acceptance rule for each. OpenAI recommends [the Responses API](https://developers.openai.com/api/docs/guides/deployment-checklist) for current model behavior, tool calling and stateful work. An API project, credentials, billing access, and a model available to that project are prerequisites. The model pages list no free-tier support; rate limits depend on usage tier. Check the account's actual access and limits before sizing a rollout.

| Work assigned to the model | Starting candidate | Starting effort | Promote or change when |
| --- | --- | --- | --- |
| Repeated extraction, classification or a structured summary with a clear answer | `gpt-6-luna` | Low for routine work; compare with its Medium default | Error rate or review time exceeds the team's threshold |
| Coding, research, tool use or a professional judgment step | `gpt-6.1-sol` | Medium default; test High for difficult cases | Representative tasks fail even with sound inputs and instructions |
| The hardest reasoning or review step where quality is decisive | `gpt-6-astra` | Compare Medium and High on the same cases | Keep only when the measured gain justifies its added cost and time |

The table turns OpenAI's [model guidance](https://openai.com/index/practical-guide-building-gpt-6/) and [model pages](https://developers.openai.com/api/docs/models/compare) into evaluation starting points. Check the API model IDs and supported effort settings: Astra and GPT-6.1 Sol support Low through Max; Luna also supports None. GPT-6.1 Sol excludes None and Minimal. OpenAI says to try Extra High or Max where supported after High falls short. Compare effort settings on the same task set because quality, duration and token use can change together.

For Standard processing and prompts up to 272,000 input tokens, the current text rates per million tokens are:

| Model | Input | Cached input | Cache write | Output |
| --- | --- | --- | --- | --- |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 |
| GPT-6.1 Sol | $2.00 | $0.10 | $2.50 | $10.00 |
| GPT-6 Astra | $10.00 | $1.00 | $12.50 | $50.00 |

Sources: the [Luna](https://developers.openai.com/api/docs/models/gpt-6-luna), [GPT-6.1 Sol](https://developers.openai.com/api/docs/models/gpt-6.1-sol) and [Astra](https://developers.openai.com/api/docs/models/gpt-6-astra) model pages. A request with more than 272,000 input tokens carries higher rates for the entire request. Other processing modes, regional processing where available, and some tools change the bill. All three pages list a 1,050,000-token context window and a 128,000-token maximum output; a large window is a capacity limit, not a reason to send every available document.

Estimate cost for the full task path: input, cached input, cache writes, output, tool charges, retries and any long-context uplift. Divide that total by tasks accepted under the same review rule. Then compare latency at the user-facing step and for the whole workflow. Unit prices alone cannot tell a team which route is cheapest per successful result.

## Specify the assignment and output before adding tools

Give each step a clear input, intended reader or downstream consumer, allowed sources and tools, constraints, and completion condition. OpenAI's guide also asks teams to state which decisions the model can make and which require a person's approval. Keep project instructions, skills and prompts consistent about those boundaries.

For machine-readable output, define the fields and valid values in advance, then use the [Structured Outputs guide](https://developers.openai.com/api/docs/guides/structured-outputs) where a schema suits the task. Treat a valid shape as one check; a field can conform to a schema and still be factually wrong. If the output is a human handoff, require the result, evidence used, checks performed and unresolved items. Review the result against the original inputs and the team's acceptance rule.

Put stable instructions and shared reference material before changing task details when evaluating [prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching). Reuse can reduce recurring input cost, but cache writes and later context must be included in the estimate. BIG CHANGE's [earlier prompt-cache guide](https://bigchange.ai/blog/gpt-6-prompt-caching-agent-context-costs) covers cache diagnostics in depth.

## Keep long-running work inspectable

The [October 2 guide](https://openai.com/index/practical-guide-building-gpt-6/) describes mid-turn steering, asynchronous tool calls and parallel subagents for independent work. A correction sent through the Responses WebSocket API is queued; it does not undo completed actions or stop a tool already running. An asynchronous tool lets unrelated work continue, but dependent work must wait for its result. Multi-agent support for GPT-6.1 Sol in the Responses API is currently beta.

For a multi-step run, persist the task ID, chosen model and effort, current stage, tool call and result IDs, approvals, and the evidence behind the final answer. Decide in advance what happens after a timeout, a failed tool call, a changed instruction or a duplicate result. When context grows, [compaction](https://developers.openai.com/api/docs/guides/compaction) can reduce what is carried forward; inspect what the continued run actually retains. OpenAI's [background mode](https://developers.openai.com/api/docs/guides/background) is another documented option for tasks that outlast a single request. Choose these controls to match the job's duration and recovery needs.

OpenAI's guide recommends a direct API or connected tool where it can perform the step, and screen interaction where it is necessary. BIG CHANGE's [Agents API browser-task guide](https://bigchange.ai/blog/openai-agents-api-computer-use-browser-guide) covers the computer-use interface and its supervision path.

## A worksheet for a decision the team can reproduce

Use the same cases and review rule for every candidate. This worksheet is a proposed evaluation method; no results have been entered or tested by BIG CHANGE.

| Record for each case and candidate | Entry to keep |
| --- | --- |
| Task and expected result | Real input ID, output requirements, allowed tools, acceptance rule |
| Configuration | API model ID, effort, processing mode, prompt version, schema or output contract |
| Outcome | Accepted, rejected or needs review; failure reason; reviewer |
| Time | End-to-end duration and duration at the user-facing step |
| Usage and charge | Input, cached input, cache-write and output tokens; tool fees; retries; long-context or regional uplift |
| Decision | Accepted tasks divided by attempted tasks; total charge divided by accepted tasks; unresolved failure types |

Include easy and difficult cases, malformed inputs and interrupted tool steps that the real workflow encounters. Keep the cases fixed while comparing models, then repeat after changing prompts or tool permissions. Review failures by type: missing evidence, wrong field, tool error, missed instruction or an answer that needs human correction. Move a step to a different model only after the same acceptance rule shows a useful gain. A cheaper model with more rework can cost more per accepted task; a slower model may be acceptable in a background stage but unsuitable for an interactive one.

Before release, check the project's actual rate and spend limits, data controls, timeouts, retry behavior, monitoring and human approval boundaries against OpenAI's [deployment checklist](https://developers.openai.com/api/docs/guides/deployment-checklist). Keep a sampled review path after release and re-run the worksheet when a model alias, prompt, tool or workload changes. Use the resulting task data to make the routing decision.

## Sources & further reading

- [OpenAI's October 2 GPT-6 family guide](https://openai.com/index/practical-guide-building-gpt-6/) sets out the vendor's model, effort, instruction and long-running workflow recommendations. It does not report BIG CHANGE test results or a best model for an individual team's workload.
- [GPT-6 Luna](https://developers.openai.com/api/docs/models/gpt-6-luna), [GPT-6.1 Sol](https://developers.openai.com/api/docs/models/gpt-6.1-sol) and [GPT-6 Astra](https://developers.openai.com/api/docs/models/gpt-6-astra) document the API IDs, supported effort, context, prices and tiered limits used here. Current account access and billing still need checking.
- [OpenAI's API deployment checklist](https://developers.openai.com/api/docs/guides/deployment-checklist) supports the recommendation to evaluate representative tasks, configure the Responses API and plan production controls. The worksheet above is BIG CHANGE's proposed method, not a vendor benchmark.
- [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs), [compaction](https://developers.openai.com/api/docs/guides/compaction) and [background mode](https://developers.openai.com/api/docs/guides/background) explain the specific interfaces referenced in the workflow.

## Sources

- [A model guide for the GPT-6 family](https://openai.com/index/practical-guide-building-gpt-6/) — Primary vendor guide for model roles, reasoning effort, instruction and output design, caching, steering, async calls and beta multi-agent support. Its recommendations do not establish local task success or cost.
- [GPT-6 Astra model documentation](https://developers.openai.com/api/docs/models/gpt-6-astra) — Official model ID, supported effort, context and maximum output, modality, Standard token rates, long-context price boundary and usage-tier limits. Rates can change.
- [GPT-6.1 Sol model documentation](https://developers.openai.com/api/docs/models/gpt-6.1-sol) — Confirms current 6.1 Sol ID, default and supported reasoning levels, Responses API tool calling, context, output, rates, regional and long-context conditions, and tiered limits.
- [GPT-6 Luna model documentation](https://developers.openai.com/api/docs/models/gpt-6-luna) — Confirms focused-task positioning, None through Max effort, Responses tool support, context/output limits, Standard rates and usage-tier limits.
- [API deployment checklist](https://developers.openai.com/api/docs/guides/deployment-checklist) — OpenAI's production recommendations for Responses, model and effort choice, representative evaluations, tool/context controls, reliability and monitoring. Does not validate the proposed worksheet.
- [Structured model outputs](https://developers.openai.com/api/docs/guides/structured-outputs) — Official output-schema interface. A conforming structure does not establish factual correctness; the article calls for separate review.
- [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) — Official cache behavior and stable-prefix guidance, referenced briefly. Prior BIG CHANGE #29 owns the detailed cache diagnostic angle.
- [Compaction](https://developers.openai.com/api/docs/guides/compaction) — Official long-conversation context reduction mechanism; the article advises checking retained state rather than assuming perfect preservation.
- [Background mode](https://developers.openai.com/api/docs/guides/background) — Documents a long-running Responses mode; used only to identify an available continuity option.