GPT-6.1 Sol Makes the Case for Measuring Cost per Accepted Task
A practical framework for evaluating GPT-6.1 Sol: include retries, review, caching and failure costs before switching production workloads.
On this page
A model upgrade can lower the invoice while making the product worse. It can also charge the same token price and make the product dramatically cheaper. The missing unit in both conversations is the completed, accepted piece of work.
What actually changed in the September release?
In its September 29 announcement, OpenAI positioned GPT-6.1 Sol near GPT-6 Astra on several evaluated tasks, at lower prices. The published standard API rates were $2 per million input tokens, $10 per million output tokens and $0.10 per million cached input tokens. OpenAI also explicitly cautioned that its research evaluation environment can differ from production product behavior. These are vendor-reported results and launch prices, checked October 6; Crashtech has not independently benchmarked the model.
The useful question is what this changes in your own operating model. If an application performs a fixed number of calls, tokens dominate the conversation. If an agent chooses its own tools, retries and stopping point, the price of a token explains much less. A superficially successful answer can also transfer expensive work to the human who must discover and repair its mistakes.
What should the denominator be?
Use accepted outcomes rather than generated responses. For a support assistant, an accepted outcome might be a correctly resolved ticket without a later reopening. For a coding assistant, it might be a change that passes the relevant checks and survives review. For document extraction, it might be a record whose required fields match the original evidence, with uncertainty surfaced rather than guessed away.
Define acceptance before running the comparison. Otherwise a team can unconsciously lower its quality bar to justify a cheaper model. A response that sounds convincing should not count as complete when the actual task required a database update, a working file or a correct calculation. Keep a separate count for safe refusals and escalations; they may be the right outcome for cases the system cannot responsibly finish.
Cost per accepted task = total evaluation-run cost divided by accepted tasks. Include the expense of failed attempts in the numerator. Also publish the acceptance rate beside this metric, because a system can appear inexpensive by abandoning the difficult cases that matter most to customers.
How can the arithmetic reverse a purchasing decision?
Consider an illustrative experiment, not a measurement of any named model. Model A costs $20 to attempt 100 tasks and yields 80 accepted results. Its model-only cost per accepted task is $0.25. Model B costs $30 and yields 95 accepted results, or about $0.32 each. A token-focused comparison favors A.
Now suppose the batch using A needs two hours of correction and the batch using B needs half an hour. At an assumed internal review cost of $40 per hour, their combined costs become $100 and $50. Dividing by accepted tasks gives $1.25 and about $0.53 respectively. The decision reverses because review dominates inference. The reviewer rate and correction time are assumptions; replace both with measurements from your own workflow.
The opposite reversal is possible too. An expensive model might produce prettier writing without reducing correction effort or improving acceptance. Paying for that capability can still be sensible for a premium editorial product, but it should be an intentional product decision rather than an unexplained infrastructure expense.
Where does caching belong in the evaluation?
Treat caching as an observed operating condition. A large reusable instruction prefix can make repeated workloads cheaper, but an evaluation consisting entirely of warm requests can overstate savings for a product with many first-time users. Record billed cached and uncached usage rather than assuming every repeated token receives the discounted rate.
Run at least two representative conditions: the first request in a workflow and a later request with reusable context. Preserve authorization boundaries when constructing context; a cache optimization must never reuse one customer’s private material for another customer. A small prompt change can alter the economics, so retain the prompt version beside the usage record.
The same discipline applies to output length. Asking for shorter answers may reduce cost while removing evidence that a reviewer needs. Measure whether the shorter format preserves the acceptance criteria. Compress repetition, not the proof that makes an answer trustworthy.
What does a useful migration experiment look like?
Choose historical tasks that resemble current traffic, remove sensitive data appropriately, and retain the original expected outcomes. Include messy documents, incomplete requests and tool failures. A collection of clean demonstration prompts will tell you whether the model can perform well, but not how often your product will need to recover when conditions are bad.
Keep the tools, permissions, timeout budget and available documents equivalent across candidates. Run repeated trials for tasks where outputs vary. Reviewers should see the result and evidence without being told which candidate produced them, where practical. Record acceptance, completion time, human correction, tool errors and any unauthorized action separately.
Start with a bounded traffic slice whose effects are reversible. Maintain an explicit rollback route and compare the live sample with the offline evaluation. If quality falls only for a particular document type or language, fix that slice rather than hiding it inside an overall average. An aggregate improvement is not permission to neglect a critical minority of requests.
When should a more capable model remain available?
Keep an escalation path for tasks with ambiguous requirements, difficult evidence reconciliation or unusually costly mistakes. Escalation should have a reason that can be audited: missing evidence, repeated tool failure, unresolved contradictions or a failed acceptance check. Do not use the model’s verbal confidence as the only routing signal.
A practical routing record contains the task category, first model, effort setting, escalation reason and final acceptance result. Review it regularly. If nearly every case escalates, the cheap first pass is probably adding latency and cost. If almost nothing escalates despite frequent reviewer corrections, your trigger is probably too permissive.
For teams building retrieval workflows, our RAG versus agentic RAG explainer covers a related decision: when additional reasoning earns its extra work. The release names will keep changing. A stable definition of accepted work is what lets an engineering team benefit from those changes without rebuilding its judgment every month.
Frequently asked questions
How should a team evaluate GPT-6.1 Sol before migration?
Evaluate representative tasks with fixed acceptance criteria and comparable tools, permissions and time budgets. Record success rate, correction time, latency and total billed usage, including retries. Begin with a reversible production sample after offline testing, and keep an escalation or rollback path for workloads that regress.
What does cost per accepted AI task mean?
It is the total expense of attempting a workload divided by the number of results that meet predefined quality requirements. Include unsuccessful attempts and relevant tool or review costs. Publish acceptance rate alongside it so that abandoning difficult requests cannot make an otherwise weak system look artificially economical.
Should prompt caching determine which AI model we choose?
Caching is one part of the operating cost, not the whole selection decision. Measure actual cached usage under both first-request and repeated-request conditions. Then compare quality, correction effort and completion time. Preserve tenant isolation and check current provider terms before making projections from a cached-token price.
/* Comments */
Comments are offline right now — we reconnect automatically, nothing is lost.