GPT-6.1 Sol has replaced GPT-6 Sol only seven days after the earlier model's release, according to an evaluation published by Artificial Analysis. The independent benchmarking firm reports that the update moved close to GPT-6 Astra on its composite intelligence measure while costing substantially less per completed benchmark task.
Artificial Analysis placed GPT-6.1 Sol one point behind GPT-6 Astra on its Intelligence Index. It said the new Sol model gained four points over GPT-6 Sol and five over GPT-5.6 Sol. The comparison is a result within the firm's own index rather than a universal measure of model quality, and scores can vary with benchmark design, model settings and workload.
The listed usage prices remain $2 per million input tokens and $10 per million output tokens, matching GPT-6 Sol. The cache-read discount increased from 90 percent to 95 percent, which Artificial Analysis said slightly reduces the blended cost for agent-style workloads that reuse cached context.
At maximum reasoning effort, the firm measured a cost of 72 cents per Intelligence Index task for GPT-6.1 Sol, compared with $3.26 for GPT-6 Astra and $1.05 for GPT-6 Sol. Artificial Analysis therefore placed all tested effort settings for the new model on its cost-efficiency frontier, meaning no cheaper model in its dataset reached the same intelligence score. Those figures represent the evaluator's standardized task mix, not a guarantee of costs in a particular production application.
The reported improvement was uneven across individual tests. Artificial Analysis recorded gains of four points on AA-Briefcase and five on GDPval-AA, both aimed at agentic knowledge work. It also reported a 12-point increase on Terminal-Bench 4.0, five points on Humanity's Last Exam and six on GDP.pdf. On the firm's AA-Omniscience test, accuracy rose eight points while the measured hallucination rate fell from 60 percent to 54 percent.
Higher output volume accompanied some of the gains. GPT-6.1 Sol used about 10 to 30 percent more output tokens than GPT-6 Sol across tested effort settings. Even so, the evaluator judged the low- and medium-effort configurations to be efficient in relation to their index scores. This distinction matters because fixed per-token prices do not by themselves determine the cost of completing a task.
Coding results followed a similar pattern. At maximum effort, the model gained three points over GPT-6 Sol on the Artificial Analysis Coding Agent Index and finished two points below GPT-6 Astra. The xhigh setting scored one point above Astra in that index at less than 15 percent of Astra's measured task cost, though it also outperformed the nominally higher max setting by three points.
The results indicate a rapid improvement in the Sol tier under this particular evaluation suite. Developers considering the update will still need to test their own prompts, latency requirements and error tolerance, especially because composite benchmarks can conceal large differences between tasks.



