Planning proxy · GPT and Claude models
Usage estimator
Compare a simple, editable cost proxy across GPT-5.6, GPT-6, GPT-6.1 Sol, and Claude Opus 5.5 models at different reasoning levels.
Compare performance
Select a row below or use Find, then choose a target model. The closest match applies only to the chosen benchmark.
Inputs & assumptions
Values are saved in this browser on edit. Prices are fixed API price assumptions shown in the table.
How the estimate works
The estimated cost is calculated for the assumed token mix. Output is the visible output token assumption, adjusted by the selected reasoning multiplier.
estimated cost = (input tokens ÷ 1,000,000 × input price) + (output tokens ÷ 1,000,000 × output price × reasoning multiplier)
Combined performance is the equal-weighted average of the available AA Coding Index, FrontierCode, and DeepSWE scores. Missing scores are excluded. The cost-per-point value uses the estimated request cost, so it changes with the editable token and reasoning assumptions.
combined performance = sum of available benchmark scores ÷ number of available benchmarks
cost per point = estimated cost ÷ combined performance
The table displays cost per point multiplied by 1,000 and rounded to three decimals. Sorting and winner selection use the unscaled value.
Lower cost-per-point values are more cost-effective. Equal raw weighting is a transparent heuristic; benchmark coverage varies by row, and the scores measure different things. This is not a validated universal performance index.
By default, every score is normalized against the cheapest of the 40 configurations, so the baseline is 1.0×. Select a row to make it the 1.0× reference and highlight the closest lower and higher combinations for the other models. Use Compare performance to select a target model and match its reasoning levels on one benchmark instead. Matching uses the smallest absolute score difference, excluding unavailable scores and incompatible AA Coding Index benchmark versions. This is an illustrative API-price proxy, not a promise about product behavior.
Default assumptions
Defaults use 900 assumed input tokens, 1,200 visible output tokens, and illustrative effort multipliers of 1.0×, 1.4×, 2.0×, 2.8×, and 4.0× from low through max. Edit them in Inputs & assumptions for your own scenario.
Standard API prices per 1M input / output tokens are fixed here as: GPT-5.6 Luna $0.20 / $1.20, Terra $2 / $12, Sol $4 / $20; Claude Opus 5.5 $4 / $20; GPT-6 Luna $0.10 / $0.50, Sol $2 / $10, and Astra $10 / $50; GPT-6.1 Sol $2 / $10.
Sources
Standard API token prices come from the OpenAI model catalog, including GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna, and Anthropic's Claude Opus 5.5 model pricing. Cached input, cache writes, Batch, Flex, Fast mode, tools, and long-context surcharges are not included.
AA Coding Index scores are manually calculated as the equal-weighted average of Terminal-Bench and SciCode results, rounded to the nearest integer. Existing ratings were captured September 7, 2026 and use Terminal-Bench v2.1. GPT-6 Sol and Luna ratings were captured September 27, 2026 using Terminal-Bench 4.0 and SciCode from the GPT-6 Sol and GPT-6 Luna release pages. Because the Terminal-Bench version changed, these GPT-6 ratings are not directly comparable to the older ratings. A dash means the paired component results could not be verified for that model and version; Claude Opus 5.5 remains blank because a matching Terminal-Bench v2.1 and SciCode pair is not published. GPT-6.1 Sol remains blank because its paired component scores have not been verified. Artificial Analysis’s “xhigh” setting uses the same “xhigh” label here.
FrontierCode scores are manually maintained snapshots of the Cognition FrontierCode 1.1 Main leaderboard (100 tasks). Existing GPT scores were captured September 22, 2026; Claude Opus 5.5's five effort-level scores were captured September 24, 2026 from “All reasoning levels” using the Claude Code harness. GPT-6.1 Sol’s five effort-level scores were captured October 3, 2026 from “All reasoning levels” using the Codex harness. The displayed percentage is Cognition's weighted rubric score, not its pass rate; runs flagged for unfair internet use receive zero.
Existing DeepSWE Pass@1 ratings are manually maintained snapshots from the DeepSWE v1.1 leaderboard (113 tasks), captured September 3, 2026 from its “All effort levels” view. The GPT-6 Sol and Luna max scores are OpenAI-reported results published September 22, 2026; their other effort scores were not listed on the public leaderboard when checked. GPT-6.1 Sol’s five effort-level scores are OpenAI-reported results from its interactive DeepSWE chart, published September 29, 2026 and checked October 3, 2026; the model was not listed on the public leaderboard when checked. Claude Opus 5.5's max score is Anthropic-reported at 74.2%, the average of five max-effort trials, published September 22, 2026; it was not listed on the public leaderboard when checked September 24. The leaderboard's “Claude Opus 5” effort rows are for the earlier model, not Opus 5.5, so no additional Opus 5.5 effort scores are available there. This dashboard uses Pass@1 scores only, excluding cost, output-token, and agent-step metrics.