Planning proxy · GPT-5.6

Usage estimator

Compare a simple, editable cost proxy across Luna, Terra, and Sol at different reasoning levels.

Estimated relative usage

Heuristic GPT-5.6 model and reasoning usage comparison
Inputs & assumptions
tokens
tokens
Reasoning multipliers editable assumptions

Values are saved in this browser on edit. Prices are fixed API price assumptions shown in the table.

How the estimate works

The estimated cost is calculated for the assumed token mix. Output is the visible output token assumption, adjusted by the selected reasoning multiplier.

estimated cost = (input tokens ÷ 1,000,000 × input price) + (output tokens ÷ 1,000,000 × output price × reasoning multiplier)

Combined performance is the equal-weighted average of the two displayed benchmark scores. The cost-per-point value uses the estimated request cost, so it changes with the editable token and reasoning assumptions.

combined performance = (AA Coding Index + DeepSWE Pass@1) ÷ 2 cost per point = estimated cost ÷ combined performance

The table displays cost per point multiplied by 1,000 and rounded to one decimal for readability. Sorting and winner selection use the unscaled value.

Lower cost-per-point values are more cost-effective. Equal raw weighting is a transparent heuristic; the benchmark scores measure different things and are not a validated universal performance index.

By default, every score is normalized against the cheapest of the 15 configurations, so the baseline is 1.0×. Select a row to make it the 1.0× reference and highlight the closest lower and higher combinations for the other models. This is an illustrative API-price proxy, not a promise about product behavior.

Default assumptions

Defaults use 900 assumed input tokens, 1,200 visible output tokens, and illustrative effort multipliers of 1.0×, 1.4×, 2.0×, 2.8×, and 4.0× from low through max. Edit them in Inputs & assumptions for your own scenario.

Model prices are fixed here as: Luna $1 / $6, Terra $2.50 / $15, and Sol $5 / $30 per 1M input / output tokens.

Sources

Coding Index scores are a manually maintained snapshot from the displayed model chart on Artificial Analysis Coding Index, captured August 3, 2026. The chart’s headline score is an equal-weighted average of Terminal-Bench v2.1 and SciCode; this dashboard copies those displayed scores rather than recalculating them. The chart’s “xhigh” setting is shown with the same “xhigh” label here.

DeepSWE Pass@1 ratings are a manually maintained snapshot from the DeepSWE leaderboard, version 1.1, covering 113 tasks and updated July 25, 2026. The values were transcribed from the leaderboard's “All effort levels” view and are shown as the published percentage. DeepSWE's “xhigh” setting uses the “xhigh” reasoning label here. This dashboard includes the Pass@1 scores only; the leaderboard's uncertainty, average cost, output-token, and agent-step metrics are not included.