Best balance
☀ Sol medium
73.01% · $0.4196/task
Sol, Luna & Astra · coding
Best balance
73.01% · $0.4196/task
Lower cost
66.59% · $0.2169/task
Highest score
75.22% · $0.6461/task
Astra max leads AutomationBench (41.4%). Sol high leads this coding snapshot at lower cost than every Astra setting.
Cheapest within 3 score points of the best for balance; within 10 for lower cost. Custom thresholds, not guarantees.
The coding range starts 10 percentage points below the highest score and ends at the cheapest setting with that highest score. Rings and bold labels mark the three picks; other coding points are gently faded. This makes unnecessary spending visible without hiding other results.
Low latency chooses the shortest reported total response among settings within 10 index points of the best Artificial Analysis Intelligence Index score. This is a separate broad quality test. It does not measure coding-task completion. Missing speed values are excluded.
Frontier means no other configuration in the same dataset is both cheaper and at least as high-scoring. This does not account for speed or statistical uncertainty.
Enter your token counts. Include reasoning tokens; they are billed as output.
Current official API prices, checked October 1, 2026. USD per million tokens; standard short context:
| Model | Input | Cache hit | Cache write | Output |
|---|---|---|---|---|
| Sol | $2.00 | $0.10 | $2.50 | $10.00 |
| Luna | $0.10 | $0.01 | $0.125 | $0.50 |
| Astra | $10.00 | $1.00 | $12.50 | $50.00 |
Cost = uncached input × input rate + cache hits × cached rate + cache writes × write rate + (answer + reasoning) × output rate, divided by one million. Cache categories partition total input, so nothing is counted twice. Above 272k input tokens, input/cache rates double and output rates multiply by 1.5 for the full request. Batch and Flex halve rates; Fast doubles them. Ultrafast is listed for Astra only, at six times standard rates. Unsupported estimates are omitted. These are API estimates, not Codex subscription charges.
Use task cost with an acceptable score. Token price alone misses how many tokens a model uses. For your own coding work, measure spend and elapsed time per accepted task, including retries and review.
Speed views use Artificial Analysis: tokens/second, first-chunk latency, or reported total response time. The leaderboard does not specify answer length for its total column. These measure API responses, not complete coding tasks. Different units cannot be averaged.
The shared-task average equally weights three benchmark families. It is a custom view, not an official coding score. Small differences may be noise; the release points do not supply confidence intervals.
No comparable ultra points are published here. Luna ultra is unsupported; Astra API supports low through max. Blank speed values are unpublished.
Saved launch results. Effort, cost, and score retained.
Independent quality, cost, and speed measurements. Saved October 1; individual measurement dates unavailable.
Community observations of resets, reasoning budgets, and capability.
Check benchmark versions and evaluation setups before comparing results.
Saved data, reviewed October 1, 2026. Reloading does not refresh it. Luna's launch computer-use results predate OpenAI's September 25 image-encoding fix.
Fast, source-reviewed updates after model releases. Then recommendations tested on real tasks.
Available now: free data API · Workflow planner · Roadmap