Results
Score by model and effort
Each dot is one run. Shades are reasoning effort, from light (min) to full colour (max).
Score against cost
Mean score per model and effort against mean inference cost per run. Hover a line to pick out a model.
Ascension 20
GPT-6 Astra at max effort on the hardest difficulty, next to its Ascension 0 runs.