GPT-6 Astra on the Model ML Composite

5 MIN. READ

GPT-6 Astra launched yesterday. Ahead of its release, we tested it against the Model ML Composite, our benchmark for AI in financial services.

The Composite measures how frontier models perform across the financial workflows our clients use the most. It is the evaluation behind our routing: it tells us which model to send a given piece of work to, where a cheaper one is sufficient, and whether a new release changes the answer.

This report covers how GPT-6 Astra performed across five categories, what it costs, and which tasks are worth routing to it.

Overview

We tested GPT-6 Astra and other closed-source frontier models across five of the Composite’s key areas of evaluation: analytical finance, financial workflows, single-document intelligence, multi-document intelligence, and PowerPoint creation.

Its main improvements from GPT-5.6 Sol were seen in the following areas:

  • Financial calculations. Higher accuracy on analytical finance and numeric workflow tasks, with more explicit component-level working that makes definitions and results easier to check.

  • Document retrieval. Recovers relevant disclosures that GPT-5.6 Sol overlooks, including information outside the main financial tables, and works them into the answer.

  • Presentation quality. Higher visual-quality scores and more complete slide compositions, with lower token use and faster median completion.

Key takeaways

When tested alongside other closed-source frontier models, GPT-6 Astra achieved the highest score in four of the five categories we tested. In three of them, it also revealed to be cheaper than the next-best model.

Against Fable 5.1 the case is straightforward: it received higher scores than it in all five categories, and a lower cost in four. However, the picture shifts against the cheaper end of the frontier across a series of categories.

On analytical finance, Gemini 3.8 Flash scores within a point of GPT-6 Astra at 86% less. On financial workflows it scores higher, at 72% less. Where GPT-6 Astra earns its premium is on the harder tasks, and the gap widens as the work gets more complex.

Here's what we found across each category:

  • On analytical finance, GPT-6 Astra, Fable 5.1 and Gemini 3.8 Flash are effectively on par, but GPT-6 Astra is the most expensive. GPT-6 Astra scores 93.21 at an estimated $6.92 per task, Fable 5.1 scores 92.54 at $5.23, and Gemini 3.8 Flash scores 92.39 at $0.97, or 86% less than GPT-6 Astra.

  • Financial workflows show GPT-6 Astra’s largest score gain over Fable 5.1, but Gemini leads overall. GPT-6 Astra scores 12.36 points higher than Fable 5.1 at $1.61 versus $1.82 per task. Gemini 3.8 Flash achieves the highest score, 76.33, at $0.45 per task. That’s 3.97 points above GPT-6 at 72% lower cost.

  • GPT-6 Astra leads on single-document retrieval, but Gemini 3.8 Flash offers an equally accurate, cost-efficient alternative. GPT-6 Astra scores 86.30 at $0.68 per task, while Gemini reaches 81.48 for roughly a third of that, with GPT-5.6 Sol in between on 82.96. Opus 5 and Fable 5.1 are beaten on both score and cost.

  • On multi-document retrieval, GPT-6 Astra is marginally more accurate than GPT-5.6 Sol at double the cost. Astra scores 78.71 at $2.11 per task against Sol's 76.77 at $0.97. Gemini 3.8 Flash is cheaper still at $0.45 but drops to 63.23, a margin wide enough to disqualify it as a substitute.

  • In PowerPoint creation, Astra offers comparable visual quality to Fable 5.1 at a substantially lower cost. GPT-6 Astra scores 80.88 against Fable 5.1’s 80.79, at $4.65 versus $12.71 per task; a 63% lower cost. Its displayed score also exceeds GPT-5.6 Sol’s 72.25, though Sol costs less at $2.01 per task.


PowerPoint creation: GPT-6 Astra vs GPT-5.6 Sol

We saw meaningful improvement across classic investment banking analysis from building AVPs to buyer profiles. Even more, GPT-6 Astra's clearest advantage is on open-ended, client-facing design work, where the task and prompts are less prescriptive and the model determines the composition itself. It demonstrates greater range on these tasks, and applies it where allowed.

Slide building was a notable improvement over GPT-5.6 Sol. GPT-6 Astra scores higher on every quality component while finishing the work in two-thirds of the time.

Measure

GPT-6 Astra

GPT-5.6 Sol

Visual quality

76.83

67.25–72.25

Instruction following

69.14

61.93

Banker writing / paired slides

86.30

80.00

Deck completion rate

95%

95%

Median elapsed time

289.8 sec

440.7 sec


Example 1: Building a product portfolio slide

When asked to build a product portfolio slide laying out five product families with revenue mix and growth, GPT-6 Astra outperformed GPT-5.6 Sol in both visual structure and level of detail.

GPT-6 Astra

On the other hand, GPT-6 Astra placed representative imagery beside each product family, separated the revenue and growth figures into clean aligned columns, and grouped the families into core home audio and expansion avenues, making the breadth and structure of the portfolio visible at a glance.

GPT-5.6 Sol

GPT-5.6 Sol pictured only one of the five product families and left the rest as text in a dense table, with the title and subtitle sitting close together.


Example 2: Building a product feature overview slide

When asked to build a product feature overview for the Rivian R1S, GPT-6 Astra built better slides than GPT-5.6 Sol in terms of layout and presentation of information.

GPT-6 Astra

On the other hand, GPT-6 Astra paired a vehicle image with concise commercial positioning and five clearly separated feature-to-benefit rows, each tied to a customer benefit, with announced but unreleased features flagged separately at the bottom.

GPT-5.6 Sol

GPT-5.6 Sol included substantial useful detail, but the subtitle crowded the title, the configuration and capability tables competed for the same space, and the commercial-relevance callout left almost no padding below its final line.


Conclusion

GPT-6 Astra presents a substantive improvement from GPT-5.6 Sol. Overall, it posts the highest score in four of the five categories tested here, and in three of them it is also cheaper than the next-best model. Financial workflows is the exception: Gemini 3.8 Flash scores 3.97 points higher at 72% less. On analytical finance, GPT-6 leads by less than a point over both Fable 5.1 and Gemini while costing seven times what Gemini does. The premium is clearest on PowerPoint creation, where GPT-6 matches Fable 5.1's visual quality at a third of the price, and demonstrates a greater range where the tasks and prompts allow.

This is why Model ML stays model-agnostic. GPT-6 raises the top score, but on analytical work a model costing a fraction as much provides an alternative. Task-level routing sends work to GPT-6 where its advantage holds, and to a cheaper model wherever one performs the task equally well.

The A to Qs 1-4

The A to Qs 1-4

The A to Qs 1-4

New York

West 38th St,
New York

San Francisco

Market St,
San Francisco

London

King's Cross,
London

Hong Kong

Stanley St Central,
Hong Kong

© 2026 Model ML. All rights reserved.

New York

West 38th St,
New York

San Francisco

Market St,
San Francisco

London

King's Cross,
London

Hong Kong

Stanley St Central,
Hong Kong

© 2026 Model ML. All rights reserved.