OpenAI Features Model ML After Testing of GPT-5.6 Pre-Release

5 MIN. READ

OpenAI has published a paper on its website detailing its work with Model ML ahead of the release of GPT-5.6.

Before launch, GPT-5.6 was evaluated on the Model ML Composite, our evaluation benchmark for AI in Financial Services.

The Composite tests models on hundreds of real-world finance tasks created by finance domain experts, across five categories: PowerPoint creation, Excel creation, analytical finance, financial workflows and document retrieval. Unlike most finance benchmarks, which focus on tasks such as filings review, it also assesses the quality of final deliverables, such as pitch books and financial models.

Every model runs inside the Model ML agent harness with the same tools, retry behaviour and context handling, so differences in results come down to the model itself. Each task is then graded in two ways: deterministic checks, such as whether an LBO's IRR falls within a set tolerance or its sources and uses table balances, and LLM judges scoring each criterion against its own rubric.

GPT-5.6 performed strongly across the Composite. Three results stood out:

  • Deck efficiency: 21% fewer tokens per PowerPoint deck than Fable 5, with equivalent performance.

  • Review-ready output: a 16.6 percentage-point lead over Opus 5 on decks ready for substantive review.

  • Workbook efficiency: 36% fewer tokens per Excel workbook than Opus 5.

There was also a change in how the model approaches a task. Given a brief for an investment committee memo and pointed to the relevant sources, GPT-5.6 first plans the content and layout of each slide. It then carries out the research and calculations, builds native charts and tables, and finishes by reviewing each slide visually. The result is an editable PowerPoint deck with traceable sources.

Earlier models could complete analyst-level tasks, but usually only when users broke the work into detailed steps and specified the format of the output. GPT-5.6 gets considerably closer to a finished deliverable on its own. By using fewer tokens, each piece of work is also faster and more cost-efficient to produce. More of its decks were also ready for review on the first pass. Professionals can therefore spend their time making judgment calls rather than cleaning up the output.

Read OpenAI's article on how they worked with Model ML here.

The A to Qs 1-4

The A to Qs 1-4

The A to Qs 1-4

New York

West 38th St,
New York

San Francisco

Market St,
San Francisco

London

King's Cross,
London

Hong Kong

Stanley St Central,
Hong Kong

© 2026 Model ML. All rights reserved.

New York

West 38th St,
New York

San Francisco

Market St,
San Francisco

London

King's Cross,
London

Hong Kong

Stanley St Central,
Hong Kong

© 2026 Model ML. All rights reserved.