
Breaking Down the Model: The Agent Harness
5 MIN. READ
Agent harnesses have attracted a lot of attention recently. This is in part because it's become clear that the product is no longer the model itself, but the system built around it. That system is the agent harness. A good harness lets a model do good work, but no model, however capable, can make up for a poor one.
In this article, we explain what an agent harness is, why it matters, and how Model ML’s agent harness is purpose-built for finance.
What is an agent harness?
On its own, a model such as GPT-5 or Claude Opus only takes text in and produces text out. The agent harness is everything built around the model that turns it into something able to take actions and complete tasks.
At its core, the harness defines a loop and handles every step within it. It builds the context, sends it to the model and reads the response. The model might ask to take an action, such as using a tool to create a PowerPoint deck. The harness is the code that carries out that action and tells the model what happened. It then starts the loop again, and repeats it until the work is done.
Why the agent harness matters, especially in finance
The harness makes the agent plan its work and stay on track. This is what separates reliable, client-ready work from an agent that wanders off halfway through a task. To support this, our harness gives the agent a file system, tools to read and write code, and sub-agents it can delegate to. It also gives the agent skills and memories, and has it ask clarifying questions instead of guessing.
The harness also curates what the model sees. It gives the model enough information to do the job, but not so much that it's overwhelmed. When this is done well, the same model can do better work with fewer tokens. When it's done badly, no model can make up for it.
In finance, the bar is higher still. Take a request to produce a 50-slide PowerPoint deck on the major players in the semiconductor industry. That one task involves thousands of steps: research, gathering data, building a narrative and constructing the slides. Every step has to follow a coherent plan, and that plan may change as the research comes in. The agent also has to work with PDFs, transcripts, Excel models, images, code and dashboards. The harness has to handle all of it.
Why use Model ML's agent harness for finance?
Three things set Model ML's agent harness apart: what shapes it, how it picks the right model for each task, and how it keeps costs down.
Built for finance, by finance professionals
Our harness is designed for finance. It's shaped by our clients and by the ex-bankers, investors and consultants who work at Model ML.
Benchmarked to use the right model for each task
The harness is informed by the Model ML Composite, the leading benchmark for AI in financial services. Model ML doesn't tie the platform to one provider. Instead, it routes every task to the model best suited to it, and the Composite is the evaluation behind that routing. It shows which models perform best on specific kinds of financial work, where a cheaper model is good enough, and when a new release actually changes the answer. We re-run it as models ship and prices change, so routing always reflects the latest models and prices.
Model agnostic: new models on day one, at lower cost
Because the harness is model agnostic, our users get new models on the day they're released. It also helps keep costs down. We often find that firms working with a single provider run its largest, most expensive model for every query, even when a model at a tenth of the price would do the job perfectly well. Our harness isn't tied to any one provider, so it can choose the best model for each task based on accuracy, latency and cost. Our entire stack is designed around the customer.
As open-source models improve, fine-tuning and self-hosting will become a growing part of our workload. By the end of the year, more than 50% of agent workloads at Model ML will run on our own models and infrastructure. Running on our own GPUs means we pay only for inference, which can be 95% cheaper than renting tokens from the big labs. That makes cost an area where we can be the best, not just something we manage. The agent harness is the only reason we can make this shift without users noticing anything but lower costs.



