Skip to content
Book demo
Back to Blog

The best value for every token: GPT-5.6 Sol is now our default in Cosmos

Jul 29, 2026
Siyu Zhan
Siyu Zhan
The best value for every token: GPT-5.6 Sol is now our default in Cosmos

In April, Matt wrote that the era of single-model engineering is over. Three models cleared the bar at the time, and he predicted there would be more by the end of the year - both better and cheaper.

It took just a few weeks for him to be proved right.

What just happened

On June 9, Anthropic shipped Claude Fable 5, a new tier above Opus. A week later, Z.ai released GLM-5.2 with open weights and a million-token context window. On June 30, Anthropic released Claude Sonnet 5, which became the default across many of its own products that same day. On July 8, Grok 4.5 followed, trained on real developer sessions from Cursor.

The next day, OpenAI made GPT-5.6 generally available in three tiers - Sol, Terra, and Luna. While most providers lead with price per million tokens, this family made its case on token efficiency per task. That is the number that ultimately determines your bill: the list price multiplied by every token the model spends reaching an outcome. A model that takes twice as many turns costs twice as much at the same list price. Sol is the most token-efficient model we have tested at its quality level.

To round out the eight weeks, we also got Kimi K3 from Moonshot on July 16, and Claude Opus 5 on July 23.

None of these releases made the default model you were already running any worse. They simply made the decision to keep running it look older than it did before. The most useful thing we can do for customers is keep that decision current, so nobody has to track all of this themselves. That means we owe you an explanation of how we make the decision!

Our approach, and why Sol

We want Cosmos to give you the best value for every token you spend. Our approach is to make the default the most token efficient model that is best suited for the real-world professional software engineering work completed in Cosmos.

Cosmos orchestrates agents through long-horizon tasks - a ticket to a reviewed PR, a migration across a repo, an incident investigation that starts from an alert and ends in a diff. These tasks involve dozens of steps. A model that loses the thread at step forty has not simply given you a bad answer; it has wasted half an hour and left you with a half-finished branch.

There is a lot of discussion around the price advantage of open weights models, and they are getting better really fast. However, we need to be careful to distinguish between token price and token efficiency. A cheap default model that fails often enough to require a second attempt has never truly been cheap. Retries cost tokens, and on a long-running task, they can cost the entire run, not just the step that went wrong. That is why simply choosing the cheapest model based on token price isn’t the best approach for choosing the default.

Instead, we hold the default to a pass-rate floor across our internal benchmarks and online testing. Among the models above that floor, we select the most token-efficient model, i.e. the one with the lowest cost per task.

Out of all the models released in the last eight weeks, GPT-5.6 Sol is the most token efficient model that clears this floor. Our own engineering team has validated this, testing all the new models on our own codebase. A recurring theme has been that Sol needs less steering on requests that weren't fully specified to begin with, and this was reflected in the cost per task.

We will be reviewing this

GPT 5.6 Sol is the default in Cosmos starting today. As a user, you’ll still have the option to pick any model you’d like. Most users leverage the Cosmos Advisor to help them pick the best model while setting up a loop.

Trying to choose the best model can feel overwhelming. You can trust us to change this default again, as new models continue to be released weekly. When new data suggests that there is an even more token efficient model that meets the bar of real-world professional software engineering work, we’ll be sure to update the default and let you know.

Getting more value from every token

Taking a step back, we are seeing another step change in how the most forward-leaning AI-teams are building software. Up until now, AI has predominately been changing how individuals work with individual agents. More recently, there is a new pattern emerging. Teams are building learning loops across the SDLC. With increased usage, token costs are no longer an afterthought.

At Augment, we want our users to get the greatest possible value from the tokens they spend. We do this in several ways:

  • We give users the optionality to adjust their default model for different use cases.
  • We introduced Prism, our model router, as we believed in combining the power of open source and frontier models in a cache efficient way to deliver the best outcomes.
  • We built the Cosmos Advisor, grounded in our knowledge base, to help users choose the right model for each task based on its complexity and requirements.

These are just a few of the ways we help our users get more value from every token. You can rest assured that we will continue investing in this area so that, whenever you build with Cosmos, you are using the best model for the task at the best possible cost.

Written by

Siyu Zhan

Siyu Zhan

Engineering Manager

Siyu Zhan is an Engineering Manager at Augment Code, where she leads efforts to automate the software development lifecycle with AI agents. With over a decade of experience building products at the intersection of complex systems and user experience, she brings deep expertise in scaling engineering teams and shipping products that solve real-world problems. Before joining Augment, Siyu spent nearly three years at Stripe leading teams that built Stripe Global Payouts and the Stripe Payments Dashboard, and served as Head of Engineering for Commercialization at Nuro, where she designed and built the entire product stack from the ground up as the company's first product engineer hire. Her career also includes backend engineering at Uber's UberPOOL team and financial software development at Bloomberg LP.

Get Started

Give your codebase the agents it deserves

Install Augment to get started. Works with codebases of any size, from side projects to enterprise monorepos.