October 7, 2026
Heard in AI

OutSystems CEO says a harness and router cut its AI token bill

OutSystems CEO Woodson Martin says the company's AI token spending peaked in June and July, then fell below forecast after it built its own harness and model router. He argues most enterprise work doesn't need frontier models.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published October 7, 2026

Woodson Martin, chief executive of OutSystems, a company whose platform large organizations use to build business applications and AI agents, says his company's AI bill has already been through a full boom and correction. Speaking with host Nathan Labenz on The Cognitive Revolution, in an episode published October 7, 2026, Martin said OutSystems' token spending peaked in June and July and has been declining since. At the time of the conversation, he said, the company was spending less than it had forecast for the third quarter.

Tokens are the small chunks of text that language models read and write. AI providers charge by the token, so a company that sends a lot of work to models sees that work directly on its invoice.

OutSystems is one of the episode's sponsors, and much of what Martin described involves the company's own products.

A year of "use models for everything"

Martin said that at the same point a year earlier, OutSystems' "sole focus" was to "shift everything to AI, use models for everything." That push, he said, taught the company a great deal and sharply accelerated how quickly it shipped software. Prefacing the numbers with "I think," he said OutSystems shipped four major features in the fourth quarter of last year, 19 in the first quarter of this year and 26 in the second. These were not small tweaks but "major new capabilities," Martin said. The cost showed up "in the token burn, but worth it for the impact."

Labenz asked how the company's token spending had changed and said he often asks guests for the ratio of token spending to payroll. Martin did not give a ratio. He described how the spending had changed over time instead.

A better harness and a router

The turning point, Martin said, was realizing that most of what the company had been doing with AI "we could do just as well, a lot cheaper with a better harness." A harness is the software wrapped around a model that supplies its instructions and context, gives it tools and runs its working loop. OutSystems built one for its engineering organization that is, in Martin's words, "full of all of our own context" and tuned to how the company works.

The company also built a gateway, which Martin called an LLM router. The logic was that not every job needed to go to the most capable and expensive model of the moment. The router sends specific jobs to lower-cost models instead. Martin said the company's spending was below forecast "because of optimizations largely."

Amazon offers a commercial example of the general idea, though not OutSystems' gateway. Its Bedrock prompt routing uses a single endpoint that predicts how well each of two models in the same family will answer an incoming prompt. Customers configure a fallback model and the quality difference required before the router chooses the other model. AWS notes that the feature is optimized for English prompts and cannot learn from an application's own performance data, so specialized workloads may need their own testing.

"Really boring" workloads

Martin then applied the lesson beyond his own company. "I think the reality for most organizations today is that they don't need frontier models for their enterprise workloads, like almost none of enterprise workloads," he said. Frontier models are the newest and most capable systems from the leading AI labs. Martin separated two kinds of work: running the day-to-day operations of a business, and creating a new piece of software. For operations, he said, models "three years old" can handle most of the job.

"It's really boring what most of the enterprise workloads where AI can make a real advance happen today," Martin said. In his view, much of that work can be done more cheaply with deterministic code, meaning conventional software that produces the same output every time, without a model at all. Much of the rest can go to a cheaper model. That could be an open-weight model, whose weights a company can download and run itself, or an earlier version of a model from one of the big labs.

Martin said OutSystems sees this in its own work and among its customers. That is why, he argued, companies should build their agentic systems on a platform that makes it easy to swap models, either because a better model arrives or because another model can do the job more cheaply. He predicted "without question" that most enterprise workloads will run at a much lower cost per token than the most advanced models charge today. He added that people at the frontier are cutting prices too. One recent example is Anthropic's Claude Opus 5.5, announced September 22, 2026, which costs $4 per million input tokens and $20 per million output tokens. Anthropic claims typical workloads cost 40% less than on Opus 5, crediting both lower prices and fewer tokens per task.

The CFO's token bill

Martin argued that companies were not "on a sustainable trajectory in the first half of this year." He said that "every CFO got the token bill in February and March," and called it "an oh shit moment for the whole world."

Companies responded with controls of various kinds, Martin said. Some capped token spending per person. Some installed routers to send work to cheaper models. Others retreated from what he called an "AI first for everything everywhere strategy." He recalled big announcements from companies that "maybe went too far too fast," naming Microsoft, Meta and Uber with an "I think."

Martin does not see that retreat as permanent. Companies can get back to "really leveraging AI and everything," he said, but they need to do it "in a way that is smart," with platforms to help them. He said that is the job OutSystems does for its customers.

Open models, Chinese models and the customer mix

Labenz pushed for specifics. He asked what model mix OutSystems recommends to customers. He guessed that developers were still upgrading to Opus 5.5 right away for coding, while simpler jobs such as converting PDFs into structured data could clearly be done more cheaply. He also asked whether Martin recommends fine-tuning for such jobs, and how customers feel about Chinese models.

Martin said he wasn't sure he had "unique insight" on Chinese models, but some OutSystems customers are "perfectly happy to go there." They run their own versions of those models, distill them, or tune them further after training. Distillation means training a smaller model to imitate a larger one. Customers usually make these choices to save money, Martin said. Geography also plays a part. Half of OutSystems' customers are in Europe and many are in Asia, and Martin said attitudes differ around the world: "Not everybody is like super excited to trust their stuff to US companies." As a result, he said, customers use a wide variety of models in the agentic systems they build on the platform.

Martin did not directly recommend that customers fine-tune models for their own workloads. Instead he described the tuning OutSystems has done itself.

Tuning beneath Mentor

That tuning sits behind Mentor, OutSystems' AI service for building applications and agentic systems on its platform. According to OutSystems' product page, Mentor turns descriptions or requirements into editable designs and presents implementation plans before making complex changes. Martin said users can work with Mentor directly in OutSystems' development environment. They can also reach it through MCP services from "any coding agent or harness of your choice."

MCP, the Model Context Protocol, is a standard way for AI applications to connect to outside servers that provide tools and context, so one tool does not need a separate custom integration for every service, according to its official documentation. In a June 2, 2026 post, OutSystems' Luis Blando described Mentor services that let outside agents hand off the generation of data structures, screens and logic to the platform.

Behind those services, Martin said, OutSystems has "picked and fine-tuned a bunch of models" to run building tasks efficiently. As a result, a customer's own model, whether a frontier model or an older one, gets "a much lighter weight job" when it works through those services, because much of the optimization has already happened "under the covers."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Connected ideas and articles

From the conversation

Podcast episodes

The Cognitive Revolution

Software That Never Breaks: OutSystems CEO Woodson Martin on Building Enterprise-Grade Apps at ...

Episode published This article draws on 1:08–2:18 and 27:15–35:32 (approximate times)

Article history

Updates to this article

Tags