As Google, OpenAI and Meta Push A.I. Agents, Enterprises Need Models That Think Less
Observer ·

The next breakthrough in enterprise A.I. may not come from models that reason more deeply, but from systems that know when to rely on established rules instead. Campfire’s John Glasgow argues that specialized models and deterministic systems can offer greater consistency at a fraction of the cost.
The race to make A.I. more capable is entering a new phase. Google’s announcement yesterday (Oct. 8) of a “universal” Gemini agent that can run in the background of apps and devices marks the latest step in a broader push to make A.I. act on users’ behalf. OpenAI introduced its always-on agents, called dots, in September, just weeks after Meta released Muse , a personal agent capable of shopping, booking travel, sending emails and making payments. The ambition has evolved from giving users a better way to generate text or analyze simple information to build systems that can do the work themselves, across both enterprise workflows and personal life.
For the past several years, A.I. has been organized around a fairly simple assumption: smarter models will produce better software. Every new generation promises stronger reasoning, longer context windows and better performance across a growing list of benchmarks. Companies, understandably, have treated those improvements as a roadmap: when a better model arrives, it’s time to upgrade.
But as A.I. moves from generating answers to executing tasks, enterprise software is beginning to expose a flaw in that logic. In many of the most valuable business workflows, the goal is not to give an A.I. model more freedom to reason. It is to give it less.
We encountered this problem building A.I. for accounting at Campfire . Accounting is an unusually unforgiving domain for probabilistic answers. A model can generate five different versions of a good sales email, and all five might be useful. But if it is categorizing the same Uber transaction 500 times, there is generally only one correct answer the company wants all 500 times. A close approximation is not close enough when the result ultimately feeds a general ledger, financial statement or audit.
This doesn’t just apply to accounting. The first wave of generative A.I. was built around tasks where variation is often a feature: writing, summarizing, brainstorming and coding. The next wave of enterprise adoption is moving into workflows where variation becomes a liability. Companies want A.I. to reconcile records, classify transactions, enforce policies, update systems and execute processes in which the answer should be reproducible and explainable. That requires a different architecture and a different way of thinking about what makes an A.I. system good.
At Campfire, one of the most important changes we made was to stop asking the model to directly “do the accounting.” Instead, we increasingly ask the model to write code that does the accounting. Rather than handing financial data to a language model and asking it to reason its way toward a result each time, the model can generate SQL or Python that applies a defined set of rules to the underlying data.
It sounds like a technical distinction, but it changes the nature of the output. A language model is probabilistic: given the same problem more than once, it can take a slightly different path or produce a different answer. Code, by contrast, is deterministic. Once the system has generated the correct query or program, that code can run repeatedly against the same inputs and produce the same result.
That difference matters most when there is one right answer, as there often is with financial data. Businesses already have well-established controls for software that behaves predictably. When a process cannot reliably be reproduced, companies need additional review and approval around it. An A.I. system that generates a deterministic process can fit more naturally into those existing controls than one that continually generates new answers.
That does not eliminate the need to check the model’s work. Generated code can contain errors, and the underlying data or business rules can be wrong. The advantage is that the resulting process can be tested, validated and rerun against defined criteria instead of relying on the model to independently reason through every transaction. Reliability, in other words, does not necessarily come from teaching the model to reason more accurately about every domain. Sometimes it comes from narrowing what we ask the model to do.
Finance provides an obvious example because numerical errors are so visible, but the principle applies to much of the back office. Imagine an A.I. system reviewing whether an expense complies with corporate policy, matching a payment to an invoice, determining which customer record should be updated or checking whether a contract contains a required provision. These are not primarily creative exercises. The enterprise value comes from getting a defined operation right, repeatedly, and maintaining an audit trail that explains what happened.
The same principle applies as more companies give agents permission to act across their software systems. An agent that can update a record, initiate a payment or change a customer account needs boundaries, predictable execution and a way to verify that the action was correct.
This shift has also changed the way we think about model selection. The prevailing assumption in A.I. software has been that the newest and largest model is probably the best default. Our experience has been more complicated. In live customer workloads, we have seen newer models perform worse than their predecessors on tasks we care about, even when their overall capabilities are theoretically greater. In one case, that deterioration was significant enough that we moved back to an earlier model.
At the same time, we post-trained a smaller open-source model specifically for our domain and found that it slightly outperforms leading commercial models on accuracy while costing roughly 2 percent as much to operate. That does not reflect a universal rule about smaller models, nor does it mean that frontier model development has stopped mattering. But it illustrates how the enterprise A.I. market is starting to mature beyond a single hierarchy of intelligence, in which every application should simply rent the smartest model available.
The A.I. industry is investing heavily in models that can tackle increasingly complex problems. For many businesses, however, the key metric is whether a model can perform a specific task accurately, consistently and affordably under real operating conditions.
A general-purpose benchmark measures how well a model performs across a broad collection of tasks. An enterprise application has a much narrower question: How reliably can this system perform my task, under my constraints, at a cost that makes the economics work? Those are very different standards.
As companies begin deploying A.I. meaningfully at scale, the economics will make this distinction harder to ignore. A model that is marginally more capable can still be the wrong choice if an application invokes it millions of times to perform a tightly bounded operation. At that volume, inference costs matter. So do latency, consistency and the engineering work required to compensate for unpredictable outputs.
There is also an opportunity cost. If a smaller, specialized model can handle routine tasks reliably, reserving a more capable model for complex exceptions may deliver better overall performance across the system. The goal, rather than using the least expensive model everywhere, is to match the level of intelligence to the work.
The result will likely be a more mixed A.I. stack than the industry’s focus on frontier models might suggest. Large, frontier models will remain extremely valuable for ambiguous, complex problems that genuinely benefit from broad reasoning. Smaller or domain-specific models can handle narrower tasks where speed, accuracy and operating costs matter. Deterministic software will do work that should never have been delegated to probabilistic reasoning in the first place. The most effective enterprise applications will combine all three.
This is also why the next chapter of A.I. adoption may be less visually impressive than the first one. Generating an image or holding a remarkably human conversation makes for a compelling demonstration. Automatically reconciling thousands of transactions and producing exactly the same answer tomorrow does not.
But that is precisely where some of the largest business value may emerge. Enterprise software has spent decades turning repeatable work into systems. Generative A.I. initially seemed to challenge that model by introducing software capable of improvising and making decisions on the spot. We are now learning that the most powerful applications may do something subtler: use A.I.’s flexibility at one layer of the system to create greater predictability everywhere else.
For businesses, the question, therefore, should not be whether they are using the most intelligent model available but whether their A.I. is designed for the actual work it has to do. Sometimes that means giving the model more room to think, but increasingly, it will mean knowing when not to.