I have argued that most of the money going into AI models is aimed at the wrong layer, because a model is becoming an operating system: two or three will matter, and most people will never know which one they are running. This is the working half of that argument. Real work does not need the best model.
The model is the least interesting layer now. What changes behavior is the harness around it: in my own account, I have run the same model and watched it behave differently when the harness changed.
Look at what people actually run. On Vercel's AI Gateway, open-weight models carried 56 percent of token volume in August 2026 but only 14 percent of spend, and price per token fell 23.2 percent in the month. [1] That is one gateway's view, and Vercel prices spend at the labs' published list prices, so it is not the whole market. But it is production traffic, not a survey of intentions, and it shows where the work goes when someone is paying for it: to the cheaper model.
I see the same thing in my own account. My data pipeline sends its bulk extraction, the mechanical work of turning raw text into structured records, to the smallest and cheapest model in a family, and reserves a larger one for analysis. It runs every day. The largest model is not used for volume work at all, because the job does not need it.
There is a name for this. Herbert Simon called it satisficing: choosing what is good enough instead of searching for the best. [2] Every choice of model is a trade among cost, speed and quality, and for most work the first two win. I can run, and I cannot run a marathon, and almost nothing I need to do requires one.
The strongest argument against this is that open models keep getting better. Linux is the precedent. Its creator announced it in 1991 as "just a hobby, won't be big and professional," and it did not stay one. [3] Epoch AI finds that the best open-weight models have trailed the closed frontier by about four months since January 2026. [4] That is close enough that a premium cannot last long for ordinary work, and far enough that the leaders still earn one on the hardest.
So the frontier keeps its premium only where quality is the one thing that cannot give. Everywhere else, sufficiency wins. An investor looking at a company that sells model access should ask what share of its customers' work is the hardest kind, and what share is the ordinary kind that a cheaper model already handles.
Sources
- AI Gateway Production Index: September 2026, Vercel
- Herbert A. Simon, "Rational Choice and the Structure of the Environment," Psychological Review 63 (1956), 129-138; background: Herbert A. Simon, Wikipedia
- Linux first announced, August 25, 1991: Torvalds' post to comp.os.minix, History of Information
- Open vs. closed model gap, Epoch AI Data Insight
The descriptions of the author's own data pipeline and of running models are the author's account and cannot be independently sourced. This is not investment advice.