Vercel’s September AI Gateway report says open-weight models processed 56% of the tokens routed through its gateway in August, while accounting for 14% of estimated spend. The split shows how usage volume and model budgets can tell different stories. It is a snapshot of one gateway’s traffic, not a measure of the whole AI market or of customers’ actual bills.
At 20:03 in episode 290 of the All-In podcast, the panel discusses the pace of open-weight and premium model releases and argues that more capable downloadable models widen access. That is the panel’s interpretation. A separate, more limited data point comes from Vercel’s September AI Gateway Production Index, which reports traffic through August.
The big change
- What changed: Open-weight models moved from a minority to 56% of tokens on Vercel AI Gateway between December and August, while taking 14% of August’s list-price-estimated spend.
- Why it matters: Teams choosing models need to separate token demand from spend, then compare quality and full operating cost on the tasks they run. The index provides no task-specific results or savings for an individual team.
- What to watch: The practical question is whether a chosen model meets a team’s quality, latency, privacy and operating requirements at its full cost. Token share alone cannot answer that; license terms, hardware and engineering time also matter.
What the 56% and 14% figures measure
Vercel says open-weight models handled 56% of the tokens routed through AI Gateway in August, the first month they took a majority of that gateway’s volume. They accounted for 14% of estimated spend. The company’s chart shows the share rising from 13% in April to 56% in August. The report is dated September 17 and uses data through August.
The gap does not mean the models are free to run. A downloadable weight can still require compute, electricity and operations. Based on the rounded shares, estimated spend per token in the open-weight group was about one-eighth that of the other models in this dataset. This calculation uses Vercel’s figures: 14% of spend divided by 56% of tokens, compared with the remaining 86% of spend divided by 44% of tokens. It describes an average across different models and token types, not the price or capability of a particular model.
The index counts input, output, reasoning, cached-input and cache-creation tokens. Those token types can have different prices, so a group’s average also reflects what its users sent and generated. Vercel estimates spend from labs’ published list prices; it says actual bills may differ. That approach lets the company compare traffic from teams bringing their own provider keys, but it is not an invoice audit or a measurement of compute costs for self-hosting.
The report also says the average estimated price per token across the gateway fell 23.2% in August. Among teams that each ran more than 10 million tokens in both July and August, the median team’s estimated price per token fell 7.6%. Vercel attributes part of the decline to the growth of open-weight use. The report does not isolate how much of the change came from that shift, from different token mixes, or from teams changing models and workloads.
One gateway is not the whole market
The index is anonymized aggregate traffic routed through Vercel AI Gateway. It offers a view of real production requests passing through that service, which is more concrete than a model leaderboard or a release announcement. But the report does not establish that its customers represent all AI users, disclose a sampling frame for the wider market, or show adoption across workloads that do not use the gateway. Its percentages should therefore be read as Vercel’s gateway mix.
Vercel also notes that its open-weight classification follows the current AI Gateway model list, which is broader than the definition used in earlier reports. Earlier figures may change as Vercel revises its methodology and receives more complete data. The month-to-month trend is informative, but the report itself flags that the category boundary shifted.
“Open-weight” is also not interchangeable with “open source.” Model weights are learned numerical parameters that can be released for download. The Open Source Initiative’s Open Source AI Definition 1.0 sets a broader standard: it calls for the freedom to use, study, modify and share an AI system, and identifies training-data information, code and model parameters as parts of the preferred form for modification. A downloadable checkpoint may still come with license restrictions or omit training code and data details. Teams need to check each model’s actual license and release materials.
What the gap means for AI buyers
Vercel’s figures support a narrow conclusion: on this gateway, open-weight models carried more token volume at a smaller share of list-price-estimated spend. They do not show that the same workload would cost less on a local server, or that a lower-cost model can replace a premium one without changes in quality, speed or reliability.
For a product team, the relevant comparison is the cost of completing the same task to the same acceptance standard. That means including input and output volume, retries, tools, routing, latency and any human review, then comparing provider billing with the hardware and operations needed to run weights directly. The September index does not provide those per-task comparisons. Its value is as a signal that open-weight models now account for substantial usage in one production gateway, while the spending mix remains concentrated elsewhere.
The All-In panel’s broader argument—that model availability is expanding—can be evaluated against this kind of usage evidence, but the Vercel index neither verifies every model claim made in the discussion nor predicts where adoption will go next. For now, the useful distinction is between how many tokens a category processes, what its estimated provider charges are, and what it costs a team to get a dependable result.
Sources & further reading
- Vercel, “Open-weight models take 56% of token volume” (September 17, 2026) — Vercel’s August gateway data, token definitions, list-price spend method, classification caveat and revision note.
- Open Source Initiative, Open Source AI Definition 1.0 — the freedoms and components the initiative uses to define open-source AI, distinct from weight availability alone.
- All-In episode 290, open-model discussion at 20:03 — the panel’s discussion of release pace and model access; it is cited as commentary, not as evidence for Vercel’s usage or spending figures.



