The chart showing "Token falls below one dollar" has gone viral online. The actual situation, however, is: the Jevons Paradox combined with RSI, as AI demand decreases, pricing becomes more explosive.
According to a note from Goldman Sachs' Delta One trading desk cited by ZeroHedge, what is proposed is a conditional stress test, rather than a factual judgment that "global computing power is already excessive."
A chart shared by the American financial platform ZeroHedge on social media from Goldman Sachs Delta One trading desk shows a sharp drop in artificial intelligence computing power Token prices, which may weaken the key support for stock market valuations based on the AI computing power themea theme that has driven the recent global stock market bull run. This chart regarding the plummeting Token prices indicates that unless the consumption of computing power continues to surge enough to offset the impact of falling prices, the effects may be detrimental. From another perspective, the decrease in unit Token costs stimulates increases in invocation volume, inference depth, and agent loops, thus opening a second curve of surging computing power demand through the RSI for continuous training.
However, as the development of large AI models, led by OpenAI and Anthropic, moves into a new phase of "recursive self-improvement (RSI)", along with recent strong performance and outlooks from Nvidia and a series of megawatt-level computing power capacity contracts driving the prosperous expansion of the AI computing power industry chain, existing fundamental data of the AI computing power industry does not support the pessimistic assumption of a "systemic surplus of computing resources across the entire industry." On the contrary, it indicates that the demand for AI computing power tends to explode further as prices drop.
This note from Goldman Sachs Delta One trading desk disseminated by ZeroHedge is not an official industry benchmark forecast published by Goldman Sachs' official research team. Its core valuation formula is based on model service revenue Token unit price Token consumption quantity. If the price index falls by 29% in a month, under the simplified assumption of a constant business structure, Token usage must increase by approximately 40.8% to keep total revenue unchanged. If factors such as chip throughput and power per watt are continuously improved, it is necessary to maintain underlying hardware demand, and the composite growth rates of Token quantities, calculations per Token, and inference durations need to be even higher.
The price drop of Tokens can stimulate broader and longer-lasting invocations through the Jevons Paradox, but it cannot automatically guarantee revenue for model vendors or returns on infrastructure; a more reasonable current market prediction suggests that the global enterprise AI computing power demand is still in a phase of strong expansion. However, older generation AI accelerators, highly leveraged cloud vendors, and data centers lacking long-term customer contracts may first experience localized, generational surpluses.
The recent hot discussions about the so-called Jevons effect, also known as the "Jevons Paradox," describe a counterintuitive economic theory: when cutting-edge technology advances the efficiency of a particular resource (such as energy, raw materials, or AI computing power infrastructure), it lowers unit costs and thereby stimulates a massive expansion of market demand, ultimately leading to an increase in the total consumption of all types of resources instead of a decrease. This concept was first proposed by British economist William Stanley Jevons in his 1865 book "The Coal Question."
As AI Token prices plummet, the surplus alert is ringing once more.
This institutional report shared by ZeroHedge on social media points out that the increasingly fierce competition among leading AI developers is significantly compressing Token costs. Zerohedge quotes warnings from Goldman Sachs Delta One trading desk that the plummeting prices of AI Tokens may weaken the valuation support for the AI trading frenzy in the stock market, unless the consumption of computing power surges enough to offset the impact of falling prices.
The report indicates that the fierce competition among the leading developers of large AI models is driving a significant reduction in the cost per Token, making the current monetization structure increasingly difficult to maintain. The Goldman Sachs trading desk specifically highlighted Meta's recent release of Muse Spark 1.3which has shown strong competitiveness in benchmarksas well as widespread speculation about the downward trajectory of Token costs for an upcoming model called Astra from OpenAI.
As inference workloads begin to shift partly towards more localized, lower-cost hardware systems, the Goldman Sachs trading desk raises questions about whether charging per Token can become a long-term business model. The report states, "Combined with the increasing shift of inference to cheaper local hardware systems, we find it hard to believe that a per-Token pricing model can become a sustainable final business model." The report emphasizes that this shift will directly impact the stock market.
Goldman Sachs further warns that the combination of broader AI adoption and lower Token sales prices does not automatically constitute a continued positive benefit for investors; if the rate of decline in the costs of generating AI outputs exceeds the overall growth rate of consumption, the entire industry may face an oversupply of AI computing power infrastructure.
The trading desk explicitly warns that unless upcoming AI models drive fundamental performance and intelligence leaps in underlying capabilities, "the emergence of periodic oversupply of computing power will become a real possibility."
With the price of one million Tokens falling below one dollar, a counterintuitive sign is that closed-source models unexpectedly lead the price drop wave. This market re-evaluation and re-pricing of Token prosperity is taking place as the overall industry prices continue to decline. Notably, leading market analytics firm Silicon Data's benchmark index, used to track the market cost of large language models, has yet to reach a bottom, recently dipping to a historic low of $0.97 per million Tokens. Goldman Sachs' research report indicates that this index plummeted by 29% just in August.
Over the past two weeks, the prices of both open-weight models and proprietary models have decreased; however, Silicon Data indicates that the significant price cuts of some high-performance frontier closed-source models are the main drivers of the overall market price drop.
The rise of RSIAI begins to participate in the creation of the next generation of large AI models! The price drop of Tokens does not indicate a retreat in computing power; the real variable is whether the increase in volume can outpace the fall in prices.
The note from Goldman Sachs Delta One trading desk cited by ZeroHedge presents a conditional stress test rather than a factual judgement of "the global computing power is already oversupplied."
Existing industry data does not yet support the idea that "the entire industry has systematically excess computing power." Nvidia's revenue for the second fiscal quarter of 2027 grew 106% year-on-year to $96.2 billion, with data center revenue growing 117% to $89 billion, and approximately 70% of the revenue growth outlook for 2028 is clearly defined as "supply-constrained outlook."
Moreover, according to recent media reports, Anthropic has signed Lambda computing power agreements worth approximately $35 billion, as well as a six-year Nscale agreement valued at about $45 billion, corresponding to around 460 megawatts of capacity. In August, South Korea's semiconductor exports surged by 209% year-on-year to $46.65 billion, setting a record high against last year's high base; Goldman Sachs' official research further projects that global AI investment will reach about $1 trillion by 2026.
AI server storage chip components remain the clearest supply bottleneck within the AI computing power industry chain. Market research firm TrendForce predicts that in the third quarter of 2026, the contract prices for traditional DRAM and NAND Flash will increase by 13%-18% and 10%-15%, respectively, compared to a high base; HBM and server RDIMM are expected to account for 51% of the global DRAM bit supply in 2026, and HBM contract prices may still rise by 70%-140% in 2027. Their more aggressive forecast is that DRAM and NAND will rise from accounting for 47% of major cloud service providers' capital expenditure in 2026 to 68% in 2027, but this ratio is also driven by surging storage prices, not purely representative of procurement growth.
On the corporate side, the real focus should be on "Cost per Successful Task" rather than just the price per million Tokens. Although Gemini 3.8 Flash maintains a price of $0.75 per million input Tokens and $3.75 per million output Tokens, the increase of approximately 30% in output Tokens and the rise in agent invocation rounds mean that the cost per benchmark task has actually increased by about 40% compared to the previous generation.
Meanwhile, a report from Goldman Sachs' official research team shows that AI agent-driven, proxy AI workflows are expected to drive Token monthly consumption to grow 24 times to 120 trillion Tokens between 2026 and 2030, and it is anticipated that chips may still be in short supply over the next 12-18 months. Therefore, while the price drop of Tokens can stimulate broader and longer invocation periods through the Jevons Paradox, it cannot automatically guarantee income for model vendors or returns on infrastructure. Overall demand for computing power remains in a phase of robust expansion, but older accelerators, highly leveraged cloud vendors, and data centers lacking long-term contracts may first face localized, generational surpluses.
The decrease in unit Token costs stimulates increases in invocation volume, inference depth, and agent loops, while RSI is opening a second curve of computing power demand for ongoing training. Recursive Self-Improvement (RSI) does not equate to models suddenly gaining self-awareness; its current verifiable form involves letting AI gradually participate in training strategy planning, algorithm generation, code writing, experimental design, result evaluation, error correction, and the next training rounds, up to several subsequent rounds of AI training.
The trend of RSI brings an important direct inference: advantages in AI computing power directly research advantages, which then directly convert into model advantages; those AI laboratories with sufficient GPUs can complete all verifications in a day, while frontier AI laboratories lacking GPUs can only complete five tasks or fewer in a day, which immediately widens the speed gap.
Google has made it clear that part of the progress of Gemini 3.8 Flash comes from the ability to "recursively assess and improve the underlying model" through lengthy intelligent agent loops. OpenAI has temporarily slowed down parts of its frontier training due to a misalignment between the pace of advancements in frontier model capabilities and safety assurance capabilities. SSI, following investment from Nvidia and support from the Vera Rubin system, plans to expand its computing power by an order of magnitude within 12 months. These signals indicate that AI-assisted AI research is in the process of becoming a reality, and this process will lead to increasingly massive demands for AI computing power. However, it remains unproven whether completely autonomous, sustainable exponential RSI has been achieved.
RSI reinforces the continuity of computing power demand; thus, regarding specific AI investments, future investors should monitor whether the growth rate of Token consumption can outpace the decline in unit costs, cluster utilization rates, and lease renewal rates, as well as the improvements in model capabilities per unit of computing power and whether model revenues can cover AI chip, AI server rack component, power chain device, storage, and financing costs.
RSI may create a "second demand curve" for computing power, independent of end-user demand: training is no longer simply "trainreleasecomplete," but instead involves models proposing algorithms, running hundreds to thousands of experiments in parallel, evaluating results, and initiating the next round of training in an ongoing closed loop; agents will also increase testing time compute through longer inference chains, repeated tool invocations, and multi-model collaboration. If leading laboratories, such as OpenAI and Anthropic, reallocate computing resources from high-revenue inference businesses back to training, short-term revenues may plummet, but overall computing power consumption is unlikely to decrease, and may even rise due to the competition for leading model status.
Related Articles

The swap market has fully priced in an interest rate hike in September, and Nomura goes further: in an extreme scenario, the Bank of Japan may resort to a rare "triple whammy" within this year.

AI venture capital is shifting from frenzy to selectiveness! A tide of cleansing has arrived, and a major reshuffle in the industry may be imminent.

Arbitrage trading retreat boosts a significant rise in the yen, while expectations for the Bank of Japan's interest rate hikes continue to heat up.
The swap market has fully priced in an interest rate hike in September, and Nomura goes further: in an extreme scenario, the Bank of Japan may resort to a rare "triple whammy" within this year.

AI venture capital is shifting from frenzy to selectiveness! A tide of cleansing has arrived, and a major reshuffle in the industry may be imminent.

Arbitrage trading retreat boosts a significant rise in the yen, while expectations for the Bank of Japan's interest rate hikes continue to heat up.

RECOMMEND





