A giant soap bubble holding the logos of OpenAI, Nvidia, Microsoft, Google, Amazon, Meta and Anthropic, resting in front of a tall stone wall.

In these past few weeks, we had some controversies involving Anthropic and even GitHub related to their access plan. From “issues” in Claude Code like routing to weaker models, to GitHub removing the subscription option for new users. Are we reaching the AI WALL? Is the Bubble about to burst??

Claude Code plans and GitHub removing subscription plans

While I was writing this article, we had some changes in GitHub Copilot. GitHub Copilot had to pause new sign-ups for some plans, tighten limits, remove expensive models from cheaper tiers, and migrate to usage-based billing. The justification is simple: coding agents do not behave like traditional chatbots. A short question and an autonomous session lasting hours cannot cost the same forever.

Claude Code went through similar tension. Anthropic tested removing Claude Code from the Pro plan for a small portion of new users, encouraged usage outside peak hours, adjusted limits during periods of higher demand, and even reduced the default reasoning effort to decrease latency and token consumption — a decision later reversed because it worsened the perception of quality.

In other words, we are running out of infinite computational power, maybe we are finally reaching that so-called Wall?

Image taken from https://metr.org/time-horizons/

But what the hell is the AI Wall?

The so-called AI wall is a point where models will start evolving at a lower frequency than now, not in the double exponential way models were improving in the past. In other words, from here on out models will not make leaps as large as the one from 3.5 to 4o or from 4o to 5, this last one was not that huge.

METR chart of the task duration models complete with a 50% chance of success, climbing from a few minutes for GPT-4 in 2023 to over six hours for GPT-5.2; a brick wall is pasted onto the right edge, cutting the curve short.

For models to keep evolving at this accelerated pace, which is that double exponential curve represented by the models in the meme image above, we need at least 4 fundamental stages.

  • Training Data: Texts, images, videos, anything created by human beings that can be used as training for the models.
  • Infrastructure: Basically, these are Datacenters. The more GPUs and servers, the more we can scale the models.
  • Algorithmic Efficiency: This is where the ability to make better models while spending fewer resources comes in — whether through more efficient architectures, better context usage, routing, compression, quantization, or inference techniques.
  • Regulations: This stage is related to limitations/rules that AI research may have.

Training Data

To train an AI model, you need DATA, a lot of data and according to Epoch AI, the effective stock of high-quality public human text is on the order of 300 trillion tokens, with a wide interval, and it projects that this stock could be fully used up between 2026 and 2032, or earlier if models continue being overtrained with many passes over similar data. There are already estimates that this data has already run out.

With that, what is left for AIs is Synthetic Data, if future models are trained indiscriminately on content generated by previous models, model collapse can happen: they lose tails of the distribution, diversity, and begin to “fail to perceive” reality.

Infrastructure

This is a problem that the United States is trying to solve for the AGI race, but they are facing several problems. One way to scale an LLM by brute force is by increasing its computational power. Epoch analyzed four major constraints to keep scaling training through 2030: electric power, chip manufacturing capacity, data scarcity, and the latency wall. It still thinks it is plausible to reach enormous training runs by 2030, but that already enters the scale of gigawatts, monstrous data centers, and network/power bottlenecks that look more like “civilization-scale construction” than “training a little model in the basement.” Besides that, AI usage is now high, so considering the recent controversies in the US over the creation of new datacenters, the problems related to Claude and GitHub’s own CoPilot, possibly if any country is going to win the scaling race, that country will be China.

Algorithmic Efficiency

This is perhaps one of the hardest stages of the AI “wall.” We can keep scaling by brute force, using more GPUs, more energy, more data, and even synthetic data, but the central question is: can we make models fundamentally more efficient and capable?

Since the article Attention Is All You Need, which consolidated the Transformer as the dominant architecture, we have had important advances: Chain-of-Thought, better prompting techniques, RLHF/DPO, inference-time scaling, Mixture of Experts, model routers, distillation, speculative decoding, KV cache optimizations, long context, and recent quantization methods like TurboQuant. However, many of these techniques improve model usage, cost, or inference, without necessarily representing a new architecture as transformative as the Transformer.

Regulations

The advance of AI does not depend only on technical capacity. Even if models become smarter, cheaper, and more efficient, they will still need to cross the regulatory wall. This wall does not appear in the same way in every country: in the US, regulation tends to be more fragmented and disputed among the federal government, states, and companies; in China, it is more centralized and strongly tied to content control, national security, and state governance; meanwhile the European Union follows a more formalized and risk-based path, with the AI Act in force since 2024 and obligations entering in phases.

This difference creates an asymmetry: whoever has more flexible rules can experiment faster, scale models with less friction, and perhaps get ahead in the race for increasingly autonomous systems. But this advantage comes with risks: deepfakes, surveillance, discrimination, commercial abuses, irresponsible automation, and models capable of acting beyond what we can audit.

In the AGI race, whoever has fewer regulations will really be able to get ahead, but will it be worth the risk of having a misaligned model to control, whether it is close to AGI or not?

And the AI Bubble? Will it burst?

Now that you understand where the Artificial Intelligence Wall is, we can analyze the scenario of a major explosion.

For the AI bubble not to burst, or at least not to suffer a major correction, we need several things to keep working in parallel: infrastructure needs to keep scaling, energy and chip costs need to be sustainable, new algorithmic innovations need to keep emerging, regulations need to remain flexible enough to allow experimentation, and synthetic data needs to be good enough to compensate for the gradual exhaustion of human data available on the internet.

That is why we are seeing a race for gigantic data centers, companies seeking their own energy sources, Big Techs investing absurd amounts in chips, governments trying to balance innovation with safety, and labs betting more and more on synthetic data, agents, specialized models, and reasoning at inference time.

Bloomberg diagram titled “How Nvidia and OpenAI Fuel the AI Money Machine”, with circles sized by market value — Nvidia at $4.5T, Microsoft at $3.9T, OpenAI at $500B — linked by arrows for hardware, investment, services and venture capital.

The problem is that the market seems to be pricing in a future where all these stages will be overcome continuously. And that is the risk. If just one of these pillars fails, growth may slow down. But if several fail at the same time — expensive energy, scarce chips, heavy regulation, bad synthetic data, few algorithmic gains, and models that do not deliver financial return proportional to the investment — then yes, we may see a bubble burst. And maybe it will not even be a normal bubble, but a true financial black hole, sucking in capital, energy, chips, and expectations. And this is where one of the most important, but least glamorous, discussions in modern AI comes in: token optimization.

Token Optimization

At the end of the day, behind every magical interface, every autonomous agent, and every intelligent “copilot,” there is a bill being paid. Every message sent, every document read, every tool called, every reasoning step, and every attempt to solve a task consumes tokens. Tokens become inference. Inference becomes GPU. GPU becomes energy. Energy becomes cost. And cost becomes the real limit between a revolutionary technology and an economically unviable product.

Maybe Tokens are the invisible oil of modern AI: nobody sees them burning, but every response consumes a little. Model routing, context compression, RAG, semantic cache, specialist models, and better prompt techniques will become increasingly important.

What changes for a developer?

With or without the bubble bursting, the technology of LLM models and AI agents is here to stay. Having said all that, it is worth continuing to study agent orchestration, but not leaving development knowledge aside, architecture, code, because you will need a lot of that in the future, especially if models start getting more expensive.

And remember, for the bubble to burst early, models do not need to fail, it is enough for it to take a little longer than predicted to pay for itself with or without AGI.

References

GitHub. Changes to GitHub Copilot Individual plans. 2026.
 GitHub. GitHub Copilot is moving to usage-based billing. 2026.
 Anthropic. An update on recent Claude Code quality reports. 2026.
 Anthropic Support. Claude March 2026 usage promotion. 2026.
 Epoch AI. Will we run out of data? Limits of LLM scaling based on human-generated data. 2024.
 Epoch AI. Can AI scaling continue through 2030? 2024.
 Shumailov et al. AI models collapse when trained on recursively generated data. Nature, 2024.
 Vaswani et al. Attention Is All You Need. 2017.
 Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. 2022.
 Rafailov et al. Direct Preference Optimization. 2023.
 Fedus et al. Switch Transformers: Scaling to Trillion Parameter Models. 2022.
 OpenAI. Learning to reason with LLMs. 2024.
 DeepSeek-AI. DeepSeek-V3 Technical Report. 2024.
 Google Research. TurboQuant: Redefining AI efficiency with extreme compression. 2026.
 European Commission. AI Act enters into force. 2024.
 DigiChina / Stanford. Internet Information Service Algorithmic Recommendation Management Provisions. 2022.
 China Law Translate. Interim Measures for the Management of Generative Artificial Intelligence Services. 2023.
 Bloomberg. A Guide to the Circular Deals Underpinning the AI Boom. 2026.
 MUFG. Complex Web of AI Corporate Cross-Holdings. 2026.