AI Token Demand Surges as Platforms Tighten Usage Limits

AI token demand is rising faster than inference costs are falling, prompting platforms to tighten usage limits and revise subscription plans. Zhipu has ended automatic renewal for its legacy unlimited-weekly-quota GLM Coding Plan and is compensating existing users with two months of a new plan. It has also launched GLM Coding Plan subscriptions on Tmall, priced at 118 yuan, 538 yuan and 1,078 yuan per month, with different credit allowances and support for more than 20 coding agents. The shift reflects growing pressure from agent-based AI, where coding, tool use and multi-step tasks consume far more tokens than conventional chatbot conversations. Kimi temporarily stopped accepting new consumer subscribers after demand approached its computing-capacity limit. Alibaba Cloud discontinued new purchases and renewals for part of its Coding Plan, while Tencent Cloud raised the price of some Hunyuan model input tokens by more than 460%. Zhipu reported first-half 2026 revenue of 954 million yuan, up 399.7% year on year. Revenue from its model-as-a-service (MaaS) and API business reached 825 million yuan, accounting for 86.5% of total revenue, compared with 15.2% a year earlier. Token usage increased more than 40-fold from the start of the year, while average API prices rose about 101% and inference costs fell 80%. MaaS gross margin reached 24.6%. For traders, the key signal is that AI companies are shifting from one-off model deployments to recurring cloud usage and outcome-based services. Falling token prices may stimulate even greater consumption, creating demand for AI infrastructure, chips and data-centre capacity. However, capacity constraints, higher pricing and subscription limits could increase volatility across AI-related technology and crypto tokens linked to the sector.
Neutral
The article is neutral for the broader cryptocurrency market because it does not report a direct change in crypto demand, regulation or capital flows. Its main subject is the pricing and capacity management of AI model services. The immediate trading impact is therefore likely to be concentrated in AI-related equities, semiconductor suppliers, cloud providers and crypto projects associated with artificial intelligence rather than across the entire digital-asset market. In the short term, tighter AI subscriptions and higher API prices could be interpreted as evidence of strong demand and scarce computing capacity. That may support sentiment around AI infrastructure and related tokens. However, limits on access, rising service costs and potential supply bottlenecks could also trigger concerns about margins, adoption and the sustainability of AI growth. Traders may therefore react selectively rather than treat the news as a broad risk-on signal. Over the long term, the shift from model sales to recurring API, MaaS and agent-based workloads is potentially constructive. Zhipu’s revenue growth, 40-fold increase in token usage and 80% decline in inference costs indicate improving monetisation and efficiency. At the same time, the Jevons-paradox dynamic means cheaper tokens can drive much higher consumption, keeping demand for chips, data centres and power elevated. Historical technology cycles show that strong usage growth can benefit infrastructure providers, but capacity expansions and valuation expectations can also create sharp corrections. The likely market outcome is differentiated performance: positive for credible AI infrastructure and revenue-generating platforms, but mixed for speculative tokens without clear utility or cash-flow exposure.