Chutes and Harvard Publish 6.12 Billion LLM Request Dataset

Chutes AI and Harvard researchers have released a public dataset covering 6.12 billion large language model (LLM) requests processed between April 2025 and April 2026. The dataset spans 9,174 models and 314,970 anonymised users, making it one of the largest collections of LLM serving metadata available. The dataset contains request timing, token usage, latency and time-to-first-token data, but no prompts or model responses. Chutes processed about 35.8 trillion input tokens and 2.52 trillion output tokens during the period. User identifiers rotate every three months to strengthen privacy. A key finding is that 99% of repeat requests occur within 15 minutes. The research suggests prefix-aware routing can deliver near-optimal cache-hit rates with only limited server load imbalances. This could reduce inference costs and improve efficiency for AI infrastructure providers. The study also found that average output lengths declined from hundreds of tokens to fewer than 100, potentially reflecting increased use of automated or agentic AI queries. Chutes operates on Bittensor Subnet 64, where decentralised GPU infrastructure supports open-source LLMs and payments use TAO. The project offered participating researchers a 25% discount during an opt-in data-contribution period.
Neutral
The immediate market impact is likely neutral. The release provides important evidence about LLM demand, caching and inference efficiency, but it does not announce new revenue, token demand, network upgrades or capital inflows for Bittensor. TAO is the only cryptocurrency directly linked to the report because Chutes operates on Bittensor Subnet 64 and uses TAO for payments. In the short term, traders may treat the dataset as a positive narrative catalyst for decentralised AI and GPU infrastructure. If market participants interpret the high level of repeat requests as evidence that inference workloads can be served more cheaply, sentiment toward Bittensor and related AI tokens could improve. However, data releases without direct changes to network activity or token economics have historically produced limited and short-lived price reactions compared with listings, incentive changes, emissions updates or major partnerships. Longer term, the findings could support demand for efficient decentralised inference networks. Lower caching and routing costs may improve the competitiveness of platforms such as Chutes, potentially increasing usage and transaction activity. The main risks are that the dataset covers one platform, metadata rather than user content, and an opt-in contribution programme. Traders should therefore monitor TAO volume, Bittensor subnet emissions, active users, inference revenue and broader AI-token momentum before treating the news as a bullish signal.