MIT-Sakana AI SIFT Cuts Coding-Agent Evaluation Costs
MIT and Sakana AI researchers have introduced SIFT, or Self-Improvement via Fast Tree-search, to reduce the cost of evaluating self-improving coding agents. Instead of running a full coding benchmark on every proposed code change, SIFT uses a large language model to compare candidate modifications head-to-head. A regularised Bradley-Terry model ranks the preferences, while asynchronous evaluation helps speed up the search.\n\nOn the Polyglot benchmark, an o3-mini coding agent using SIFT achieved 35.1% after 30 expansion steps. The earlier Darwin Gödel Machine reached 30.7% after 80 nodes. A Qwen3-Coder-30B configuration completed its search in 224 CPU hours, with about $34 in API costs—roughly one-tenth of the resources used by DGM.\n\nSIFT also improved results on TerminalBench 2.1, from 29.2% to 36.7%, and on SWE-60, from 40.0% to 52.1%. The framework could make autonomous coding-agent research more accessible by lowering compute and evaluation costs. However, its performance depends heavily on the quality of the AI judge. A flawed judge could rank inferior code changes more efficiently rather than identify genuine improvements. For crypto traders, the development is indirectly relevant to AI-agent, decentralised computing and blockchain automation narratives, but it contains no direct cryptocurrency or token catalyst.
Neutral
The market impact is neutral because SIFT is a research framework and does not involve a cryptocurrency, token launch, blockchain network upgrade or direct capital-flow catalyst. In the short term, the findings could strengthen the broader AI-agent narrative and encourage speculative interest in crypto projects associated with autonomous agents, decentralised computing or AI infrastructure. However, without a named token, partnership or commercial deployment, any move in related assets would likely be sentiment-driven and difficult to sustain.\n\nLonger term, cheaper evaluation could accelerate the development of autonomous coding systems and increase demand for compute, data and AI infrastructure. That may indirectly support crypto projects offering decentralised GPU networks, agent payments or blockchain-based automation. Similar to past AI research breakthroughs, initial trading reactions would probably be strongest in high-beta AI-linked tokens, while major assets such as BTC and ETH would receive limited direct benefit. Traders should watch for follow-up evidence, including open-source adoption, enterprise deployment, funding announcements and partnerships with blockchain platforms. The main risk is that poor AI judging could limit real-world performance, reducing the durability of the bullish narrative.