Xiaomi Unveils HySparse2 for Faster Long-Context AI

0
24

Xiaomi’s MiMo team has unveiled HySparse2, a core architecture for their upcoming MiMo-V3 models, designed to significantly boost the efficiency and accuracy of long-context AI agents. This new architecture promises to handle extensive data like web pages, documents, code, and tool call histories with reduced computational cost and improved retrieval of critical information from increasingly complex task histories.

Optimizing for Long-Range, Multi-Turn Agents

The HySparse2 architecture is specifically engineered for next-generation agents that require processing and recalling information across extended interactions. It aims to reduce Prefill computation, minimize KV Cache size, and enhance the precision of long-context retrieval. This builds upon Xiaomi’s previous work, including the MiMo-V2 series which introduced the Hybrid SWA (Sliding Window Attention) architecture, combining full and sliding window attention for a balance of performance and efficiency.

Building on the foundation laid by the earlier HySparse architecture, HySparse2 introduces further refinements. It features an upgraded KV sharing mechanism and a more sophisticated sparse selection strategy. This allows models to process lengthy contexts more efficiently and retrieve relevant information with greater accuracy.

The Demands of Agent Attention Architectures

Modern AI agents often deal with substantial inputs, such as entire web pages or lengthy execution logs, after a single tool call. As agents progress through tasks, they need to sift through a growing history to pinpoint the necessary information for their next action. The KV Cache stores representations of processed content for future use, and the process of ingesting new information and building these caches is known as Prefill. For multi-turn agent tasks, each new piece of information requires re-processing. Therefore, an effective architecture must meet three key requirements:

  • Efficient Input Processing: Minimize the computation needed for long inputs, allowing the model to quickly move to the next generation step.
  • Memory Efficiency: Reduce the KV Cache’s memory footprint, enabling the retention of longer histories.
  • Accurate Retrieval: Precisely locate relevant evidence within extensive histories for accurate cross-turn information integration.

The first-generation HySparse architecture took a step forward by using a small number of Full Attention layers to provide KV Cache and select important positions, which subsequent Sparse Attention layers could then reuse. This reduced attention computation and caching overhead. However, Prefill still required processing all layers, and block-level selection had room for improvement in long-distance retrieval scenarios. HySparse2 addresses these limitations by introducing two-level KV sharing, token-level sparse selection, and a unified approach to accessing local information.

Key Innovations in HySparse2

Two-Level KV Sharing

Inspired by YOCO, HySparse2 divides the model into two parts: a Self-Decoder (front half) and a Cross-Decoder (back half). The Self-Decoder uses a hybrid structure of Full Attention and SWA, while the Cross-Decoder combines Full Attention with Sparse Attention. KV sharing is implemented at two levels: KV Bridging and KV Reuse.

  • KV Bridging: This bridges the front and back parts of the model. Each Full Attention layer in the Cross-Decoder builds its KV Cache from the hidden states of the corresponding Full Attention layer in the Self-Decoder. This allows the KV caches for the Cross-Decoder’s Full Attention layers to be prepared earlier, without waiting for input to pass through all preceding layers.
  • KV Reuse: Within the same Hybrid Block (comprising one Full Attention layer followed by Sparse Attention layers), the KV Cache and selection results from the Full Attention layer are directly reused by the subsequent Sparse Attention layers. This maintains the core HySparse design of using a few full attention layers for global information and selection, with multiple sparse layers efficiently utilizing this information.

Token-Level Selection and Unified Local Access

Moving beyond the block-level selection of the original HySparse, HySparse2 employs token-level selection. This allows for a more granular allocation of attention budget to specific important tokens, even if they are scattered across different parts of the context. This finer granularity has shown improvements in tasks like RULER-v2, multi-turn retrieval (MRCR-v2), and graph reasoning (GraphWalks).

To ensure the model consistently accesses recent context, HySparse2 enforces the selection of the most recent 128 tokens. In addition to these local tokens, it selects 1,024 global tokens from outside this window. Both local and global information are accessed from the shared KV Cache provided by the Full Attention layers. This integration eliminates the need for separate SWA branches and local KV caches previously used in HySparse. Consequently, local and global information can be processed within a single sparse attention computation, leveraging KV data prepared during the Self-Decoder phase.

Reduced Prefill and Deployment Benefits

The combined effect of two-level KV sharing and the merged local window approach means that all necessary KV Caches for the Cross-Decoder can be constructed from the Self-Decoder’s hidden states. Once the Self-Decoder computation and KV Bridging projections are complete, the Prefill process can conclude. For a 49-layer model, Prefill only requires executing the first 25 Self-Decoder layers and the KV Bridging projections, involving just one Full Attention layer. This significantly reduces the computational path for long inputs. In deployment scenarios separating Prefill and Decode, the Prefill nodes require only these Self-Decoder components, effectively halving the required model weight storage.

Performance Gains: Cost, Speed, and Long-Context Capabilities

Xiaomi compared HySparse2 against Hybrid SWA and the original HySparse on an 80B-A3B MoE model. The results demonstrated substantial improvements:

  • Computational and Cache Reduction: With million-token contexts, HySparse2 reduced Prefill computation to 1/5 compared to Hybrid SWA and 1/3 compared to HySparse. KV Cache size dropped from 12GB (Hybrid SWA) and 6.7GB (HySparse) to just 2.7GB.
  • Enhanced Long-Context Performance: After equivalent fine-tuning, HySparse2 showed marked improvements in long-context benchmarks up to 256k tokens. It achieved higher scores on MRCR-v2 and RULER-v2 (measuring retrieval and QA in long contexts) and lower scores on AgentPPL and LongPPL (measuring prediction accuracy for agent trajectories and long-distance dependencies). Compared to HySparse, HySparse2 saw an average increase of 11.30 and 19.81 percentage points on MRCR-v2 and RULER-v2, respectively.

These findings indicate that reduced KV Cache and shorter Prefill times can coexist with superior long-context retrieval capabilities, establishing HySparse2 as a more efficient architectural foundation for multi-turn, long-context agents.

Accelerating Agent Interactions

Xiaomi’s previous efforts, such as MiMo-UltraSpeed, focused on optimizing the decoding (output generation) phase through techniques like low-bit quantization and speculative decoding. HySparse2 complements these efforts by drastically lowering the Prefill and caching overhead associated with processing long inputs. By addressing both input processing and output generation, Xiaomi aims to deliver faster and more efficient AI task completion for users. The company plans to integrate these advancements into MiMo-V3, striving for faster Prefill and Decode speeds within the same model scale, enabling users to accomplish more complex tasks more quickly.

Source: https://www.ithome.com/1/006/839.htm

LEAVE A REPLY

Please enter your comment!
Please enter your name here