ByteDance's Seed team revealed the technical reasons for performance fluctuations in large language models when processing ultra-long text in their latest research paper. The research team pointed out that this phenomenon mainly stems from the "phase sensitivity" caused by the block-based KV cache compression technology.
Key Mechanism Causing Retrieval Bias
This technology compresses continuous token windows into fewer entries by fixed strides to reduce memory consumption during long context reasoning. However, this compression mechanism introduces a new positional coordinate—the "phase" of a token relative to the boundaries of the compressed window.
Experiments show that when facing identical information, different phases can cause dramatic changes in retrieval difficulty. In some large open-source models, the difference in long-text retrieval accuracy caused by phase differences can be as high as 40 percentage points.

Periodic Weaknesses Require Optimization
This discovery highlights the limitations of traditional average benchmark tests. Under the seemingly excellent average scores of some models, serious periodic weaknesses may be hidden. As the application of large models for long text becomes increasingly widespread, this research provides important theoretical basis for future model architecture optimization and improvement of long-context reasoning stability.
Join Now