AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors.
- ▪Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate eff
- ▪Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings.
- ▪To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Artificial Intelligence arXiv:2607.19363 (cs) [Submitted on 5 Jun 2026] Title:AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally Authors:Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian Li View a PDF of the paper titled AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally, by Shaowen Wang and 6 other authors View PDF HTML (experimental) Abstract:Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv.org.