LayerRoPE: Treating Transformer Depth Growth as a Positional Code Podcast Por  arte de portada

LayerRoPE: Treating Transformer Depth Growth as a Positional Code

LayerRoPE: Treating Transformer Depth Growth as a Positional Code

Escúchala gratis

Ver detalles del espectáculo

Exclusivo para miembros Prime | $0.99/mes por 4 meses + $20 de crédito en Audible

$8.99 al mes después de 4 meses. Consulta términos y condiciones.
This episode examines LayerRoPE, a 2026 preprint proposing that the steady growth of hidden-state norms with depth in Transformers is a learned depth-positional code rather than a defect to suppress. It lays out the background: residual-stream accumulation under Pre-Norm, massive activations, high-norm tokens in ViTs, and the Curse of Depth. It also covers remedies such as Layer-Norm Scaling, Peri-Norm, ReZero and LayerScale, plus stream-restructuring approaches like Hyper-Connections. The core idea is to apply RoPE-style rotation along layer index instead of token position, motivated by an analysis of 16 open LLMs in which about 99 percent of norm-weight dimensions share one depth trajectory and about 1 percent supply a rotation. The discussion stays skeptical. Attention-block norms and some model families show no clear depth code, and the observed arc could just be smooth drift between neighbouring layers. A rival explanation is that RMSNorm resets input scale, so growing gains need no positional signal at all. Listeners get a clear look at the competing explanations for norm growth and at the scaling-law methodology (58M to 1.3B parameters at about 80 tokens per parameter) used to judge whether the approach helps. Sources: 1. LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition — Shikhar Srivastava, Christopher Kanan, 2026 http://arxiv.org/abs/2610.09179 2. Root Mean Square Layer Normalization — Biao Zhang, Rico Sennrich, 2019 https://scholar.google.com/scholar?q=Root+Mean+Square+Layer+Normalization 3. RoFormer: Enhanced Transformer with Rotary Position Embedding — Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, Yunfeng Liu, 2021 https://scholar.google.com/scholar?q=RoFormer%3A+Enhanced+Transformer+with+Rotary+Position+Embedding 4. The Curse of Depth in Large Language Models — Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu, 2025 https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models 5. Peri-LN: Revisiting Layer Normalization in the Transformer Architecture — Jeonghoon Kim et al., 2025 https://scholar.google.com/scholar?q=Peri-LN%3A+Revisiting+Layer+Normalization+in+the+Transformer+Architecture 6. Massive Activations in Large Language Models — Mingjie Sun, Xinlei Chen, J. Zico Kolter, Zhuang Liu, 2024 https://scholar.google.com/scholar?q=Massive+Activations+in+Large+Language+Models 7. Vision Transformers Need Registers — Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski, 2024 https://scholar.google.com/scholar?q=Vision+Transformers+Need+Registers 8. Hyper-Connections (and Manifold-Constrained Hyper-Connections, mHC) — Defa Zhu et al. (Hyper-Connections, ByteDance, 2024-25); Zhenda Xie et al. (mHC, DeepSeek, 2025), 2024-2025 https://scholar.google.com/scholar?q=Hyper-Connections+%28and+Manifold-Constrained+Hyper-Connections%2C+mHC%29 9. ReZero is All You Need: Fast Convergence at Large Depth — Bachlechner, Majumder, Mao, Cottrell, McAuley, 2020/2021 https://scholar.google.com/scholar?q=ReZero+is+All+You+Need%3A+Fast+Convergence+at+Large+Depth 10. Going Deeper with Image Transformers (CaiT / LayerScale) — Touvron, Cord, Sablayrolles, Synnaeve, Jégou, 2021 https://scholar.google.com/scholar?q=Going+Deeper+with+Image+Transformers+%28CaiT+%2F+LayerScale%29 11. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models — Wang, Zhu, Fang, Li, Shen, Zhong, 2026 https://scholar.google.com/scholar?q=Negligible+in+Size%2C+Significant+in+Effect%3A+On+Scale+Vectors+in+Large+Language+Models 12. On the Residual Scaling of Looped Transformers: Stability and Transferability — Wang, Li, Zhang, Huang, Yan, Li, 2026 https://scholar.google.com/scholar?q=On+the+Residual+Scaling+of+Looped+Transformers%3A+Stability+and+Transferability 13. Parcae: Scaling Laws for Stable Looped Language Models — Prairie, Novack, Berg-Kirkpatrick, Fu, 2026 https://scholar.google.com/scholar?q=Parcae%3A+Scaling+Laws+for+Stable+Looped+Language+Models 14. On Layer Normalization in the Transformer Architecture — Xiong et al., 2020 https://scholar.google.com/scholar?q=On+Layer+Normalization+in+the+Transformer+Architecture 15. Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks (Depth-μP) — Yang, Yu, Zhu, Hayou, 2024 https://scholar.google.com/scholar?q=Tensor+Programs+VI%3A+Feature+Learning+in+Infinite-Depth+Neural+Networks+%28Depth-%CE%BCP%29 16. Efficient Streaming Language Models with Attention Sinks — Xiao, Tian, Chen, Han, Lewis, 2023 https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks 17. Dissecting Outlier Dynamics in LLM NVFP4 Pretraining — Dong et al., 2026 https://scholar.google.com/scholar?q=Dissecting+Outlier+Dynamics+in+LLM+NVFP4+Pretraining 18. Hyper-Connections / mHC: Manifold-Constrained Hyper-Connections — Zhu et al. 2025; Xie et al. 2025, 2025 https://scholar.google.com/scholar?q=Hyper-Connections+%2F+mHC%3A+...
adbl_web_anon_alc_button_suppression_t1
Todavía no hay opiniones