AI Post Transformers Podcast Por mcgrof arte de portada

AI Post Transformers

AI Post Transformers

De: mcgrof
Escúchala gratis

Exclusivo para miembros Prime | $0.99/mes por 4 meses + $20 de crédito en Audible

$8.99 al mes después de 4 meses. Consulta términos y condiciones.
AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.
Episodios
  • Birkhoff Geometry of Manifold-Constrained Hyper-Connections: Width Fades With Depth
    Oct 10 2026
    This episode examines a theory paper on manifold-constrained hyper-connections (mHC), which widen the residual stream into four parallel streams and force each layer's mixing matrix to be doubly stochastic using Sinkhorn normalization. It explains why the composite of many mixers looks nearly uniform, and why that suggests the extra width fades with depth. The key argument is that the average of the streams passes through with gain exactly one, like an ordinary residual network. All the added capacity sits in a "difference channel" that shrinks each layer by the mixer's second singular value. In the paper's worked example, which assumes a smallest entry of 0.02, that shrinkage compounds to under 0.091 over thirty layers. The episode also covers a rigidity result: only permutation matrices avoid this shrinkage. It then starts on the training dynamics, where backprop through Sinkhorn gives the Fisher-Rao gradient and permutation corners sit at infinity in logit space. Listeners get a clear account of what DeepSeek's extra width is worth, and of why such mixers become "sticky" near permutation matrices. Sources: 1. The Birkhoff Geometry of Manifold-Constrained Hyper-Connections: Two Channels, Vertex Viscosity, and Sinkhorn as a Retraction — Xiaoyu Li, Zhizhou Sha, Chiwun Yang, 2026 http://arxiv.org/abs/2610.06653v1 2. Concerning Nonnegative Matrices and Doubly Stochastic Matrices — Richard Sinkhorn, Paul Knopp, 1967 https://scholar.google.com/scholar?q=Concerning+Nonnegative+Matrices+and+Doubly+Stochastic+Matrices 3. Sinkhorn Distances: Lightspeed Computation of Optimal Transport — Marco Cuturi, 2013 https://scholar.google.com/scholar?q=Sinkhorn+Distances%3A+Lightspeed+Computation+of+Optimal+Transport 4. Learning Latent Permutations with Gumbel-Sinkhorn Networks — Gonzalo Mena, David Belanger, Scott Linderman, Jasper Snoek, 2018 https://scholar.google.com/scholar?q=Learning+Latent+Permutations+with+Gumbel-Sinkhorn+Networks 5. mHC: Manifold-Constrained Hyper-Connections — Zhenda Xie et al. (DeepSeek-AI), 2026 https://scholar.google.com/scholar?q=mHC%3A+Manifold-Constrained+Hyper-Connections 6. Problem Complexity and Method Efficiency in Optimization — Arkadi Nemirovski, David Yudin, 1983 https://scholar.google.com/scholar?q=Problem+Complexity+and+Method+Efficiency+in+Optimization 7. Natural Gradient Works Efficiently in Learning — Shun-ichi Amari, 1998 https://scholar.google.com/scholar?q=Natural+Gradient+Works+Efficiently+in+Learning 8. Mirror Descent and Nonlinear Projected Subgradient Methods for Convex Optimization — Amir Beck, Marc Teboulle, 2003 https://scholar.google.com/scholar?q=Mirror+Descent+and+Nonlinear+Projected+Subgradient+Methods+for+Convex+Optimization 9. Mirror Descent Policy Optimization — Manan Tomar, Lior Shani, Yonathan Efroni, Mohammad Ghavamzadeh, 2020 https://scholar.google.com/scholar?q=Mirror+Descent+Policy+Optimization 10. Hyper-Connections — Defa Zhu et al., 2025 https://scholar.google.com/scholar?q=Hyper-Connections 11. mHC-lite: You don't need 20 Sinkhorn-Knopp iterations — Yongyi Yang, Jianyang Gao, 2026 https://scholar.google.com/scholar?q=mHC-lite%3A+You+don%27t+need+20+Sinkhorn-Knopp+iterations 12. Accelerating Birkhoff projection for manifold-constrained hyper-connections — Chenrui Wang, Yixuan Qiu, 2026 https://scholar.google.com/scholar?q=Accelerating+Birkhoff+projection+for+manifold-constrained+hyper-connections 13. oHC: Orthogonal hyper-connections on SO(4) via quaternions — Haoqiang Guo et al., 2026 https://scholar.google.com/scholar?q=oHC%3A+Orthogonal+hyper-connections+on+SO%284%29+via+quaternions 14. JPmHC: Dynamical isometry via orthogonal hyper-connections — Biswa Sengupta, Jinhua Wang, Leo Brunswic, 2026 https://scholar.google.com/scholar?q=JPmHC%3A+Dynamical+isometry+via+orthogonal+hyper-connections 15. Beyond the Birkhoff Polytope: Spectral-Sphere-Constrained Hyper-Connections — Zhaoyi Liu, Haichuan Zhang, Ang Li, 2026 https://scholar.google.com/scholar?q=Beyond+the+Birkhoff+Polytope%3A+Spectral-Sphere-Constrained+Hyper-Connections 16. How does mHC use its residual streams? Selective routing and near-identity mixing — Pengxiang Zhao et al., 2026 https://scholar.google.com/scholar?q=How+does+mHC+use+its+residual+streams%3F+Selective+routing+and+near-identity+mixing 17. Analyzing stream collapse in hyper-connections: From diagnosis to mitigation — Ekaterina Alimaskina, Gleb Molodtsov, Aleksandr Beznosikov, 2026 https://scholar.google.com/scholar?q=Analyzing+stream+collapse+in+hyper-connections%3A+From+diagnosis+to+mitigation 18. Ablate and rescue: A causal analysis of residual stream hyper-connections — William Peng et al., 2026 https://scholar.google.com/scholar?q=Ablate+and+rescue%3A+A+causal+analysis+of+residual+stream+hyper-connections 19. Manifold optimization over the set of doubly stochastic matrices: A second-order geometry — Ahmed Douik, Babak Hassibi, 2019 https://scholar.google.com/scholar?q=Manifold+...
    Más Menos
    Menos de 1 minuto
  • LayerRoPE: Treating Transformer Depth Growth as a Positional Code
    Oct 9 2026
    This episode examines LayerRoPE, a 2026 preprint proposing that the steady growth of hidden-state norms with depth in Transformers is a learned depth-positional code rather than a defect to suppress. It lays out the background: residual-stream accumulation under Pre-Norm, massive activations, high-norm tokens in ViTs, and the Curse of Depth. It also covers remedies such as Layer-Norm Scaling, Peri-Norm, ReZero and LayerScale, plus stream-restructuring approaches like Hyper-Connections. The core idea is to apply RoPE-style rotation along layer index instead of token position, motivated by an analysis of 16 open LLMs in which about 99 percent of norm-weight dimensions share one depth trajectory and about 1 percent supply a rotation. The discussion stays skeptical. Attention-block norms and some model families show no clear depth code, and the observed arc could just be smooth drift between neighbouring layers. A rival explanation is that RMSNorm resets input scale, so growing gains need no positional signal at all. Listeners get a clear look at the competing explanations for norm growth and at the scaling-law methodology (58M to 1.3B parameters at about 80 tokens per parameter) used to judge whether the approach helps. Sources: 1. LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition — Shikhar Srivastava, Christopher Kanan, 2026 http://arxiv.org/abs/2610.09179 2. Root Mean Square Layer Normalization — Biao Zhang, Rico Sennrich, 2019 https://scholar.google.com/scholar?q=Root+Mean+Square+Layer+Normalization 3. RoFormer: Enhanced Transformer with Rotary Position Embedding — Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, Yunfeng Liu, 2021 https://scholar.google.com/scholar?q=RoFormer%3A+Enhanced+Transformer+with+Rotary+Position+Embedding 4. The Curse of Depth in Large Language Models — Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu, 2025 https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models 5. Peri-LN: Revisiting Layer Normalization in the Transformer Architecture — Jeonghoon Kim et al., 2025 https://scholar.google.com/scholar?q=Peri-LN%3A+Revisiting+Layer+Normalization+in+the+Transformer+Architecture 6. Massive Activations in Large Language Models — Mingjie Sun, Xinlei Chen, J. Zico Kolter, Zhuang Liu, 2024 https://scholar.google.com/scholar?q=Massive+Activations+in+Large+Language+Models 7. Vision Transformers Need Registers — Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski, 2024 https://scholar.google.com/scholar?q=Vision+Transformers+Need+Registers 8. Hyper-Connections (and Manifold-Constrained Hyper-Connections, mHC) — Defa Zhu et al. (Hyper-Connections, ByteDance, 2024-25); Zhenda Xie et al. (mHC, DeepSeek, 2025), 2024-2025 https://scholar.google.com/scholar?q=Hyper-Connections+%28and+Manifold-Constrained+Hyper-Connections%2C+mHC%29 9. ReZero is All You Need: Fast Convergence at Large Depth — Bachlechner, Majumder, Mao, Cottrell, McAuley, 2020/2021 https://scholar.google.com/scholar?q=ReZero+is+All+You+Need%3A+Fast+Convergence+at+Large+Depth 10. Going Deeper with Image Transformers (CaiT / LayerScale) — Touvron, Cord, Sablayrolles, Synnaeve, Jégou, 2021 https://scholar.google.com/scholar?q=Going+Deeper+with+Image+Transformers+%28CaiT+%2F+LayerScale%29 11. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models — Wang, Zhu, Fang, Li, Shen, Zhong, 2026 https://scholar.google.com/scholar?q=Negligible+in+Size%2C+Significant+in+Effect%3A+On+Scale+Vectors+in+Large+Language+Models 12. On the Residual Scaling of Looped Transformers: Stability and Transferability — Wang, Li, Zhang, Huang, Yan, Li, 2026 https://scholar.google.com/scholar?q=On+the+Residual+Scaling+of+Looped+Transformers%3A+Stability+and+Transferability 13. Parcae: Scaling Laws for Stable Looped Language Models — Prairie, Novack, Berg-Kirkpatrick, Fu, 2026 https://scholar.google.com/scholar?q=Parcae%3A+Scaling+Laws+for+Stable+Looped+Language+Models 14. On Layer Normalization in the Transformer Architecture — Xiong et al., 2020 https://scholar.google.com/scholar?q=On+Layer+Normalization+in+the+Transformer+Architecture 15. Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks (Depth-μP) — Yang, Yu, Zhu, Hayou, 2024 https://scholar.google.com/scholar?q=Tensor+Programs+VI%3A+Feature+Learning+in+Infinite-Depth+Neural+Networks+%28Depth-%CE%BCP%29 16. Efficient Streaming Language Models with Attention Sinks — Xiao, Tian, Chen, Han, Lewis, 2023 https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks 17. Dissecting Outlier Dynamics in LLM NVFP4 Pretraining — Dong et al., 2026 https://scholar.google.com/scholar?q=Dissecting+Outlier+Dynamics+in+LLM+NVFP4+Pretraining 18. Hyper-Connections / mHC: Manifold-Constrained Hyper-Connections — Zhu et al. 2025; Xie et al. 2025, 2025 https://scholar.google.com/scholar?q=Hyper-Connections+%2F+mHC%3A+...
    Más Menos
    Menos de 1 minuto
  • HBF Sucks? Why Faster Flash Slows KV-Centric LLM Serving
    Oct 8 2026
    This episode examines "HBF Sucks?", a simulation study of what happens when High-Bandwidth Flash (stacked 3D NAND with a wide package-local interface) replaces the SSD tier in a Mooncake-style KV-offloading stack, and why serving gets slower. It first covers the KV cache, prefix reuse, and the near-tier/far-tier split, then HBF's tradeoffs: flash-scale capacity and aggregate bandwidth, but microsecond read latency and costly page-program writes. The paper's central argument is that a faster far tier only pays off under three conditions. Exposed read I/O must be the bottleneck, reads must outweigh writes (useful bytes read per byte written of at least about one), and the delivered bandwidth must be sustainable thermally and for endurance. Transient KV cache is argued to violate all three. The episode also covers the HBF-1 and HBF-2 packaging roadmaps, which cost GPU near-tier memory, and the trace-driven evaluation on Alibaba Qwen-Bailian traces across five models, modeled H100 and B200 systems, and a capacity-matched SSD baseline. It is useful for anyone weighing flash-based memory tiers for LLM serving. Sources: 1. HBF Sucks? A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving — Zhuoran Li, Zhuohang Bian, Xin Huang, Yibo Zhao, Guangyu Sun, Youwei Zhuo, 2026 http://arxiv.org/abs/2608.11668v4 2. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu, Yihua Cheng, et al., 2025 https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference 3. Strata: Hierarchical Context Caching for Long Context Language Model Serving — Zhiqiang Xie, Ziyi Xu, Mark Zhao, et al., 2025 https://scholar.google.com/scholar?q=Strata%3A+Hierarchical+Context+Caching+for+Long+Context+Language+Model+Serving 4. Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving — Shi Qiu, Yifan Hu, et al., 2026 https://scholar.google.com/scholar?q=Tutti%3A+Making+SSD-Backed+KV+Cache+Practical+for+Long-Context+LLM+Serving 5. KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider — Jiahao Wang, Jinbo Han, Xingda Wei, et al., 2025 https://scholar.google.com/scholar?q=KVCache+Cache+in+the+Wild%3A+Characterizing+and+Optimizing+KVCache+Cache+at+a+Large+Cloud+Provider 6. H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference — Minho Ha, Euiseok Kim, Hoshik Kim, 2026 https://scholar.google.com/scholar?q=H3%3A+Hybrid+Architecture+Using+High+Bandwidth+Memory+and+High+Bandwidth+Flash+for+Cost-Efficient+LLM+Inference 7. FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference — Xinyu Wang, Yalong Xue, et al., 2026 https://scholar.google.com/scholar?q=FlashAccel%3A+Leveraging+High-Bandwidth+Flash+for+High-Throughput+LLM+Inference 8. HBFSim: Fast and Faithful Simulation of High-Bandwidth Flash Under Real GPU Execution — Yanpeng Hu, Yiwei Yang, et al., 2026 https://scholar.google.com/scholar?q=HBFSim%3A+Fast+and+Faithful+Simulation+of+High-Bandwidth+Flash+Under+Real+GPU+Execution 9. Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges — Dowon Son, Yonggon Park, et al. (with Onur Mutlu), 2026 https://scholar.google.com/scholar?q=Exploring+High-Bandwidth+Flash+for+Modern+LLM+Inference%3A+Opportunities+and+Challenges 10. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention — Bin Gao, Zhuomin He, et al., 2024 https://scholar.google.com/scholar?q=Cost-Efficient+Large+Language+Model+Serving+for+Multi-turn+Conversations+with+CachedAttention 11. Data Retention in MLC NAND Flash Memory: Characterization, Optimization, and Recovery — Yu Cai, Yixin Luo, Erich Haratsch, Ken Mai, Onur Mutlu, 2015 https://scholar.google.com/scholar?q=Data+Retention+in+MLC+NAND+Flash+Memory%3A+Characterization%2C+Optimization%2C+and+Recovery Interactive Visualization: HBF Sucks? Why Faster Flash Slows KV-Centric LLM Serving
    Más Menos
    Menos de 1 minuto
adbl_web_anon_alc_button_suppression_t1
Todavía no hay opiniones