标签: Linear-Attention
所有带有此标签的文章 "Linear-Attention".
-
Kimi Linear: An Expressive, Efficient Attention Architecture
发布于:4,873 字约 18 分钟Kimi Linear,有比较详细的实验 & Scale Up。有 Linear Attention 可以去掉 RoPE 这个结论还是比较惊喜的。
-
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
发布于:15,596 字约 55 分钟AI Lab 关于” 广义 “LLM 推理加速的工作,包括 Linear Attention,Sparse Attention,Diffusion LLM,Applications 等。
-
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
发布于:3,215 字约 12 分钟DeltaNet