标签: 模型量化
所有带有此标签的文章 "模型量化".
-
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
发布于:1,219 字约 5 分钟LLM 的 Interger-Only PTQ 量化工作。
-
I-ViT: Integer-only Quantization for Efficient Vision Transformer Inference
发布于:1,272 字约 5 分钟对 ViT 的纯整型量化,W8A8,中科院 2023 ICCV
-
Efficient and Effective Methods for Mixed Precision Neural Network Quantization for Faster, Energy-efficient Inference
发布于:1,352 字约 5 分钟EAGL,声称只要用 CPU 在 3 秒内就能完成对 ResNet 的量化,效率远高于 HAWQ 等其他传统的方法
-
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
发布于:1,378 字约 6 分钟谷歌的,第一篇完整跑通 interger-only 量化推理流程的工作。
-
HAWQ: Hessian Aware Quantization of Neural Networks with Mixed-Precision
发布于:1,795 字约 7 分钟模型量化经典方法,基于黑森矩阵,一种二阶信息的量化方法。