搜索:GPTQ

共命中 5 条(服务端检索)
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost · 9-19 阅读 41·访客 38
一篇读懂模型量化:70B 是怎么塞进一张消费级显卡的
70B 模型 FP16 要 140 GB 显存,量化到 4-bit 只剩 35 GB——一张 RTX 4090 就装得下,代价是平均每个权重从 2 字节压到 0.5 字节后的精度损失。本文拆解 GPTQ 的二阶误差补偿、AWQ 的激活感知保护、GGUF k-quants 的命名规则,以及「量化伤推理」争议的实测边界。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 1·访客 1
模型量化的精度账:从 FP16 到 INT4 该怎么选
从 FP16 压到 INT4,体积与显存降到约四分之一,量化是大模型部署的标配。本文讲清量化原理、PTQ 与 QAT 两条路线、RTN/GPTQ/AWQ 的思想差异,以及按显存逐级下降的选型决策表。
原创 大模型 精选 · 原创 · 3天前 阅读 5·访客 5
G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large lan…
大模型 HuggingFace Daily Papers · 9-25 阅读 9·访客 9
Softmax Reparameterization for Output-Head Quantization
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparamet…
行业动态 HuggingFace Daily Papers · 9-25 阅读 4·访客 4