awq-quantization
Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you…
- Industry
- software-engineering
- License
- Unverified
- Source repo
- Orchestra-Research/AI-Research-SKILLs · ★ 10,718
- Source file
- 10-optimization/awq/SKILL.md
不会安装?看中文图文教程 →