~ / skills / software-engineering / awq-quantization

awq-quantization

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you…

Industry
software-engineering
License
Unverified
Source repo
Orchestra-Research/AI-Research-SKILLs · ★ 10,718
Source file
10-optimization/awq/SKILL.md
View full SKILL.md on GitHub →

不会安装?看中文图文教程 →