Qwen3.8 27B on 16GB VRAM: benchmarking 21 quantized variants
A benchmark of 21 quantized Qwen3.8 27B variants on an RTX 5080 with 16GB VRAM shows how much GGUF choice matters. The bartowski IQ4_XS variant offers the lowest mean KL divergence and highest same-top-p agreement, while more aggressive quants prove ...