What is quantization and how much does it degrade the result?
FAQQuantization shrinks a model by storing its internal numbers at lower precision. A model with thirty billion parameters takes about sixty gigabytes at full precision and just under twenty at four-bit, so it fits on a single graphics card. The quality loss at four bits is usually small and at more aggressive levels noticeable. The exact figure for your task cannot be estimated, only measured on a test set, and that is part of our work.