Separate precision from packaging
A filename can mix the base model, size, training variant, numerical representation and file format. Read them as separate choices. GGUF is not a quality level. Q4 is not a complete model name. Identify the underlying model and confirm runtime support for its architecture and format before downloading.
Hugging Face describes GGUF as a binary format containing tensors and standardised metadata. It can hold values in different representations. The extension identifies packaging, but does not prove every application can execute the file.
FP16: sixteen-bit floating point
FP16 is a sixteen-bit floating-point representation. BF16 is also sixteen bits but has a different layout and numerical range; the labels are not interchangeable instructions. Software can store one representation and use different arithmetic for some operations. Follow the model's documented loading configuration rather than selecting precision by name alone.
A dense model with eight billion stored parameters would need roughly 16 billion bytes for sixteen-bit weights: 8 billion × 16 ÷ 8. That is approximately 16 GB in decimal units, or 14.9 GiB. This is weight-only arithmetic, not a claim it fits an advertised 16 GB graphics card.
Q8 and Q4: lower-precision variants
In common quantised-model names, Q8 and Q4 indicate representations around eight and four bits respectively for relevant weights. Less precision saves space. The llama.cpp quantizer documentation explains potential accuracy loss. Its example measurements are not universal quality guarantees.
The simplified weight-only estimates for the same example are 8 GB and 4 GB. Real schemes include scaling information, blocks and sometimes higher-precision tensors, so actual files differ. Suffixes such as _K and variant letters matter; consult the conversion author's description of the exact file.
Smaller does not automatically mean faster
Lower precision can make a model easier to fit. Speed also depends on supported kernels, processor bandwidth, context and CPU/GPU transfers. Quantisation is a trade-off, not a promise Q4 beats every Q8 configuration. Hugging Face explains quantisation concepts.
Choose a compatible variant with headroom for runtime work. Compare outputs on representative tasks: your writing format, extraction fields or code requirements. Save the inputs for repeatability. Judge errors that matter to the job rather than relying on one attractive reply.
Keep a record of the choice
Record the base-model identifier, variant filename, source, licence, runtime version and context setting together. If the app rejects a file, check format and architecture support first. If it runs but gives poor results, inspect prompt formatting and task suitability as well as precision.
Avoid mixing a new model, new runtime and new precision in one comparison. An unchanged baseline helps identify the cause of a difference. This page explains labels and estimates; it supplies no executed comparison or endorsement of a particular quantised model.
Sources
Primary references checked 11 October 2026. This explanation is not a benchmark of a particular model or computer.