ApiaryActiveLive
Try: pause · settings · learn · wipe
← Community / Reading Room
WI
craft · 13 min read

What Is a Quantized Model (Q4, Q8) in Plain English

A quantized model is a smaller copy of an AI model, made by storing its numbers with less precision.

By Austin Little

When you download a model in Ollama or LM Studio, you run into labels like Q4_K_M, Q5_K_S, and Q8_0. They look like part numbers. They're really a short description of how much the model was compressed, and once you can read them, picking the right download stops being a guess.

AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events.

Short answer (read this first)

A quantized model is a smaller copy of an AI model, made by storing its numbers with less precision.

  • Q8 means about 8 bits per number. It's bigger, and closer to the original.
  • Q4 means about 4 bits per number. It's much smaller and easier to run on an ordinary computer, and gives up more fidelity.
  • The letters after the number (_K_M, _0, _S, _L) describe how the compression was done. More on those below.

LM Studio's own docs put it simply. The download options are "all copies of the same model, provided in varying degrees of fidelity," and the Q stands for quantization, "which roughly means compressing model files in size, while giving up some degree of quality." Their advice: "Choose a 4-bit option or higher if your machine is capable enough for running it."

If you just want a starting point: try a Q4_K_M version first. It's the example llama.cpp's own docs use. If it runs comfortably and you want a bit more fidelity, try Q5 or Q6, or Q8 if you have the memory. If it won't fit, go smaller, or pick a smaller model.

The rest of this page explains why.

The idea, with no math

An AI language model is, underneath, a giant pile of numbers called weights. When the model writes a sentence, your computer does a huge amount of arithmetic with those weights. A model with "8B" in its name has about eight billion of them.

When a model is first trained, each weight is usually stored at high precision, for example 16 or 32 bits per number. That's like writing every measurement as 3.14159265 when 3.14 would usually do.

Quantization rounds those numbers off. Instead of storing every weight in 16 bits, you store it in 8, or 4, or even fewer. The file gets much smaller. Your computer has less data to move around, which often makes it faster. And because the rounding is smart (more on that below), the model usually still works well.

Here's a picture that helps. Imagine a photo saved as a huge uncompressed file versus a JPEG. The JPEG is a fraction of the size and, for most uses, you can't tell the difference. Push the compression too far and you start to see blocky artifacts. Quantization is the same deal for AI models. There's a range where it's nearly free, and a point past which you start noticing.

The analogy isn't perfect. A model doesn't get "blurry" the way a photo does. When quantization goes too far, the model tends to make more small mistakes, lose track of instructions, or produce slightly worse writing. But the basic trade, smaller file in exchange for some fidelity, is the same.

What the official docs actually say

We wanted to explain this without making up numbers, so here's what the tools themselves document. Checked October 1, 2026.

llama.cpp is the open-source engine that defined the GGUF file format most local models use. Its quantize tool README says quantization "reduces the precision of model weights (e.g., from 32-bit floats to 4-bit integers), which shrinks the model's size and can speed up inference." It also says this "may introduce some accuracy loss which is usually measured in Perplexity (ppl) and/or Kullback–Leibler Divergence (kld)," and that the loss "can be minimized by using a suitable imatrix file."

LM Studio describes the download options as copies of the same model "in varying degrees of fidelity," and suggests a 4-bit option or higher if your machine can handle it.

Ollama notes in its import docs that it "does not quantize GGUF models during import," and that you should prepare and quantize them first with a tool such as llama.cpp's llama-quantize. In practice, most people never quantize anything themselves. They download a version someone already made.

Notice what none of them say: "Q4 loses exactly X percent of quality." That's on purpose. How much a given quantization hurts depends on the model, the task, and how it was quantized. Anyone who gives you one tidy percentage for "Q4 quality loss" across all models is oversimplifying.

How big is the difference? One real, documented example

llama.cpp's README includes a measured table for one specific model, Meta's Llama 3.1 8B. We're quoting it because it's documented by the people who build the tool, not because it applies to every model. Your model and your hardware will give different numbers.

From that table:

FormatBits per weight (as listed)File size (as listed)
F16 (unquantized 16-bit)16.000514.96 GiB
Q8_08.50087.95 GiB
Q6_K6.56336.14 GiB
Q5_K_M5.70365.33 GiB
Q4_K_M4.89444.58 GiB
Q3_K_M3.99603.74 GiB
Q2_K3.15932.95 GiB

A few things jump out:

  • Q8 is roughly half the size of the 16-bit original, and Q4_K_M is under a third of it, for this model.
  • "Q4" doesn't mean exactly 4 bits. Q4_K_M is listed at about 4.89 bits per weight, because some parts of the model are kept at higher precision and each block of numbers carries a little extra information to help with rounding. The number in the name is a family, not a precise measurement.
  • The same README shows speed varies too. For this model on the hardware they tested, text generation was faster on the quantized versions than on F16. Speed depends heavily on your machine, so treat that as "smaller often means faster," not as a promise.

The same README also gives a size comparison for Llama 3.1 at three sizes, in its memory and disk section. The 8B model goes from 32.1 GB original to 4.9 GB at Q4_K_M. The 70B goes from 280.9 GB to 43.1 GB. The 405B goes from 1,625.1 GB to 249.1 GB.

Reading the label: what Q4_K_M actually means

Let's take one label apart: Q4_K_M.

  • Q means quantized.
  • 4 is the bit family, around 4 bits per weight.
  • K means it's a "k-quant." That's a newer quantization method in llama.cpp that works on blocks of weights and mixes precision across different parts of the model. llama.cpp's README links the original k-quants work for anyone who wants the details.
  • M means medium. Within the k-quant family you'll often see S (small), M (medium), and L (large). These are variants that keep more or fewer parts of the model at higher precision. In llama.cpp's table for Llama 3.1 8B, Q4_K_S is listed at 4.36 GiB and Q4_K_M at 4.58 GiB, so M is a little bigger.

Other labels you'll see:

  • Q8_0 and Q4_0: the _0 marks an older, simpler style of quantization. Q8_0 is still very common, because at 8 bits the simple method works well.
  • IQ labels, like IQ4_XS or IQ2_M: "i-quants," another family in llama.cpp, often used for very small files. llama.cpp's table lists several, down to IQ1_S at about 2 bits per weight.
  • F16 or BF16: not quantized in the usual sense. These are the 16-bit versions, the closest to the original weights in most downloads.

The trade-off, plainly

Every quantization level is a trade between three things:

  1. Memory. The model has to fit in your computer's memory to run well: system RAM, or GPU memory (VRAM) if you have a graphics card. Smaller quantizations fit in less.
  2. Speed. Smaller files usually mean less data to move, which often helps speed. But it depends on your hardware.
  3. Fidelity. How closely the quantized model behaves like the original. Lower bits mean more rounding, which means more drift from the original.

There's no free lunch at the low end. llama.cpp's own option list warns that re-quantizing an already-quantized model "can severely reduce quality compared to quantizing from 16bit or 32bit." Most people won't ever do that, but it's a useful reminder that the rounding adds up.

A rough, honest mental model:

  • Q8: about as close to the original as most people need. Big.
  • Q5 and Q6: a middle ground many people like when they have the room.
  • Q4 (especially Q4_K_M): the common default for running on everyday machines. LM Studio's docs suggest staying at 4-bit or higher.
  • Q3 and Q2: for when nothing else fits. Expect the model to feel noticeably less sharp. Try it before relying on it.

Bigger model at Q4, or smaller model at Q8?

This is the question people actually get stuck on. You have room for roughly one size of file. Should you take a bigger model squeezed down to Q4, or a smaller model kept at Q8?

There's no single rule, and we won't pretend there is. A lot of people find that a larger model at a moderate quantization (like Q4_K_M) does better on reasoning and writing than a much smaller model at high precision. That's one reason Q4 is so popular. But it varies by model family and task, and very aggressive quantization (Q2, Q3) can erase the advantage.

The practical answer: try both on your real work. Give each one the same three or four prompts you actually care about, like an email you need to write, a document you need summarized, or a question in your field. Read the answers side by side.

How to pick one for your computer

Here's a calm, no-guessing process.

Step 1: Know your memory

Find out how much RAM your computer has, and whether it has a separate graphics card with its own memory (VRAM).

Step 2: Compare file size to memory

As a rough guide, the model file has to fit in memory, with room left over for your operating system, your other apps, and the model's working memory for the conversation. That working memory is called the context, and it grows with longer chats.

llama.cpp's README notes that, for its quantize tool, "memory and disk requirements are the same" because models are loaded fully into memory. For running a model, the file size is a decent first estimate of the memory it needs, plus extra for context.

Ollama's FAQ adds a useful detail: running more requests in parallel or using a longer context window increases memory use. If a model barely fits, a long conversation can push it over.

We're not going to give you a table of "8 GB RAM runs X."

Step 3: Start at Q4_K_M and adjust

  • Runs smoothly, with memory to spare? Try a Q5, Q6, or Q8 version of the same model and see if you notice better answers.
  • Runs, but slowly or with the computer struggling? Stay at Q4, or try a smaller model.
  • Won't load, or crashes? Go to a smaller model before you go below Q4. A smaller model at Q4 is often a better experience than a big model at Q2.

Step 4: Check where it's actually running

In Ollama, ollama ps shows loaded models and a Processor column. Per Ollama's FAQ, "100% GPU" means the model fit entirely in GPU memory, "100% CPU" means it's running in system memory, and a split like "48%/52% CPU/GPU" means it's running partly on each. If you have a GPU and see a split, a smaller quantization might let it fit fully on the GPU, which is usually faster.

Where you'll see these labels

In Ollama

Ollama's model library lists different versions of a model as tags. Many tags include a quantization label in the name.

In LM Studio

LM Studio's Discover tab shows several download options for each model, with labels like Q3_K_S and Q8_0.

On Hugging Face

Many people publish GGUF versions of popular models on Hugging Face, often with a whole list of files like model-Q4_K_M.gguf, model-Q5_K_M.gguf, and model-Q8_0.gguf. Same idea: same model, different compression. Stick to publishers you recognize or that the tool's own search surfaces, and read the model card.

Quantizing a model yourself (optional)

You don't need to do this. Most people never will. But if you're curious, it's free, and llama.cpp documents the process:

  1. Convert the original model to GGUF at high precision (for example BF16), using llama.cpp's convert_hf_to_gguf.py script.
  2. Run llama-quantize on that file with the type you want. The README's example uses Q4_K_M:
./build/bin/llama-quantize model-bf16.gguf model-Q4_K_M.gguf Q4_K_M
  1. To use it in Ollama, create a Modelfile with FROM /path/to/model-Q4_K_M.gguf and run ollama create my-model, per Ollama's import docs.

You'll need enough disk space for both the big original and the smaller output. The llama.cpp README is blunt that intermediate files can be large. There's also a Hugging Face space, linked from the same README, called GGUF-my-repo, that can build quants for you without any local setup.

Common myths, gently corrected

"Quantized models are a cheap knockoff." They're the same model, stored more compactly. The model's training, knowledge, and behavior come from the same weights, just rounded. At moderate levels, many people can't tell the difference in everyday use.

"Q8 is twice as good as Q4." No. The number is about storage precision, not quality on a scale. Twice the bits doesn't mean twice the quality. Often the difference is small. Sometimes, on some tasks, it's noticeable.

"There's one correct quantization." There isn't. The right one is the one that fits your machine and does your job well enough. That's why tools offer several.

"Lower bits always means faster." Often, not always. llama.cpp's measured table shows speeds that vary across formats in ways that don't line up neatly with size. Your hardware matters.

"I have to pay for a bigger model to get good results." No. Free local models at sensible quantizations handle a lot of everyday writing, summarizing, and question-answering. Free-tier cloud options exist too. Start free, and only move up when you can name a specific job that failed.

A note on the "KV cache" setting you might stumble on

If you dig into Ollama's settings, you may see a separate option, OLLAMA_KV_CACHE_TYPE, with values like f16, q8_0, and q4_0. That's not the model's quantization. It's the quantization of the model's short-term working memory for the conversation, the "K/V cache."

Ollama's FAQ says q8_0 uses "approximately 1/2 the memory of f16" with "a very small loss in precision," and q4_0 uses "approximately 1/4 the memory of f16" with "a small-medium loss in precision that may be more noticeable at higher context sizes." Ollama says this requires Flash Attention and that the impact depends on the model and task. Leave it at the default unless you're running out of memory on long conversations.

A plain checklist

  • [ ] Know your RAM (and VRAM, if you have a separate graphics card)
  • [ ] Pick a model size that fits comfortably
  • [ ] Start with a Q4_K_M version
  • [ ] Test it on three or four of your real tasks
  • [ ] If there's room, compare a Q5, Q6, or Q8 version on the same tasks
  • [ ] If it doesn't fit, try a smaller model before going below Q4
  • [ ] Use ollama ps (Ollama) to check whether it's running on GPU, CPU, or both
  • [ ] Keep notes on what worked, so you don't have to rediscover it

FAQ

What does Q4 mean on an AI model? It means the model's weights are stored at about 4 bits each, instead of the 16 or 32 bits used in training. The file is much smaller and easier to run, with some loss of fidelity.

Is Q8 better than Q4? Q8 stays closer to the original model and is bigger. Whether you'll notice the difference depends on the model and your tasks. If Q8 fits and runs well on your machine, it's a safe choice. If it doesn't, Q4 is a common, sensible default.

What does the K_M in Q4_K_M mean? K means it's a "k-quant," a llama.cpp method that compresses in blocks and mixes precision across the model. M means the medium variant, which is a bit larger than S (small) and smaller than L (large).

Does quantization change what the model knows? It doesn't add or remove training. It rounds the numbers the model uses. Heavy rounding can make the model less reliable, but it's still the same model.

Which quantization does LM Studio recommend? LM Studio's docs say to choose a 4-bit option or higher if your machine is capable of running it.

Can I quantize a model myself for free? Yes. llama.cpp's llama-quantize tool is open source, and its README walks through the steps. Most people just download a version someone already quantized.

Sources

Checked October 1, 2026:

  • llama.cpp, quantize tool README: https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md
  • Ollama FAQ (GPU/CPU loading, K/V cache quantization): https://docs.ollama.com/faq
  • Ollama, Importing a Model: https://docs.ollama.com/import
  • LM Studio, Download an LLM: https://lmstudio.ai/docs/app/basics/download-model
  • LM Studio, Import Models: https://lmstudio.ai/docs/app/advanced/import-model
Frequently asked
What does Q4 mean on an AI model?
It means the model's weights are stored at about 4 bits each, instead of the 16 or 32 bits used in training. The file is much smaller and easier to run, with some loss of fidelity.
Is Q8 better than Q4?
Q8 stays closer to the original model and is bigger. Whether you'll notice the difference depends on the model and your tasks. If Q8 fits and runs well on your machine, it's a safe choice. If it doesn't, Q4 is a common, sensible default.
What does the K_M in Q4_K_M mean?
K means it's a "k-quant," a llama.cpp method that compresses in blocks and mixes precision across the model. M means the medium variant, which is a bit larger than S (small) and smaller than L (large).
Does quantization change what the model knows?
It doesn't add or remove training. It rounds the numbers the model uses. Heavy rounding can make the model less reliable, but it's still the same model.
Which quantization does LM Studio recommend?
LM Studio's docs say to choose a 4-bit option or higher if your machine is capable of running it.
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room