ApiaryActiveLive
Try: pause · settings · learn · wipe
← Community / Reading Room
RQ
craft · 13 min read

Running Qwen or Mistral Locally for Blog Drafts

People ask me which local model to use for writing as if there's one right answer. There isn't. There's a right answer for your machine, your patience, and…

By Austin Little

You don't need a frontier subscription to get a usable first draft. A mid-size open model on your own machine can outline, expand, and tighten English prose well enough to save real time, as long as you pick one that fits your hardware and you never let it be the last pair of eyes.

AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events.

People ask me which local model to use for writing as if there's one right answer. There isn't. There's a right answer for your machine, your patience, and the kind of drafting you do. This guide narrows it down to two families that are easy to run through Ollama and that ship under open licenses: Qwen, from Alibaba's Qwen team, and Mistral, from Mistral AI.

I'll tell you what the official pages actually say about each model: sizes, context windows, licenses. I'll tell you how to choose based on your machine. And I'll give you a drafting workflow and prompts that keep the model useful without letting it invent things.

What I won't do is hand you a leaderboard. I haven't run a controlled benchmark for this article, and I'm not going to pretend otherwise. Vendor benchmark claims appear on model pages, and I'll note them as claims. Your own test on your own writing beats anybody's chart.

Why run a model locally for drafting

Privacy. Drafts are where you're messy. Half-formed opinions, client names, rough notes. A local model never sends any of it anywhere.

No meter. Once the model is downloaded, drafting costs electricity. No subscription, no per-token bill, no free-tier cap that cuts you off mid-draft.

Stability. Cloud models get swapped, retuned, or retired. A model file on your disk behaves the same next month as it did today, unless you update it.

Learning. Running models locally teaches you what they're actually good and bad at. You stop treating them like oracles.

The tradeoffs are real. Local models are smaller than the biggest cloud models, and they run at whatever speed your hardware allows. For first drafts, outlines, and tightening, that's often fine. For hard reasoning or deep research, it may not be.

The setup in brief

If you haven't installed Ollama yet, the Apiary guide on using AI without paying walks through it. The short version: download Ollama for your operating system from ollama.com, install it, open a terminal, and run a model:

ollama run qwen3:4b

The first run downloads the model. After that, it starts from disk. Ollama also serves a local API on localhost:11434, which you can use from scripts and editors.

One caution: Ollama also offers cloud-hosted models through its own service.

How to think about model size

Model names include a parameter count like 4b, 8b, or 14b. That's billions of parameters. Bigger models generally know more and follow complicated instructions better, but they need more memory and run slower.

The Ollama library lists a download size for each tag. That size is a useful rough guide to how heavy a model is, but it isn't the full memory you'll need while running. Context length adds memory on top.

A practical way to choose:

  1. Start with the smallest model in the family that isn't a toy. For drafting, that's usually around 4B.
  2. Draft one real piece with it.
  3. If it loses the thread, ignores instructions, or writes mush, step up one size.
  4. If your machine starts freezing, swapping, or the model crawls, step down.

To check whether a model is running on your GPU, CPU, or split between them, Ollama's FAQ points to ollama ps. The PROCESSOR column shows "100% GPU," "100% CPU," or a split. If a model is mostly on CPU on a modest machine, a smaller model will usually feel much better.

The Qwen family

Qwen3

The Ollama library describes Qwen3 as the latest generation in the Qwen series at the time the page was written, "offering a comprehensive suite of dense and mixture-of-experts (MoE) models." Here are the tags as listed on the Ollama library page when we fetched it:

The Qwen3-8B model card on Hugging Face lists it under the Apache 2.0 license, with 8.2B parameters and a context length of 32,768 tokens natively, extendable to 131,072 tokens with a technique called YaRN.

Thinking mode. Qwen3's model card says it supports switching between a "thinking mode" for complex reasoning and a "non-thinking mode" for general dialogue, within one model, and that thinking is enabled by default in the Transformers chat template. In thinking mode, the model writes a reasoning trace before its answer.

For blog drafting, you usually don't want that trace. It's slower and you have to skip past it. Ollama's thinking docs describe a think field on API requests: true requests thinking output, false requests no thinking output if the model permits it. Through the API:

curl http://localhost:11434/api/chat -d '{
  "model": "qwen3:8b",
  "messages": [{"role": "user", "content": "Tighten this paragraph: ..."}],
  "think": false,
  "stream": false
}'

**Vendor claims.** The Ollama readme for Qwen3 includes the Qwen team's claims, including that Qwen3 has "superior human preference alignment, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following," and that even Qwen3-4B "can rival the performance of Qwen2.5-72B-Instruct." Those are the developer's claims, not our measurements. Test it yourself.

### Qwen3.5

The Ollama library also lists **Qwen3.5**, described as "a family of open-source multimodal models." Tags listed at our fetch included `qwen3.5:0.8b` (1.0GB), `qwen3.5:2b` (2.7GB), `qwen3.5:4b` (3.4GB), `qwen3.5:9b` (6.6GB, the `latest` tag), `qwen3.5:27b` (17GB), `qwen3.5:35b` (24GB), and `qwen3.5:122b` (81GB), each listed with a 256K context window and text-plus-image input. The page also lists MLX variants for Apple Silicon.

Multimodal means it can look at images as well as read text.

### Qwen2.5

Older, but still on the shelf. The Ollama library lists Qwen2.5 tags from `qwen2.5:0.5b` (398MB) to `qwen2.5:72b` (47GB), with `qwen2.5:7b` (4.7GB) as `latest`, and a 32K context window per tag. If a tutorial you're following names Qwen2.5, it'll still work. For new setups, I'd start with a newer family.

## The Mistral family

### Mistral 7B (`mistral`)

The plain `mistral` tag on Ollama is Mistral 7B, updated to version 0.3. The library page lists it at 4.4GB with a 32K context window and says it's "distributed with the Apache license" and available in instruct and text-completion variants. The version table on that page dates v0.3 to May 2024.

It's old by AI standards, but it's small, well understood, and still perfectly capable of turning an outline into readable paragraphs. If your machine is modest, it's a reasonable first try.

The library readme repeats Mistral AI's original claims, like outperforming Llama 2 13B on all benchmarks. Those claims compare it to older models. Treat them as history, not a buying guide.

### Mistral NeMo (`mistral-nemo`)

The Ollama library describes Mistral NeMo as a 12B model built by Mistral AI in collaboration with NVIDIA. The `mistral-nemo:12b` tag is listed at 7.1GB. The readme says it offers "a large context window of up to 128k tokens."

For blog drafting, NeMo is a nice middle step: bigger than 7B, smaller than the 24B models. If Mistral 7B feels thin and your machine has headroom, try NeMo next.

### Mistral Small (`mistral-small`, `mistral-small3.2`)

The `mistral-small` tags on Ollama include `22b` (13GB) and `24b` (14GB, the `latest` tag). The readme describes Mistral Small 3 as having 24B parameters, released under Apache 2.0, and says it can fit "in a single RTX 4090 or a 32GB RAM MacBook once quantized." That's Mistral's claim, quoted from the readme, not our test.

There's also `mistral-small3.2:24b`, listed at 15GB with a 128K context window and text-plus-image input. Its readme says Small 3.2 improves instruction following, produces fewer "infinite generations or repetitive answers," and has a more robust function-calling template compared to Small 3.1.

For drafting, "fewer repetitive answers" and "better instruction following" are exactly the things you care about. If your hardware can carry a 24B model, Small 3.2 is worth a try.

### Ministral 3 (`ministral-3`)

The Ollama library lists Ministral 3 in `3b` (3.0GB), `8b` (6.0GB, the `latest` tag), and `14b` (9.1GB), each with a 256K context window and text-plus-image input. The readme describes the family as designed for edge deployment and lists an Apache 2.0 license.

## So which one should you pick?

I won't rank them. I haven't run the tests to back a ranking. Instead, here's how to choose by situation.

**"My laptop is modest, and I just want to try it."**
Start with `qwen3:4b` or `mistral` (7B). Both are on the small end of useful. Draft one real piece. If the output is mushy or ignores instructions, step up.

**"I have some headroom and want a better default."**
Try `qwen3:8b` or `mistral-nemo`. Compare them on the same outline. Keep whichever sounds less like a brochure.

**"I have a strong machine and care about long drafts."**
Try `mistral-small3.2` or `qwen3:14b`. Bigger models tend to hold a long outline together better. Watch your memory with `ollama ps`.

**"I want one model for drafting and looking at screenshots."**
Look at the multimodal tags: `qwen3.5`, `ministral-3`, or `mistral-small3.2`. Check the size against your hardware first.

**"I need it to be fast more than smart."**
Go small. A fast small model you actually use beats a big slow one you avoid.

## A bake-off you can run in an hour

Here's how to choose for real, on your own writing.

1. **Pick one outline** for a piece you actually need to write. Five to seven headings, a few bullet notes under each.
2. **Pull two or three candidate models** that fit your hardware.
3. **Run the same prompt** (below) on each, with the same outline.
4. **Read the outputs blind if you can.** Have someone paste them into a doc labeled A, B, C.
5. **Score each on four things:**
   - Did it stick to my outline, or wander?
   - Did it invent facts, numbers, names, or quotes?
   - How much editing would it take to sound like me?
   - Did it repeat itself?
6. **Pick the winner.** Note the model tag and date in a text file so you remember.

Repeat this every few months, or whenever a new model shows up in the library. Model names rotate fast.

## The drafting workflow

### 1. You write the outline

Don't ask the model to decide what the piece is about. You decide. Write the headings and the key points. Put the facts you know under each heading, with sources.

### 2. The model expands one section at a time

Long one-shot generations drift. Smaller chunks stay on track and are easier to check.

### 3. You check every claim

Anything the model adds that wasn't in your outline is suspect until proven otherwise.

### 4. The model tightens

Once you've got real content, a model is good at cutting wordiness. Feed it a paragraph and ask for a tighter version that keeps the meaning.

### 5. You put your voice back

Models smooth everything toward the same neutral tone. Your job is to make it sound like a person.

## Prompts that keep a local model honest

**Section expansion**

You are helping draft a blog section in plain, direct English. Heading: <heading> My notes: <notes>

Write 2–4 short paragraphs using ONLY the information in my notes.

  • Do not add statistics, studies, quotes, names, product names, or prices.
  • If my notes are too thin to fill a paragraph, write [THIN: needs more]

instead of padding.

  • No intro like "In today's world". Start with the point.

**Tightening**

Tighten this paragraph. Cut filler and repetition. Keep every fact and my meaning. Do not add anything new. Return only the revised paragraph.

<paragraph>


**Voice pass**

Here are three paragraphs I wrote myself, for voice reference: <your writing>

Rewrite the draft below to match that voice: sentence length, word choice, directness. Do not change facts. Do not add facts.

<draft>


**Fact flagging**

List every factual claim in the text below: numbers, names, dates, product features, prices, cause-and-effect claims. For each one, say whether it appears in my notes (provided) or not. Do not judge whether claims are true; just tell me which ones came from my notes.

NOTES: <notes> TEXT: <draft>


That last prompt is surprisingly useful. Local models aren't good fact checkers, but they're decent at comparing two texts and listing what's in one and not the other. It gives you a checklist of claims to verify.

## Context length, and why your draft got cut off

If you paste a long outline plus reference notes and the model seems to forget the beginning, you've probably hit the context limit.

Ollama's docs describe context length as the maximum number of tokens the model has access to in memory. You can change it:

- In the Ollama app, with a context-length slider in settings, per the context-length doc.
- With the `OLLAMA_CONTEXT_LENGTH` environment variable when serving.
- In an interactive session with `/set parameter num_ctx <number>`, per the FAQ.

A larger context uses more memory. The model's advertised maximum (like 128K for NeMo or 256K for some Qwen tags) is the ceiling, not what you'll get by default, and not necessarily what your machine can hold.

The simpler fix: work section by section. Paste only the heading, its notes, and maybe the previous section for continuity.

## Things local models do badly in drafting

Be honest with yourself about these:

- **Facts.** They'll produce plausible-sounding numbers, studies, and quotes. Never trust a fact the model added.
- **Current events.** A local model knows nothing after its training data. It can't look anything up unless you wire in a tool.
- **Your voice.** Out of the box, they write like a polite brochure. You have to drag them toward you.
- **Long-range structure.** Smaller models lose the thread on long outputs. Go section by section.
- **Repetition.** Some models loop. Mistral's own readme for Small 3.2 lists fewer repetitive answers as an improvement, which tells you it was a known issue.
- **Knowing when to stop.** Asked for 1,000 words from 200 words of notes, a model will invent 800 words. Ask for [THIN] markers instead.

## Licensing, briefly

Most of the models above list Apache 2.0 on their official pages: Qwen3-8B on Hugging Face, Mistral 7B on the Ollama library page, Mistral NeMo on Hugging Face, and Mistral Small 3 and Ministral 3 in their Ollama readmes. Apache 2.0 is a permissive open-source license.

Check the license on the specific model you use, though, especially before using a model's output in a commercial product.

## A minimal daily setup

Once you've picked a model, here's what a drafting session looks like:

1. Open your outline in your editor.
2. Open a terminal: `ollama run <your-model>`.
3. For each section: paste the expansion prompt with that section's notes. Copy the result into your draft.
4. Search the draft for [THIN]. Fill those yourself or cut them.
5. Run the fact-flagging prompt on the full draft. Verify each claim that didn't come from your notes, or delete it.
6. Run the tightening prompt on any flabby paragraph.
7. Read the whole thing out loud. Fix anything that doesn't sound like you.

## Short answers

**Is Qwen or Mistral better for English blog writing?** I haven't run a controlled test to say. Both families have models that can draft readable English. Run the one-hour bake-off on your own outline and keep the one you edit least.

**What's the smallest model worth using for drafts?** For most people, around 4B: `qwen3:4b` is listed at 2.5GB on Ollama. Below that, models get noticeably weaker at following detailed instructions. Try it and see.

**Do I need a GPU?** No, but it helps a lot. Ollama runs on CPU too. Use `ollama ps` to see where your model is running.

**Can I turn off Qwen3's thinking output?** Ollama's API supports a `think` field; set it to `false` to request no thinking output if the model permits it.

**Are these models free to use?** They're free to download and run locally through Ollama. Several list the Apache 2.0 license on their official pages. Check the specific model's license for your use.

## The bottom line

Pick the smallest Qwen or Mistral model your machine runs comfortably. Start around 4B to 8B. Write the outline yourself, expand one section at a time, tell the model to flag gaps instead of filling them, and check every claim it adds. Run a quick bake-off on your own writing instead of trusting anyone's leaderboard, including mine, since I didn't give you one. The model drafts. You publish.

## References

- Ollama library: qwen3 — https://ollama.com/library/qwen3 (fetched 2026-10-01)
- Ollama library: qwen3.5 — https://ollama.com/library/qwen3.5 (fetched 2026-10-01)
- Ollama library: qwen2.5 — https://ollama.com/library/qwen2.5 (fetched 2026-10-01)
- Ollama library: mistral — https://ollama.com/library/mistral (fetched 2026-10-01)
- Ollama library: mistral-nemo — https://ollama.com/library/mistral-nemo (fetched 2026-10-01)
- Ollama library: mistral-small — https://ollama.com/library/mistral-small (fetched 2026-10-01)
- Ollama library: mistral-small3.2 — https://ollama.com/library/mistral-small3.2 (fetched 2026-10-01)
- Ollama library: ministral-3 — https://ollama.com/library/ministral-3 (fetched 2026-10-01)
- Qwen/Qwen3-8B model card — https://huggingface.co/Qwen/Qwen3-8B (fetched 2026-10-01)
- mistralai/Mistral-Nemo-Instruct-2407 model card — https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407 (fetched 2026-10-01)
- Ollama FAQ — https://docs.ollama.com/faq (fetched 2026-10-01)
- Ollama context length — https://docs.ollama.com/context-length (fetched 2026-10-01)
- Ollama thinking — https://docs.ollama.com/capabilities/thinking (fetched 2026-10-01)
Frequently asked
What is Running Qwen or Mistral Locally for Blog Drafts about?
People ask me which local model to use for writing as if there's one right answer. There isn't. There's a right answer for your machine, your patience, and…
What should you know about why run a model locally for drafting?
Privacy. Drafts are where you're messy. Half-formed opinions, client names, rough notes. A local model never sends any of it anywhere.
What should you know about the setup in brief?
If you haven't installed Ollama yet, the Apiary guide on using AI without paying walks through it. The short version: download Ollama for your operating system from ollama.com, install it, open a terminal, and run a model:
What should you know about how to think about model size?
Model names include a parameter count like 4b , 8b , or 14b . That's billions of parameters. Bigger models generally know more and follow complicated instructions better, but they need more memory and run slower.
What should you know about qwen3?
The Ollama library describes Qwen3 as the latest generation in the Qwen series at the time the page was written, "offering a comprehensive suite of dense and mixture-of-experts (MoE) models." Here are the tags as listed on the Ollama library page when we fetched it:
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room