ApiaryActiveLive
Try: pause · settings · learn · wipe
← Community / Reading Room
OO
craft · 14 min read

Ollama on a Mac Mini for Home Writing: Realistic Setup

I like the Mac mini for this job because it's small, it's quiet, and it's the kind of computer people already have on a desk or can share in a household.…

By Austin Little

A Mac mini on a shelf can be a quiet, private writing assistant that never sends your drafts anywhere. It won't be a frontier model, and it doesn't have to be. This is the setup I'd hand a friend who writes for a living or for love: official steps, honest limits, and no made-up benchmark numbers.

AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events.

I like the Mac mini for this job because it's small, it's quiet, and it's the kind of computer people already have on a desk or can share in a household. Ollama is free, open-source software for running language models on your own machine. Put the two together and you get a writing helper that works without a subscription and without sending your manuscript to a server.

I'm going to stick to what the official Ollama docs and the Ollama model library say today. You'll notice I don't give you a "this Mac mini runs this model at this many words per second" table. I haven't measured it on your machine, and neither has the forum post that claims to know. I'll show you how to measure it yourself in about five minutes instead.

What you're actually building

The setup has three parts:

  1. The Ollama app on macOS, which runs a small local server on your Mac.
  2. One or two models downloaded from the Ollama library and stored on your disk.
  3. A way to talk to it: the terminal, the Ollama app, or a writing tool that can connect to a local Ollama server.

Everything happens on the Mac. By default, Ollama listens only on the Mac itself, not on your home network. The FAQ says Ollama binds to 127.0.0.1 on port 11434 by default. That means other devices in your house can't reach it unless you change that on purpose.

What it's good at for writers

Local models are good at the everyday parts of writing:

  • Tightening a paragraph you already wrote
  • Offering three alternate headlines or openings
  • Turning rough notes into an outline
  • Summarizing your own long notes so you can find the thread
  • Reading a draft back to you as "what is this piece arguing?"
  • Catching repeated words and clunky sentences

What it's not good at

  • Facts. A local model doesn't know what happened last week, and it can state wrong things confidently. Don't use it as a source.
  • Citations. Never trust a reference a model gives you without checking it. Apiary has a separate guide on how to check AI citations.
  • Your voice, out of the box. It'll write in a generic voice until you show it yours. More on that below.

Before you start: check the requirements

The Ollama macOS docs list two system requirements:

  • macOS Sonoma (version 14) or newer
  • Apple M series (CPU and GPU support) or x86 (CPU only)

So an Apple silicon Mac mini can use its GPU for models. An older Intel Mac mini can still run Ollama, but on the CPU only, which the docs make clear. To check your Mac, open the Apple menu and choose About This Mac. It shows the chip and the macOS version.

Memory: the honest answer

Here's where most setup guides make up a number. I won't. The Ollama docs don't publish a "you need this much memory for this model" table for Macs, and I'm not going to invent one.

What the docs do tell you is useful:

  • The FAQ explains that the ollama ps command shows each loaded model's size and whether it's running on the GPU, the CPU, or split between them. That's how you find out what fits.
  • The FAQ also explains that more parallel requests and longer context use more memory, and that required memory scales with the number of parallel requests times the context length.

So the realistic approach is: start with a small model, load it, and look at ollama ps. If it says the model is fully on the GPU and your Mac feels responsive, you have room.

Disk: plan for more than you think

The macOS docs say plainly that after installing Ollama you'll need additional space for models, which "can be tens to hundreds of GB in size." If your Mac mini's internal drive is tight, you can store models elsewhere. The FAQ says to set the OLLAMA_MODELS environment variable to a different folder. On macOS, models live in ~/.ollama/models by default.

Step 1: Install Ollama the official way

The macOS docs say the preferred method is to mount ollama.dmg and drag Ollama into the system-wide Applications folder.

  1. Download Ollama from the official site, ollama.com.
  2. Open the downloaded .dmg.
  3. Drag the Ollama app into Applications.
  4. Open Ollama from Applications.

On first start, the docs say the Ollama app checks whether the ollama command-line tool is in your PATH. If it isn't, it asks permission to create a link in /usr/local/bin. Say yes if you want to use the terminal. You'll probably want to.

To confirm it worked, open Terminal and run:

ollama ls

That lists the models you've downloaded. Right after installing, it should be empty.

Should it start at login?

The FAQ says Ollama on macOS registers as a login item when you install it. If your Mac mini is a dedicated writing box, leaving it on is convenient. If you'd rather start it by hand, open System Settings, search for Login Items, find Ollama under Allow in the Background, and turn it off. The FAQ says Ollama respects that setting across upgrades.

Updates

The FAQ says Ollama on macOS downloads updates automatically. Click the menu bar icon and choose Restart to update to apply one. You can also download the latest version by hand.

Step 2: Turn on local-only mode (optional, recommended for private drafts)

Ollama now has cloud features too: cloud-hosted models and web search. If your goal is that your drafts never leave your Mac, you can turn those off.

The FAQ describes it two ways:

  • Set "disable_ollama_cloud": true in ~/.ollama/server.json
  • Or set the environment variable OLLAMA_NO_CLOUD=1

Then restart Ollama. The FAQ says the logs will show Ollama cloud disabled: true once it's working. You'll lose access to cloud models and web search, which is the point.

On privacy, the FAQ says: "Ollama runs locally. We don't see your prompts or data when you run locally." It also describes what happens if you do use cloud models: prompts and responses are processed to provide the service, but per the FAQ they're not stored or logged and aren't used for training. If you want zero doubt, use local-only mode.

Setting environment variables on a Mac

This trips people up. The FAQ says that when Ollama runs as the macOS app, environment variables should be set with launchctl, then Ollama should be restarted. For example:

launchctl setenv OLLAMA_NO_CLOUD 1

Then quit and reopen the Ollama app. Setting a variable in your terminal profile alone won't reach the app.

Step 3: Pick a first model by reading the library page

Open the Ollama library. It's a long list, and it changes often. On the day I checked, several entries had been updated within the past day or week. Don't pick from a list in an article, including this one. Learn to read the page.

How to read a model listing

Each entry shows:

  • Name, like gemma4, qwen3, llama3.2, or mistral
  • Tags that tell you what it can do: tools, thinking, vision, embedding, or cloud
  • Sizes, shown as parameter counts like 1b, 4b, 8b, 27b. Smaller numbers mean smaller models.
  • Pull counts and when it was last updated

Click into a model and you'll see each tag's file size and context window. For example, the gemma4 page on the day I checked listed tags from small edge models up to larger workstation models, each with a download size and a context window of 128K or 256K tokens.

Two cautions:

  • File size isn't the same as memory needed to run it. A model also needs room for its working context. That's why ollama ps is the real test.
  • Tags marked cloud run on Ollama's servers, not your Mac. If you've turned on local-only mode, skip them.

A sensible first pick

For writing, start with a small general-purpose instruct model from a family you recognize, at one of the smaller sizes. Pull it:

ollama pull <model>:<tag>

Then run it:

ollama run <model>:<tag>

Use the exact name and tag from the library page. Many model pages also list a latest tag and show which size it points to, so check that before you pull a bare name.

Start small even if you think your Mac can handle more. You're learning the workflow first. A smaller model that answers quickly is more useful for drafting than a large one you wait on.

Step 4: Measure what your Mac can actually do

This is the step that replaces every made-up speed table on the internet.

  1. Run your model and ask it to rewrite a paragraph you wrote.
  2. While it's still loaded, open a second Terminal window and run:
ollama ps

The FAQ shows that this lists each loaded model with its NAME, ID, SIZE, PROCESSOR, and UNTIL columns. The PROCESSOR column tells you:

  • 100% GPU: the model fit entirely in GPU memory. Good.
  • 100% CPU: it's running in system memory on the CPU.
  • A split, like 48%/52% CPU/GPU: part of it didn't fit on the GPU.

On an Apple silicon Mac mini, you want 100% GPU for the snappiest experience. If you see a split, try a smaller tag. If you're on an Intel Mac mini, you'll see CPU, because that's what the docs say is supported.

  1. Notice how the Mac feels. Can you still switch apps? Does the fan spin up? That matters more than any number.

Write down what you tried and what ollama ps said. That's your benchmark.

Step 5: Make it a writing assistant, not a chatbot

Out of the box, a model answers like a generic assistant. A few small changes make it a much better writing partner.

Give it a standing brief with a Modelfile

Ollama lets you build a custom version of a model with a short text file called a Modelfile. It doesn't retrain anything. It just saves your settings and instructions so you don't retype them.

Here's a simple one for editing:

FROM <model>:<tag>
PARAMETER temperature 0.6
SYSTEM """You are a careful line editor for a working writer. Keep the writer's voice. Prefer short sentences and plain words. Never add facts, names, numbers, or quotes that are not in the original. When you suggest a change, show the revised text and one sentence on why."""

Save it as Modelfile, then run:

ollama create my-editor -f Modelfile
ollama run my-editor

The Modelfile docs describe each piece:

  • FROM sets the base model. It's required.
  • PARAMETER temperature controls how adventurous the output is. The docs describe higher values as more creative and lower values as more coherent.
  • SYSTEM sets the standing instruction.

The docs also describe MESSAGE, which lets you include example exchanges. That's a powerful way to show a model your voice: put in a short paragraph you wrote and a sample of how you'd want it edited.

Context length: how much it can read at once

For writers, context length matters. It's how much text the model can consider at once.

The FAQ says Ollama's default context window is 4096 tokens, and that you can change it:

  • For everything, with the OLLAMA_CONTEXT_LENGTH environment variable
  • In a chat session, with /set parameter num_ctx <number>

Bigger context uses more memory. The FAQ is explicit that memory scales with context. If you raise it and ollama ps starts showing a CPU/GPU split, bring it back down. For long manuscripts, it's usually better to work chapter by chapter anyway. You'll get more focused feedback.

Keep the model warm while you write

The FAQ says models stay loaded in memory for 5 minutes by default after use, then unload. If you write in bursts, that means the first request after a break can take longer while the model reloads.

You can change this with the OLLAMA_KEEP_ALIVE environment variable, or with the keep_alive setting in API calls. The FAQ says it accepts a duration like "10m" or "24h", a number of seconds, a negative number to keep the model loaded indefinitely, or 0 to unload right away. To free memory when you're done:

ollama stop <model>

On a shared family Mac, I'd keep the default or something modest, so the model isn't holding memory while someone else is using the computer.

Step 6: Connect a writing tool (optional)

You don't have to write in the terminal. The Ollama API docs list a local server at http://localhost:11434/api, and an OpenAI-compatible endpoint at http://localhost:11434/v1. Local requests don't need an API key.

That matters because many writing apps, editor plugins, and note tools let you point them at a custom endpoint. If a tool supports "OpenAI-compatible" or "Ollama" as a provider, you can usually give it http://localhost:11434/v1 and the model name, and it'll work entirely on your Mac.

A few cautions:

  • Read the tool's privacy policy. Some apps send text to their own servers before calling your model. The point of local is lost if the wrapper isn't local too.
  • Browser extensions need permission. The FAQ says Ollama allows cross-origin requests from 127.0.0.1 and 0.0.0.0 by default, and that browser extensions need their origins allowed with OLLAMA_ORIGINS. Only allow extensions you trust.
  • Don't expose it to the internet. The FAQ shows how to change the bind address with OLLAMA_HOST, and how to use tunnels. For a home writing setup, you almost certainly don't need any of that. If you want another computer in the house to use the Mac mini's model, change it deliberately and understand that anything on your network can then reach it.

A realistic writing day with a local model

Here's how I'd actually use this, in order.

Morning: outline. Paste your notes and ask for three possible outlines. Pick one and rewrite it yourself. The model's outline is a menu, not a decision.

Midday: draft. Write the draft yourself. This is the part that's yours. If you get stuck, ask the model for five ways to start the next paragraph, and then write your own sixth.

Afternoon: edit. Run your my-editor model on one section at a time. Accept changes that sound like you. Reject the rest. Watch especially for anything it added that wasn't in your draft, like a fact, a name, a number, or a quote. Delete it unless you can verify it.

Evening: read-back. Ask the model, "In two sentences, what is this piece arguing, and where does it lose the thread?" It's a cheap second reader. Then read it out loud yourself, which is still the best editor there is.

Housekeeping

See what you have, remove what you don't use

From the CLI reference:

ollama ls          # list downloaded models
ollama ps          # list running models
ollama rm <model>  # remove a model

Models are large. If you tried three and only use one, remove the other two.

Where things live

The macOS docs list:

  • ~/.ollama: models and configuration
  • ~/.ollama/logs: logs. app.log is the app, and server.log is the server.

If something isn't working, server.log is the first place to look.

Uninstalling cleanly

The macOS docs give a full list of files and folders to remove, including the app in Applications, the CLI link in /usr/local/bin, several Library folders, and ~/.ollama. If you ever want to remove Ollama, use that list from the docs rather than guessing. Deleting ~/.ollama deletes your downloaded models.

Safety lines for a home setup

  • Local doesn't mean correct. Treat anything factual a model says as a lead to check.
  • Don't paste passwords, financial account numbers, or medical records into any model you haven't thought carefully about, even a local one.
  • Be careful with "agent" tools. Some integrations can run commands on your computer. A chat box that edits text is a different risk from a tool that can change files. Know which one you installed.
  • Mind the household. If the Mac mini is shared, remember that chat history in some front-end apps may be visible to whoever uses the computer.

FAQ

Can a Mac mini run Ollama?

Yes, if it meets the requirements in the official macOS docs: macOS Sonoma (14) or newer, with Apple M series (CPU and GPU) or x86 (CPU only).

Which model should I use for writing?

Start with a small, general-purpose instruct model from the Ollama library, at one of the smaller sizes. Load it, check ollama ps, and step up only if it fits fully on the GPU and you want better output.

How much memory do I need?

The official docs don't publish a Mac memory table, so I won't give you one. Use ollama ps to see each model's loaded size and whether it fits on the GPU. Longer context and more parallel requests use more memory.

Will my writing be sent to the internet?

Not when you run local models. The Ollama FAQ says it doesn't see your prompts or data when you run locally. For extra certainty, turn on local-only mode with OLLAMA_NO_CLOUD=1 or disable_ollama_cloud in server.json.

Why is the first answer after a break slow?

Models unload after 5 minutes of inactivity by default, per the FAQ. The first request reloads it. You can change that with OLLAMA_KEEP_ALIVE.

Can I use it from my laptop in another room?

You can, by changing OLLAMA_HOST, but then anything on your network can reach it. For most home writers, keeping it local to the Mac mini is the safer default.

The short version

Install Ollama from the official .dmg and drag it to Applications. Check that you're on macOS Sonoma or newer. Turn on local-only mode if privacy is the point. Pick a small model by reading the library page, not a forum. Use ollama ps to see what actually fits. Save your editing instructions in a Modelfile. Keep facts out of its job description.

A quiet Mac mini and a small model won't write your book. They'll help you finish it.

Frequently asked
Can a Mac mini run Ollama?
Yes, if it meets the requirements in the official macOS docs: macOS Sonoma (14) or newer, with Apple M series (CPU and GPU) or x86 (CPU only).
Which model should I use for writing?
Start with a small, general-purpose instruct model from the Ollama library, at one of the smaller sizes. Load it, check `ollama ps`, and step up only if it fits fully on the GPU and you want better output.
How much memory do I need?
The official docs don't publish a Mac memory table, so I won't give you one. Use `ollama ps` to see each model's loaded size and whether it fits on the GPU. Longer context and more parallel requests use more memory.
Will my writing be sent to the internet?
Not when you run local models. The Ollama FAQ says it doesn't see your prompts or data when you run locally. For extra certainty, turn on local-only mode with `OLLAMA_NO_CLOUD=1` or `disable_ollama_cloud` in `server.json`.
Why is the first answer after a break slow?
Models unload after 5 minutes of inactivity by default, per the FAQ. The first request reloads it. You can change that with `OLLAMA_KEEP_ALIVE`.
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room