By Austin Little
Fine-tuning sounds like the grown-up move: train the model on your stuff and it'll finally "get" you. Sometimes that's true. Most of the time, for a small operator with a local model, a better prompt and a folder of your own documents gets you most of the way, in an afternoon instead of a month. Here's how to tell which situation you're in.
AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events.
I run small things: a publishing site, some writing, some advice for neighbors about bees and garage doors. When local AI got good enough to be useful, the first question everyone asked me was "should I fine-tune it?" The honest answer is "maybe later." This guide explains why, and how to recognize when "later" has arrived.
Everything here uses free, local tools: Ollama for running models, and open-source training tools if you do get to the fine-tuning step. Where I quote a tool's documentation, I link it. Where a claim is a tool maker's opinion rather than a settled fact, I say so.
Three ways to change what a model does
People lump these together, and they're very different jobs.
1. Prompting
You change what you say to the model. That includes instructions, examples, a standing "system" brief, and settings like temperature. Nothing about the model changes. It takes minutes, it costs nothing, and you can undo it instantly.
2. Retrieval (often called RAG)
You change what the model can see when it answers. You keep your own documents (notes, policies, past articles, product sheets) in a searchable index. When you ask a question, the system finds the relevant passages and hands them to the model alongside your question. The model itself still doesn't change. It's reading your documents in the moment.
3. Fine-tuning
You change the model itself by training it further on examples. Afterward, the new behavior is built in, even with a short prompt. It takes data preparation, compute, testing, and maintenance.
Here's the key difference in one line. Prompting and retrieval change the input. Fine-tuning changes the model. Changing the input is cheap and reversible. Changing the model is neither, at least not without effort.
The decision ladder
I use a ladder. Start at the bottom. Only climb when you can name what the current rung can't do.
Rung 1: Write a better prompt. Rung 2: Save it as a reusable model profile. Rung 3: Add examples. Rung 4: Add retrieval over your documents. Rung 5: Try a bigger or different base model. Rung 6: Fine-tune.
Most small operators stop at Rung 3 or 4 and are happy. Let's walk up.
Rung 1: Write a better prompt
Most "the model doesn't understand me" problems are prompt problems. Before you change anything else, make your prompt do four things.
- Say the job. "Edit this paragraph for clarity" is better than "make this better."
- Say the audience. "For a beekeeping club newsletter, readers are mostly retirees, plain language."
- Say the rules. "Keep it under 150 words. Don't add facts that aren't in the original. Don't use the word 'delve.'"
- Say the format. "Give me the revised paragraph, then one line on what you changed."
Try the same task five times with your improved prompt. If four of the five are usable, you're done. You don't need to fine-tune.
Settings matter too
Ollama exposes generation settings you can change without training anything. The Modelfile reference lists, among others:
temperature: the docs describe higher values as more creative and lower values as more coherent.top_kandtop_p: control how adventurous word choices are.repeat_penalty: discourages repetition.seed: the docs say a fixed seed makes the model generate the same text for the same prompt, which is great for testing whether a change actually helped.num_ctx: the context window size, meaning how much text the model can consider at once.
For editing work, a lower temperature often helps. For brainstorming, a higher one. Use a fixed seed while you test prompts so you're comparing like with like.
Rung 2: Save it as a reusable model profile
Once a prompt works, stop retyping it. Ollama lets you package a base model plus your instructions and settings into a named model with a Modelfile. This isn't fine-tuning. The weights don't change. It just saves your setup.
FROM <base-model>:<tag>
PARAMETER temperature 0.5
SYSTEM """You edit for a small beekeeping newsletter. Plain words, short sentences, friendly but not cute. Never add facts, numbers, names, or quotes that are not in the original."""
Then:
ollama create club-editor -f Modelfile
ollama run club-editor
The docs describe FROM as required, SYSTEM as the system message, and PARAMETER as run settings. You can view any model's Modelfile with ollama show --modelfile <model>, which is a good way to see how a model's template and defaults are set up.
A lot of people who think they need fine-tuning actually need this. One named model per job, like club-editor, headline-helper, or invoice-reply, each with its own brief.
Rung 3: Add examples
Models learn a lot from a few good examples placed right in the prompt. This is sometimes called few-shot prompting.
Ollama's Modelfile has a MESSAGE instruction for exactly this. The docs describe it as a way to specify message history, using the roles system, user, and assistant, so the model will "answer in a similar way."
FROM <base-model>:<tag>
SYSTEM """Rewrite customer questions into short, warm replies in our house style."""
MESSAGE user Do you service torsion springs on weekends?
MESSAGE assistant Yes, we do weekend spring calls. Tell us your address and the door size, and we'll confirm a time.
MESSAGE user Can I adjust my own spring?
MESSAGE assistant We don't recommend it. Torsion springs hold a lot of tension and can cause serious injury. We're happy to come take a look.
Three to five examples that show the voice and the edge cases often do more than pages of instructions. Pick examples that show:
- Your normal tone
- One case where the right answer is "no"
- One case where the model should refuse to guess
If the outputs now sound like you most of the time, you're done. If they still drift, keep climbing.
Rung 4: Add retrieval over your documents
If the problem is that the model doesn't know your stuff (your prices, your policies, your past articles, your hive records), that's usually a retrieval problem, not a training problem.
How local retrieval works
- Split your documents into chunks, like paragraphs or short sections.
- Turn each chunk into an embedding, a list of numbers that represents its meaning.
- Store the embeddings in a simple index.
- When you ask a question, embed the question, find the closest chunks, and paste them into the prompt with your question.
Ollama can do the embedding in steps 2 and 4 on your own machine. The embeddings docs say embeddings "turn text into numeric vectors you can store in a vector database, search with cosine similarity, or use in RAG pipelines." They list recommended embedding models (embeddinggemma, qwen3-embedding, and all-minilm) and an endpoint:
curl -X POST http://localhost:11434/api/embed \
-H "Content-Type: application/json" \
-d '{"model": "embeddinggemma", "input": "Your text here"}'
Two tips straight from the docs:
- Use cosine similarity for most semantic search.
- Use the same embedding model for both indexing and querying. If you switch embedding models, rebuild the index.
Why retrieval often beats fine-tuning for facts
- It's current. Change a document, re-embed it, and the next answer reflects it. A fine-tuned model knows only what it was trained on.
- It's checkable. You can show which passages the answer came from. With a fine-tuned model, there's no paper trail.
- It's cheap. Embedding a folder of documents on a home computer is an afternoon project, not a training run.
- It's reversible. Delete a document from the index and it's gone from answers.
A fair note on the other side
Not everyone frames it this way. The Unsloth fine-tuning guide, from a team that builds fine-tuning tools, says fine-tuning can "inject and learn new domain-specific information," and states that "fine-tuning can replicate all of RAG's capabilities, but not vice versa." It also calls the idea that fine-tuning can't teach new knowledge a misconception.
That's a real position from people who do this work every day, and it's worth knowing. My view for small operators is narrower: even if fine-tuning can teach facts, retrieval is usually the faster, cheaper, and more checkable way to get your current facts into an answer. You can always add fine-tuning later on top of retrieval.
Rung 5: Try a different base model
Before training anything, try another model. The Ollama library lists many families and sizes, with tags for features like tools, thinking, vision, and embedding. A different model, or a larger size of the same one, might simply be better at your task.
Run the same five-task test you used on Rung 1 with the same prompt and seed. Compare. This is the cheapest "upgrade" there is, and it's free.
Watch memory as you go. The Ollama FAQ explains that ollama ps shows each loaded model's size and whether it's on the GPU, the CPU, or split. If a bigger model doesn't fit well, it may be slower than it's worth.
Rung 6: Fine-tune (when you've earned it)
You've climbed the ladder. Prompts are tight, examples are in, retrieval handles your facts, and you've tried other models. You still have a specific, repeatable problem. Now fine-tuning might be the right tool.
Good reasons to fine-tune
- A consistent format or style the model keeps breaking, even with examples. For example, a strict output structure for every reply, every time.
- A narrow, repetitive task at volume, like classifying hundreds of messages into the same few categories, where a short prompt and a small specialized model beat a long prompt on a big one.
- You want shorter prompts. A fine-tuned model can carry behavior that would otherwise need a long system prompt and many examples on every call.
- You have good data. Hundreds of clean, consistent examples of exactly the input and output you want.
Bad reasons to fine-tune
- "It doesn't know our prices." That's retrieval.
- "It sometimes gets facts wrong." Fine-tuning won't make a model reliable on facts, and it can make it confidently wrong in your voice.
- "It doesn't sound like me." Try examples first. Voice is very responsive to three good samples.
- "Everyone says fine-tuning is the real thing." That's not a requirement.
What fine-tuning actually involves
The modern home-scale approach is usually LoRA or QLoRA. Instead of changing every weight in the model, you train small add-on pieces called adapters.
- The Hugging Face PEFT docs describe parameter-efficient fine-tuning as adapting large pretrained models "without fine-tuning all of a model's parameters," which they say significantly decreases compute and storage costs and makes training more accessible on consumer hardware.
- The Unsloth guide describes LoRA as training a small set of added low-rank adapter weights while the base model stays frozen, and QLoRA as combining LoRA with 4-bit precision. It recommends starting with QLoRA, starting with a small instruct model, and warns that jumping straight to full fine-tuning is a common mistake.
Here's the work, step by step.
- Build a dataset. Usually pairs of input and ideal output. The Unsloth guide stresses that quality and amount largely determine the result, and that a well-structured set of question-and-answer pairs beats a raw dump of documents for most uses.
- Hold some back for testing. The Unsloth guide suggests setting aside part of your data, for example 20%, to test on.
- Train. Watch the training loss. The Unsloth guide notes that a loss dropping toward zero can mean overfitting, where the model memorizes your examples instead of learning the pattern.
- Evaluate. Compare the fine-tuned model against your best Rung 4 setup on tasks it hasn't seen. If it isn't clearly better, you've learned something valuable cheaply.
- Package and run it. Convert or export so you can run it locally.
- Maintain it. When your style or needs change, you'll need new data and another training run.
Free, local ways to do it
On a Mac: the mlx-lm LoRA guide from the MLX project (the ml-explore organization on GitHub) describes fine-tuning with LoRA or QLoRA:
- Install the training extras with
pip install "mlx-lm[train]". - The main command is
mlx_lm.lora. Runmlx_lm.lora --helpto see options. - Training data is a
train.jsonlfile, with atest.jsonlfor testing and an optionalvalid.jsonlfor validation loss. - If you point it at a quantized model, training uses QLoRA. Otherwise it uses regular LoRA.
- Adapters are saved to an
adapters/folder by default, and the guide includes a step to fuse adapters into the model.
I'm not giving memory or time numbers for training. They depend heavily on the model size, the method, the context length, and your hardware, and I haven't measured your setup. Start with the smallest model the tool supports and a small dataset, and see what happens on your machine.
Getting a fine-tuned model into Ollama
The Ollama import docs describe two routes:
- Safetensors: write a Modelfile with
FROM /path/to/safetensors/directory, thenollama create my-model. - GGUF: write a Modelfile with
FROM /path/to/file.gguf. The docs note that Ollama doesn't quantize GGUF models during import, so you'd prepare and quantize first with a GGUF tool like llama.cpp'sllama-quantize.
So a common path is: train LoRA adapters, fuse them into the model, export to Safetensors or GGUF, then import into Ollama with a Modelfile.
How to test any rung honestly
Whatever you try, test it the same way so you're not fooled by one lucky answer.
- Write ten real tasks from your actual work. Not toy examples.
- Write what a good answer looks like for each, briefly.
- Run each setup on all ten, with the same seed where you can.
- Score each answer: usable as-is, usable with light edits, or not usable.
- Keep the scores in a simple table. Compare rungs.
A setup that moves three tasks from "not usable" to "usable with light edits" is a real gain. A setup that feels smarter on one task and worse on two others isn't.
Also check for two failure types that matter for small operators:
- Invented facts. Did it add any number, name, price, or quote that isn't in your source? That's a fail, no matter how good the prose is.
- Voice drift. Does it still sound like you on the tenth task, not just the first?
The real costs, side by side
| Prompting | Model profile + examples | Retrieval | Fine-tuning | |
|---|---|---|---|---|
| Changes the model? | No | No | No | Yes |
| Time to first result | Minutes | Under an hour | An afternoon or so | Days to weeks, including data prep |
| Data you need | None | A few good examples | Your documents | Many clean input/output pairs |
| Easy to update? | Instantly | Instantly | Re-embed changed docs | Retrain |
| Shows its sources? | No | No | Yes, if you display retrieved passages | No |
| Free and local? | Yes | Yes, with Ollama | Yes, with Ollama embeddings | Yes, with open-source tools on capable hardware |
The time estimates are rough guidance from the shape of the work, not measurements. Your mileage will vary with your documents and your hardware.
Three small-operator scenarios
The newsletter writer who wants "my voice"
Start at Rung 3. Put three of your best paragraphs into MESSAGE examples with a short brief. Most people are surprised how far this goes. If after a few weeks you have dozens of before-and-after edits you've made by hand, you have the start of a fine-tuning dataset. Keep them in a file.
The small service business with policies and prices
Start at Rung 4. Your prices and policies change, and you need answers to be checkable. Put your policy documents into a local retrieval setup. Have the model answer only from retrieved passages and say "I don't know" otherwise. Fine-tuning would freeze today's prices into the model, which is the opposite of what you want.
The club secretary sorting member emails
This might be a Rung 6 job eventually. Sorting emails into the same handful of categories, hundreds of times, is the kind of narrow, repetitive task where a small fine-tuned model can shine. Start with a prompt and examples anyway. If it's already accurate on your ten-task test, you've saved yourself a project.
Safety lines
- Don't train on private data you don't have the right to use. Customer emails, member records, and client documents carry obligations. Ask before you train.
- Don't train on secrets. A fine-tuned model can repeat parts of its training data. Passwords, account numbers, and personal details don't belong in a dataset.
- Keep retrieval indexes private too. An index of your documents is a copy of your documents.
- Label AI-assisted output where your readers would want to know.
FAQ
Is fine-tuning necessary to use a local model well?
No. For most small operators, a clear prompt, a saved Modelfile profile, a few examples, and retrieval over your documents cover most needs. Fine-tuning is an option for specific, repeatable problems those steps can't solve.
Is a Modelfile the same as fine-tuning?
No. A Modelfile saves a base model with your instructions, examples, and settings under a new name. The model's weights don't change.
Can I fine-tune on a Mac?
The mlx-lm project documents LoRA and QLoRA fine-tuning with its mlx_lm.lora command. How large a model you can train depends on your hardware, so start small.
Will fine-tuning stop the model from making things up?
Don't count on it. Retrieval with "answer only from these passages" instructions, and your own checking, are better guards against invented facts.
What's RAG in one sentence?
You search your own documents for passages relevant to the question and give those passages to the model with the question, so it answers from your material.
How do I know it's time to fine-tune?
When you can name a specific, repeatable task that still fails after better prompts, examples, retrieval, and a different base model, and you have plenty of clean examples of the right answer.
The short version
Start with the cheapest change that could work. Write a clearer prompt. Save it as a named model. Add a few examples. Put your documents behind retrieval. Try another base model. Only then, with a real dataset and a real test, fine-tune.
Most of the time you'll stop early, and that isn't settling. It's choosing the tool that matches the job, and keeping your afternoon.