ApiaryActiveLive
Try: pause · settings · learn · wipe
← Community / Reading Room
HT
craft · 19 min read

How to Prompt Local Models So They Stop Writing Brochure Guff

Everyone who has used a language model knows the sound. It is the voice of a hotel lobby pamphlet. It goes like this:

By Austin Little

You asked your local model for a short note about a weekend hive inspection, and it gave you three paragraphs about "the vibrant, ever-evolving world of beekeeping." This guide is about making it stop. None of it costs money, and almost all of it is about how you ask.

AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events.

What "brochure guff" actually is

Everyone who has used a language model knows the sound. It is the voice of a hotel lobby pamphlet. It goes like this:

  • An opening that restates your question in grander words. "In today's fast-paced world, effective communication has never been more important."
  • Adjectives that do no work: vibrant, seamless, robust, cutting-edge, holistic, game-changing, transformative.
  • A list of three things where one would do, and all three are abstract.
  • A closing paragraph that summarizes what you just read and then invites you to "embark on your journey."
  • No numbers, no names, no specific nouns. Nothing you could check or argue with.

It is not that the model is stupid. It is that this register is the safest, most average-sounding text it knows. A great deal of the public web is marketing copy, press releases, and SEO filler, and models learn from text that looks a lot like that. When you give a vague request, you get the average answer, and the average answer sounds like a brochure.

The fix is not a magic phrase. The fix is to remove the vagueness that invites the average. You do that in four places: the job you describe, the material you hand over, the examples you show, and the settings you run with. We will go through each.

I will use Ollama for the settings examples because it is free, runs on your own machine, and its Modelfile reference documents exactly which knobs exist. The prompting ideas work the same way in LM Studio or any other local tool, and in free-tier cloud chat boxes too.

Why local models drift into guff more easily

A small local model has less room to hold nuance than a giant hosted one. That is not a knock on local AI. It just means two things in practice:

  1. Small models lean harder on defaults. If you do not tell a small model what voice to use, it reaches for the most common one. The most common one is generic and upbeat.
  2. Small models follow concrete instructions better than abstract ones. "Be less flowery" is abstract. "No adjectives unless they carry a measurement or a color" is concrete. Small models do much better with the second kind.

So the craft here is turning your taste into rules and examples concrete enough that a modest model can follow them. If you can do that, you will also find that bigger models get better, because the same clarity helps them too.

One honest caveat. Different models behave differently, and they change between versions. Pick a model your machine can hold comfortably, from the Ollama library, and test the techniques below on it.

Fix one: describe the job, not the vibe

Most guff starts with a request like "Write a post about our honey sale." That sentence contains a topic and nothing else. The model has to guess the reader, the length, the purpose, the tone, and the facts. It guesses "generic marketing reader," and here comes the brochure.

A job description answers five questions:

  • Who is reading? "Neighbors on our street email list, most of whom already know us."
  • What should they do or know afterward? "Know the sale is Saturday 9 to noon at the garage, cash or card."
  • How long? "Under 80 words."
  • What format? "Plain email body. No subject line. No headings."
  • What voice? "Like a neighbor talking over the fence. Friendly, not excited."

Put those together and the request becomes:

Write a plain email body for neighbors on our street email list.
Most of them already know us. They need to know:
- honey sale this Saturday, 9 to noon
- in our garage at the end of the block
- cash or card
- we have about 30 jars this year (do not change this number)
Under 80 words. No headings, no subject line, no exclamation marks.
Sound like a neighbor talking over the fence, not a store.

Notice what is happening. Every line removes a decision the model would otherwise make by averaging. The more decisions you make, the less brochure you get.

The "who is reading" line does the most work

If you only add one thing, add the reader. "Explain this to a ninth-grader who is good at math but has never heard of bees" produces completely different prose than "explain this." A named reader gives the model a target. Generic copy is what you get when the target is "everyone."

Give the purpose in a verb

"So they can decide whether to come" is a purpose. "To engage the community" is not; it is guff about guff. When you write the purpose, use a plain verb the reader would do: decide, buy, fix, show up, reply, stop worrying.

Fix two: hand over the facts, and forbid new ones

The second big source of brochure prose is a model with nothing to say. If you give it a topic but no material, it has to fill space, and filler is adjectives.

So feed it material. Paste your notes, your bullet points, your rough draft, the customer's question, the measurements, the dates. Then tell it plainly:

Use only the facts in my notes below. If something is missing,
write [MISSING: what] instead of making it up.

This does two jobs at once. It reduces guff, because the model is busy arranging your real details instead of inventing pleasant ones. And it reduces made-up facts, which matter far more than tone. A model that is told it is allowed to leave a gap is less likely to fill the gap with a confident fiction. It will not be perfect, so you still read the output, but it is a real improvement in practice.

The bracket trick is also useful for you. When the draft comes back with [MISSING: start time], you see exactly what you forgot.

Rough notes beat polished prompts

People often spend a long time polishing the instructions and then provide two vague sentences of content. Flip that. Your instructions can be short and blunt. Your notes should be specific and messy. "Hive 2 queen seen, eggs yes, capped brood patchy, two frames honey, temper calm, 68F overcast" is wonderful material. A model can turn that into a clean paragraph. It cannot turn "the inspection went well" into anything but fluff.

Fix three: ban the guff by name

Telling a model "don't be generic" rarely works, because "generic" is not a thing it can check. Naming specific words and moves it should avoid works much better, because those are concrete.

Here is a ban list I find useful. Adjust it to your own pet peeves:

Do not use these words: vibrant, seamless, robust, leverage,
holistic, cutting-edge, game-changer, unlock, elevate, delve,
journey, landscape, tapestry, testament, in today's world.
Do not open by restating the question.
Do not end with a summary or a call to "embark" or "explore".
Do not use exclamation marks.
No more than one adjective per sentence.

Two notes on ban lists.

First, keep them short enough to matter. A ban list of 200 words will get partly ignored, especially by small models with limited context. Ten to twenty of your worst offenders is plenty. If a new one shows up repeatedly, add it and drop one that has stopped appearing.

Second, a ban list alone can make prose stiff, because the model will reach for a near-synonym. That is why you pair it with a positive instruction about what to do instead, which is the next fix.

Fix four: show, do not describe

The single strongest technique for voice is an example. Describe a voice in adjectives and the model hears more adjectives. Show it two short samples of the voice you want, and it can imitate the rhythm, sentence length, and vocabulary directly.

A simple pattern:

Here are two examples of the voice I want.

Example 1:
"We checked the hives Saturday. Hive 2 is fine. Hive 3 is short on
stores, so we'll feed it this week. Nobody got stung."

Example 2:
"The door was off the track on the left side. I didn't touch the
spring. I called the company and they're coming Tuesday."

Now write in that voice about the following notes:
[your notes]

What the examples teach, without you having to say it: short sentences, concrete nouns, first person, no hype, and an ending that just stops when the information is done.

Use your own past writing

The best examples are paragraphs you wrote yourself on a good day. Find an email or a note you liked, strip out anything private, and paste it in as the voice sample. The model will drift toward you rather than toward the internet's average.

Show a bad example, carefully

You can also show the model what you do not want, labeled clearly:

BAD (do not write like this):
"Embark on an unforgettable journey into the vibrant world of
beekeeping, where nature's tiny marvels await!"

GOOD (write like this):
"Bees are calmer on warm, still days. Inspect then if you can."

Contrast pairs help some models a lot. With very small models, though, a bad example can occasionally leak into the output, because the words are now in front of it. If you see that, drop the bad example and keep only the good one.

Making it stick: system prompts and Modelfiles in Ollama

Typing all of this every time gets old. Ollama lets you save it into a custom model so the rules are always on. This is all in the official Modelfile reference.

The SYSTEM instruction

A Modelfile is a small text file. The FROM line names the base model, and the SYSTEM line sets a system message, which is the standing instruction the model sees before every conversation. The documented syntax for a multi-line system message uses triple quotes:

FROM your-base-model
SYSTEM """
You write plain, specific prose for a neighbor or a coworker.
Short sentences. Concrete nouns. First person when it fits.
Use only facts the user provides. If something is missing,
write [MISSING: what].
Do not use: vibrant, seamless, robust, leverage, holistic,
cutting-edge, unlock, elevate, delve, journey, landscape.
Do not restate the question. Do not end with a summary.
No exclamation marks.
"""

Replace your-base-model with a model you already pulled from the library. Then, per the docs, you save the file as Modelfile and run:

ollama create plainwriter -f Modelfile
ollama run plainwriter

Now plainwriter is your guff-resistant writer, and the base model stays untouched. You can have several: one for emails, one for article outlines, one for kid-friendly explanations.

If you want to see how a model you already have is set up, the docs show ollama show --modelfile followed by the model name. It prints the Modelfile, including the template and stop sequences. Reading it is a good way to learn what is already baked in.

The MESSAGE instruction for built-in examples

The Modelfile reference also documents a MESSAGE instruction, which lets you add example conversation turns with the roles user and assistant. That means you can bake your voice examples straight into the model:

MESSAGE user Notes: hive 2 queen seen, eggs yes, calm, 68F.
MESSAGE assistant Checked hive 2. Saw the queen and fresh eggs. Calm bees. About 68 degrees.

Two or three of those pairs teach the voice by demonstration every time you open the model. Keep them short and true to how you want it to sound.

Temperature, and why turning it down is not the whole answer

The Modelfile reference lists temperature as a parameter and describes it this way: increasing the temperature makes the model answer more creatively. The docs list a default of 0.8. You set it with a line like:

PARAMETER temperature 0.5

People often assume guff is a temperature problem. It mostly is not. A low temperature makes the model pick its most likely words more consistently, and for an underspecified prompt, the most likely words are often exactly the brochure clichés. So lowering temperature can make output more predictable without making it less generic.

My practical rule:

  • For factual rewrites, summaries, and anything where you supplied the facts, try a somewhat lower temperature so it sticks to your material.
  • For drafting where you want some variety in phrasing, leave it near the default or a bit higher, and rely on examples and ban lists to steer the voice.
  • Change one thing at a time and compare. Do not trust anyone's magic number, including mine. I am not giving you one because I have not measured it on your model.

You can also test a value inside a running chat without making a new model.

Repetition penalty for the "robust, robust, robust" problem

Small models sometimes fall into a loop where the same word or phrase keeps showing up. The Modelfile reference documents repeat_penalty, which sets how strongly to penalize repetitions, and repeat_last_n, which sets how far back the model looks. The docs describe the default repeat_penalty as 1.0, meaning disabled, with a higher value such as 1.1 penalizing repetition more.

If your drafts keep repeating the same filler word, a small bump is worth trying. Do not crank it up a lot. Penalize repetition too hard and the model starts avoiding words it legitimately needs, like your product name or "the," and the prose turns odd in a different way.

Context size, and why it matters for voice

Your system prompt, your examples, your notes, and the conversation so far all have to fit in the model's context window. If they do not, the oldest material falls out, and often that is your carefully written voice instruction. The model then slides back to its default register halfway through a long session.

The Ollama docs cover setting context size through num_ctx in a Modelfile, /set parameter num_ctx in a session, or the OLLAMA_CONTEXT_LENGTH environment variable for the server. Larger context uses more memory, so raise it only as far as your machine handles comfortably.

The simpler fix for long sessions: start a fresh chat for each new piece of writing. Your custom model carries the system prompt and examples, so a fresh chat gets the full rules again without the clutter of the last job.

A reusable prompt template

Here is a template you can keep in a notes file and paste into any local model. Fill in the brackets.

JOB: Write [format] for [specific reader].
PURPOSE: After reading, they should [plain verb].
LENGTH: Under [N] words.

VOICE: Plain and specific. Short sentences. Concrete nouns.
Sound like [a neighbor / a coworker / a patient teacher].

RULES:
- Use only the facts in NOTES. If something is missing, write
  [MISSING: what].
- Do not use: [your 10 to 20 worst words].
- Do not restate the question. Do not end with a summary.
- No exclamation marks.

EXAMPLE OF THE VOICE:
"[one or two sentences you wrote yourself]"

NOTES:
[paste your messy notes]

It looks long, but most of it never changes. Once your custom Modelfile has the voice and rules baked in, your day-to-day prompt shrinks to the JOB, PURPOSE, LENGTH, and NOTES lines.

The two-pass method: draft, then de-guff

Sometimes the first draft still has fluff in it. Instead of rewriting by hand, run a second, narrower pass. Narrow jobs are easier for small models than broad ones.

Pass one: write the draft with the template above.

Pass two: paste the draft back with a strict editing instruction.

Edit the text below. Keep every fact. Change nothing about meaning.
1. Delete any sentence that does not add a new fact or step.
2. Replace vague words with the specific thing they refer to.
   If you don't know the specific thing, write [SPECIFY].
3. Cut any opening sentence that restates the topic.
4. Cut any closing sentence that summarizes.
Return only the edited text.

TEXT:
[paste draft]

The [SPECIFY] marker is the important part. It flags spots where the draft was vague because you never gave it the detail. That tells you what to go find, which is the real reason the draft was soft.

A third pass for reading aloud

If you are writing something people will hear, like a talk or a voicemail script, one more pass helps:

Rewrite this so it sounds natural read aloud by a regular person.
Shorter sentences. No words a person wouldn't say out loud.
Keep every fact.

Then actually read it aloud. Your mouth is a better guff detector than any model. If you would feel silly saying a sentence to a neighbor, cut it.

Common mistakes that bring the brochure back

Asking it to "make it engaging." "Engaging" is a marketing word, and asking for it summons marketing prose. Ask for "clear," "specific," or "short" instead. If you want warmth, show a warm example.

Asking for a "blog post." The phrase drags in a whole genre: a catchy intro, headers, a conclusion, a call to action. If you want a plain explanation, ask for "a plain explanation in four paragraphs." Name the shape, not the genre.

Stacking contradictory rules. "Be concise but comprehensive, casual but professional." The model splits the difference, and the difference is beige. Pick one per pair.

Too many instructions for a small model. A small model with thirty rules will follow some and drop others unpredictably. Prioritize. Put the three rules you care about most first and in plain words.

Long chats that wander. As we covered, voice instructions can fall out of context. Start fresh for each piece.

Trusting the facts because the tone is good. This is the dangerous one. Once the prose sounds plain and human, it is easy to forget the model can still be wrong. Plain-sounding wrong facts are worse than guff, because they are more believable. Check every number, date, and name against your notes or a primary source.

Worked example: a hive-inspection note, before and after

Here is the kind of thing a vague prompt tends to produce. This is an illustration I wrote to show the pattern, not output copied from a specific model.

Vague prompt: "Write about my hive inspection today."

Typical guff: "Today's hive inspection was a truly rewarding experience that highlighted the incredible resilience of these remarkable pollinators. As I carefully opened the hive, I was greeted by a vibrant hum of activity, a testament to the thriving colony within..."

It goes on. It contains no information about the hive.

Better prompt:

JOB: A short log entry for my own beekeeping notebook.
PURPOSE: So next week I remember what to check.
LENGTH: Under 60 words.
VOICE: Terse, like a field note. No adjectives unless measured.
RULES: Use only my notes. Mark missing info as [MISSING: what].
NOTES: Hive 2. Queen seen. Eggs yes. Capped brood patchy.
Two frames honey. Calm. Overcast, 68F. Plan: recheck brood pattern.

The kind of result you want: "Hive 2, overcast, about 68F. Queen seen, eggs present. Capped brood patchy, so recheck the pattern next visit. Two frames of honey. Bees calm. [MISSING: date]"

That is the entire improvement in one picture. Nothing magical, just a defined reader, a purpose, a length, a voice, and real notes. And the [MISSING: date] tag is the model being honest instead of guessing.

Prompting for different jobs

The core method does not change, but the emphasis shifts by task.

Emails and messages

Lead with the reader and the one thing they must do. Set a hard word limit. Ban greetings like "I hope this email finds you well" if they bother you. Paste your own past email as the voice example.

Explanations for learners

Name the learner precisely: age, what they already know, what they do not. Ask for one concrete example per idea. Ask it to define any term the first time it appears. A good addition: "If a sentence would only make sense to someone who already understands this, rewrite it."

Product or service descriptions

This is where guff is most tempting, because the genre itself is brochure-shaped. Push against it with facts: dimensions, materials, what it does, what it does not do, who it is not for. "Who it is not for" is an excellent de-guff instruction, because marketing copy never says it.

Summaries of long documents

Ask for a summary that keeps numbers, names, and dates, and drops adjectives. Ask it to quote rather than paraphrase anything that sounds like a commitment or a deadline. Then check the quotes against the source, because a summary can still misstate things.

Article outlines

Ask for headings that are questions a real person would type or say, not clever titles. "How long does a hive inspection take?" beats "The Rhythm of the Hive." Then write the body yourself or section by section with your notes.

When a local model is not enough

Sometimes the model you can run comfortably just will not hold a voice across a long piece, even with good prompts. That is a hardware and size limit, not a failure on your part. You have free options:

  • Work in smaller pieces. Ask for one section at a time with the same custom model. Small models do better on short jobs.
  • Try a different model from the library that your machine can still hold. Models vary in how well they follow style instructions. Test with the same prompt and notes so the comparison is fair.
  • Use a free-tier cloud model for the draft and local for the edit, or the reverse, as long as the material is not private. Check the provider's own pages for current limits, because free tiers change.
  • Write the first draft yourself and use the model only for the de-guff pass. This is underrated. Editing is a narrower job than writing, and narrow jobs suit small models.

You do not need a paid subscription to get plain prose. You need clear jobs, real material, a few examples, and a model you have tuned a little with a Modelfile.

A short checklist to tape near your screen

  1. Who is reading, specifically?
  2. What should they do afterward, as a plain verb?
  3. How many words, at most?
  4. Did I paste real notes with numbers, names, and dates?
  5. Did I tell it to mark missing facts instead of inventing them?
  6. Did I show one or two examples of the voice?
  7. Did I ban my worst ten words?
  8. Is this a fresh chat?
  9. Did I run a de-guff pass?
  10. Did I check every fact myself?

If the answer to all ten is yes, the brochure is mostly gone. If a little is left, cut it by hand. That last edit is yours, and it is the part that makes the writing sound like you.

The bottom line

Brochure prose is what a model writes when it has to guess. Every guess you remove, about the reader, the purpose, the length, the facts, and the voice, pushes the output toward something plain and useful. Local models reward this more than big hosted ones, because they rely on your instructions more. Bake your rules into an Ollama Modelfile with SYSTEM and MESSAGE, adjust temperature and repetition gently and one at a time, start fresh chats, and keep checking facts. It is free, it is on your machine, and it sounds like you.

References

  • Ollama Modelfile Reference: https://docs.ollama.com/modelfile (fetched 2026-10-01)
  • Ollama FAQ: https://docs.ollama.com/faq (fetched 2026-10-01)
  • Ollama download: https://ollama.com/download (fetched 2026-10-01)
  • Ollama model library: https://ollama.com/library
Frequently asked
What is How to Prompt Local Models So They Stop Writing Brochure Guff about?
Everyone who has used a language model knows the sound. It is the voice of a hotel lobby pamphlet. It goes like this:
What should you know about what "brochure guff" actually is?
Everyone who has used a language model knows the sound. It is the voice of a hotel lobby pamphlet. It goes like this:
What should you know about why local models drift into guff more easily?
A small local model has less room to hold nuance than a giant hosted one. That is not a knock on local AI. It just means two things in practice:
What should you know about fix one: describe the job, not the vibe?
Most guff starts with a request like "Write a post about our honey sale." That sentence contains a topic and nothing else. The model has to guess the reader, the length, the purpose, the tone, and the facts. It guesses "generic marketing reader," and here comes the brochure.
What should you know about the "who is reading" line does the most work?
If you only add one thing, add the reader. "Explain this to a ninth-grader who is good at math but has never heard of bees" produces completely different prose than "explain this." A named reader gives the model a target. Generic copy is what you get when the target is "everyone."
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room