ApiaryActiveLive
Try: pause · settings · learn · wipe
← Community / Reading Room
FT
craft · 20 min read

Free Transcription to Draft Workflow Without a Card

A lot of good writing starts as talking. You're walking the dog, driving to a job site, standing at a hive with your gloves on, and the idea finally comes out…

By Austin Little

You talk faster than you type. The trick is getting from a rambling voice memo to a clean draft without handing a card number, or your raw thoughts, to a subscription you'll forget to cancel.

AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events.

A lot of good writing starts as talking. You're walking the dog, driving to a job site, standing at a hive with your gloves on, and the idea finally comes out in a whole sentence. If you wait until you're at a keyboard, the sentence is gone. So you hit record on your phone, ramble for four minutes, and end up with a file you never open again.

This guide is about fixing that, for free. Not "free for seven days." Not "free once you add a payment method." Free in the plain sense: software that runs on a computer you already own, or tools that do not ask for a card at all. The stack is boring on purpose: your phone's recorder, a converter called ffmpeg, an open speech-recognition model called Whisper (run through OpenAI's Python package or the lighter whisper.cpp port), and a local language model through Ollama to turn the transcript into a first draft. Then you, editing.

That last step isn't optional. The machine gets you from noise to rough text. You get from rough text to something with your name on it.

What this workflow is, and what it isn't

Let's name the pipeline so we can talk about each piece:

  1. Capture. Record a voice note on whatever device is in your pocket.
  2. Move. Get the audio file onto a computer.
  3. Convert. Turn it into a format the transcriber wants.
  4. Transcribe. Speech to raw text, locally.
  5. Shape. A local language model turns the transcript into an outline or rough draft.
  6. Edit. You fix facts, cut filler, and put your voice back in.
  7. File. Save the audio, transcript, and draft together so you can trace any line back to what you actually said.

What it isn't:

  • It isn't a dictation app that types as you talk. This is a batch workflow. You record, then process. Real-time dictation exists, but it's a different tool and a different set of tradeoffs.
  • It isn't a meeting recorder for other people's voices. More on consent below. Whisper's own model card warns against transcribing recordings of people taken without their consent.
  • It isn't a fact checker. If you said a wrong number into your phone, the draft will have a wrong number in it, now with nicer grammar.

Why local first

There are three reasons I push the local path before any cloud option.

Your voice notes are more private than you think. People say things into a phone they'd never type into a web form: a client's name, a medical appointment, a half-formed opinion about a coworker, the address of the job. If transcription happens on your machine, that audio never has to leave it.

No meter is running. Local transcription costs electricity and patience. There's no per-minute bill, no free-tier cap you'll hit halfway through a long recording, and no account to get locked out of.

It keeps working when a free plan changes. Free cloud tiers move. Limits tighten, features get paywalled, products get shut down. A model file on your own disk doesn't change unless you change it.

The cost of local is setup time and hardware. If your computer is old or low on memory, a big model will feel slow. The fix is to use a smaller model, not to give up. We'll get into sizes.

Step 1: Capture without overthinking it

Use the voice recorder that came with your phone. Every mainstream phone has one.

A few habits make the transcript dramatically better:

  • Say the job at the top. "This is a draft for the Apiary piece on fall feeding. Working title: Feeding Bees in October." That sentence becomes your file label and tells the language model what it's looking at later.
  • Talk in paragraphs. Pause a beat between ideas. You don't have to be polished, but a breath between thoughts helps you, the transcriber, and the editor.
  • Say "new section" out loud. It looks silly. It works. You can tell the language model later to treat "new section" as a heading break.
  • Flag the facts you're unsure about. Say "check this" right after a number or a name. "The hive weighed about sixty pounds, check this." The phrase lands in the transcript and you can search for it.
  • Get out of the wind. Wind noise, road noise, and a running faucet hurt any transcriber. If you're outdoors, turn your back to the wind or cup the phone.
  • Keep recordings short-ish. Several five-minute notes are easier to manage than one forty-minute monologue. You'll find things faster, and if one file is garbled, you lose less.

A note on consent

If your recording includes other people, ask first. That's basic decency, and in some places recording laws are strict about it.

Whisper's model card is direct about this: the developers caution against using the models to transcribe recordings of individuals taken without their consent. Take that seriously. This workflow is for your own voice.

Step 2: Move the file to your computer

Use whatever you already use: a cable, a local file share, emailing it to yourself, or a cloud folder you already have. The point is to get the audio file onto the machine that'll do the transcribing.

One caution: if the reason you're going local is privacy, routing the file through a third-party cloud folder first undoes some of that. For sensitive notes, use a cable or a local network transfer.

Make a folder structure now, before you have fifty files:

voice-drafts/
  2026-10-01-fall-feeding/
    audio-original.m4a
    audio-16k.wav
    transcript.txt
    draft-v1.md
    draft-final.md

One folder per piece. Date first so they sort. Keep the original audio forever, or at least until the piece is published. If a reader ever asks "did you really say that?", you can check.

Step 3: Install the free tools

You need three things: ffmpeg, a Whisper implementation, and Ollama. All three are free. None of them asks for a card.

ffmpeg

Whisper's README lists ffmpeg as a requirement and gives install commands for common package managers: apt on Ubuntu or Debian, pacman on Arch, Homebrew on macOS, and Chocolatey or Scoop on Windows. For example, on Ubuntu:

sudo apt update && sudo apt install ffmpeg

On a Mac with Homebrew:

brew install ffmpeg

Whisper: pick one of two routes

Route A: OpenAI's Python package (openai-whisper). This is the reference implementation. The README installs it with:

pip install -U openai-whisper

The README says the code was trained and tested with Python 3.9.9 and PyTorch 1.10.1 and is expected to work with Python 3.8 to 3.11 and recent PyTorch versions. It also notes you might need Rust installed if the tiktoken dependency doesn't have a prebuilt wheel for your platform.

Route B: whisper.cpp. This is a C/C++ port of Whisper. Its README describes it as a plain C/C++ implementation without dependencies, with support for CPU-only inference plus acceleration on Apple Silicon (via Metal and Core ML), NVIDIA GPUs, Vulkan, and more. It's MIT-licensed. If you're comfortable running a few build commands, this is often the lighter option on a modest machine.

The quick start from the README:

git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
sh ./models/download-ggml-model.sh base.en
cmake -B build
cmake --build build -j --config Release
./build/bin/whisper-cli -f samples/jfk.wav

That last line transcribes a sample file that ships with the repo. If you see text come out, it's working.

Which route should you pick? If you already have Python set up and you're comfortable with pip, Route A is fewer moving parts. If Python installs tend to fight you, or you want to squeeze more out of an older laptop, try Route B. Both use Whisper models. Both are free. Both run locally.

Ollama

Ollama runs language models locally. Install it from ollama.com, then pull a model. The Apiary guide on using AI without paying covers the install in more detail. For this workflow, you just need one small-to-mid model that's good at cleaning up English text. More on choosing one in Step 6.

Step 4: Convert the audio

Phones record in various formats. The Python Whisper package uses ffmpeg under the hood to read audio, so it can usually take common formats directly.

whisper.cpp is pickier. Its README says the whisper-cli example currently runs only with 16-bit WAV files and gives this conversion command:

ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav

In plain words: take the input file, resample it to 16 kHz, make it mono, and write 16-bit PCM WAV. Swap input.mp3 for whatever your phone produced.

I'd do this conversion even if you're using the Python route. A clean, mono, 16 kHz WAV is a predictable input, and keeping it next to the original means you always know exactly what the transcriber heard.

Step 5: Transcribe locally

Picking a Whisper model size

This is where most people go wrong: they grab the biggest model because it sounds best, and then their laptop sits there for ages. Start small.

The Whisper README lists six model sizes, four of which have English-only versions. Here's the table as published there. The VRAM figures and relative speeds come from OpenAI's README, and the README notes the relative speeds were measured transcribing English speech on an A100 GPU. Real-world speed varies a lot depending on your hardware, the language, and how fast you talk.

SizeParametersEnglish-onlyMultilingualRequired VRAM (per README)Relative speed (per README)
tiny39 Mtiny.entiny~1 GB~10x
base74 Mbase.enbase~1 GB~7x
small244 Msmall.ensmall~2 GB~4x
medium769 Mmedium.enmedium~5 GB~2x
large1550 MN/Alarge~10 GB1x
turbo809 MN/Aturbo~6 GB~8x

Three notes from the README that matter for English voice notes:

  • The .en English-only models tend to do better for English, especially at the tiny.en and base.en sizes. The README says the gap shrinks at small.en and medium.en.
  • turbo is described as an optimized version of large-v3 with faster transcription and minimal accuracy loss. It's the default model in the CLI example.
  • turbo isn't trained for translation. If you need to translate non-English speech into English, the README says to use a multilingual model like medium or large.

whisper.cpp publishes its own memory table for its converted models:

ModelDisk (per whisper.cpp README)Memory (per whisper.cpp README)
tiny75 MiB~273 MB
base142 MiB~388 MB
small466 MiB~852 MB
medium1.5 GiB~2.1 GB
large2.9 GiB~3.9 GB

My suggestion: for English voice notes, start with base.en on a modest machine or small.en if you have some headroom. Run one real note through it. If the transcript is good enough to edit, stop there. If it's mangling words you need, step up a size. You're looking for "good enough to edit," not perfect.

Running it: Python route

From the README, the basic command is:

whisper audio.wav --model turbo

To pick an English-only model and set the language explicitly:

whisper audio-16k.wav --model small.en --language English

Run whisper --help to see every option, including output formats.

Running it: whisper.cpp route

Download the model you want with the included script, then point whisper-cli at your WAV file. The CLI help lists -m for the model path, -f for the input file, -l for the spoken language (defaulting to en), and -otxt to write a plain text file:

sh ./models/download-ggml-model.sh small.en
./build/bin/whisper-cli -m models/ggml-small.en.bin -f ~/voice-drafts/2026-10-01-fall-feeding/audio-16k.wav -otxt

There's also -osrt if you want a subtitle file with timestamps, which is handy for finding exactly where in the audio a phrase came from.

whisper.cpp also supports voice activity detection (VAD) with the --vad flag and a separate VAD model. VAD detects speech segments so the transcriber can focus on the parts where someone is talking. For voice notes with long pauses, it's worth trying. The README walks through downloading the Silero VAD model.

What to expect from the raw transcript

The raw transcript will be a wall of text with your "ums," false starts, and repeated phrases. That's normal. It's raw material.

It may also contain errors you won't notice unless you look. Whisper's model card is candid: because the models are trained on large-scale noisy data, the predictions may include text that wasn't actually spoken in the audio. The card calls this hallucination. It also notes the models can generate repetitive text, and that these problems may be worse in lower-resource languages.

So the rule is simple: any line you're going to publish, especially a number, a name, or a quote, should be checked against the audio. That's why we keep the original file.

Step 6: Shape the transcript with a local model

Now you have transcript.txt. The next job is turning it into something that looks like a draft. This is where a local language model through Ollama earns its keep.

Choosing a model

Open the Ollama library and pick something that fits your machine. Two families that work well for English cleanup:

  • Qwen3. The Ollama library lists tags from qwen3:0.6b up through much larger sizes, with qwen3:4b listed at 2.5GB and qwen3:8b (the latest tag) at 5.2GB as of our fetch.
  • Mistral. The mistral tag on Ollama is Mistral 7B, updated to v0.3, listed at 4.4GB with a 32K context window, distributed under the Apache license per the library page.

The download size is not the same as the memory you'll need while it's running. Leave headroom.

The companion piece on running Qwen or Mistral locally goes deeper on picks. For this workflow, a 4B to 8B model is plenty for cleanup and outlining.

Mind the context window

A long transcript can exceed what the model is set to read at once. Ollama's docs say context length is the maximum number of tokens the model has access to in memory, and that larger context increases memory use. Ollama's FAQ says you can change it in an interactive session with /set parameter num_ctx and through the API with the num_ctx option.

The practical workaround if a transcript is long: split it. Process one five-minute voice note at a time, or cut the transcript at your "new section" markers. Smaller chunks also make it easier to see if the model drops something.

Prompts that work

Here's the most important rule: tell the model not to add facts. Language models love to helpfully fill in gaps. In a transcript cleanup job, that's the last thing you want.

Prompt 1: Clean transcript (minimal edit)

Below is a raw transcript of me talking. Clean it up:
- Remove filler words (um, uh, like, you know) and false starts.
- Fix obvious transcription errors only if the intended word is clear.
- Keep my wording and order. Do not summarize.
- Do NOT add any facts, numbers, names, or examples that are not in the transcript.
- Where the transcript says "check this", keep the words [CHECK] in place.
- Where the transcript says "new section", start a new paragraph.

TRANSCRIPT:
<paste>

Prompt 2: Outline

From the cleaned transcript below, make an outline with H2 and H3 headings.
Use only ideas present in the transcript. Under each heading, list the
transcript sentences that belong there. If something does not fit, put
it under "Leftovers". Do not invent new points.

Prompt 3: Rough draft in my voice

Turn this outline into a rough draft. Plain, direct sentences. Short
paragraphs. Second person where natural. Do not add statistics, quotes,
product names, or claims that are not in the outline. Keep every [CHECK]
marker. If a section is thin, write "[THIN: needs more]" instead of
padding it.

That last instruction, "write [THIN] instead of padding," is the one I'd tattoo on the inside of my eyelids. A model asked to write 1,000 words from 300 words of material will invent the other 700. You want it to tell you where the gaps are, not fill them with confident guesses.

Running a prompt from the terminal

The quick way is interactive: ollama run qwen3:8b, then paste. For longer transcripts, the local API is cleaner. Ollama's docs show it listening on localhost:11434. Here's a simple pattern with curl:

curl http://localhost:11434/api/chat -d '{
  "model": "mistral",
  "messages": [{"role": "user", "content": "PASTE PROMPT AND TRANSCRIPT HERE"}],
  "stream": false
}'

If you're using a model with a "thinking" mode (Qwen3 has one, per its model card), the output may include a reasoning trace. Ollama's thinking docs describe a think field on chat requests: false asks for no thinking output if the model allows it. For cleanup jobs, you usually just want the answer.

Step 7: Edit like the byline is yours, because it is

This step takes the longest, and it should. The machine saved you the typing. It didn't save you the thinking.

The edit checklist

Go through the draft with the transcript and audio open beside it.

  1. Search for [CHECK]. Verify every flagged fact against a real source. If you can't verify it, cut it or say plainly that it's your estimate.
  2. Search for [THIN]. Decide whether to record another voice note on that section, write it by hand, or cut it.
  3. Read every number aloud and compare to the audio. Transcription errors love numbers. "Fifteen" and "fifty" sound alike in wind.
  4. Check every proper noun. Names of people, places, products. Whisper may spell them phonetically. The language model may then "correct" them into a different real name.
  5. Hunt for things you didn't say. If a sentence sounds polished but you don't remember saying anything like it, find it in the transcript. If it isn't there, delete it.
  6. Put your voice back. Models smooth everything into the same polite paste. Restore your phrases, your examples, your rough edges.
  7. Cut the throat-clearing. Models love to open with "In today's fast-paced world." You don't talk like that. Delete it.

Signs the model went too far

  • A statistic appears that you never said.
  • A quote is attributed to someone, and you didn't quote anyone.
  • A "study" or "expert" shows up.
  • A product recommendation appears that you didn't make.
  • The draft is much longer than your transcript, and you didn't ask for expansion.

Any one of these means: go back to the cleaned transcript and re-run a stricter prompt, or edit by hand.

Step 8: File it so you can trace it later

When the piece is done, your folder should hold:

  • The original audio
  • The 16 kHz WAV you transcribed
  • The raw transcript
  • The cleaned transcript
  • Each draft version
  • The final

That's a paper trail from your mouth to the published page. If you ever need to show where a line came from, you can. On a site like Apiary, where we put an AI disclosure on drafted pages, that trail is part of being honest about how the piece was made.

Free cloud options, and why they're second choice here

Sometimes your computer just can't run even the small models comfortably. Then you're looking at cloud options.

Be strict about the "without a card" part. Lots of transcription and AI services advertise free plans that turn out to need a payment method at signup, or free trials that auto-convert. Before you upload anything:

  • Read the pricing page yourself. Not a roundup article. The provider's own page.
  • Check whether signup asks for a card. If it does, it isn't free for our purposes.
  • Check the data policy. Does the service keep your audio? Use it to train models? For how long? If you can't find a clear answer, assume the worst and don't upload anything sensitive.
  • Check the limits. Free tiers usually cap minutes or file size. A cap you hit halfway through a recording is a cap you'll hit at the worst moment.

Your computer or phone may also have built-in dictation, which some people use as a free option.

I'm not listing specific cloud services here on purpose. Free tiers change too often, and a list would be stale before you read it. The local stack above won't change on you.

Troubleshooting

"The transcript is gibberish." Check the audio first. Play the 16 kHz WAV. If it sounds bad to you, it'll be worse for the model. Re-record somewhere quieter. If the audio sounds fine, try a larger model or set the language explicitly.

"whisper-cli fails on my file." Remember that the whisper.cpp README says whisper-cli currently only takes 16-bit WAV. Run the ffmpeg conversion command above.

"pip install fails." Read the error. If it mentions Rust or setuptools_rust, the Whisper README covers it: install Rust, or pip install setuptools-rust. If it's a Python version issue, the README says the code is expected to work with Python 3.8 to 3.11.

"The transcript repeats the same line over and over." The Whisper model card says the architecture is prone to repetitive text. Try VAD in whisper.cpp to skip silence, split the file into shorter pieces, or try a different model size.

"Ollama is really slow." Use a smaller model. Check ollama ps. Ollama's FAQ says that command shows whether a model loaded on GPU, CPU, or split between them. If it's mostly CPU on a weak machine, a smaller model will feel much better.

"The draft has stuff I didn't say." Tighten the prompt. Add "Do NOT add any facts" in capital letters. Use the cleanup prompt first, then the outline prompt, then the draft prompt, as separate steps. Smaller steps mean fewer chances to wander.

A sample week

Here's what this looks like in practice once it's set up.

Monday through Thursday: Record voice notes whenever an idea shows up. Three to six minutes each. Say the working title at the top.

Friday morning, 20 minutes: Move the week's files to your computer. Run the conversion and transcription. Let it run while you make coffee.

Friday, 30 minutes: Run the cleanup prompt on each transcript. Skim. Group the notes that belong to the same piece.

Friday, 30 minutes: Run the outline and rough-draft prompts on the best group.

Next week: Edit that draft by hand, a little each day. Record a voice note to fill any [THIN] section.

That's roughly an hour and a half of machine-assisted work to get from a week of rambling to a real draft.

What this workflow won't fix

  • Thin ideas. If you didn't have much to say, cleanup won't add substance. It'll just make the thinness easier to read.
  • Wrong facts. The tools faithfully transcribe and politely rephrase what you said, including your mistakes.
  • Your voice. It'll get flattened if you let the model write the final. Edit it back in.
  • Other people's privacy. Local tools make it easy to transcribe anything. That doesn't make it right to. Get consent.

Short answers

Can I transcribe voice notes for free without a credit card? Yes. Whisper and whisper.cpp are open source and run locally. Whisper's README says its code and model weights are released under the MIT License, and whisper.cpp is MIT-licensed too.

Which Whisper model should I use for English? Start with base.en or small.en. The Whisper README says English-only models tend to do better for English, especially at the smaller sizes. Step up only if the transcript isn't editable.

Do I need a GPU? No. whisper.cpp supports CPU-only inference. It's slower, but it works. Smaller models help a lot.

Can a local model turn my transcript into a blog post? It can turn it into a rough draft. You still edit. Tell it not to invent facts, and check every number.

Is Whisper perfect? No. Its own model card warns that it can produce text that wasn't spoken. Check anything that matters against the audio.

The bottom line

The free path from voice to draft is real: phone recorder, ffmpeg, Whisper or whisper.cpp, a small local model in Ollama, and you. No card, no meter, no upload of your private rambling to someone else's server. Start with the smallest model that gives you an editable transcript. Tell the language model to flag gaps instead of filling them. Keep the audio so you can check every line. Then edit until it sounds like you, because it should.

References

  • openai/whisper — https://github.com/openai/whisper (fetched 2026-10-01)
  • Whisper model card — https://github.com/openai/whisper/blob/main/model-card.md (fetched 2026-10-01)
  • ggml-org/whisper.cpp — https://github.com/ggml-org/whisper.cpp (fetched 2026-10-01)
  • whisper.cpp CLI — https://github.com/ggml-org/whisper.cpp/tree/master/examples/cli (fetched 2026-10-01)
  • Ollama FAQ — https://docs.ollama.com/faq (fetched 2026-10-01)
  • Ollama context length — https://docs.ollama.com/context-length (fetched 2026-10-01)
  • Ollama thinking — https://docs.ollama.com/capabilities/thinking (fetched 2026-10-01)
  • Ollama library: qwen3 — https://ollama.com/library/qwen3 (fetched 2026-10-01)
  • Ollama library: mistral — https://ollama.com/library/mistral (fetched 2026-10-01)
Frequently asked
What is Free Transcription to Draft Workflow Without a Card about?
A lot of good writing starts as talking. You're walking the dog, driving to a job site, standing at a hive with your gloves on, and the idea finally comes out…
What should you know about what this workflow is, and what it isn't?
Let's name the pipeline so we can talk about each piece:
What should you know about why local first?
There are three reasons I push the local path before any cloud option.
What should you know about step 1: Capture without overthinking it?
Use the voice recorder that came with your phone. Every mainstream phone has one.
What should you know about a note on consent?
If your recording includes other people, ask first. That's basic decency, and in some places recording laws are strict about it.
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room