By Austin Little
A Raspberry Pi is a small, low-power computer that a lot of us already own, sitting in a drawer or running a home project. People keep asking whether it can run an AI model locally, with no cloud and no subscription. The honest answer is yes, within limits that matter. This guide walks through those limits using the official specs, without speed numbers we haven't measured ourselves.
AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events.
The short answer
- Yes, a Raspberry Pi can run a small local language model. Both Ollama and llama.cpp can run on 64-bit ARM Linux, which is what a modern Pi runs.
- The model has to be small. "Small" here means roughly the 1-billion to few-billion parameter range, not the giant models behind big cloud chat services.
- RAM is the first wall. The model has to fit in memory along with the operating system. More RAM means more room for a bigger model.
- Speed is the second wall. A Pi will generally be slower than a modern laptop or desktop at this. How much slower depends on the model, the Pi, cooling, and settings. We are not quoting a tokens-per-second figure because we have not benchmarked one.
- There's now an official add-on for generative AI. Raspberry Pi's AI HAT+ 2 carries its own accelerator and 8GB of onboard RAM and is designed to run LLMs and vision-language models on a Raspberry Pi 5.
- It's a great learning box, a fine private helper for short tasks, and a poor replacement for a frontier chat assistant.
What "local AI model" means here
When people say "run AI on a Pi," they usually mean one of three things:
- A small chat model (an LLM) that answers questions, rewrites short text, or summarizes a paragraph, all offline.
- A vision model that detects objects in a camera feed, like a bird-feeder camera that notices when something lands.
- Speech tools like speech-to-text or text-to-speech.
This guide focuses mainly on number 1, local language models, because that's what most people mean by "can it run ChatGPT." Vision models on a Pi are a well-trodden path with their own official accessories; we touch on that below.
What the official specs say
We checked Raspberry Pi's own product pages on 2026-10-01.
Raspberry Pi 5
According to the official Raspberry Pi 5 product page:
- CPU: Broadcom BCM2712, 2.4GHz quad-core 64-bit Arm Cortex-A76, with 512KB per-core L2 caches and a 2MB shared L3 cache.
- GPU: VideoCore VII, supporting OpenGL ES 3.1 and Vulkan 1.3.
- RAM options: LPDDR4X-4267 SDRAM in 1GB, 2GB, 4GB, 8GB, and 16GB variants.
- Storage and expansion: microSD slot, plus a PCIe 2.0 x1 interface for fast peripherals (which needs a separate M.2 HAT or other adapter), so you can attach an M.2 SSD.
- Power: 5V/5A DC via USB-C with Power Delivery support. The FAQ recommends a high-quality 5V 5A USB-C supply and notes you may have problems with an under-powered supply.
- Cooling: The FAQ says that, like most general-purpose computers, the Pi 5 performs best with active cooling.
Raspberry Pi 4
The official Raspberry Pi 4 Model B spec page lists a Broadcom BCM2711 quad-core Cortex-A72 (ARM v8) 64-bit SoC at 1.8GHz, with 1GB, 2GB, 3GB, 4GB, or 8GB of LPDDR4-3200 SDRAM depending on model.
A Pi 4 can run small models too, but it's an older, slower chip than the Pi 5. If you're buying new specifically for local AI, the Pi 5 is the more sensible starting point. If you already have a Pi 4 with 4GB or 8GB, it's worth trying the smallest models before spending anything.
Older and smaller boards
Older boards with less memory, or 32-bit operating systems, are a much harder fit. Ollama's Linux install docs offer an ARM64 package, which points to 64-bit ARM as the supported path.
Why RAM is the first wall
A language model is, at heart, a very large file of numbers. To answer you, the computer has to load that file into memory and work through it for every word it produces.
Two things shrink that file:
- Fewer parameters. A 1-billion-parameter model is much smaller than a 7-billion or 70-billion one.
- Quantization. This stores the numbers at lower precision. The llama.cpp project lists support for 1.5-bit through 8-bit integer quantization "for faster inference and reduced memory use."
To make this concrete, look at Ollama's own model library pages, which list download sizes. On 2026-10-01:
gemma3:270mwas listed at 292MB.gemma3:1bat 815MB.llama3.2:1bat 1.3GB.llama3.2:3b(the defaultllama3.2tag) at 2.0GB.gemma3:4bat 3.3GB.gemma3:12bat 8.1GB.
Those are download sizes, not exact memory requirements. Running a model usually needs more memory than the file size, because the program also has to hold working memory for the conversation (the "context").
So what does that mean in practice?
- A 1GB or 2GB Pi is a very tight fit for a language model. The very smallest models might load, but you'll have little room to spare.
- A 4GB Pi gives you room for the small end of the model range.
- An 8GB or 16GB Pi gives you more headroom, either for a somewhat larger model or for more breathing room with a small one.
More RAM doesn't make the CPU faster. It lets you load a bigger model, and a bigger model on the same CPU is usually slower. That's the trade-off you'll keep bumping into: bigger and smarter, or smaller and quicker.
Why speed is the second wall
On a laptop with a strong GPU, the graphics chip does most of the heavy math. On a stock Raspberry Pi, the model runs on the CPU.
llama.cpp's README describes the project as aiming for LLM inference "with minimal setup and state-of-the-art performance on a wide range of hardware," and it lists ARM NEON optimization (in the context of Apple silicon) and many backends, including Vulkan. The Pi 5's GPU supports Vulkan 1.3 according to the official spec.
What we can say without numbers:
- Expect waiting. For short questions with a small model, many people find the speed usable for patient tasks. For long answers or big models, it can feel slow.
- The first response takes longest. Loading the model from storage into memory takes time, especially from a microSD card.
- Heat matters. Running a model keeps the CPU busy. Raspberry Pi's own FAQ says the Pi 5 performs best with active cooling. A Pi that's overheating will slow itself down.
- Power matters. An under-powered supply can cause instability. The official recommendation is a quality 5V 5A USB-C supply.
The AI HAT+ 2: what it changes
Raspberry Pi sells AI accelerator boards that sit on top of a Pi 5. There are two families, and the difference matters for language models.
According to Raspberry Pi's official documentation:
- AI HAT+ uses a Hailo-8L (13 TOPS, INT8) or Hailo-8 (26 TOPS, INT8) chip. It uses the Pi 5's memory, and the documentation's comparison table lists LLM support as not supported. It's built for vision tasks like object detection, pose estimation, and camera post-processing.
- AI HAT+ 2 uses a Hailo-10H accelerator rated at 40 TOPS (INT4), with 8GB of its own onboard RAM. The documentation lists LLM and vision-language model support as supported, and describes it as able to run LLMs and VLMs up to about 6 billion parameters.
Raspberry Pi's announcement post for the AI HAT+ 2 makes the scale honest: it says edge LLMs on the AI HAT+ 2, sized to fit the onboard RAM, typically run at 1 to 7 billion parameters, and that these smaller models are not designed to match the knowledge of the much larger cloud models. That's exactly the right expectation to carry.
Some practical notes:
- The AI HAT+ 2 is for the Raspberry Pi 5, connecting over the Pi 5's PCIe interface.
- Raspberry Pi says the software for compatible generative AI models is in Hailo's GitHub repositories, and Hailo provides a set of sample models. That means you're generally choosing from models prepared for the Hailo chip, not running any model file you like.
- It's an extra purchase. It's optional.
Do you need an accelerator?
No. You can run small models on the Pi's CPU with free software. An accelerator is something to consider later if you've tried the CPU route, you know what you want to build, and you've checked that the models you need are supported on it.
Software: Ollama and llama.cpp on a Pi
Both of these are free and open source.
Ollama
Ollama's official Linux documentation gives a one-line install script and also a manual path, including a specific ARM64 install package (ollama-linux-arm64.tar.zst). It also explains how to run Ollama as a background service with systemd, how to update it, view logs, and uninstall it.
A typical session looks like:
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma3:1b
Before you run any install script from the internet, it's reasonable to open it in a browser and look at it first, or use the manual install steps in Ollama's docs instead.
llama.cpp
llama.cpp is the engine many local AI tools are built on. It's plain C/C++ "without any dependencies" according to its README, and it supports a wide range of hardware. Building it yourself on a Pi gives you the most control over settings, like how many CPU threads to use and how much context to allow.
It's more hands-on than Ollama. If you're comfortable with a terminal and like tinkering, it's worth a look. If you just want to chat with a small model, start with Ollama.
Storage tip
Models are large files, and loading them from a slow microSD card takes time. The Pi 5's PCIe interface supports an M.2 SSD through a separate HAT or adapter, which the official page describes as giving "speedy data transfer." If you plan to swap models often, faster storage helps with load times. It doesn't change how fast the model "thinks" once loaded.
A realistic first project
Here's a gentle way to find out what your Pi can do, without buying anything new.
Step 1: Check what you have
- Which Pi model? How much RAM? (On Raspberry Pi OS, running
free -hin a terminal shows memory.) - Is the OS 64-bit? (
uname -mshould sayaarch64on a 64-bit ARM system.) - Do you have decent cooling and the right power supply?
- How much free storage? (
df -h.)
Step 2: Start with the smallest model
Pick a model listed in the hundreds of megabytes or around 1GB on Ollama's library. Ask it something simple: "Rewrite this sentence to be friendlier," or "Give me three name ideas for a bee garden."
Step 3: Notice how it behaves
- How long until the first word appears?
- Does the answer make sense?
- Does the Pi get hot?
- Does anything else on the Pi slow down?
Step 4: Step up one size, carefully
If the small model works and you have RAM to spare, try the next size up. Watch memory use (htop is a friendly tool for this). If the system starts swapping to disk heavily, you've gone too big.
Step 5: Decide what it's for
The best Pi AI projects are narrow:
- A private "rewrite this kindly" helper on your home network.
- A small offline assistant for a classroom or a cabin with spotty internet.
- A learning lab for understanding how models, prompts, and quantization work.
- A summarizer for short local notes you don't want to send to the cloud.
What a Pi model is good at, and what it isn't
Reasonable jobs
- Short rewrites and tone changes. "Make this text message more polite."
- Simple brainstorming. Names, lists, ideas.
- Basic summarizing of short passages you paste in.
- Learning and experimenting. Seeing how changing prompts or models changes output.
- Privacy-first tasks where "slow but local" beats "fast but cloud."
Poor fits
- Facts you need to rely on. Small models are more likely to get facts wrong, and they can sound sure while doing it. Always check facts against a real source.
- Long documents. Big inputs need more memory and more time.
- Coding help on complex projects. Possible with small coding models, but expect limits.
- Anything safety-critical: medical, legal, financial, electrical. That's true for every AI model, not just Pi-sized ones.
- Replacing a frontier chat assistant. Raspberry Pi's own AI HAT+ 2 announcement says these smaller models aren't designed to match the knowledge set of the larger cloud models.
Honest trade-offs compared to other free options
You don't need a Pi to use AI for free. Here's how it compares.
A Pi vs. your existing laptop or desktop
If you have a reasonably modern laptop or desktop, it will usually run local models more comfortably than a Pi, especially if it has a lot of RAM or a capable GPU. Ollama runs on Mac, Windows, and Linux.
A Pi vs. free cloud chat
Free cloud chat tools run much bigger models and will usually be faster and more knowledgeable. The trade-off is privacy and dependence: your text goes to someone else's server, under their terms and limits. A Pi keeps text at home.
A Pi vs. a free API tier
Some providers offer free API tiers with rate limits. Those can be great for projects that need a stronger model.
Apiary's general rule holds: bring your own model when you can, and pick the smallest tool that does the job honestly.
Safety and housekeeping
- Keep the OS updated. A Pi on your network is a computer on your network. Update Raspberry Pi OS regularly.
- Don't expose the model API to the internet. If you open Ollama's API to your local network, keep it on your local network. Don't port-forward it to the public internet without understanding the risks.
- Mind what you paste. Local is more private, but a Pi that others can access is not a vault. Keep passwords and one-time codes out of prompts everywhere.
- Check model licenses. Open models come with licenses that set rules for use. Read them for anything beyond personal tinkering.
- Watch the heat. Use the recommended cooling, especially for long sessions.
Myths to drop
"A Pi can run ChatGPT." It can run a small open model in the same general family of technology. It's not the same model, and it doesn't have the same knowledge.
"More RAM makes it fast." More RAM lets you load bigger models. Bigger models on the same CPU usually go slower.
"You need the AI HAT to run any language model." No. Small models run on the CPU with free software. The AI HAT+ 2 is optional acceleration with its own model support list. The original AI HAT+ doesn't support LLMs at all, per Raspberry Pi's documentation.
"TOPS tells you how fast chat will be." TOPS is a hardware rating. Real chat speed depends on the model, the software, memory, and how the work is split between chips. Don't translate a TOPS number into a promise.
"If it runs, it's accurate." Running and being right are different things. Small models are useful helpers and unreliable authorities.
Quick answers
"Can a Raspberry Pi 5 run a local LLM?" Yes, small ones, on the CPU with Ollama or llama.cpp. Expect it to be patient rather than fast. Get good cooling and a proper power supply.
"How much RAM does a Raspberry Pi need for AI?" More is better. A 4GB board can handle the smallest models; 8GB or 16GB gives more headroom. Actual needs depend on the model and settings. Check model download sizes on Ollama's library and leave room for the OS.
"Does Ollama work on Raspberry Pi?" Ollama's Linux docs include an ARM64 install package, which is the architecture a 64-bit Raspberry Pi OS uses.
"Is the Raspberry Pi AI HAT+ 2 worth it?" It's the official way to accelerate generative AI on a Pi 5, with its own 8GB of RAM and support for models up to about 6 billion parameters, per Raspberry Pi's documentation. It's worth it if you've outgrown the CPU route and the models you want are supported. It's not required.
"Can a Raspberry Pi Zero or Pi 3 run AI?" Very small vision or speech tasks may be possible on older boards, but a language model is a tough fit with that little memory. Start with a Pi 4 or Pi 5 if you have one.
"How fast is it?" We're not quoting a number we haven't measured. Try the smallest model on your own Pi and see whether it's fast enough for your task.
Bottom line
A Raspberry Pi can run a local AI model, and that's genuinely delightful: a small, quiet computer answering you with no cloud and no bill. Go in with honest expectations. Choose a small model, give the Pi cooling and proper power, and use it for short, private, low-stakes jobs. If you want more, the official AI HAT+ 2 exists, and so does the laptop you may already own. Start with what's in the drawer. That's the most Apiary way to do it.
References
- Raspberry Pi 5 product page and specifications. https://www.raspberrypi.com/products/raspberry-pi-5/ (fetched 2026-10-01)
- Raspberry Pi 4 Model B specifications. https://www.raspberrypi.com/products/raspberry-pi-4-model-b/specifications/ (fetched 2026-10-01)
- Raspberry Pi AI HAT+ 2 product page. https://www.raspberrypi.com/products/ai-hat-plus-2/ (fetched 2026-10-01, via search)
- Raspberry Pi news: Introducing the Raspberry Pi AI HAT+ 2. https://www.raspberrypi.com/news/introducing-the-raspberry-pi-ai-hat-plus-2-generative-ai-on-raspberry-pi-5/ (fetched 2026-10-01, via search)
- Raspberry Pi documentation, AI HAT+ about page. https://github.com/raspberrypi/documentation/blob/master/documentation/asciidoc/accessories/ai-hat-plus/about.adoc (fetched 2026-10-01, via search)
- Ollama Linux documentation. https://docs.ollama.com/linux (fetched 2026-10-01)
- Ollama model library: gemma3, llama3.2. https://ollama.com/library/gemma3, https://ollama.com/library/llama3.2 (fetched 2026-10-01)
- llama.cpp README. https://github.com/ggml-org/llama.cpp (fetched 2026-10-01)