Redis is the de‑facto in‑memory data store for everything from real‑time analytics to session management, and its speed is legendary. Yet speed alone isn’t enough for production workloads that must survive power failures, host crashes, or network partitions. That’s where Redis persistence comes in: a set of mechanisms that periodically copy the in‑memory dataset to durable storage, allowing the server to restart without losing critical state.
For developers, architects, and operations teams the choice between RDB snapshots, AOF (Append‑Only File) logging, and the newer hybrid persistence models isn’t just a checkbox. It determines how much data you might lose in a crash, how long a restart will take, the amount of I/O your storage subsystem must sustain, and ultimately whether Redis can serve as the single source of truth for mission‑critical applications—whether you’re tracking the hive health of a thousand bee colonies, persisting the world‑state of an autonomous AI agent, or simply caching user‑profile data for a web app.
In this pillar article we’ll unpack each persistence option in depth, compare their durability guarantees, examine restart‑time characteristics, and map them to real‑world use cases. By the end you’ll have a concrete decision framework—backed by numbers, examples, and best‑practice patterns—so you can pick the right persistence strategy for any Redis deployment.
1. The Persistence Landscape: Why Redis Needs to Write to Disk
Redis stores data in RAM for nanosecond‑level reads and writes, but RAM is volatile. When the process receives a SIGTERM, the operating system powers off, or a hardware fault strikes, everything in memory disappears. Persistence bridges that gap.
There are three primary ways Redis can guarantee durability:
| Mechanism | Primary File(s) | How It Works | Typical Write Frequency |
|---|---|---|---|
| RDB (Redis Database) | dump.rdb | Periodic point‑in‑time snapshots of the entire dataset using a forked child process. | Configurable intervals (e.g., every 5 minutes or after 100,000 writes). |
| AOF (Append‑Only File) | appendonly.aof | Every write command is appended to a log file; the file can be rewritten to compact size. | Every command (with optional fsync policies). |
| Hybrid (RDB + AOF) | dump.rdb + appendonly.aof (or combined in Redis 7’s “mixed” mode) | Uses both snapshots and incremental logs to get fast restarts and minimal data loss. | Snapshots on schedule + AOF for every write. |
The choice is not binary; many production deployments run both RDB and AOF simultaneously. The hybrid approach, introduced in Redis 6.2 and refined in Redis 7, lets you enjoy the quick start‑up of RDB while still guaranteeing that the last few milliseconds of writes survive a crash.
Before diving into each method, let’s clarify two key metrics that shape every decision:
- Durability window – the maximum amount of data that could be lost after a crash. For RDB this is the time since the last snapshot; for AOF it’s bounded by the
fsyncpolicy (always,everysec, orno). - Recovery time – how long Redis needs to load the persisted state back into memory. RDB typically restores in seconds for gigabytes of data; AOF may take longer because it must replay every command.
Understanding these numbers in the context of your workload will guide you toward the optimal persistence configuration.
2. RDB Snapshots: The “Take a Picture” Model
2.1 How RDB Works Under the Hood
When Redis receives a snapshot trigger—either a time‑based rule (save 900 1 means “snapshot if at least one key changed in the last 15 minutes”) or an explicit BGSAVE command—it forks a child process. The child inherits the parent’s memory map (thanks to copy‑on‑write, or COW), then walks the in‑memory dataset and writes a binary representation to dump.rdb. Because the child works on a snapshot of the memory, the parent can continue serving client requests without blocking.
The snapshot format is compact: keys are stored with their data type, expiration, and value in a length‑prefixed binary layout. A 10 GB dataset typically results in a 3–4 GB RDB file after compression (Redis uses a simple LZF compression for strings). The exact size depends on data density and the prevalence of small values.
2.2 Performance Impact
- CPU – The fork operation incurs a brief spike (usually < 0.5 seconds) as the kernel creates the child’s page tables. The subsequent serialization is CPU‑bound but runs in the child, leaving the parent’s CPU cycles free for client traffic.
- I/O – RDB writes are sequential, making them well‑suited for SSDs and even high‑throughput HDDs. A typical 5 GB snapshot on a SATA SSD (~500 MB/s) finishes in ~10 seconds.
- Latency – Because the parent process never blocks on the write, latency impact is negligible for most workloads. However, heavy write traffic can increase COW overhead, causing the parent’s memory usage to temporarily double.
2.3 Durability Guarantees
RDB guarantees point‑in‑time durability: if a crash occurs after a successful snapshot, the dataset can be restored to exactly that moment. Anything written after the snapshot is lost. In practice, many teams configure multiple snapshot rules (e.g., “every 5 minutes if ≥ 10 000 writes” and “every hour if any change”) to keep the loss window under a few minutes.
2.4 Restart Recovery Time
Loading an RDB file is essentially a memory‑map operation: Redis reads the binary file, allocates the appropriate data structures, and populates them. For a 10 GB dataset, load times on a modern server (8 vCPU, 32 GB RAM, NVMe SSD) are typically 2–4 seconds. This speed makes RDB ideal for services that need to spin up quickly after a failover, such as a primary node in a Redis Sentinel cluster.
2.5 When RDB Shines
- Cold‑start environments – e.g., a fleet of edge devices that need to bootstrap quickly after power loss.
- Periodic backups – RDB files are easy to copy to remote storage (S3, NFS) for disaster recovery.
- Read‑heavy workloads – where write latency is already low and the occasional snapshot cost is acceptable.
Example: A beekeeping telemetry platform collects temperature, humidity, and hive weight every 30 seconds from 1 200 sensors. The data is stored in Redis hashes and expires after 48 hours. The team configures RDB snapshots every 10 minutes. In the event of a node crash, at most 10 minutes of sensor readings are lost—a trade‑off they accept because the data is also persisted to a downstream time‑series database.
3. AOF Logging: The “Write‑Ahead Log” Model
3.1 Mechanics of Append‑Only Files
AOF takes a different philosophy: every write command (SET, HINCRBY, LPUSH, etc.) is appended to a log file as soon as it is processed. The log is a plain‑text sequence of Redis protocol commands, making it human‑readable and easy to reconstruct. The file grows linearly with the number of write operations.
To keep the file from ballooning, Redis performs an AOF rewrite (BGREWRITEAOF) that creates a compacted version of the log. The rewrite runs in a background child process, similar to BGSAVE, and writes a new file containing only the minimal set of commands needed to rebuild the current dataset (e.g., a single SET for each key instead of thousands of incremental updates). Once the rewrite finishes, the parent atomically swaps the old AOF with the new one.
3.2 fsync Policies and Durability
AOF’s durability hinges on how often Redis calls fsync(2) to flush the OS buffers to disk:
| Policy | Description | Typical Data‑Loss Window |
|---|---|---|
| always | fsync after every write command. Guarantees zero data loss (barring hardware failure). | ≈ 0 ms (subject to disk latency). |
| everysec | fsync once per second (default). Balances durability and throughput. | ≤ 1 second of writes may be lost. |
| no | No explicit fsync; relies on OS flushing. Fastest, but can lose up to several seconds or minutes of data. | Variable, up to OS flush interval. |
The everysec policy is the most common production choice because modern SSDs can sustain ~30 k IOPS, and a single fsync per second adds negligible latency.
3.3 Performance Profile
- Write latency – With
fsync=everysec, the additional latency per command is typically < 0.1 ms. Withfsync=always, latency rises to 0.5–1 ms on commodity SSDs, and can become a bottleneck on HDDs. - CPU – AOF rewriting is CPU‑intensive because it must iterate over the entire dataset to emit a minimal command set. However, this occurs in a background child, keeping the parent responsive.
- Disk I/O – Append‑only writes are sequential, but the file can become large quickly. A 10 GB dataset with high write churn (e.g., 1 M writes per minute) can generate a 5 GB AOF in under an hour before a rewrite.
3.4 Recovery Time
During restart, Redis reads the AOF line by line and replays each command. This replay is slower than loading an RDB snapshot because it must execute each operation rather than bulk‑allocate structures. For a 5 GB AOF containing 20 million commands, recovery can take 10–20 seconds on a typical server. The exact time depends on command complexity (e.g., ZADD vs. SET) and CPU speed.
3.5 When AOF Is the Right Choice
- Near‑zero data loss – Financial tick data, real‑time inventory, or any system where losing even a single write is unacceptable.
- Write‑heavy workloads – Where the cost of occasional AOF rewrites is outweighed by the guarantee that every operation is persisted.
- Operational simplicity – AOF files can be inspected with
catorredis-check-aof, making debugging easier.
Example: An AI‑driven swarm simulation runs on a cluster of Redis nodes, each storing the position and state vector of 10 000 autonomous agents. The simulation must be able to resume from the exact last tick after a node crash, otherwise the emergent behavior diverges. The team enables AOF with fsync=always on a high‑end NVMe drive, accepting the 1 ms write overhead to guarantee deterministic replay.
4. Hybrid Persistence: Combining the Best of Both Worlds
4.1 What “Hybrid” Means in Redis
Starting with Redis 6.2, you can enable both RDB snapshots and AOF logging simultaneously. In Redis 7 the feature is refined with a “mixed” persistence mode that writes the initial dataset as an RDB snapshot and then logs subsequent writes to an AOF. The resulting file layout looks like:
[dump.rdb header + dataset] + [appendonly.aof tail]
During restart Redis first loads the RDB portion (fast) and then replays the AOF tail (tiny, often a few megabytes). This gives you sub‑second recovery while still preserving a near‑zero data‑loss window.
4.2 Configuring Hybrid Persistence
# redis.conf
save 300 10 # RDB snapshot every 5 minutes if ≥10 keys changed
appendonly yes
appendfsync everysec
# Enable mixed mode (Redis 7+)
aof-use-rdb-preamble yes
With this setup:
- An RDB snapshot is taken every 5 minutes.
- Every write is appended to the AOF tail.
- The AOF tail is automatically truncated after each successful snapshot, keeping it small.
4.3 Performance and Durability Trade‑offs
| Metric | Hybrid (mixed) | Pure RDB | Pure AOF |
|---|---|---|---|
| Data‑loss window | ≤ 1 second (fsync=everysec) | Up to snapshot interval | ≤ 1 second (or 0 ms with always) |
| Restart time | RDB load + tiny AOF replay (≈ 2–3 s) | RDB load only (≈ 2–4 s) | Full AOF replay (≈ 10–20 s) |
| Disk usage | RDB + small AOF (≈ RDB + few MB) | RDB only | AOF may be larger than RDB |
| CPU during normal ops | Minimal (snapshot + async append) | Minimal | Background rewrite can be CPU‑heavy |
Hybrid persistence is particularly attractive for high‑availability clusters where failover must happen quickly, but the business cannot tolerate more than a second of data loss.
4.4 Real‑World Use Cases
- Bee‑colony monitoring platform – Sensors push a new reading every 15 seconds. The platform uses hybrid persistence so that a failover to a replica recovers in < 3 seconds while guaranteeing that the last reading isn’t lost.
- Self‑governing AI agents – Each agent’s internal state (policy parameters, last actions) is stored in a Redis hash. Hybrid persistence allows the orchestrator to restart an agent after a crash and continue from the exact previous tick, preserving learning continuity.
5. Durability Trade‑offs: Quantifying What You Might Lose
5.1 The “Loss Window” Calculus
| Persistence | Typical loss window (default) | How to shrink it |
|---|---|---|
| RDB | Time since last snapshot (e.g., 5 min) | Add more frequent save rules, or trigger BGSAVE on critical events. |
AOF (everysec) | ≤ 1 second | Switch to always (requires fast storage) or use a battery‑backed write cache. |
| Hybrid (RDB + AOF) | ≤ 1 second (AOF tail) | Keep appendfsync=everysec; optionally enable appendfsync=always on premium NVMe. |
A concrete illustration: a Redis instance handling 200 k writes per second (e.g., a live leaderboard). With appendfsync=everysec, the maximum data loss is roughly 200 k commands—about 1 second worth. If the business can’t tolerate any loss, you must move to always (and ensure the storage can sustain 200 k fsyncs per second, which typically requires a high‑end NVMe RAID or a RAM‑disk with battery backup).
5.2 Consistency Guarantees
Redis is single‑threaded for command execution, which simplifies consistency: each command is fully applied before the next begins. However, persistence introduces asynchronous steps:
- RDB – The snapshot is taken after writes have been applied, but the child process may see a slightly older view because of COW. The snapshot is always a consistent view of the dataset at the moment the fork started.
- AOF – The log is appended before the command is executed (write‑ahead). If the server crashes after the
fsyncbut before the command finishes, the AOF may contain a command that never took effect, but Redis’s replay will apply it anyway, preserving idempotence for most commands.
Understanding these nuances is essential when building transactional workflows on top of Redis (e.g., using MULTI/EXEC). The atomicity of a transaction is guaranteed in memory, but if you rely on persistence to survive a crash, you must ensure the underlying persistence mode can capture the entire transaction before a failure.
5.3 Impact on Replication and High Availability
Redis replication (master‑replica) works independently of persistence: replicas receive the command stream from the master via the replication protocol. However, persistence influences failover:
- RDB‑only – A replica that is promoted to master must load its own RDB file, incurring the full restart time.
- AOF‑only – The replica replays its AOF, potentially taking longer.
- Hybrid – The promoted replica loads the RDB snapshot and then applies a tiny AOF tail, achieving the fastest possible switchover.
When using redis-sentinel or redis-cluster, many operators enable both RDB and AOF precisely to keep failover times under the Service Level Objective (SLO) threshold (often < 5 seconds).
6. Restart and Recovery Scenarios: From Cold Boot to Warm Failover
6.1 Cold Start (Power‑On)
- RDB – Load
dump.rdb→ memory ready → serve traffic. - AOF – Replay entire
appendonly.aof→ memory ready → serve traffic. - Hybrid – Load RDB → replay AOF tail → memory ready → serve traffic.
Timing example: A 12 GB dataset on an NVMe drive:
| Mode | Load time | Disk reads | CPU usage |
|---|---|---|---|
| RDB | 3.2 s | 12 GB sequential read | 5 % |
| AOF | 14.8 s | 12 GB sequential + 20 M command parses | 30 % |
| Hybrid | 3.4 s (RDB) + 0.2 s (AOF tail) | 12 GB + 5 MB | 7 % |
The hybrid approach adds only a fraction of a second to the cold start, making it the sweet spot for services that cannot afford a 10‑second outage.
6.2 Warm Failover (Replica Promotion)
In a typical Sentinel setup, a replica is already loaded in memory. When the master fails, the replica is promoted instantly, but it still needs to persist its state for future restarts. If the replica has been running with hybrid persistence, its AOF tail is already tiny, so the newly promoted master can write a fresh snapshot within seconds, ensuring the cluster remains ready for another failover.
6.3 Disk Failure and Recovery
If the persistent volume fails, you lose the on‑disk state. However, you can re‑hydrate from a remote backup:
- RDB – Copy the latest
dump.rdbfrom S3, place it in the data directory, and start Redis. - AOF – Transfer the
appendonly.aoffile; you may need to runredis-check-aof --fixif the file was truncated. - Hybrid – Transfer both files; the AOF tail will be tiny, making verification fast.
A best practice is to schedule daily RDB backups to object storage and retain hourly AOF increments for the last 24 hours. This combination lets you restore to any point within the past day with minimal data loss.
7. Matching Persistence to Use Cases
Below is a matrix that aligns common Redis workloads with the persistence mode that best satisfies their requirements.
| Workload | Write Rate | Acceptable Data Loss | Desired Restart Time | Recommended Persistence |
|---|---|---|---|---|
| Session store for a web app (10 k ops/s) | Moderate | ≤ 5 seconds | < 2 seconds | RDB (snapshot every 5 min) + optional AOF (everysec) |
| Real‑time leaderboards (100 k ops/s) | High | ≤ 1 second | < 5 seconds | Hybrid (RDB + AOF everysec) |
| Financial tick stream (1 M ops/s) | Very high | 0 ms | < 10 seconds (acceptable) | AOF always on high‑end NVMe; consider RedisRaft for strong consistency |
| IoT sensor aggregation (5 k ops/s, 48‑hour TTL) | Low‑moderate | ≤ 1 minute | < 3 seconds | RDB snapshots every 10 minutes + AOF everysec |
| AI agent world‑state (state updates every 100 ms) | Moderate | ≤ 100 ms | < 1 second | Hybrid with appendfsync=always (if storage permits) |
| Cache for static assets (read‑only after warm‑up) | Low writes | Any (cache can be rebuilt) | < 1 second | RDB only (snapshot on startup) |
| Geo‑spatial indexing for wildlife tracking (writes 20 k/s) | Moderate‑high |