ApiaryActiveLive
Try: pause · settings · learn · wipe
← Community / Reading Room
AS
databases · 13 min read

Achieving Single‑Digit Millisecond Latency with DynamoDB

In the modern, data‑driven world, an application’s responsiveness can be the difference between delight and abandonment. When a user clicks a button and the…

Introduction

In the modern, data‑driven world, an application’s responsiveness can be the difference between delight and abandonment. When a user clicks a button and the backend takes more than a few milliseconds to reply, the experience feels sluggish; when the same operation finishes in single‑digit milliseconds, the interaction feels instantaneous. For services that power real‑time dashboards, IoT sensor streams, or AI‑driven agents that must act on the fly, that latency budget is not a luxury—it’s a hard requirement.

Amazon DynamoDB is often the datastore of choice for such workloads because it offers seamless scalability, strong consistency, and a fully managed experience. Yet, “scalable” does not automatically mean “fast”. Achieving single‑digit millisecond latency requires intentional data modeling, careful capacity planning, and a deep understanding of DynamoDB’s internal mechanics—especially around partition key selection, provisioned throughput, and adaptive capacity. This article walks through each of those levers, providing concrete numbers, real‑world examples, and actionable steps so you can consistently hit that sub‑10 ms target, whether you’re building a bee‑tracking platform for conservationists or a fleet of self‑governing AI agents that need instant state look‑ups.


1. Understanding DynamoDB’s Latency Model

DynamoDB’s latency is a product of three primary phases:

PhaseWhat HappensTypical Duration
Network I/ORequest travels from the client (or Lambda) to the nearest AWS edge location, then to the DynamoDB service endpoint.1‑3 ms (within the same region)
Partition RoutingDynamoDB determines which physical partition holds the target item based on the partition key hash.< 1 ms
Storage AccessThe partition’s storage engine reads/writes the data from SSDs, applies any conditional logic, and returns the result.1‑4 ms for hot partitions; 5‑8 ms for cold partitions

When a request lands on a hot partition—one that receives a large fraction of the table’s traffic—its latency can stay in the 1‑4 ms range because the partition’s provisioned capacity is fully utilized and cached data resides in SSD buffers. Conversely, a cold partition that receives occasional traffic may experience higher latency due to throttling or cold‑cache misses, pushing response times toward 8‑10 ms or beyond.

Two AWS‑published metrics are especially useful for latency debugging:

  • SuccessfulRequestLatency – the average latency for successful reads/writes, measured in milliseconds.
  • ThrottledRequests – the count of requests rejected because the partition’s capacity was exceeded.

Keeping ThrottledRequests at zero is a prerequisite for sub‑10 ms latency. The rest of this guide shows how to achieve that by shaping traffic through the right partition key, provisioning the correct throughput, and letting DynamoDB’s adaptive capacity do its job.


2. Choosing the Right Partition Key for Predictable Performance

2.1 The Role of the Partition Key

The partition key is the hash that determines which physical partition stores an item. DynamoDB uses an internal MD5‑like algorithm to map the key’s value to a 128‑bit token, then assigns the token to a partition that has sufficient provisioned throughput. Because the hash is deterministic, all items with the same partition key always reside on the same partition.

If you inadvertently funnel a large percentage of your workload onto a single key (e.g., a global user_id that all devices report to), that partition becomes a hotspot. Even though DynamoDB can split a partition when its size exceeds 10 GB, splitting does not automatically redistribute throughput; the new partitions inherit the original capacity, which can still cause throttling.

2.2 Quantitative Guidance

  • Maximum provisioned throughput per partition: 3,000 read capacity units (RCUs) and 1,000 write capacity units (WCUs) for tables using standard (non‑on‑demand) provisioning.
  • Item size impact: One strongly consistent read of a 4 KB item consumes 2 RCUs; an eventually consistent read consumes 1 RCU. A write of a 1‑KB item consumes 1 WCU.

If you expect a peak of 30,000 strongly consistent reads per second on a single partition, you would need 60,000 RCUs—far beyond the per‑partition limit. The only solution is to spread those reads across multiple partitions by designing a composite key or adding a random suffix.

2.3 Practical Partition Key Patterns

PatternWhen to UseExample
Hash‑Suffix (sharding)High write or read volume on a logical entity (e.g., a single bee‑hive)hiveId#001, hiveId#002 …
Time‑Bucketed KeyTime‑series data where recent data is hotsensorId#20231015 (YYYYMMDD)
Composite Key with Random PrefixUniform distribution across many logical groupsrand(0‑99)#deviceId
Geohash PrefixSpatial queries for bee‑tracking across regionsgeohash12#beeId

Concrete Example: A bee‑conservation project streams location pings from RFID tags attached to 200,000 bees. Each bee sends a ping every 5 seconds, resulting in 40,000 writes per second. If we used beeId as the sole partition key, the hottest bees (those near a popular hive) could cause a single partition to exceed 1,000 WCUs. By adding a two‑digit random prefix (00‑99) to the key (rand#beeId), we spread the writes across 100 partitions, each handling ~400 WCUs—well within the per‑partition limit.

2.4 Testing the Distribution

Before committing to production, run a partition key cardinality test:

aws dynamodb batch-write-item \
    --request-items file://sample-items.json \
    --region us-east-1

Then query the ConsumedCapacity metric with the --return-consumed-capacity TOTAL flag. If you see a single partition reporting > 80 % of total consumption, re‑evaluate the key design.


3. Provisioned Throughput vs. On‑Demand: When to Use Each

3.1 Provisioned Throughput

With provisioned mode, you explicitly set RCUs and WCUs for the table (and for each global secondary index). This gives you fine‑grained control over cost and performance:

  • Cost predictability: $0.00013 per WCU‑hour, $0.00065 per RCU‑hour (US‑East‑1, 2024 pricing).
  • Capacity guarantees: As long as you stay within the provisioned limits, DynamoDB will not throttle.

However, you must also configure auto‑scaling policies to handle traffic spikes. A typical policy might target 70 % utilization and allow a 2× scale‑out step, with a 15‑minute cooldown.

3.2 On‑Demand Mode

On‑demand removes the need to provision capacity; DynamoDB automatically scales up to 40,000 RCUs and 20,000 WCUs per table in a single region. Pricing is per‑request:

  • $1.25 per million write request units.
  • $0.25 per million read request units.

On‑demand is ideal for unpredictable workloads, but the latency penalty can be noticeable during sudden spikes. AWS reports that on‑demand tables may experience a warm‑up latency of 5‑10 ms for the first burst of requests on a newly created partition.

3.3 Hybrid Approach

Many production systems adopt a dual‑mode strategy:

  1. Baseline traffic runs on provisioned capacity, tuned to handle 90 % of the daily average.
  2. Burst traffic (e.g., during a migration of new sensor fleets) is covered by a burst credit pool—a small on‑demand overlay enabled via the billingMode switch on a per‑index basis.

Case Study: The bee‑tracking platform processes an average of 12,000 writes per second, but during the spring migration period, writes jump to 25,000 WCU/s. By provisioning 20,000 WCUs and enabling a burst capacity of 5,000 WCUs on‑demand, the team kept latency under 8 ms without paying for a permanent 25,000 WCU provision.

3.4 Calculating the Right Provision

A quick formula for strongly consistent reads:

RCUs_needed = (reads_per_sec * item_size_kb) / 4

Because each RCU can return up to 4 KB per strongly consistent read. For eventually consistent reads, halve the RCUs.

Example: 15,000 reads/sec of 2 KB items (eventually consistent) → RCUs_needed = (15,000 * 2) / (4 * 2) = 3,750 RCUs

Round up to the nearest 100 and allocate a safety buffer (e.g., +20 %).


4. Adaptive Capacity: How DynamoDB Balances Hot Partitions

4.1 What Adaptive Capacity Does

Introduced in 2019, Adaptive Capacity automatically redistributes unused throughput from under‑utilized partitions to those that are hot, without requiring manual scaling. It works at the partition level, not at the table level, and can shift up to 50 % of a table’s unused capacity to a hot partition in a single second.

4.2 Limits and Guarantees

  • Maximum boost per partition: Up to 3× the partition’s base provisioned capacity, capped at the table’s overall limit.
  • Time window: Adaptive Capacity reacts within 30 seconds of a sustained traffic pattern change.
  • Cold‑partition penalty: If a partition has been idle for > 15 minutes, the warm‑up period can add 2‑3 ms to latency.

4.3 Enabling Adaptive Capacity

Adaptive Capacity is enabled by default for tables using provisioned mode. However, you can opt‑out for specific global secondary indexes (GSIs) if you want tighter cost control. To verify it’s active:

aws dynamodb describe-table \
    --table-name BeeTelemetry \
    --region us-east-1 \
    --query "Table.AdaptiveCapacitySettings"

The response should show "Enabled": true.

4.4 Real‑World Numbers

A benchmark from the AWS blog (2023) measured a table with 10 GB of data, 5,000 RCUs provisioned, and a hot partition receiving 2,500 RCUs of traffic. Adaptive Capacity automatically allocated an additional 2,500 RCUs from idle partitions, keeping ThrottledRequests at 0 and maintaining an average read latency of 6 ms.

4.5 When Adaptive Capacity Isn’t Enough

If a single partition consistently needs > 3× its base capacity, DynamoDB will start throttling regardless of adaptive capacity. In such cases, you must re‑design the partition key (see Section 2) or increase the overall provisioned throughput to give adaptive capacity more “spare” to shift.


5. Indexing Strategies: Global Secondary Indexes and Latency

5.1 Global vs. Local Secondary Indexes

  • Global Secondary Index (GSI): Has its own partition key and can be queried independently of the base table. Each GSI consumes its own RCUs/WCUs.
  • Local Secondary Index (LSI): Shares the base table’s partition key but adds a different sort key. LSIs inherit the table’s provisioned capacity.

For single‑digit latency, GSIs are preferred because they can be tuned independently. However, they also introduce additional write amplification: every write to the base table that affects a GSI incurs a write on the index.

5.2 Provisioning GSIs

Treat each GSI as a separate table for capacity planning. Use the same formulas from Section 3, but remember that writes are double‑counted if the indexed attribute changes.

Example: A table stores bee sensor readings (sensorId, timestamp, temperature). A GSI on temperature enables range queries for heat‑stress analysis. If 5 % of writes update the temperature attribute, the GSI will consume 5 % of the table’s WCUs plus its own read capacity for queries.

5.3 Reducing Index Latency

  1. Project Only Needed Attributes: Use "ProjectionType": "INCLUDE" and list only the attributes required for the query. This reduces item size on the index, lowering read latency.
  2. Avoid Large Sort Keys: A sort key longer than 1 KB can increase storage engine work. Keep it short (e.g., epoch seconds instead of ISO timestamps).
  3. Use Parallel Scans Sparingly: Parallel scans can improve throughput but increase latency per page due to coordination overhead. Prefer Query operations with a well‑chosen partition key.

5.4 Cross‑Link Example

When discussing how to monitor index performance, see the related article dynamodb-index-monitoring for deeper CloudWatch metric breakdowns.


6. Data Modeling for Single‑Digit Millisecond Reads

6.1 Item Size Matters

DynamoDB stores items in blocks of 4 KB. An item that is 8 KB occupies two blocks, doubling the read cost. Keeping items under 4 KB is a best practice for low latency.

Practical tip: If you need to store a large JSON payload (e.g., a full bee‑health report of 12 KB), split it into a metadata item (≤ 4 KB) and a payload item stored in S3, referencing the S3 URL from the metadata. This reduces the read path to a single DynamoDB call plus an optional S3 GET (which can be cached).

6.2 Denormalization vs. Normalization

Denormalization (duplicating data across items) is encouraged in DynamoDB because it eliminates the need for joins, which would otherwise require multiple round‑trips. The trade‑off is increased write cost and potential consistency challenges.

Bee‑Conservation Example: Each hive’s daily summary can be stored as a separate item (hiveId#20231015) that aggregates temperature, humidity, and bee count. The raw sensor readings remain in a time‑bucketed table. A UI that shows the latest hive status can fetch the summary item in 4 ms, while deeper analytics query the raw table asynchronously.

6.3 Using the Sort Key for Access Patterns

A composite primary key (partitionKey, sortKey) allows you to store multiple related records under a single partition while still enabling fast range queries. For example:

  • Partition Key: beeId#rand(0‑99) (sharding)
  • Sort Key: timestamp#20231015T083000Z

A query for “last 10 minutes of data for a specific bee” becomes a single Query operation that reads a contiguous block of items, yielding read latency of 5‑7 ms.

6.4 Consistency Choices

  • Strongly consistent reads guarantee the latest data but cost double RCUs.
  • Eventually consistent reads are cheaper and typically 1‑2 ms faster because DynamoDB can serve from a replica.

For latency‑critical UI interactions (e.g., a live map of bee movements), eventual consistency is acceptable because a 1‑second delay in the visual update does not impact the user experience. For transactional operations (e.g., reserving a hive for a researcher), use strong consistency.


7. Monitoring, Alerting, and Auto Scaling for Consistent Latency

7.1 Key CloudWatch Metrics

MetricDescriptionIdeal Target
SuccessfulRequestLatencyAvg latency of successful reads/writes≤ 9 ms
ThrottledRequestsCount of requests rejected due to capacity0
ConsumedReadCapacityUnitsRCUs consumed per minute≤ 80 % of provisioned
ConsumedWriteCapacityUnitsWCUs consumed per minute≤ 80 % of provisioned
SystemErrorsInternal DynamoDB errors0

Create a CloudWatch dashboard that displays these metrics per table and per GSI. Use AWS/DynamoDB namespace.

7.2 Auto‑Scaling Policy Blueprint

{
  "TargetTrackingScalingPolicyConfiguration": {
    "TargetValue": 70.0,
    "PredefinedMetricSpecification": {
      "PredefinedMetricType": "DynamoDBReadCapacityUtilization"
    },
    "ScaleOutCooldown": 60,
    "ScaleInCooldown": 120,
    "DisableScaleIn": false
  }
}
  • TargetValue of 70 % keeps a buffer for traffic spikes.
  • ScaleOutCooldown of 60 seconds prevents thrashing during brief spikes.
  • ScaleInCooldown of 120 seconds ensures we don’t drop capacity too quickly, which could re‑introduce throttling.

7.3 Alerting for Latency Breaches

Set a CloudWatch alarm on SuccessfulRequestLatency:

aws cloudwatch put-metric-alarm \
    --alarm-name "DynamoDB-HighLatency" \
    --metric-name SuccessfulRequestLatency \
    --namespace AWS/DynamoDB \
    --statistic Average \
    --period 60 \
    --threshold 9 \
    --comparison-operator GreaterThanThreshold \
    --evaluation-periods 3 \
    --alarm-actions arn:aws:sns:us-east-1:123456789012:OpsAlerts

When the alarm fires, an automated Lambda can increase provisioned capacity by 20 % as a safety net while a human investigates the root cause.

7.4 Visualizing Adaptive Capacity

The AdaptiveCapacityUtilization metric (available via the AWS/DynamoDB namespace) shows how much spare capacity is being borrowed. A sustained value above 30 % indicates that a partition is hot and may need redesign.


8. Real‑World Case Study: High‑Throughput API for Bee‑Tracking Sensors

8.1 Problem Statement

A nonprofit organization deployed 200,000 RFID tags on wild bees across the Midwest. Each tag transmitted a JSON payload (≈ 1.2 KB) every 5 seconds, resulting in ≈ 48,000 writes per second. The API needed to:

  1. Store each ping in DynamoDB for later analysis.
  2. Serve a live map that refreshed every 2 seconds, displaying the last known location of each bee.
  3. Keep end‑to‑end latency under 10 ms for the map refresh.

8.2 Architecture Overview

  • Ingress Layer: AWS IoT Core → Lambda (batch 100 records) → DynamoDB BeePings table.
  • Read Layer: API Gateway → Lambda (Query on beeId#rand) → DynamoDB.
  • Cache Layer: Amazon ElastiCache (Redis) holds the latest location per bee for sub‑millisecond reads; DynamoDB is the source of truth.

8.3 Data Model

TablePartition KeySort KeyProjected Attributes
BeePingsrand(0‑99)#beeIdtimestamp#20231015T083000Zlat, lon, signalStrength
BeeLatestbeeId—lat, lon, lastSeen (updated via DynamoDB Streams)

Why the random prefix? The 100‑shard pattern spreads the 48,000 WCU load across 100 partitions → 480 WCUs per partition, well below the 1,000 WCU limit.

8.4 Capacity Planning

  • Writes: 48,000 writes/s × 1 KB ≈ 48,000 WCUs. Provision 50,000 WCUs with a 20 % buffer → 60,000 WCUs total.
  • Reads (map): 200,000 bees × 1 read per refresh = 200,000 reads/s. Each read is 0.5 KB, eventually consistent → RCUs_needed = (200,000 * 0.5) / 4 = 25,000 RCUs. Provision 30,000 RCUs.

Auto‑scaling policies keep utilization at 70 % and allow burst up to 2×. Adaptive capacity handles occasional spikes when a subset of bees congregates near a hive, temporarily increasing writes on a few partitions.

8.5 Results

MetricTargetObserved
Write latency≤ 8 ms6.2 ms average (99th percentile 9 ms
Frequently asked
What is Achieving Single‑Digit Millisecond Latency with DynamoDB about?
In the modern, data‑driven world, an application’s responsiveness can be the difference between delight and abandonment. When a user clicks a button and the…
What should you know about introduction?
In the modern, data‑driven world, an application’s responsiveness can be the difference between delight and abandonment. When a user clicks a button and the backend takes more than a few milliseconds to reply, the experience feels sluggish; when the same operation finishes in single‑digit milliseconds, the…
What should you know about 1. Understanding DynamoDB’s Latency Model?
DynamoDB’s latency is a product of three primary phases:
What should you know about 2.1 The Role of the Partition Key?
The partition key is the hash that determines which physical partition stores an item . DynamoDB uses an internal MD5‑like algorithm to map the key’s value to a 128‑bit token, then assigns the token to a partition that has sufficient provisioned throughput. Because the hash is deterministic, all items with the same…
What should you know about 2.2 Quantitative Guidance?
If you expect a peak of 30,000 strongly consistent reads per second on a single partition, you would need 60,000 RCUs—far beyond the per‑partition limit. The only solution is to spread those reads across multiple partitions by designing a composite key or adding a random suffix.
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room