Verified against Microsoft Learn on September 29, 2026. Every limit and formula on this sheet comes from the pages below, and every worked example has been recalculated.
Service quotas and limits · Autoscale throughput · Scaling provisioned throughput · Optimize request cost
Limits to memorize#
| Resource | Limit |
|---|---|
| Physical partition | 10,000 RU/s and 50 GB (30 GB for API for Cassandra) |
| Logical partition (one partition key value) | 20 GB and 10,000 RU/s |
| Item size | 2 MB (UTF-8 length of the JSON) |
| RU/s per container or shared database | 1,000,000 (raise it with a support ticket) |
| Logical partitions and storage per container | Unlimited |
| Minimum throughput per GB stored | 1 RU/s (manual), 10 RU/s of autoscale max |
| Containers in a shared throughput database | 25; add more with dedicated throughput |
| Serverless container | Starts at 5,000 RU/s, up to 5,000 RU/s per physical partition, single region account |
| Free tier (one account per subscription) | 1000 RU/s and 25 GB free for the life of the account; opt in at creation |
What operations cost#
| Operation | Request units |
|---|---|
| Point read of a 1 KB item (id and partition key) | 1 RU |
| Point read of a 100 KB item | 10 RU |
| Insert a 1 KB item with no indexing | About 5.5 RU |
| Replace an item | Twice the cost of inserting it |
| Query returning a thousand 1 KB items | About 1,000 RU |
| Each physical partition a query checks | At least about 2.5 RU, even with no matches |
| Read at strong or bounded staleness | Twice the RU of the same read at session or weaker |
A point read beats a query. A query that filters on one item's id and partition key is still a query; only the SDK or REST point read gets the 1 RU price. Include the partition key in filters so a query checks one partition instead of every one.
Provisioned throughput is set in increments of 100 RU/s. Index fewer paths to lower the RU cost of every write.
Choose a throughput model#
| Model | How it works | Best for |
|---|---|---|
| Manual (standard) | A fixed RU/s you change yourself | Steady, predictable traffic |
| Autoscale | Scales between 10% of the max and the max you set | Variable or unpredictable traffic, new apps, dev and test |
| Shared database throughput | Up to 25 containers share one RU/s budget, with no per-container guarantee | Many small, low-traffic containers |
| Serverless | No provisioned throughput; billed per RU consumed | Intermittent or light traffic in a single region |
Throughput set on the container, manual or autoscale, is the recommended approach for most workloads. A container’s model is fixed: move the data to a new container to switch between dedicated and shared throughput.
Minimum throughput formulas#
| Resource | Minimum | Worked example |
|---|---|---|
| Manual container | MAX(400, storage GB × 1, highest RU/s ever ÷ 100) | Scaled to 50,000 RU/s with 20 GB: MAX(400, 20, 500) = 500 RU/s. At 2,000 GB: 2,000 RU/s |
| Autoscale container (max RU/s) | MAX(1000, storage GB × 10, highest max ever ÷ 10), rounded up to the nearest 1000 | Scaled to 50,000 with 20 GB: MAX(1000, 200, 5000) = 5,000 RU/s. At 2,000 GB: 20,000 RU/s |
| Manual shared database | The manual container terms, and 400 + MAX(containers − 25, 0) × 100 | 30 containers: 400 + 5 × 100 = 900 RU/s |
| Autoscale shared database (max RU/s) | The autoscale terms, and 1000 + MAX(containers − 25, 0) × 1000, rounded up to the nearest 1000 | 27 containers: 1000 + 2 × 1000 = 3,000 RU/s |
Scaling up raises the floor for later. After a container reaches 100,000 RU/s, the lowest manual throughput you can set is 1,000 RU/s.
Autoscale in numbers#
- Range: throughput T scales so that 0.1 × Tmax ≤ T ≤ Tmax. A max of 20,000 RU/s scales between 2,000 and 20,000.
- Entry point: Tmax 1000 RU/s, scaling from 100 to 1000 RU/s. Set Tmax in increments of 1000.
- Billing: each hour, you pay for the highest RU/s the system scaled to in that hour. Autoscale saves money when you use the full Tmax for 66% or fewer of the hours in a month.
- Rates: single write region accounts pay the autoscale rate; accounts with multiple write regions pay the same multi-region write rate with no extra charge for autoscale.
- Storage: a resource stores up to 0.1 × Tmax GB. Beyond that, the max rises automatically. A 50,000 RU/s max holds 5,000 GB; at 6,000 GB the max becomes 60,000 RU/s.
- Dynamic scaling: when it is enabled on the account, each physical partition and region scales independently based on its own usage.
Partition split math#
Is the new RU/s at or below physical partitions × 10,000?
Yes The change completes instantly on the current partitions.
No Partitions split until there are ROUNDUP(RU/s ÷ 10,000). The change runs asynchronously, typically in 4 to 6 hours, and reads and writes continue throughout.
Should storage stay evenly spread after the split?
Yes Scale to 10,000 × current partitions × the next power of two, so every partition splits, then lower the RU/s.
Uneven is acceptable Scale straight to the target.
| Scenario | Result |
|---|---|
| 3 partitions at 30,000 RU/s, scaled to 45,000 | ROUNDUP(45,000 ÷ 10,000) = 5 partitions: two of the three split |
| 2 partitions, 80 GB, 20,000 RU/s, scaled to 30,000 | Only one partition splits: one holds 40 GB and two hold 20 GB each |
| The same container, scaled to 40,000 and then lowered to 30,000 | Both partitions split: 4 partitions of 20 GB, each with 7,500 RU/s |
| Responding to 429 errors | First raise RU/s to partitions × 10,000, which is instant, and check whether that is enough |
Pre-provision partitions before a large migration#
- Initial throughput per new physical partition: 6,000 RU/s with manual throughput, 10,000 RU/s with autoscale or a shared database.
- Example: 1 TB at 40 GB per partition needs 1000 ÷ 40 = 25 partitions.
- Manual: create the container at 25 × 6,000 = 150,000 RU/s, then raise it instantly to 250,000 RU/s for the load.
- Autoscale or shared: create it at 25 × 10,000 = 250,000 RU/s.
- Lower the throughput after the load; the new floor is the highest RU/s ÷ 100 (manual).

