↓ Skip to main content
  1. Certifications/
  2. DP-420 Study Hub/

DP-420 Video Guide: John Savill on Cosmos DB Optimization

John Savill, a colleague here at Microsoft, has been teaching Azure to the community for years, and his free videos are among the best preparation anywhere. His Cosmos DB Optimization deep dive is an hour on the ideas that decide both the cost and the performance of an Azure Cosmos DB design: request units, free tier, manual and autoscale throughput, serverless, physical and logical partitions, partition key cardinality, document size, global secondary indexes, hierarchical partition keys, and storage-heavy and write-heavy workloads.

Cosmos DB Optimization by John Savill's Technical Training · watch on YouTube to like, comment, and subscribe

How to use this guide

This is a deep dive on optimization rather than a full exam cram, so pair it with the domain pages for the AI, security, and change feed skills. Watch it once early in your studies, then use the sections below to revisit topics. Select any timestamp to jump the player to that point. The badges show which exam objectives a section supports and link to those skills on the domain pages. John also shares the whiteboard from the video, which makes an excellent one-page review sheet.

Why it matters for DP-420
#

Objective 1.1 asks you to evaluate request unit consumption and choose the right throughput, and objective 1.3 asks you to choose partition keys and estimate workload cost. Objective 2.3 builds on both when it asks you to optimize query performance. John’s explanation of how requested throughput becomes physical partitions, and why a high-cardinality partition key that matches your most common query is the single most important design choice, is the reasoning behind many of those questions.

The video, section by section
#

The key points below are a study summary of what John explains in each section, written to help you find and review topics. Where a detail has changed since the video was recorded, the summary follows the current Microsoft Learn documentation. It is not a substitute for watching him explain it.

Account structure and request units

  • John frames the video around cost, explaining that Azure Cosmos DB spend depends on the capacity model you choose, the number of request units (RUs) you configure, and how your data model shapes the way you interact with data.
  • The Azure resource you create is an account, the account contains databases, and each database contains containers that hold the JSON documents making up your data set.
  • Some settings belong to the account and most throughput and partitioning choices belong to the container, with the database level used only occasionally, so knowing where each setting lives is central to optimizing cost.
  • He describes an RU as the compute currency of Azure Cosmos DB, where a point read of a small item costs about one RU while inserts, upserts, deletes, and queries cost more depending on the operation, how the data is distributed, and the number of regions.
  • Provisioned throughput is a per-second rate that is billed hourly, so the published monthly price for 100 RU/s already reflects that rate sustained for the whole month rather than a cost you multiply by every second.

Free tier and manual throughput

  • The Azure Cosmos DB free tier provides the first 1,000 RU/s and 25 GB of storage at no charge on one account per subscription, which suits development, testing, and prototyping.
  • Production workloads usually move to provisioned throughput, and John notes that manual and autoscale are both provisioned options, so the fixed option is more accurately called manual rather than simply provisioned.
  • With manual throughput you set a fixed RU/s value and pay for it every hour whether you use it or not, so it fits highly predictable, constant workloads whose usage stays close to the provisioned amount.
  • You also pay for consumed storage, and the provisioned throughput applies in every region you add, so cost scales with the number of regions, while multiple write regions and a configurable consistency level let you balance write latency against stronger guarantees.
  • Reserved capacity with one-year or three-year terms lowers the effective price, and larger commitments earn deeper discounts, which can justify manual throughput even when some capacity sits idle.

Autoscale throughput and billing

  • John describes autoscale as the right choice for the vast majority of workloads: you set a maximum RU/s and Azure Cosmos DB scales throughput between 10% of that maximum and the full value based on demand.
  • Billing follows actual scaling rather than the configured maximum, with each hour charged for the highest RU/s the system scaled to and never less than the 10% floor.
  • For accounts with a single write region the autoscale rate is 1.5 times the manual rate, which puts the break-even point at roughly 66% average utilization, and below that level autoscale usually costs less, while accounts with multiple write regions pay the same rate for either option.
  • With dynamic scaling, which is on by default for newer accounts, each physical partition in each region scales and bills independently, so a busy primary region does not raise the charge for quieter replica regions.
  • Both manual and autoscale enforce a hard maximum, and John highlights this ceiling for vendors who need predictable costs when they bill their own customers for different service tiers.

Configuring throughput and account limits

  • The capacity mode is chosen at the account level when you create an Azure Cosmos DB account, which is either provisioned throughput or serverless.
  • In a provisioned account you pick manual or autoscale and set the RU/s on each container, which John calls the recommended path for most designs.
  • You can optionally provision throughput on a database so its containers share it, but sharing offers no guaranteed distribution during contention and can lead to noisy-neighbor effects, so it is best kept for dev and test scenarios.
  • He demonstrates that a container given its own dedicated throughput sits outside the database allocation entirely, so its RU/s do not need to fit within the shared value, while containers without their own setting keep sharing the database throughput.
  • The free tier discount applies to a provisioned account, so only usage above the first 1,000 RU/s is billed, and the account throughput limit caps the total RU/s provisioned across the account, with autoscale resources counting their maximum and free tier accounts able to stay within the free allowance by using a 1,000 RU/s cap.

Physical partitions and choosing a capacity mode

  • John explains that physical partitions power every Azure Cosmos DB container behind the scenes, with each physical partition supporting up to 10,000 RU/s and 50 GB of storage and replicated for durability.
  • At container creation the service knows only the requested RU/s, so that value sets the initial number of physical partitions, and partitions later split as storage or throughput grows, redistributing logical partitions across them.
  • Autoscale keeps the physical partitions needed for its maximum even while running near its 10% floor, which helps explain its price premium, and every region holds the same partition layout while utilization can differ from region to region.
  • Serverless bills per million RUs actually consumed with no per-second allocation, which suits sporadic, bursty usage, while John's rough math shows steady workloads cost several times more on serverless than on provisioned throughput.
  • Serverless currently runs in a single region and its maximum throughput grows with the number of physical partitions, and therefore with stored data, without guaranteed throughput, so John's guidance is autoscale for nearly everything, manual above about 66% utilization, and serverless for truly sporadic use.

Planning for RU efficiency with the partition key

  • Planning starts with the seasonality of the workload, which guides the choice between autoscale and serverless, and with an RU estimate, where autoscale leaves generous room to adjust and throttling signals that the maximum should rise.
  • John reminds viewers that Azure Cosmos DB is a NoSQL database without joins or foreign keys, so RU efficiency comes from the partition key, the queries you run, the attributes and size of each document, and how you separate bounded from unbounded data.
  • You set the partition key when you create the container and cannot change it in place afterward, and every document you write must include the partition key property.
  • The partition key value is hashed to place each document in a logical partition, each logical partition lives in exactly one physical partition, and a logical partition can hold up to 20 GB.
  • Because most workloads are read heavy, queries should filter on the partition key so they route to a single partition and avoid cross-partition fan-out, as in a user profile store partitioned by user ID.

Cardinality, document size, and related data

  • In Azure Cosmos DB a strong partition key has high cardinality, meaning many distinct values, so hashing spreads logical partitions evenly across physical partitions.
  • A key with only a few possible values, such as three, limits data to a handful of logical partitions and can concentrate storage and requests unevenly when one value dominates.
  • John recommends pairing high cardinality with the property you query most often, so requests stay targeted to one partition while data and throughput remain evenly distributed.
  • The maximum item size is 2 MB, yet he advises keeping documents small because larger items consume more RUs every time you read or write them.
  • Embed small bounded sets, such as a customer's two or three addresses, in the same document, and store unbounded data, such as blog comments, as separate documents that share the article's partition key so they stay in the same logical partition.

Global secondary indexes

  • When a frequent query filters on a property other than the partition key, a global secondary index maintains a copy of the data under a different partition key, and one container can have several of them.
  • Azure Cosmos DB implements each global secondary index as a separate read-only container synchronized through the change feed, replacing the custom container and change feed setup that teams previously had to build and maintain.
  • The index container uses autoscale throughput and can have a lower maximum than the source when it is queried less often, but you pay for the duplicate storage, the RUs to keep it in sync, and its own throughput.
  • John frames the choice as a cost comparison: use query diagnostics to see how much the cross-partition queries on the second key cost and how often they run, then weigh that against the cost of maintaining the index.
  • Single-partition queries on the index can also lower latency, although at scale the RU trade-off usually decides the matter, and John has suggested to the product group that Azure Advisor could someday recommend when an index pays off.

Storage-heavy and write-heavy partitioning

  • Because each Azure Cosmos DB logical partition holds up to 20 GB, storage-heavy data where a single partition key value outgrows that size calls for a hierarchical partition key with up to three levels.
  • The complete combination of levels determines the logical partition, so each full key value must stay within 20 GB, while queries that supply only the first level, or the first two, still route to a subset of partitions.
  • For write-heavy workloads such as IoT telemetry from vehicles or elevators keyed by device ID, John describes adding a GUID as a second level to raise cardinality and spread writes evenly.
  • If the data is not queried by device ID in Azure Cosmos DB and another service processes it downstream, the partition key can simply be a GUID, giving maximum cardinality for a high-volume landing zone.
  • The summary reinforces understanding how you will interact with your data, choosing autoscale unless usage is truly sporadic, picking a high-cardinality partition key, and considering a global secondary index for frequent queries on other attributes.

Keep going
#