1. Choose the distribution unit from access patterns
A positions database can hold millions of documents while concentrating most load on one partition-key value. High cardinality helps but does not prove request distribution: one busy fund can dominate reads and writes. Before choosing the key, describe critical requests, available filters and the intended transaction unit. Distinguish logical partitions, defined by key value, from service-managed physical partitions. Increasing total capacity does not automatically fix concentration that hits local limits. Estimate data growth and RU consumption per key, including daily closing and retry behavior. Low average account utilization does not rule out throttling on a hot subset. A useful review compares the busiest keys and representative queries rather than relying on the total number of documents.
2. Treat key changes as migration
An item’s partition-key value is not an attribute that replace can change in place. Moving an item requires creating it under the new value and removing the original; across different logical partitions, those operations do not automatically form one transaction. If a fund-owner-based key changes when responsibility transfers, the design must account for that lifecycle. Migration needs a stable business identifier, temporary coexistence, duplicate detection, validation and recovery from interruption. Do not present a data-model change as simply increasing RU capacity. Include queries, integrations and references still looking under the old value before deleting the earlier copy. Decide how a client behaves while both representations exist or while the transfer is incomplete.
3. Bound atomicity before choosing a batch
Transactional batch groups point operations sharing a partition-key value in the same container. They all succeed or the batch rolls back. Do not confuse this with bulk execution, whose purpose is throughput rather than joint atomicity. When a request stores a position and its control record, placing both items appropriately can support that local contract. If they occupy different partitions, another coordination design or a revised transaction unit is needed. When an operation fails, its individual result identifies the cause while others can show 424 failed dependency. A 409 on creation can identify an existing item. Diagnosis should find the initiating cause rather than treat every 424 as an independent failure requiring separate writes to be retried.
4. Preserve read-your-writes across instances
Session consistency provides guarantees within a session, and the SDK tracks partition-bound session tokens. In a multi-instance application, a write can occur on one instance and the next read on another using a different client. Do not assume the token was shared merely because both use the same Cosmos DB account. Decide how to preserve the required context and test the real load-balanced path. Preserve the token without interpreting or modifying it, and use the relevant partition’s context. It is a minimum version barrier, not a request for an exact historical snapshot. A new instance without that context does not demonstrate reading another instance’s previous write. If every read requires the latest global version, compare that requirement with the selected guarantee and the cost of stronger alternatives.
5. Distinguish consistency from lost-update protection
Two teams can read the same version and calculate different changes. A consistent read does not prevent the second write from overwriting the first team’s work if the application fails to check the version it read. With optimistic concurrency, the application sends If-Match containing the observed ETag; if it no longer matches, the operation is rejected with 412. The next step depends on business rules: reread, reconcile and propose the update again, or surface the conflict for a decision. Removing If-Match simply to make a retry succeed abandons the intended protection. Do not confuse 412 with 429 throttling, since increasing capacity does not fix a stale version. The local model demonstrates this contract with a fictional integer; it models neither real ETag format nor cross-Region conflict policies.
6. Accept business outcomes and observed costs
Acceptance should combine representative requests, hot keys, concurrency and failover where required. Record RU use, latency, retries, conflicts and final operation outcomes, using fictional identifiers in training material. An application can show low average response time yet fail to confirm a recent write when the request moves between instances. Another can support aggregate peak demand while throttling one fund. Define separate evidence for distribution, atomicity, read consistency and concurrency. For production, give support a way to distinguish 409, 412, 424 and 429 within the operation’s context. The aim is to select action from the violated contract rather than increase capacity or repeat work when state reconciliation is required. Each signal should have a known owner and a bounded response.
record = {"version": 3, "position": 100}
def conditional_update(expected, new_value):
if expected!= record["version"]:
return 412
record.update(version=record["version"] + 1, position=new_value)
return 200
reader_a = record["version"]
reader_b = record["version"]
assert conditional_update(reader_a, 105) == 200
assert conditional_update(reader_b, 99) == 412
assert record["position"] == 105
assert record["version"] == 4
assert conditional_update(record["version"], 103) == 200
print("five fictional version checks passed; no Cosmos DB request performed")
Fictional case: two operators correct the same position. The first writes with the current ETag. The second receives 412. Instead of removing the condition, the application rereads the position, exposes the conflict and applies the approved reconciliation rule.
Common pitfalls
Confusing cardinality with actual distribution; treating bulk as a transaction; losing session tokens across instances; retrying a 412 write after removing its version condition.
Related topics: Data modeling · Integration and outbox · Capacity and observability
Partitioning, atomicity, consistency and concurrency are related decisions, but each needs its own contract.
Reference: Cosmos DB partitioning · AZ-305 objectives 2026-04-17