A practical guide to Amazon DynamoDB — how NoSQL data modelling actually works, how keys and queries relate to each other, and what the pricing model means for your application.
DynamoDB is Amazon's fully managed NoSQL database. If you have worked with a traditional relational database like MySQL or PostgreSQL before, DynamoDB requires a genuine shift in thinking rather than just learning a new syntax. Relational databases are built around tables with fixed schemas, joins between tables, and flexible queries written in SQL. DynamoDB abandons most of that in exchange for extreme scalability and consistent, predictable performance regardless of how large your dataset grows.
A DynamoDB table does not enforce a fixed schema in the way a SQL table does. Each item — the DynamoDB term for what would be a row in a relational database — can have a different set of attributes, as long as it includes the key attributes that identify it. This flexibility is powerful but also means the database will not catch structural mistakes for you the way a rigid schema would; the discipline has to come from how you design your application.
The defining characteristic of DynamoDB, and the thing that trips up almost everyone coming from a SQL background, is that queries are built around the keys you define at table creation time, not around arbitrary conditions on any attribute. In a relational database you can write a query filtering on any column and the database figures out how to execute it efficiently, sometimes slowly. In DynamoDB, efficient queries only work well against the specific keys the table was designed around. This means data modelling has to happen upfront, based on how your application will actually query the data, rather than being something you can bolt on later.
DynamoDB is fully managed, meaning there is no server to configure, patch, or scale manually. AWS handles replication, backups, and scaling automatically. This is part of why DynamoDB pairs so naturally with Lambda in serverless architectures — neither one requires you to manage any underlying infrastructure, and both scale automatically in response to demand.
Every DynamoDB table has a primary key, and that primary key is either a simple partition key or a composite key made up of a partition key plus a sort key. Understanding the difference is the single most important concept in DynamoDB.
A partition key alone identifies a single, unique item in the table. If you have a table of users with userId as the partition key, each userId maps to exactly one item. This is straightforward and works well when each entity naturally has one record.
A composite key, using both a partition key and a sort key, allows multiple items to share the same partition key while being distinguished by different sort key values. This is where DynamoDB gets genuinely powerful. Consider a table storing orders, using userId as the partition key and orderId as the sort key. All orders belonging to the same user share a partition key, which means you can efficiently retrieve every order for a specific user in a single query, sorted by orderId, without scanning the entire table.
Here is what creating a table with a composite key looks like using boto3:
In DynamoDB terminology, HASH refers to the partition key and RANGE refers to the sort key — older naming that predates the current terms but still shows up throughout the SDK and documentation. The AttributeDefinitions section only needs to list attributes used in keys or indexes; every other attribute an item might have does not need to be declared upfront.
Choosing a good partition key is a design decision with real consequences. DynamoDB distributes data across multiple physical partitions behind the scenes based on the partition key value, and it spreads read and write traffic accordingly. A partition key with poor distribution — for example, using a status field with only three possible values as your partition key — concentrates all the traffic for each value onto a small number of partitions, which can create performance bottlenecks under heavy load. A partition key with high cardinality, like a userId or a UUID, spreads traffic much more evenly.
DynamoDB core components: https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/HowItWorks.CoreComponents.htmlDynamoDB offers a handful of core operations for reading and writing data, and understanding when to use each one matters for both performance and cost. GetItem retrieves a single item by its exact primary key — this is the fastest and cheapest way to read data, and should be your default choice whenever you know the full key of the item you want.
Here is a basic read and write example using boto3:
Query is the operation you use when you want multiple items that share the same partition key, optionally filtered further by the sort key. Following the earlier orders example, a Query for all orders belonging to user123 looks like this:
Scan is the operation to be cautious with. It reads every single item in the table, regardless of any key, and then applies filters afterward if you provide them. Scan works, but it is slow and expensive on large tables because DynamoDB has to examine every item rather than jumping directly to the ones you need. If you find yourself relying on Scan regularly in production code, it is usually a sign that your table design does not match how your application actually needs to query data, and it is worth reconsidering the key structure or adding a Global Secondary Index.
Global Secondary Indexes, or GSIs, let you query the same data using a different partition key and sort key than the table's primary key. This is how you handle the common situation where you need to query the same data in more than one way — for example, looking up orders by userId in one part of your application, and looking up orders by status in another. Rather than scanning, you create a GSI with status as its partition key, and DynamoDB maintains that index automatically as items are written to the base table.
UpdateItem lets you modify specific attributes of an existing item without replacing the whole item, and DeleteItem removes an item by its key. Both operations, along with PutItem, can include conditions that must be true for the operation to succeed — useful for things like preventing an update from happening if the item has changed since you last read it.
DynamoDB offers two capacity modes, and choosing between them is one of the first decisions you make when creating a table. On-Demand mode, shown as PAY_PER_REQUEST in the earlier example, charges based on the actual number of reads and writes your table receives, with no capacity planning required. This is the simpler option and the right default for most new projects, especially anything with unpredictable or low traffic, since you pay only for what you use and never have to think about throttling from under-provisioned capacity.
Provisioned mode requires you to specify a fixed number of read and write capacity units the table should support, and you pay for that reserved capacity whether you use it or not. This mode is cheaper per unit than On-Demand, but only pays off when your traffic is predictable and consistently high enough to justify the reservation. Provisioned mode also supports auto-scaling, which adjusts capacity within limits you define, softening the tradeoff somewhat, but for smaller and newer projects On-Demand almost always remains the more sensible starting point.
Storage costs in DynamoDB are billed per gigabyte per month, and are generally a small part of the total bill compared to read and write costs for most applications. The free tier includes 25 GB of storage and enough read and write capacity for a small application to run entirely free, and unlike some AWS free tier offers, aspects of the DynamoDB free tier do not expire after twelve months, making it genuinely useful for long-running personal projects with modest traffic.
DynamoDB Streams is worth knowing about even if you do not need it immediately. When enabled, a stream captures a time-ordered sequence of changes made to items in a table — every insert, update, and delete — which can trigger a Lambda function automatically. This pattern is common for keeping a search index in sync with your main table, sending notifications when specific data changes, or replicating data into another system without your application code needing to explicitly handle that replication.
The practical habit that matters most with DynamoDB is designing your access patterns before you design your table. Write down every way your application needs to query the data — by user, by date, by status, by any other dimension — before you decide on a partition key and sort key. Retrofitting a table design after the application is built is possible using GSIs, but it is considerably more work than getting the core structure right from the start, and understanding your access patterns early is what separates a DynamoDB table that scales cleanly from one that requires expensive Scans and constant workarounds.