← Back to Tutorials

AWS S3

A practical guide to Amazon S3 — how buckets and objects work, how permissions are structured, and how to use storage classes to keep costs under control.

4 Units Price: Free

Unit 1
What Is S3 and How Is It Organised?

Amazon S3, short for Simple Storage Service, is object storage. That distinction matters because it is different from the kind of storage most people are used to on their own computer. A traditional file system organises data in a hierarchy of folders and subfolders. S3 does not really work that way — internally it is a flat structure where every object lives inside a bucket, identified by a unique key rather than a nested folder path. The folder-like structure you see in the AWS Console is a visual convenience built on top of that flat design, using the key naming convention to simulate folders.

A bucket is the top-level container in S3. Bucket names have to be globally unique across all of AWS, not just within your own account, because S3 uses the bucket name as part of the URL that identifies your data. Inside a bucket, you store objects — files, essentially, though S3 calls them objects because they can be virtually anything: images, videos, backups, log files, static website assets, or application data. Each object is identified by a key, which is the full "path" to that object, such as photos/2026/summer/beach.jpg. Despite looking like a folder path, this is really just a single string that S3 uses as an identifier.

S3 is designed for durability rather than speed of access to a single file. AWS engineers S3 to what is described as eleven nines of durability, meaning the mathematical probability of losing an object is astronomically small. This is achieved by automatically replicating data across multiple physical facilities within a region. You do not configure this replication yourself — it happens by default the moment you upload an object.

Because buckets and objects are accessible over HTTP, S3 is also commonly used to host static websites, serve downloadable files directly to users, and act as the storage backend for applications that need somewhere durable to keep uploaded content. A Lambda function that processes user uploads, for example, will often read from and write to an S3 bucket as part of its normal workflow.

Unit 2
Uploading, Organising, and Accessing Objects

Getting a file into S3 can be done in a few different ways depending on the context. The AWS Console lets you drag and drop files directly into a bucket through the browser, which is fine for occasional manual uploads. For anything programmatic, the AWS CLI and SDKs are the standard approach. Here is a simple example using boto3, the AWS SDK for Python, to upload a file:

import boto3 s3 = boto3.client("s3") s3.upload_file( Filename="local-report.pdf", Bucket="my-app-bucket", Key="reports/2026/local-report.pdf" )

This uploads the local file and stores it under the key reports/2026/local-report.pdf inside the bucket. Even though this looks like a nested folder structure, remember that S3 is storing it as a single flat key — the slashes are just part of the string. Reading a file back works the same way in reverse, using download_file or get_object depending on whether you want the file saved locally or its contents returned directly in your code.

Naming conventions for keys matter more than people expect, particularly for applications with high request volume. A common pattern is to prefix keys with something that distributes requests evenly, such as a hash or a reversed timestamp, rather than everything starting with the same date prefix. This is less of a concern for smaller applications, but worth knowing as your usage grows.

Every object in S3 can carry metadata alongside its content — things like content type, cache control headers, and custom key-value pairs you define yourself. Setting the content type correctly matters if you are serving files directly to a browser, since it tells the browser how to interpret the file, whether that is an image, a PDF, or plain text.

S3 also supports versioning at the bucket level. When enabled, every time you overwrite or delete an object, S3 keeps the previous version rather than discarding it. This is extremely useful for protecting against accidental deletion or overwrite, though it does mean storage costs accumulate for every version kept, so it is worth pairing versioning with a lifecycle policy that eventually removes old versions after a set period.

AWS S3 documentation: https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html
Unit 3
Permissions and Bucket Policies

By default, every S3 bucket and every object inside it is private. Only the account that created the bucket, and identities with explicit permission, can access it. This default has not always been the case historically — S3 permissions used to be a common source of security incidents when buckets were accidentally made public — which is why AWS now blocks public access by default at the account level and requires you to deliberately opt out of that protection if you actually need a public bucket.

Access to S3 is controlled through a combination of IAM policies and bucket policies. IAM policies, which you learned about in the IAM course, are attached to users and roles and describe what that identity is allowed to do across AWS, including S3. Bucket policies work differently — they are attached directly to the bucket itself and describe who is allowed to access that specific bucket, regardless of which AWS account they come from. This makes bucket policies the right tool when you need to grant access to a different AWS account or to the public internet, while IAM policies are usually the right tool for controlling what your own users and applications can do.

Here is a bucket policy that allows public read access to objects in a bucket, commonly used for static website hosting:

{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": "*", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::my-website-bucket/*" } ] }

Notice the Principal field, which is set to "*", meaning anyone. This is exactly the kind of statement that should be used carefully and only when you genuinely intend for the content to be public — a static website's assets, for example, rather than a bucket holding user data or backups.

For most application backends, the better pattern is to keep the bucket entirely private and grant access only to the specific IAM role that needs it, such as a Lambda execution role. If end users need to upload or download files directly without going through your backend, S3 supports presigned URLs — temporary, time-limited URLs that grant access to a specific object without making the bucket public. This is the standard way to let a user upload a profile picture directly to S3 from their browser without exposing broader access to the bucket.

Unit 4
Storage Classes, Pricing, and Practical Tips

Not all data needs to be accessed at the same speed, and S3 pricing reflects that through storage classes. S3 Standard is the default and most expensive per gigabyte, designed for data accessed frequently. S3 Standard-Infrequent Access costs less per gigabyte but charges a retrieval fee, making it suitable for data you keep around but rarely read, like older backups. S3 Glacier and Glacier Deep Archive are far cheaper still, designed for long-term archival where retrieval can take minutes to hours rather than being instant — a good fit for compliance records or historical data you are required to keep but essentially never access.

Choosing the right storage class for each type of data can meaningfully reduce your bill, especially once you are storing large volumes. S3 Intelligent-Tiering is worth mentioning here too — it automatically moves objects between access tiers based on actual usage patterns, which removes the need to manually decide and adjust storage classes yourself, at the cost of a small monitoring fee per object.

Lifecycle policies let you automate storage class transitions and object expiration without writing any code. You define a rule — for example, move objects to Infrequent Access after 30 days, then to Glacier after 90 days, then delete them entirely after a year — and S3 handles the rest automatically. This is one of the most effective cost-control tools available in S3 and is worth setting up for any bucket that accumulates data over time, such as logs or backups.

S3 pricing itself has three components: storage cost per gigabyte per month, request cost (a small charge per PUT, GET, or other API call), and data transfer cost for data leaving AWS to the internet. Storage within the same region as your other services and data transferred within AWS is often free or heavily discounted, while data transferred out to end users over the internet is where costs typically accumulate for content-heavy applications.

A few habits worth building early: enable versioning only where it genuinely protects against a real risk, since it doubles storage costs if not paired with a lifecycle policy to clean up old versions. Set a lifecycle policy on any bucket that accumulates logs or temporary data, since forgotten storage is one of the more common sources of slowly creeping AWS bills. And use presigned URLs rather than making buckets public whenever the access needed is temporary or user-specific rather than genuinely public content. Getting these habits right from the start makes S3 both cheaper and more secure as your usage grows.