← Back to Tutorials

AWS CloudWatch

A practical guide to monitoring your AWS infrastructure — from reading logs to setting up alarms that tell you when something is going wrong before your users notice.

4 Units Price: Free

Unit 1
What Is CloudWatch?

When you deploy something on AWS — a Lambda function, an EC2 instance, an API Gateway, a DynamoDB table — things happen. Requests come in, functions run, errors occur, performance changes over time. Without some way to observe all of this, you are essentially running your application blind. You would have no idea whether your function is failing silently, whether your database is getting slower under load, or whether someone is hammering your API endpoint with requests. CloudWatch is the AWS service that solves this problem.

At its core, CloudWatch is a monitoring and observability platform. It collects data from AWS services automatically and gives you tools to view that data, set thresholds, and get notified when something crosses a line you care about. Almost every AWS service sends data to CloudWatch by default — you do not have to configure anything to start seeing basic metrics for Lambda invocations or API Gateway request counts. The data just starts flowing in as soon as you start using those services.

CloudWatch organises its data into three main categories. Logs are text output from your applications and services — the same kind of output you would see in a terminal if you were running something locally. Metrics are numerical measurements collected over time, like the number of requests per minute or the average duration of a Lambda function. Alarms watch those metrics and trigger actions when a value crosses a threshold you define, like sending you an email when your error rate goes above one percent.

There is also a fourth piece worth knowing about: CloudWatch Dashboards. These let you build custom views that pull together metrics from multiple services into a single screen. Rather than clicking around to find the data you care about, you build a dashboard once and have everything visible at a glance. It sounds like a small convenience, but in practice it is the difference between knowing what is happening in your system and having to go dig for it every time something feels off.

Unit 2
Logs and Metrics

Logs in CloudWatch are organised into log groups and log streams. A log group is the top-level container, usually representing a single application or service. Inside a log group, each individual source of log output — like a specific Lambda function instance — writes to its own log stream. When you open a log group in the CloudWatch console, you see a list of streams, and inside each stream you see the individual log entries in chronological order.

For Lambda functions, CloudWatch Logs is enabled automatically. Every time your function prints something — in Python that means calling print(), in Node.js it means console.log() — that output ends up in a log stream under a log group named after your function. This is one of the most useful things to know when you are debugging: if your Lambda function is behaving unexpectedly, the first place to look is its log group in CloudWatch.

Log entries have timestamps, and CloudWatch gives you tools to filter and search through them. The Logs Insights feature lets you run queries against your logs using a simple query language. Here is an example query that finds all log entries containing the word "error" from the last hour:

fields @timestamp, @message | filter @message like /error/ | sort @timestamp desc | limit 50

This kind of query is much faster than scrolling through raw log streams, especially when your function is handling thousands of requests per hour and you need to find the handful of entries that contain a specific pattern.

Metrics work differently from logs. Rather than raw text, a metric is a named numerical value tied to a timestamp, and CloudWatch stores a time series of these values. AWS services publish metrics automatically — Lambda publishes metrics like Invocations, Errors, Duration, and Throttles. API Gateway publishes Count, Latency, and 4XXError. DynamoDB publishes ConsumedReadCapacityUnits and ConsumedWriteCapacityUnits. You can view these metrics in the CloudWatch console by navigating to the Metrics section and browsing by service.

You can also publish your own custom metrics from your application code. This is useful when you want to track something specific to your business logic — the number of successful payments processed, the time spent in a particular part of your code, or the size of a queue you are managing yourself. Custom metrics are sent using the AWS SDK and show up in CloudWatch alongside the standard AWS metrics.

CloudWatch Logs Insights documentation: https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/AnalyzingLogData.html
Unit 3
Alarms and Notifications

Metrics are only useful if someone is looking at them, and no one can stare at a dashboard all day. This is what CloudWatch Alarms are for. An alarm watches a single metric over a time period you define, compares it against a threshold you set, and changes state depending on whether the metric is above or below that threshold. When the state changes, the alarm can trigger an action — most commonly, sending a notification through Amazon SNS so you get an email or a text message.

Setting up an alarm is straightforward. You choose the metric you want to watch, define the evaluation period (how many minutes of data to look at before deciding the alarm state), set the threshold value, and configure what happens when the alarm triggers. The three possible alarm states are OK, meaning everything is within the expected range; ALARM, meaning the threshold has been crossed; and INSUFFICIENT_DATA, which usually means not enough data has been collected yet to make a determination.

A practical example: suppose you have a Lambda function that handles API requests, and you want to know immediately if it starts throwing errors. You would create an alarm on the Errors metric for that function, set the threshold to something like "greater than 5 errors in any 5-minute window," and connect it to an SNS topic that sends you an email. The moment your function starts failing, you get notified rather than finding out hours later when a user complains.

Billing alarms are another type worth knowing about, especially early in your AWS journey. You can create an alarm on the EstimatedCharges metric in CloudWatch — it lives under the Billing namespace — and set it to notify you when your estimated monthly bill crosses a certain amount. This does not prevent charges from accumulating, but it gives you an early warning so you can investigate before the bill gets out of hand. Setting up a billing alarm is one of the first things worth doing in any new AWS account, and it takes about two minutes.

CloudWatch Alarms can also trigger more complex actions than just sending a notification. You can connect an alarm to an Auto Scaling policy, so that when CPU usage on a group of EC2 instances goes above 80 percent, AWS automatically adds more instances. Or you can connect an alarm to a Systems Manager action to run a remediation script automatically. For most projects starting out, the email notification path is the right one, but it is worth knowing that alarms can do more than just tell you about problems.

Creating a billing alarm: https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/AlarmThatSendsEmail.html
Unit 4
Summary and Practical Tips

CloudWatch is one of those services that feels optional until the moment you actually need it, at which point you wish you had set it up properly from the start. The pattern is familiar to anyone who has run production software: something breaks, you go to investigate, and the information you need either does not exist or is buried in a format that makes it hard to use. Setting up logging and monitoring early — before something goes wrong — is what makes the difference between a five-minute fix and a two-hour investigation.

The most immediately useful thing you can do right now, if you have any Lambda functions running, is open the CloudWatch console and look at their log groups. You will almost certainly find output you did not know was there — error messages, timing information, print statements from your code. Getting familiar with what is already being logged before you need it in an emergency is a good habit to build.

Log retention is something worth configuring early. By default, CloudWatch keeps logs forever, which sounds convenient but can lead to surprisingly large storage costs over time. For most development and testing scenarios, a retention period of 7 to 30 days is plenty. You can set this per log group in the console or through infrastructure-as-code, and it is one of those small configurations that pays for itself quickly.

When it comes to alarms, start with the basics: an error rate alarm on any Lambda functions that handle real traffic, and a billing alarm on your account. These two alone will catch the most common problems — functions that start failing silently, and unexpected cost spikes. Add more specific alarms as you learn more about where your application's weak points are.

CloudWatch Dashboards are worth the small time investment once you have a few services running. Build a dashboard that shows the metrics you actually care about — request count, error rate, latency, database throughput — and you will find yourself opening it reflexively when you deploy a change or investigate a report of something behaving oddly. A good dashboard does not replace proper alerting, but it gives you situational awareness that is hard to get any other way.

Finally, CloudWatch integrates with almost everything else in AWS. Logs from EC2, ECS, API Gateway, Lambda, RDS, and most other services can all end up in the same place, queryable with the same tools. As your infrastructure grows, having that unified view becomes increasingly valuable. Starting with CloudWatch from the beginning means you already have that foundation in place when you need it.