← Back to Tutorials

AWS Bedrock

A practical guide to Amazon Bedrock — how to access foundation models, make your first API call, and build AI-powered features into your applications without managing any ML infrastructure.

4 Units Price: Free

Unit 1
What Is Amazon Bedrock?

Building AI features into an application used to require either a significant machine learning background or a heavy dependency on a single provider's API. Amazon Bedrock sits somewhere in the middle — it is a managed AWS service that gives you access to a range of foundation models from different providers through a single, consistent API, without any model hosting or infrastructure to manage on your end.

A foundation model is a large pre-trained model capable of handling a broad range of tasks — generating text, answering questions, summarising documents, writing code, and more. Rather than training one from scratch, which requires enormous amounts of compute and data, you call one of these models with a prompt and get a response back. Bedrock makes this possible within your existing AWS environment, meaning you can combine it naturally with Lambda, DynamoDB, S3, and the other services you are already using.

The model catalogue in Bedrock includes offerings from Amazon itself — the Nova and Titan families — as well as models from Anthropic (Claude), Meta (Llama), Mistral, Cohere, and others. Each model has different strengths, different pricing, and different context window sizes. Some are better at following complex instructions; others are faster and cheaper for simpler tasks. The fact that they all share the same API surface means you can swap between them without rewriting your integration code.

Bedrock also includes features beyond simple text generation. Agents let you build systems where a model can take multi-step actions, calling tools and APIs autonomously to complete a task. Knowledge Bases connect a model to your own documents using retrieval-augmented generation, so the model can answer questions based on content it was never trained on. These are more advanced use cases, but they are built on the same foundation as a basic API call, so understanding the core works first makes everything else easier to pick up.

Unit 2
Enabling Model Access and Setting Up Permissions

Before you can call any model through Bedrock, you have to explicitly request access to it. This is not automatic — AWS requires you to go into the Bedrock console, navigate to Model Access, and enable each model you want to use. For most models the approval is instant; for a few it requires a short form describing your use case. This step catches people off guard the first time because the API exists and your credentials are valid, but every call returns an access denied error until you have enabled the model.

The process is straightforward. Open the AWS Console, search for Bedrock, and look for the Model Access section in the left sidebar. You will see a list of available models grouped by provider. Select the ones you want, submit the request, and within a few minutes — sometimes immediately — they will show as Access Granted. You only need to do this once per AWS account per region.

Regions matter with Bedrock. Not every model is available in every region, and the region where you enable access is the region where you can call the model. If you are building in eu-west-1 (Ireland), check which models are available there before deciding on one. Amazon's own models tend to have the broadest regional availability; third-party models are sometimes limited to us-east-1 initially and expand over time.

Permissions work the same way as any other AWS service. The identity calling Bedrock — whether that is an IAM user, a Lambda execution role, or an EC2 instance profile — needs the bedrock:InvokeModel permission attached to it. A minimal IAM policy for a Lambda function that calls one specific model looks like this:

{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "bedrock:InvokeModel", "Resource": "arn:aws:bedrock:eu-west-1::foundation-model/amazon.nova-micro-v1:0" } ] }

Locking the Resource to a specific model ARN rather than using a wildcard is good practice — it means the function can only call that one model, not any model in your account. If you need to call multiple models, list each ARN in the Resource array.

Model access documentation: https://docs.aws.amazon.com/bedrock/latest/userguide/model-access.html
Unit 3
Making Your First API Call

Once you have model access and the right permissions, calling Bedrock from Python is straightforward. The AWS SDK for Python, boto3, includes a Bedrock runtime client with a converse method that provides a consistent interface across all supported models regardless of their underlying differences. Here is a complete working example that sends a message and prints the response:

import boto3 client = boto3.client( "bedrock-runtime", region_name="eu-west-1" ) response = client.converse( modelId="amazon.nova-micro-v1:0", messages=[ { "role": "user", "content": [{"text": "Explain what AWS Bedrock is in two sentences."}] } ], inferenceConfig={ "temperature": 0.7, "maxTokens": 512 } ) reply = response["output"]["message"]["content"][0]["text"] print(reply)

The converse method takes a modelId, a messages array, and an optional inferenceConfig. The messages array follows a simple alternating pattern — user messages and assistant messages in sequence, which lets you send multi-turn conversations rather than just single prompts. The inferenceConfig controls the model's behaviour: temperature affects how creative or predictable the output is (lower values produce more consistent responses, higher values introduce more variation), and maxTokens sets a hard limit on how long the response can be.

The response object contains the model's reply nested inside output.message.content, which is an array because some models can return multiple content blocks. For text responses you almost always want the first element's text field, as shown in the example above.

Multi-turn conversations work by passing the full message history each time. If you want the model to remember what was said earlier in a conversation, you collect all previous messages and include them in the next request. Bedrock itself is stateless — it does not remember previous calls — so maintaining that history is your responsibility. In a web application you would typically store the conversation in a database or in the user's session and reconstruct the messages array on each request.

Unit 4
Pricing, Model Selection, and Practical Tips

Bedrock pricing is per-token, billed separately for input tokens (the prompt and conversation history you send) and output tokens (the response the model generates). The exact rate depends on which model you choose, and the differences are significant. Amazon's Nova Micro, for example, is one of the cheapest options available — fractions of a cent per thousand tokens — while larger, more capable models from Anthropic or other providers cost considerably more per call. Choosing the right model for the task at hand is one of the most effective ways to control costs.

As a rough rule: use a smaller, cheaper model for simple tasks like classification, short summaries, or extracting structured data from text. Use a larger model when you need nuanced reasoning, complex instruction following, or high-quality long-form output. The cheapest model that produces acceptable results for your use case is almost always the right choice, and the only way to find that threshold is to test a few options against real examples from your application.

Token counting matters more than it might seem. Every request you send to Bedrock includes all the previous messages in the conversation, not just the latest one. A long conversation accumulates tokens quickly, and those tokens are charged on every subsequent request. If your application supports long conversations, you will want to think about how to manage this — truncating older messages, summarising earlier parts of the conversation, or setting a maximum context length that you enforce in your code before sending the request to Bedrock.

There is no free tier for Bedrock — every call is billed from the first token. For development and testing this is usually fine because the per-token cost is low enough that typical experimentation costs almost nothing. But it is worth setting a billing alarm in CloudWatch before you start, just as a safeguard, especially if you are building something that might get unexpected traffic or where a bug could cause your code to call the model in a loop.

One of the most useful patterns when building with Bedrock inside a Lambda function is to keep the system prompt — the instructions that define how the model should behave — separate from the user messages. System prompts can be stored in an environment variable or loaded from S3, which means you can update the model's behaviour without redeploying your function. This separation also makes it easier to test different prompts and compare the results without changing any application logic.

Bedrock fits naturally into the rest of the AWS ecosystem. A Lambda function triggered by API Gateway handles incoming requests, calls Bedrock to generate a response, optionally stores the conversation history in DynamoDB, and returns the result — all within the same AWS account, using IAM roles for permissions and CloudWatch for logs and monitoring. Once you are comfortable with each of those services individually, combining them into a working AI-powered backend is mostly a matter of wiring them together correctly.