Large language models can answer questions, summarise documents, write code, and interact with external systems. But building a reliable AI application requires more than sending a prompt and displaying the response.

A production-ready application must manage conversation history, provide relevant context, use tools safely, handle different response types, and evaluate whether the generated output is useful.

In this tutorial, we’ll build ShopHelper, a customer-support assistant for an imaginary online shop. By the end, ShopHelper will be able to:

  • Answer general questions in a consistent tone
  • Remember what a customer said earlier
  • Look up order statuses by calling a function in your code
  • Handle Claude’s multi-block responses safely
  • Process support tickets using workflows
  • Evaluate whether prompt changes improve results

Each section adds one piece, so you can follow along in your own editor.

Prerequisites

You should have:

  • Basic Python knowledge
  • Python 3.9 or later
  • An Anthropic API key
  • Familiarity with functions and JSON

How to Set Up the Project and Keep Your API Key Secure

Create a virtual environment and install the Anthropic Python SDK:

python -m venv .venv
source .venv/bin/activate
pip install anthropic python-dotenv

On Windows:

.venv\Scripts\activate

Create a .env file:

ANTHROPIC_API_KEY=your_api_key_here

An API key is a secret credential. Never place it in browser JavaScript, mobile-app code, or client-side configuration. Never commit it to a repository:

echo ".env" >> .gitignore

If you add a web interface later, keep the key on your backend:

Browser → Your backend → Claude API

Create app.py:

import os

from anthropic import Anthropic
from dotenv import load_dotenv

load_dotenv()

MODEL = "claude-sonnet-5"

client = Anthropic(
    api_key=os.environ["ANTHROPIC_API_KEY"]
)

load_dotenv() loads the value from .env. The MODEL constant means you only need to change the model name in one place. Confirm that the model identifier is available to your account before running the example.

How to Make Your First Request

response = client.messages.create(
    model=MODEL,
    max_tokens=500,
    messages=[
        {
            "role": "user",
            "content": "Explain what an API is in simple terms."
        }
    ],
)

answer = "".join(
    block.text
    for block in response.content
    if block.type == "text"
)

print(answer)

A request contains three important parts:

  • model selects the Claude model that handles the request. Models can differ in capability, speed, and cost.
  • max_tokens limits the maximum amount of text Claude can generate. A smaller value can reduce latency, but Claude may stop before completing its answer.
  • messages contains the conversation. Each message has a role and content. The role is usually user or assistant.

For example, a one-off request contains one user message. A multi-turn conversation contains earlier user and assistant messages.

Claude returns response.content, which is a list of typed content blocks. Common blocks include:

Block typeMeaning
textGenerated text
tool_useA request for your application to call a tool
thinkingReasoning content when enabled

The example collects text blocks instead of assuming response.content[0] is always text.

You can inspect usage information for monitoring:

print(response.usage.input_tokens)
print(response.usage.output_tokens)

How to Manage Conversation History

Claude doesn’t automatically remember separate API requests. Send relevant history with every request:

messages = [
    {
        "role": "user",
        "content": "What is your returns policy?"
    },
    {
        "role": "assistant",
        "content": "Items can be returned within 30 days."
    },
    {
        "role": "user",
        "content": "How long do I have?"
    },
]

response = client.messages.create(
    model=MODEL,
    max_tokens=300,
    messages=messages,
)

The assistant message records Claude’s earlier answer, allowing the final question to be interpreted in context.

A simple chat function can maintain the history:

def chat(history, user_text):
    history.append({
        "role": "user",
        "content": user_text,
    })

    response = client.messages.create(
        model=MODEL,
        max_tokens=500,
        messages=history,
    )

    reply = "".join(
        block.text
        for block in response.content
        if block.type == "text"
    )

    history.append({
        "role": "assistant",
        "content": reply,
    })

    return reply


history = []

print(chat(history, "What is your returns policy?"))
print(chat(history, "How long do I have?"))

Each call adds the new user message, sends the complete history, and stores Claude’s response for the next turn. In production, store histories by customer or session ID.

How to Manage History as it Grows

Unlimited history increases input size and may make it harder for Claude to focus. One option is to retain only recent messages:

def trim_history(history, max_messages=10):
    trimmed = history[-max_messages:]

    while trimmed and trimmed[0]["role"] != "user":
        trimmed.pop(0)

    return trimmed

Another option is to summarise older turns while keeping recent messages:

def summarise_history(history, keep_last=6):
    old = history[:-keep_last]
    recent = history[-keep_last:]

    transcript = "\n".join(
        f"{message['role']}: {message['content']}"
        for message in old
    )

    response = client.messages.create(
        model=MODEL,
        max_tokens=250,
        messages=[{
            "role": "user",
            "content": (
                "Summarise this conversation in under 100 words. "
                "Keep order numbers and unresolved issues.\n\n"
                f"<conversation>{transcript}</conversation>"
            ),
        }],
    )

    summary = "".join(
        block.text
        for block in response.content
        if block.type == "text"
    )

    return summary, recent

Keep the summary as separate application state and include it as context in the next request. Don’t insert it as an additional user message before recent, because that can create invalid consecutive user messages.

Sensitive information should also be redacted before storage or transmission:

import re

def redact(text):
    return re.sub(
        r"\b(?:\d[ -]?){13,16}\b",
        "[REDACTED CARD]",
        text,
    )

How to Structure Prompts with Clear Boundaries

XML-style tags are ordinary text, not special API commands. They make each part of a prompt explicit:

<customer_reviews>
The product is comfortable, but the available colours are limited.
Customers also describe it as durable.
</customer_reviews>

<sales_data>
January: 120 units
February: 150 units
March: 98 units
</sales_data>

<task>
Compare the reviews with the sales data.
Identify possible relationships and state uncertainty.
</task>

Here, <customer_reviews> identifies reference material, <sales_data> identifies the data, and <task> identifies the instruction. Use similar boundaries for policies, user-generated content, examples, and output requirements.

How to Use a System Prompt

A system prompt defines ShopHelper’s general behaviour:

system_prompt = """
You are ShopHelper, a friendly customer-support assistant.

Keep answers concise and clear.
Do not invent prices, policies, or order details.
If information is missing, ask for it.
"""

Pass it separately from the conversation:

response = client.messages.create(
    model=MODEL,
    max_tokens=500,
    system=system_prompt,
    messages=[
        {"role": "user", "content": "Where is my order?"}
    ],
)

Because the customer didn’t provide an order number, ShopHelper should ask for one instead of guessing.

How to Add Tools

Tools let Claude request data from your application at runtime. Define each tool as a JSON schema:

tools = [
    {
        "name": "get_order_status",
        "description": (
            "Returns the current status of an order. "
            "Call this whenever the customer asks about an order."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "order_id": {
                    "type": "string",
                    "description": "The order ID, for example ORD-12345.",
                },
            },
            "required": ["order_id"],
        },
    }
]

Pass the tools list to each request:

response = client.messages.create(
    model=MODEL,
    max_tokens=500,
    system=system_prompt,
    tools=tools,
    messages=[
        {"role": "user", "content": "Where is order ORD-12345?"}
    ],
)

Claude reads the tool descriptions and decides whether to call a tool or reply directly.

How to Handle a Tool-Use Response

When Claude wants to call a tool, response.stop_reason is "tool_use" and the content includes a tool_use block. Your application must call the function, then send the result back to Claude:

import json

# Simulated order database
orders = {
    "ORD-12345": {"status": "Shipped", "eta": "2 days"},
    "ORD-67890": {"status": "Processing", "eta": "5 days"},
}

def get_order_status(order_id):
    return orders.get(order_id, {"error": "Order not found"})

def run_tool(tool_name, tool_input):
    if tool_name == "get_order_status":
        return get_order_status(tool_input["order_id"])
    return {"error": f"Unknown tool: {tool_name}"}

def process_response(response, messages):
    while response.stop_reason == "tool_use":
        tool_results = []

        for block in response.content:
            if block.type == "tool_use":
                result = run_tool(block.name, block.input)
                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": json.dumps(result),
                })

        messages.append({
            "role": "assistant",
            "content": response.content,
        })
        messages.append({
            "role": "user",
            "content": tool_results,
        })

        response = client.messages.create(
            model=MODEL,
            max_tokens=500,
            system=system_prompt,
            tools=tools,
            messages=messages,
        )

    return "".join(
        block.text
        for block in response.content
        if block.type == "text"
    )

The loop continues until Claude stops requesting tools and returns a final text response.

Claude Responses Can Contain Multiple Blocks

A single response may contain a thinking block, a text block, and one or more tool_use blocks. Always iterate over response.content and check block.type rather than accessing a fixed index:

for block in response.content:
    if block.type == "text":
        print("Text:", block.text)
    elif block.type == "tool_use":
        print("Tool:", block.name, block.input)
    elif block.type == "thinking":
        print("Thinking:", block.thinking)

When you add an assistant turn back to the message history, pass the full response.content list, not just the extracted text. This preserves all blocks, including tool_use blocks that the API requires when a subsequent tool_result references them.

Workflows vs Agents

An agent gives Claude tools and lets it decide what to do next. This is flexible but less predictable.

A workflow is a sequence of steps you control. Claude performs each step, but your code decides the order. Workflows are easier to test, debug, and monitor in production.

For ShopHelper, use a workflow to process a support ticket:

  1. Classify the ticket (refund, shipping, technical, other)
  2. Extract the order ID if present
  3. Look up the order status if an ID was found
  4. Draft a reply
def process_ticket(ticket_text):
    # Step 1: classify
    classification_response = client.messages.create(
        model=MODEL,
        max_tokens=50,
        messages=[{
            "role": "user",
            "content": (
                "Classify this support ticket into one category: "
                "refund, shipping, technical, or other. "
                "Reply with the category name only.\n\n"
                f"<ticket>{ticket_text}</ticket>"
            ),
        }],
    )
    category = "".join(
        block.text
        for block in classification_response.content
        if block.type == "text"
    ).strip().lower()

    # Step 2: extract order ID
    extraction_response = client.messages.create(
        model=MODEL,
        max_tokens=50,
        messages=[{
            "role": "user",
            "content": (
                "Extract the order ID from this ticket. "
                "Reply with the ID only, or 'none' if absent.\n\n"
                f"<ticket>{ticket_text}</ticket>"
            ),
        }],
    )
    order_id = "".join(
        block.text
        for block in extraction_response.content
        if block.type == "text"
    ).strip()

    # Step 3: look up order status
    order_info = ""
    if order_id.lower() != "none":
        status = get_order_status(order_id)
        order_info = f"Order status: {json.dumps(status)}"

    # Step 4: draft reply
    context = f"Category: {category}\n{order_info}"
    reply_response = client.messages.create(
        model=MODEL,
        max_tokens=300,
        system=system_prompt,
        messages=[{
            "role": "user",
            "content": (
                f"<context>{context}</context>\n\n"
                f"<ticket>{ticket_text}</ticket>\n\n"
                "Write a helpful reply to this support ticket."
            ),
        }],
    )
    reply = "".join(
        block.text
        for block in reply_response.content
        if block.type == "text"
    )

    return {
        "category": category,
        "order_id": order_id,
        "order_info": order_info,
        "reply": reply,
    }

Each step is a separate API call with a focused prompt. This makes it straightforward to log, test, or replace individual steps without changing the rest of the pipeline.

Chaining, Parallelisation, Routing, and Evaluator-Optimizer

Chaining

Chaining passes the output of one step as the input to the next. The ticket workflow above is an example. Use chaining when later steps depend on earlier results.

Parallelisation

When steps are independent, run them at the same time:

import concurrent.futures

def analyse_review(review):
    sentiment_response = client.messages.create(
        model=MODEL,
        max_tokens=10,
        messages=[{
            "role": "user",
            "content": f"Sentiment of this review (positive/negative/neutral): {review}",
        }],
    )
    sentiment = "".join(
        block.text for block in sentiment_response.content
        if block.type == "text"
    ).strip()

    topic_response = client.messages.create(
        model=MODEL,
        max_tokens=20,
        messages=[{
            "role": "user",
            "content": f"Main topic of this review in 3 words: {review}",
        }],
    )
    topic = "".join(
        block.text for block in topic_response.content
        if block.type == "text"
    ).strip()

    return {"sentiment": sentiment, "topic": topic}

reviews = [
    "Great product, fast shipping!",
    "The colour faded after one wash.",
    "Good value for the price.",
]

with concurrent.futures.ThreadPoolExecutor() as executor:
    results = list(executor.map(analyse_review, reviews))

for review, result in zip(reviews, results):
    print(f"{review[:40]!r}: {result}")

Routing

Route each request to a specialised handler based on its category:

def handle_refund(ticket):
    response = client.messages.create(
        model=MODEL,
        max_tokens=200,
        system="You are a refund specialist. Be empathetic and clear about the refund process.",
        messages=[{"role": "user", "content": ticket}],
    )
    return "".join(
        block.text for block in response.content
        if block.type == "text"
    )

def handle_shipping(ticket):
    response = client.messages.create(
        model=MODEL,
        max_tokens=200,
        system="You are a shipping specialist. Provide tracking information and delivery estimates.",
        messages=[{"role": "user", "content": ticket}],
    )
    return "".join(
        block.text for block in response.content
        if block.type == "text"
    )

def handle_general(ticket):
    response = client.messages.create(
        model=MODEL,
        max_tokens=200,
        system=system_prompt,
        messages=[{"role": "user", "content": ticket}],
    )
    return "".join(
        block.text for block in response.content
        if block.type == "text"
    )

handlers = {
    "refund": handle_refund,
    "shipping": handle_shipping,
}

def route_ticket(ticket_text):
    result = process_ticket(ticket_text)
    category = result["category"]
    handler = handlers.get(category, handle_general)
    return handler(ticket_text)

Evaluator-Optimizer

An evaluator-optimizer loop generates a response, scores it, and regenerates if the score is too low:

def evaluate_response(ticket, response_text):
    eval_response = client.messages.create(
        model=MODEL,
        max_tokens=100,
        messages=[{
            "role": "user",
            "content": (
                "Score this support response from 1–10. "
                "Consider accuracy, tone, and completeness. "
                "Reply with a number only.\n\n"
                f"<ticket>{ticket}</ticket>\n"
                f"<response>{response_text}</response>"
            ),
        }],
    )
    score_text = "".join(
        block.text for block in eval_response.content
        if block.type == "text"
    ).strip()
    try:
        return int(score_text)
    except ValueError:
        return 5

def generate_with_quality_check(ticket, min_score=7, max_attempts=3):
    for attempt in range(max_attempts):
        response = client.messages.create(
            model=MODEL,
            max_tokens=300,
            system=system_prompt,
            messages=[{"role": "user", "content": ticket}],
        )
        response_text = "".join(
            block.text for block in response.content
            if block.type == "text"
        )

        score = evaluate_response(ticket, response_text)
        print(f"Attempt {attempt + 1}: score {score}")

        if score >= min_score:
            return response_text

    return response_text  # return best attempt after max tries

How to Evaluate Prompt Quality

Testing prompt changes systematically prevents regressions. Define test cases with expected keywords, then compare two prompt versions:

test_cases = [
    {
        "input": "Where is my order ORD-12345?",
        "expected_keywords": ["shipped", "2 days"],
    },
    {
        "input": "I want a refund for my broken item",
        "expected_keywords": ["return", "refund", "30 days"],
    },
    {
        "input": "Do you ship internationally?",
        "expected_keywords": ["international", "ship"],
    },
]

def evaluate_prompt(system_prompt, test_cases):
    scores = []
    for test in test_cases:
        messages = [{"role": "user", "content": test["input"]}]

        if "ORD-" in test["input"]:
            reply = process_response(
                client.messages.create(
                    model=MODEL,
                    max_tokens=300,
                    system=system_prompt,
                    tools=tools,
                    messages=messages,
                ),
                messages,
            )
        else:
            response = client.messages.create(
                model=MODEL,
                max_tokens=300,
                system=system_prompt,
                messages=messages,
            )
            reply = "".join(
                block.text for block in response.content
                if block.type == "text"
            )

        keywords_found = sum(
            1 for kw in test["expected_keywords"]
            if kw.lower() in reply.lower()
        )
        score = keywords_found / len(test["expected_keywords"])
        scores.append(score)
        print(f"Input: {test['input'][:50]}")
        print(f"Score: {score:.0%} ({keywords_found}/{len(test['expected_keywords'])} keywords)\n")

    return sum(scores) / len(scores)

prompt_v1 = """
You are ShopHelper, a friendly customer-support assistant.
Keep answers concise and clear.
Do not invent prices, policies, or order details.
If information is missing, ask for it.
"""

prompt_v2 = """
You are ShopHelper, a friendly customer-support assistant for an online shop.

Guidelines:
- Always greet the customer warmly
- Be specific about timeframes (e.g., "within 30 days" not "soon")  
- For order issues, always mention the order ID in your response
- End with an offer to help further
- Do not invent prices, policies, or order details
"""

print("Evaluating prompt v1:")
score_v1 = evaluate_prompt(prompt_v1, test_cases)
print(f"Overall score: {score_v1:.0%}\n")

print("Evaluating prompt v2:")
score_v2 = evaluate_prompt(prompt_v2, test_cases)
print(f"Overall score: {score_v2:.0%}\n")

print(f"Winner: {'v2' if score_v2 > score_v1 else 'v1'}")

Conclusion

You’ve now built ShopHelper from a single API call into a structured customer-support assistant. The application manages multi-turn conversation history, uses XML-style tags to keep prompts unambiguous, calls external tools and processes their results, routes tickets through specialised handlers, and measures prompt quality with automated test cases.

These patterns — history management, tool use, workflow chaining, parallelisation, routing, and evaluation — apply to any production AI application, not just customer support. Start with the simplest approach that meets your requirements, then add complexity only where it’s needed.