Back to blog
Artificial Intelligence

Claude API: A Practical Guide for Developers in 2026

Bruno Bracaioli
Claude API: A Practical Guide for Developers in 2026

Why the Claude API deserves your attention

In 2026, Anthropic has cemented the Claude family as a benchmark for applications that need reliable reasoning, long context (1M tokens on Opus 4.6) and tool use. Unlike SDKs that hide important decisions, the Claude API exposes clean primitives: messages, tools, streaming and cache. Knowing how to wield those primitives directly is what separates an AI app that scales from a prototype that crumbles.

This guide covers the minimum viable path: creating a key, making your first call, adding tool use, enabling streaming, and turning on prompt caching to cut cost.

60-second setup

Create an account on Anthropic Console, generate a key in Settings → API Keys, and export it as an environment variable. Never paste the key in your code.

export ANTHROPIC_API_KEY="sk-ant-..."
pip install anthropic
# or
npm install @anthropic-ai/sdk

A best practice few people follow: create separate keys per environment (dev, staging, prod) and per service. That way a leak doesn't compromise everything, and you can attribute consumption to a specific route.

First call (Python)

from anthropic import Anthropic

client = Anthropic()

resp = client.messages.create(
    model="claude-opus-4-6",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Summarize Clean Architecture in 3 bullets."}
    ],
)

print(resp.content[0].text)

Three things the docs emphasize:

  • max_tokens is required — protects your wallet.
  • role: "system" is now a top-level parameter (system="..."), not inside messages.
  • content can be a list of blocks (text, image, tool_result), not just a string.

Tool use: letting the model act

Tool use is what turns a chatbot into an agent. You define functions in JSON schema, the model decides when to call them, you execute, and feed the result back.

tools = [
    {
        "name": "get_weather",
        "description": "Returns the current temperature for a city.",
        "input_schema": {
            "type": "object",
            "properties": {
                "city": {"type": "string", "description": "City"},
            },
            "required": ["city"],
        },
    }
]

resp = client.messages.create(
    model="claude-opus-4-6",
    max_tokens=1024,
    tools=tools,
    messages=[{"role": "user", "content": "What's the temperature in Curitiba?"}],
)

When stop_reason comes back as tool_use, you run the local function, append the result as {"role": "user", "content": [{"type": "tool_result", ...}]} and call the API again. The model sees the result and forms a final response.

Golden rule: descriptions are everything. A precise description and good schema examples can reduce wrong calls by up to 80%.

Streaming: UX that feels like magic

For chat and any long output, use streaming. The user sees text appearing token by token and perceived latency drops dramatically.

with client.messages.stream(
    model="claude-opus-4-6",
    max_tokens=2048,
    messages=[{"role": "user", "content": "Write a short story about AI."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

In production, forward chunks to the frontend via Server-Sent Events (SSE) or WebSockets. The SDK also emits more granular events (content_block_start, content_block_delta, message_stop) — useful when you need to intercept tool calls in real time.

Prompt caching: brutal cost cuts

Apps with large prompts (long instructions, RAG context, few-shot examples) reuse context across calls. Prompt caching lets you mark parts of the prompt as cacheable, and Anthropic charges only 10% of normal price on subsequent hits.

resp = client.messages.create(
    model="claude-opus-4-6",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": "You are an assistant specialized in Brazilian tax law...",
            "cache_control": {"type": "ephemeral"},
        }
    ],
    messages=[{"role": "user", "content": "What is Simples Nacional?"}],
)

For apps with 5k+ token instructions, this can cut total cost by 60-80%. Check usage.cache_read_input_tokens in the response to confirm the cache hit.

Costly mistakes

  1. Not setting low max_tokens in dev — one bad call can burn money.
  2. Reusing the same key in prod and local scripts — when it leaks, you don't know where from.
  3. Ignoring stop_reason — end_turn is different from tool_use and max_tokens. Treating them all as success can mask truncated responses.
  4. Not validating input before sending it to the model — a malicious user can inject instructions into the system prompt.
  5. Streaming without heartbeat — connections drop in proxies (Cloudflare, nginx) after 30s of silence. Send periodic SSE comments.

Next steps

  • Read the official docs: docs.claude.com
  • Try the Agent SDK for multi-step tasks with native loops.
  • For serious Node.js apps, set up observability (Helicone, Langfuse) from day one.
  • Define console budgets to avoid end-of-month surprises.

The Claude API is powerful precisely because it's simple. A 50-line file already gives you a working agent. Use it well.

Compartilhar:

Fique por dentro

Receba novos artigos sobre IA, desenvolvimento e tecnologia direto no seu email.