
Why the Claude API deserves your attention
In 2026, Anthropic has cemented the Claude family as a benchmark for applications that need reliable reasoning, long context (1M tokens on Opus 4.6) and tool use. Unlike SDKs that hide important decisions, the Claude API exposes clean primitives: messages, tools, streaming and cache. Knowing how to wield those primitives directly is what separates an AI app that scales from a prototype that crumbles.
This guide covers the minimum viable path: creating a key, making your first call, adding tool use, enabling streaming, and turning on prompt caching to cut cost.
60-second setup
Create an account on Anthropic Console, generate a key in Settings → API Keys, and export it as an environment variable. Never paste the key in your code.
export ANTHROPIC_API_KEY="sk-ant-..."
pip install anthropic
# or
npm install @anthropic-ai/sdk
A best practice few people follow: create separate keys per environment (dev, staging, prod) and per service. That way a leak doesn't compromise everything, and you can attribute consumption to a specific route.
First call (Python)
from anthropic import Anthropic
client = Anthropic()
resp = client.messages.create(
model="claude-opus-4-6",
max_tokens=1024,
messages=[
{"role": "user", "content": "Summarize Clean Architecture in 3 bullets."}
],
)
print(resp.content[0].text)
Three things the docs emphasize:
max_tokensis required — protects your wallet.role: "system"is now a top-level parameter (system="..."), not insidemessages.contentcan be a list of blocks (text, image, tool_result), not just a string.
Tool use: letting the model act
Tool use is what turns a chatbot into an agent. You define functions in JSON schema, the model decides when to call them, you execute, and feed the result back.
tools = [
{
"name": "get_weather",
"description": "Returns the current temperature for a city.",
"input_schema": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City"},
},
"required": ["city"],
},
}
]
resp = client.messages.create(
model="claude-opus-4-6",
max_tokens=1024,
tools=tools,
messages=[{"role": "user", "content": "What's the temperature in Curitiba?"}],
)
When stop_reason comes back as tool_use, you run the local function, append the result as {"role": "user", "content": [{"type": "tool_result", ...}]} and call the API again. The model sees the result and forms a final response.
Golden rule: descriptions are everything. A precise description and good schema examples can reduce wrong calls by up to 80%.
Streaming: UX that feels like magic
For chat and any long output, use streaming. The user sees text appearing token by token and perceived latency drops dramatically.
with client.messages.stream(
model="claude-opus-4-6",
max_tokens=2048,
messages=[{"role": "user", "content": "Write a short story about AI."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
In production, forward chunks to the frontend via Server-Sent Events (SSE) or WebSockets. The SDK also emits more granular events (content_block_start, content_block_delta, message_stop) — useful when you need to intercept tool calls in real time.
Prompt caching: brutal cost cuts
Apps with large prompts (long instructions, RAG context, few-shot examples) reuse context across calls. Prompt caching lets you mark parts of the prompt as cacheable, and Anthropic charges only 10% of normal price on subsequent hits.
resp = client.messages.create(
model="claude-opus-4-6",
max_tokens=1024,
system=[
{
"type": "text",
"text": "You are an assistant specialized in Brazilian tax law...",
"cache_control": {"type": "ephemeral"},
}
],
messages=[{"role": "user", "content": "What is Simples Nacional?"}],
)
For apps with 5k+ token instructions, this can cut total cost by 60-80%. Check usage.cache_read_input_tokens in the response to confirm the cache hit.
Costly mistakes
- Not setting low
max_tokensin dev — one bad call can burn money. - Reusing the same key in prod and local scripts — when it leaks, you don't know where from.
- Ignoring
stop_reason—end_turnis different fromtool_useandmax_tokens. Treating them all as success can mask truncated responses. - Not validating input before sending it to the model — a malicious user can inject instructions into the system prompt.
- Streaming without heartbeat — connections drop in proxies (Cloudflare, nginx) after 30s of silence. Send periodic SSE comments.
Next steps
- Read the official docs: docs.claude.com
- Try the Agent SDK for multi-step tasks with native loops.
- For serious Node.js apps, set up observability (Helicone, Langfuse) from day one.
- Define console budgets to avoid end-of-month surprises.
The Claude API is powerful precisely because it's simple. A 50-line file already gives you a working agent. Use it well.

