
What changed (and what didn't)
In 2023, "prompt engineering" became a meme. In 2024, it was declared dead. In 2026, with more capable but also more expensive models with absurd context windows, the discipline has never been more important — it just changed shape. It's no longer about finding the magic phrase that unlocks the model. It's about structuring context so the model returns consistent, auditable, and cheap output.
This article distills the techniques that still work in 2026, based on what Anthropic, OpenAI and Google publish in their official docs and what shows up in real production code.
Principle 1: Be specific, not polite
Models don't need "please" and "thank you". They need clear instructions about format, scope and success criteria.
Bad:
Help me write an email to a customer.
Good:
Write a formal English email to a B2B customer informing them that
invoice #1234 is 5 days overdue. Tone: cordial but firm. Max 120 words.
Include a clear CTA for payment.
The difference isn't politeness — it's the number of decisions the model has to make alone. The more implicit decisions, the more output variability.
Principle 2: Use XML structure
Anthropic recommends XML to delimit sections because Claude's tokenizer treats tags with special attention. OpenAI and Gemini also handle them well.
<context>
The company sells B2B software for law firms.
</context>
<task>
Generate 3 blog post titles about GDPR.
</task>
<rules>
- Each title is max 60 characters
- No clickbait
- Focus on practical benefit
</rules>
<output_format>
JSON with array "titles"
</output_format>
Why it works: markers eliminate ambiguity between instructions and data, and give the model cognitive "hooks" to organize reasoning.
Principle 3: Few-shot > zero-shot
Whenever the task has a pattern, show 2-5 examples before the real input. Accuracy goes up dramatically.
Classify sentiment as POSITIVE, NEGATIVE or NEUTRAL.
Text: "Loved the product, arrived fast."
Sentiment: POSITIVE
Text: "Service was ok, nothing special."
Sentiment: NEUTRAL
Text: "Broke on day two. Garbage."
Sentiment: NEGATIVE
Text: "{real_input}"
Sentiment:
Rule: cover edge cases in your examples. If you only show positives and negatives, the model will force one of them on neutral cases.
Principle 4: Controlled chain-of-thought
"Think step by step" became cliché, but it works — as long as you direct the thinking. Models like Claude have a native thinking mode where reasoning is separate from the final answer.
Before answering, in <scratchpad>:
1. List the relevant information from context
2. Identify what's being asked
3. Check for contradictions
4. Formulate the answer
Then, in <answer>, give only the final answer in 2 sentences.
You get two things: more correct answers (because the model "drafts") and clean output (because you only render the <answer> tag in the product).
Principle 5: Role priming with substance
"You are an X expert" alone does nothing. What works is describing the kind of reasoning you want.
Empty:
You are a security expert.
Useful:
You are a security auditor in OWASP style. When you receive code, you
first identify attack vectors (injection, XSS, deserialization), then
propose the simplest possible fix. Don't generalize — cite specific
lines and functions.
The difference is between "act as" and "reason like".
Principle 6: Use structured output
If you need JSON, ask for JSON with a schema. Modern models support native structured output (Claude has tool_use, OpenAI has response_format, Gemini has responseSchema). Going from "manual markdown parsing" to "validated JSON" cuts bugs by 90%.
tools = [{
"name": "save_classification",
"input_schema": {
"type": "object",
"properties": {
"category": {"type": "string", "enum": ["A", "B", "C"]},
"confidence": {"type": "number", "minimum": 0, "maximum": 1},
"rationale": {"type": "string", "maxLength": 200},
},
"required": ["category", "confidence", "rationale"],
},
}]
You "trick" the model into always returning the right format because it thinks it's calling a function.
Principle 7: Test like code
The biggest mistake when starting prompt engineering is treating prompts as throwaway text. Treat them like code: versioned, with regression tests, with metrics.
- Save prompts in versioned files (not hardcoded strings).
- Keep a dataset of 20-50 representative cases.
- For each change, run the dataset and compare results.
- Use tools like
promptfooor implement a simple evaluator.
Without this, every optimization is guesswork.
What NOT to do anymore
- Don't use "creative" jailbreaks — modern models are trained against them and you're wasting time.
- Don't call the model 5 times to refine when one well-structured call solves it.
- Don't trust "temperature: 0" as synonymous with determinism — there's still variance between runs.
- Don't dump 200k tokens when 10k well-selected ones work better.
Conclusion
In 2026, prompt engineering became an engineering discipline, not an art. Those who treat it as code, measure results and iterate methodically extract the most from models at the lowest cost. Those still hunting for "the magic prompt" spend 10x more and ship less.

