HACKS VITAE

TOOLS & TECH · SOURCES SHOWN

Prompt Engineering 2026 Prompting Manual
Architecting High-Performance AI Outputs

PRICES, VERSIONS AND FACTS AS OF SEPTEMBER 2026

What Anthropic's, OpenAI's and Google's own prompting guides recommend in September 2026: clear structure over persuasion, whether “think step by step” still helps, long context, tools, self-review, and a template to adapt.

8 SECTIONS · HOVER A POINT TO JUMP
Published
June 13, 2026
Updated
September 29, 2026
Facts as of
September 2026
Read
10 min
Sections
8

BACKGROUND · JACQUES LOUIS DAVID, ANTOINE LAURENT LAVOISIER AND MARIE ANNE LAVOISIER, 1788 · THE MET, OPEN ACCESS

THE SHORT VERSION

  1. Some older prompting tricks are now unnecessary. OpenAI's guide says telling its reasoning models to “think step by step” is unnecessary and can sometimes hinder them; the phrase comes from a 2022 paper, when it helped a great deal.
  2. Anthropic, OpenAI and Google all recommend marking out the parts of a prompt — instructions, source material, output format — with XML-style tags or headings.
  3. The models named here list context windows of around a million tokens, but longer is not automatically better: Anthropic says accuracy and recall degrade as the context grows.
  4. The evidence on asking a model to critique its own draft is mixed; a review that can check the draft against something outside it — test results, a tool, a source — is the version with more support.
  5. No prompt stops a model from making things up. Letting it say it does not know reduces errors; Anthropic says its techniques “don't eliminate them entirely”.

THE ARTICLE · 10 MIN

Prompt Engineering 2026 Prompting Manual

Prompting advice ages quickly. A lot of what still circulates was written for models that did not reason before answering, could read only a few pages at once, and had no tools. Current models from the large labs do all three, and their makers’ own guides now say some of the older tricks are unnecessary.

This page sets out what those guides recommend, where the research behind a tactic is mixed, and a template you can adapt. It was written in June 2026 and brought up to date in September 2026. Model behaviour changes within months, so treat it as a snapshot and test anything here against the model you use.

1. The Core Infrastructure of Modern LLMs

Three changes explain why older prompting habits matter less than they did.

Models reason before they answer. The phrase “Let’s think step by step” comes from a 2022 paper by Takeshi Kojima and colleagues. Adding it before each answer raised one large model’s accuracy on the MultiArith arithmetic test from 17.7% to 78.7%. Reasoning models now do that work internally. OpenAI’s guide for its reasoning models says prompting them to “think step by step” or “explain your reasoning” is unnecessary, and that the instruction “may not enhance performance (and can sometimes hinder it)”. Anthropic’s guide says a general instruction such as “think thoroughly” often produces better reasoning than a hand-written step-by-step plan.

Models can hold very long inputs. Anthropic lists a 1M-token context window for Claude Opus 5.5, the model it suggests starting with (its fastest model, Claude Haiku 4.5, has 200k). OpenAI lists 1.05M tokens for GPT-6 Astra, its flagship model, and Google lists an input limit of 1,048,576 tokens for Gemini 3.8 Flash. Google’s prompting guide puts 100 tokens at roughly 60 to 80 words, so a million tokens is somewhere around 600,000 to 800,000 words.

Longer is not automatically better. Anthropic’s documentation says that as the token count grows, “accuracy and recall degrade”, which it calls context rot. A 2023 study by Nelson Liu and colleagues, “Lost in the Middle”, found that models used information at the start or end of a long input much better than information in the middle.

Models can use tools. Anthropic, OpenAI and Google each offer a code-execution tool that lets the model write and run code in a separate environment (Anthropic and OpenAI describe theirs as a sandbox), and web search is offered as a tool too. A prompt can now say when those tools should be used.

2. The Structural Layout: Separate the Parts of a Prompt

The three vendors’ guides agree on one thing more than any other: mark out the different parts of a prompt so the model can tell instructions from material.

  • Anthropic: XML tags “help Claude parse complex prompts unambiguously, especially when your prompt mixes instructions, context, examples, and variable inputs.”
  • OpenAI: “Use delimiters like markdown, XML tags, and section titles to clearly indicate distinct parts of the input.”
  • Google: “XML-style tags (e.g., <context>, <task>) or Markdown headings are effective. Choose one format and use it consistently within a single prompt.”

For long inputs (around 20,000 tokens or more), Anthropic also suggests placing the documents near the top of the prompt, above the question and instructions. It reports that questions at the end “can improve response quality by up to 30 percent in tests, especially with complex, multidocument inputs.”

The layout below is one way to put that into practice. The five block names are this page’s own labels, not a standard; any clear, consistent labels do the same job.

### SYSTEM PERSONA
The expertise, perspective, tone and limits you want the model to work within.

### INPUT DATA
<source_context>
[Paste the source material here: documents, data, specifications, notes]
</source_context>

### TASKS
1. First instruction, stated plainly.
2. Second instruction, stated plainly.

### GUARDRAILS
- Limits on what the answer may do (for example, "Do not assume facts that are not in the source").
- When to use tools (for example, "Use web search to check anything that may have changed since your training data").

### OUTPUT FORMAT
The exact format you want back (for example, a Markdown table or JSON).

3. Three Tactics, and What Backs Them

Re-grounding a long conversation

In a long chat, earlier material can drift out of the model’s effective attention, which is what context rot and the “Lost in the Middle” results describe. One practical response is to ask the model to restate where things stand before a new phase of work:

“Before we continue, review our conversation so far. Summarise the current status, the decisions we have made and the open questions in three bullet points. Write nothing else.”

This is a working habit rather than a tested technique; no study of this exact prompt was found. The same idea is built into some products: Anthropic’s compaction feature “automatically summarizes earlier parts of the conversation” so a long session can continue. Another option is to start a fresh conversation and paste the summary in.

Making the model use a tool for calculations

The code-execution tools exist for this: OpenAI describes its code interpreter as a way for models “to write and run Python code in a sandboxed environment” for problems in “data analysis, coding, and math”.

  • Vague: “Calculate the compound growth rate in this spreadsheet.”
  • Specific: “Write and run a Python script to compute the compound growth rate from the attached data. Do not estimate the figure. Show the script and its output alongside the final number.”

Asking for a review before the final answer

Asking a model to critique and revise its own draft has mixed evidence. A 2023 paper, “Self-Refine” (Aman Madaan and colleagues), reported that outputs improved by about 20% in absolute terms on average across seven tasks when a model gave itself feedback and revised. A later paper by Jie Huang and colleagues, presented at ICLR 2024, found that models “struggle to self-correct their responses without external feedback” on reasoning problems, and that performance sometimes got worse after self-correction.

The two results sit together more comfortably when the review has something outside the model to check against. A 2024 survey of self-correction research by Ryo Kamoi and colleagues, published in Transactions of the Association for Computational Linguistics, found that earlier studies often used “unfair evaluations that over-evaluate self-correction”, and concluded that self-correction “works well in tasks that can use reliable external feedback”. Anthropic’s guide suggests adding a line such as “Before you finish, verify your answer against [test criteria]”, and notes that one of its models, Claude Opus 5, checks its own work well without being asked. A review prompt that gives the model something outside itself to check against looks like this:

“Before giving your final answer, check your draft against [the test results / the source documents / the output of the code you ran]. List anything in the draft they do not support, revise the draft, then show only the revised version.”

4. Older Habits and What the Guides Say Now

Older habitWhat current guides and research say
Adding “Let’s think step by step” to every questionIt helped older models a great deal (Kojima and colleagues, 2022). OpenAI says it is unnecessary for its reasoning models and can sometimes hinder them.
Being very polite, or very rude, to get a better answerThe research is mixed. A 2024 study in English, Chinese and Japanese found that impolite prompts often did worse. A small 2025 study (50 questions, one model, ChatGPT-4o) found very rude prompts scored slightly higher: 84.8%, against 80.8% for very polite ones. The vendors’ guides ask for clear, direct instructions rather than any particular tone.
Emphasis in capitals: “CRITICAL: you MUST…”Anthropic says its newer models may now overreact to that kind of language and suggests plain wording such as “Use this tool when…”. Google’s guide says to avoid “overly persuasive language”.
Cutting source material down to fit a small limitThe models named above accept around a million tokens, so the full text often fits. Choosing what goes in still matters, because accuracy and recall degrade as the context grows.
Taking the first answer as finishedA review step can help, especially when it checks the draft against something outside the model, such as test results, a tool’s output or sources.

5. A Template to Adapt

Copy it, change the parts in brackets, and test it on the model you use.

### SYSTEM PERSONA
You are a senior systems architect. Write directly and clearly, without introductions or pleasantries.

### INPUT DATA
<technical_specification>
[Paste your code, notes or documentation here]
</technical_specification>

### TASKS
1. Read the <technical_specification> block and identify the three most serious bottlenecks.
2. For each bottleneck, outline one concrete fix.
3. Use web search to check whether any recent security guidance affects these fixes, and name the source for each point.

### GUARDRAILS
- No introductory phrases, filler or closing summaries.
- Base every statement on the <technical_specification> block or on sources you found with web search, and name them.
- If the material is not enough to complete a task, write "Data Insufficient" instead of guessing.

### OUTPUT FORMAT
Markdown, with ### headings for each part. Put the findings in a three-column table:

| Bottleneck | Fix | Relevant guidance and source |
| :--- | :--- | :--- |

The “Data Insufficient” line follows advice in Anthropic’s guide on reducing hallucinations: give the model explicit permission to say it does not know.

Frequently Asked Questions (FAQ)

What has changed in prompting since the early chatbots?

Reasoning models work through problems before answering, so the step-by-step phrasing that helped older models is often unnecessary. Context windows grew to around a million tokens on many current models, and models can run code and search the web. The vendors’ guides now put the weight on clear structure and specific instructions.

Why use tags like <source_context> in a prompt?

Tags or headings mark where source material ends and instructions begin. Anthropic, OpenAI and Google all recommend some form of delimiter for this. The tag names themselves are up to you; consistency matters more than the exact words.

Should I still summarise documents before pasting them in?

Often you can paste the full text, since many current models accept around a million tokens. Anthropic’s documentation still notes that accuracy and recall degrade as the context grows, so leaving out material that is not needed can help, and placing long documents above the question is Anthropic’s suggested layout.

Can prompting stop a model from making things up?

No prompt removes the problem. Anthropic’s guide says its techniques “significantly reduce hallucinations” but “don’t eliminate them entirely”, and advises checking critical information. Letting the model say it does not know, grounding it in supplied documents, and making it use tools for calculations all help. For more on why models invent answers, see Why AI makes things up.

Putting It to Use

The habit that holds up whichever model you use is to test prompts on your own work. Anthropic’s guide makes the same point about its own advice: treat a tip as measured on one model, and re-check it against your own tests before relying on it with another. In practice that can be as simple as a short list of real questions with answers you already know, run again whenever you change the prompt or the model changes, changing one thing at a time. For the terms used on this page, see the AI glossary; for how these models work underneath, see How large language models work.

Sources

Checked September 2026.

  • ai models

SHARE & CITE

Hacks Vitae. "Prompt Engineering 2026 Prompting Manual: Architecting High-Performance AI Outputs." June 13, 2026. https://www.hacksvitae.com/life-hack/prompt-engineering-2026-prompting-manual-architecting-high-performance-ai-outputs

That's what we found. The rest is your call.

118 articles, each with its sources listed. Spotted something off? hacksvitae@gmail.com

Open the library