How to Write Technical Documents That Put People First, Even in the Age of LLMs

Last Friday, just before leaving work, a Slack message popped up. A junior developer had dumped our entire API documentation into ChatGPT and asked "How do I authenticate with this endpoint," and the model had confidently walked them through an OAuth flow that had been deprecated two versions ago. Rain was falling on the River Clyde outside the Glasgow office window, and I stared at my monitor for a good while, thinking. The documentation wasn't wrong. I just hadn't sufficiently considered what happens when documentation is read by a machine.

Over the following months, we overhauled our team's documentation structure quite a bit. Here's what I learned along the way. The core idea is simple: first, create documentation that reads well for humans, then layer on structural cues that AI search systems can parse correctly. If you reverse that order, you ruin both.

Break documentation into units of "one question"

The most common mistake when writing technical documentation is cramming too much context into a single page. If one "Authentication" page contains everything (OAuth, API keys, service accounts, token renewal, error handling) humans get tired of scrolling, and LLMs miss or jumble content in the middle.

Here's what I actually put into practice.

  • Limit each page to answering exactly one question. "How do I issue an API key?" and "How does API key authentication work?" are separate pages.
  • Place a one-sentence summary at the very top of each page. When this summary gets picked up as a chunk in an LLM's RAG pipeline, it needs to contain the key information.
  • Explicitly link related pages at the end of the body text. Something like "Next step: See how to renew tokens." For humans, it's navigation; for crawlers, it's a relationship signal.

After breaking things apart this way, even within our own team, the time spent hunting for "where was that again?" dropped noticeably. When the documentation structure is clear, everyone benefits: search engines, LLMs, and humans alike.

Whiteboard diagram showing documentation pages each focused on a single question

Embed version and context directly in the body text

The root cause of that Slack incident was that the deprecated authentication method hadn't been fully removed from the documentation. It was still sitting inside a "Legacy" tab. To human eyes, the tab UI makes the distinction clear, but when an LLM scrapes the page text, the tab boundaries disappear. The current method and the deprecated method end up side by side in the same blob of text.

Here are some principles I changed afterwards.

  • Separate deprecated content onto its own page, and on the very first line of that page, write: "This authentication method has been deprecated since v2.3. For the current method, see (link)." As body text, not a UI banner.
  • Annotate code examples with API version numbers, not dates, as comments. Something like // Works with API v3.1+. LLMs tend to interpret "v3.1+" as a condition more accurately than "as of 2024."
  • State prerequisites in plain text at the beginning of the body. "This guide assumes you are using Python 3.10 or higher and our SDK v4.0 or higher." For humans, it saves time; for LLMs, it sets the scope.

Don't hide critical information behind UI elements like tabs, accordions, or toggles. This is the most basic hygiene rule for technical documentation in the LLM era.

Keep code examples at a "copy-and-run" level

I'm the kind of person who actually runs the code examples in documentation. When I proposed to the team that we add documentation code tests to our CI pipeline, the initial reaction was that it was overkill. But once we set it up, the frequency at which it caught broken examples was staggering. Every quarter, there were always three or four.

Working code examples matter for humans, but they matter especially when an LLM cites that code and presents it to a user. If an LLM confidently recommends a broken example, the user loses trust in the documentation itself.

Here are some practical tips.

  • Write one sentence directly above the code block describing what the code does. "The code below sends a GET request to the /users endpoint using a Bearer token." This sentence is decisive in helping the LLM accurately grasp the purpose of the code.
  • Use clear placeholder formats like YOUR_API_KEY for environment variables and placeholders. xxx or ... can be misinterpreted by LLMs as actual values or ignored entirely.
  • Place an example output directly below the code block. When the expected result is stated explicitly, it's easier for humans to debug, and it gives the LLM context that "running this code produces this result."

Code editor showing a documented code example with expected output

The invisible layer of structured metadata

This is where we enter territory that is mostly invisible to human readers but makes a huge difference for AI systems.

Our team added several fields to each documentation page's frontmatter: the page's purpose (tutorial, reference, how-to, explanation), the target audience level (beginner, intermediate, advanced), the relevant API version, and the date of last verification. This information doesn't show up on the rendered page, but crawlers and RAG systems can use it for filtering and prioritisation when indexing the documentation.

After introducing this, the accuracy of our internal AI chatbot's answers noticeably improved. We didn't measure exact numbers, but cases where it pulled answers from irrelevant pages clearly decreased. The document type field was particularly effective, because it made it possible to distinguish between a query like "find it in tutorials" and "find it in the reference."

Trying to design a complex schema from the start is exhausting. I started with just three fields (type, audience, and api_version) and that was enough.

In truth, there's one question that runs through all of these principles: "Can this document be understood if you rip it out of context and read it on its own?" Whether human or machine, a reader encountering the documentation for the first time arrives at a single page without having read the pages before or after it. That one page needs to contain the question, the prerequisites, the answer, and the next steps. This was already what made good documentation before LLMs came along, and it's even more true now.

The rain rarely lets up in Glasgow, and there's always a documentation editor open on my monitor. The next time you open a document, try re-reading that one-sentence summary on the first line. That single sentence is the first impression, for humans and machines alike.

Comments