TS-25: Technical Documentation
Documentation is a love letter to your future self.
– Julio Biason
This technical standard covers general principles for writing and maintaining technical documentation — in all its forms. It is concerned with the information architecture and lifecycle of documentation, but not with sentence-level writing style, tone of voice, terminology, citations, formatting, and other conventions. For that, see TS-26.
Types of documentation
Technical documentation takes numerous forms, each with a different audience, owner, lifecycle, and update trigger.
Different forms of technical documentation SHOULD be captured in distinct artifacts, located separately from one another.
Conflating multiple forms of documentation leads to verbose, complex docs with no clear ownership or maintenance lifecycle. Readers with different needs are poorly served by the same text, volatile content goes stale alongside stable content, and records that should be immutable get edited in place, losing the history of why things changed.
At a minimum, we MUST distinguish between the following types of documentation.
- User documentation. Describes how to use a software product. MUST be kept strictly in sync with new releases that change the public interface or other surface characteristics of the software. Reference documentation is a subset of user documentation, alongside installation and usage tutorials, and other guidance oriented around the user’s tasks and goals. Embedded documentation, which is embedded into the user interface of a software product, is another subcategory of user documentation.
- Architecture and design documentation. Explains how major components fit together, how they communicate, and how data flows through the system. Describes the as-is product system. Updated only when key building blocks, data flows, and deployment topologies change.
- Process documentation. Records point-in-time decisions. Examples include ADRs, RFCs, and design docs. These docs are typically treated as immutable once written, but can be deprecated or superseded by records that capture new decisions that override old ones.
- Operational documentation. Runbooks, incident playbooks, and deployment guides. Written for on-call engineers under time pressure. Scannability is prioritized over depth.
- Onboarding documentation. Oriented toward newcomers who lack context that existing contributors take for granted. Higher tolerance for explaining "obvious" things.
- Historical documentation. Changelogs and release notes.
Audiences
Each unit of documentation MUST be written for a specific audience. Mixing audiences in one document tends to under-serve all of them.
The typical audiences to consider are the following.
- Maintainers. Contributors who work on the codebase. They need enough context to change the code safely. They tolerate technical depth and internal jargon.
- Codeowners. A subset of maintainers who are ultimately responsible for approving all code integrations to the trunk.
- Contributors. People contributing code and configuration changes who are external to the organization or team that owns and maintains the codebase.
- Newcomers. A subset of maintainers and contributors who lack the shared context that existing maintainers and contributors take for granted. Onboarding documentation is for them. It should make fewer assumptions than regular reference documentation, for example.
- Operators. People running or supporting the software in production. They need operational procedures, not implementation detail.
- Users. Developers who use the software as a dependency, library, or API. They need to know the public surface and how to use it correctly, but they should not need to understand internals. They may be external to the organization that owns and runs the software, and they may not even be known to the maintainers.
A common mistake is mixing maintainer/contributor documentation with user documentation in a project’s README. This is particularly common in public open-source projects — understandably, as external contributors are often also the users of the software. It is RECOMMENDED instead to use a CONTRIBUTING file as the entry point for contributors and maintainers, leaving the README to focus on installation, usage, and API reference documentation for consumers. This applies equally to public open-source and private closed-source projects.
Determining value
Judgment about what documentation to create and maintain, and what not to, is usually made by the maintainers who are closest to the code. This so happens to be the audience that is least likely to need the documentation.
This is why software developers habitually undervalue documentation. Indeed, the more senior the engineer, and the more experience they have with a particular codebase, the more likely they are to neglect important documentation artifacts. The more fluent you are at reading source code — which is the source of truth for what a system does — the less value you will place on the effort involved in maintaining separate documentation that covers much the same ground.
It is true that not everything needs a separate document. Much information can be communicated through self-documenting code: clear naming, small functions, obvious hierarchical structure, good code comments, type signatures, data schemas, and tests. But, no matter how well a codebase achieves that noble goal, there will always be questions that cannot be answered by reading source code. And even where questions can be answered that way, it will often be quicker and less error-prone to find the answers by reading distilled information in other artifacts and formats, such as diagrams.
And, of course, not all audiences can read source code.
There’s a good test to determine whether a new unit of documentation would be valuable enough to justify its cost. The test is to ask a relevant stakeholder — a maintainer, a user, or support staff, say — a question about the system that they ought to know the answer to. If they cannot answer it, and if they can’t quickly find the answer for themselves, then documenting the answer will probably be worthwhile. The more frequently the question is asked by the relevant audience members, the more valuable the documentation will be.
Choosing what documentation to maintain then becomes a matter of asking the right questions. Those questions can uncovered from multiple sources. The classic source is code review. If a reviewer cannot answer their own question, you have a gap in your documentation. If the question keeps recurring in code review, you’ll get a good return on investing in filling that gap.
The onboarding process is another excellent opportunity to patch gaps in documentation. Capture the questions asked by newcomers, and add them to your product backlog. Support channels, such as help desks and issue trackers, are another source of gaps in user-oriented documentation. Every customer interaction is a signal for potential gaps to close.
The general principle is to treat recurring questions as defects in your documentation, and to log those defects for resolution, and to prioritize them by their value, just as you would handle a bug or feature request.
Location
There’s nothing like good documentation that lives alongside the work it relates to.
Out-of-band documentation tends to become stale as the code evolves separately from it. By contrast, code-level documentation — eg. inline comments, API specifications generated from code, and READMEs maintained alongside the source — tends to stay correct. When the code changes, this kind of documentation is naturally updated alongside it, often in the same revisions.
Therefore, in choosing where to locate our technical documentation, we should always err on the side of communicative code with good inline documentation, plus adjacent documentation captured in the same repository as the source code.
Docs-as-code
If documentation is kept close to the code, it follows that the lifecycle of documentation can be deeply integrated with that of the code. Docs-as-code is the practice of treating documentation with the same tools and workflows as source code. This means authoring in plain-text markup, storing in version control, reviewing through pull requests, and building and publishing through CI.
The goal is to merge the workflows for development and documentation, such that documentation changes are made alongside the code changes they describe, by the same people, in the same review loop. This is the most effective way to keep code and its documentation synchronized.
A docs-as-code workflow typically includes the following.
- Documentation is authored in a lightweight markup language — eg. Markdown, AsciiDoc, or reStructuredText — rather than a binary format or a proprietary CMS.
- Documentation is stored in the same version control repository as the source code, or at least in a nearby repository that shares the same review and release workflow.
- Documentation changes are reviewed through the same pull-request process as code changes. Documentation is subject to the same scrutiny and the same approval gates as the code.
- Documentation is built and published automatically by CI on merge, so that the published documentation always reflects the current state of the project.
- Documentation is tested where practical — eg. link checking, linting, and validation of markup and cross-references — as part of the CI pipeline. Broken documentation fails the build just as broken code does.
DocOps extends docs-as-code further by applying DevOps practices — eg. automation, monitoring, and continuous delivery — to the documentation pipeline.
Publications
The documentation that we embed within, or maintain directly alongside, the code is the source content for the documentation that we publish.
A single unit of source documentation may be published in multiple output formats, such as API reference guides, man pages, --help output, or online tutorials.
For each type of documentation, there SHOULD be a single authoritative source from which all publications are generated.
Whatever the format, published documentation SHOULD have the following characteristics.
- Discoverable. Documentation that cannot be found might as well not exist. Published docs SHOULD funnel users through all the likely pathways they might take to find the information they need, so that no matter where a reader starts, they find a pointer to the documentation within one or two steps. The documentation need not live in every one of these places, but it SHOULD be reachable from each of them.
- The project’s README, website, package registry page, and repository description SHOULD point to the documentation.
- Search, navigation menus and indexes, and cross-references from related content SHOULD lead readers to it.
- CLI tools SHOULD print a pointer to their documentation in
--helpoutput. - API surfaces SHOULD link from generated reference to narrative documentation and vice versa.
- Error messages and logs MAY link directly to the relevant troubleshooting page.
- Addressable. Documentation that cannot be referenced precisely cannot be discussed. Provide addresses that link directly to content at a granular level, so that readers can bookmark, share, and reference specific sections, in bug reports, pull requests, support tickets, and conversations. The more granular and easier to access, the better. This keeps the documentation at the center of the work rather than beside it.
- Every section heading SHOULD be a link target.
- Published documentation SHOULD provide stable URLs that do not break when content is reorganized.
- Deep links SHOULD survive version changes where possible, so that old links resolve to the current equivalent rather than 404.
- Cumulative. Order content so prerequisite concepts come first. A reader who arrives with partial knowledge and begins reading partway through should be able to rewind to earlier content to fill gaps, eg. to gain knowledge of technical terms that appear in the documentation.
- Complete. Within each publication, documentation content SHOULD NOT be truncated from its source, or otherwise made less than fully available.
- Accessible. Published documentation, especially that for users, MUST be accessible to readers through assistive technologies such as screen readers, as well as readers who navigate by keyboard or who need magnification or high contrast. This means using proper semantic market, including alternative text descriptions for all graphics, not using color along to convey meaning, and so on.
- Beautiful. Visual style should be intentional and aesthetically pleasing. Even text-only documentation has visual style in its spacing and capitalization. Aesthetics are not important to every reader, but some will struggle to find comfort in documentation that ignores them.
Tooling
If the validation and compilation of documentation into its various published formats is to be deeply integrated into the regular change management procedures around source code, then the choice of documentation tooling — including chosen markup languages, static site generators, build tools, and hosting platforms — becomes a requirement of the devops infrastructure.
Broadly, two separate toolchains are needed. One for the generation of API reference docs from code and inline annotations. And one for more free-flowing docs written in prose.
Documentation generators — tools that extract API reference from source code comments — are useful for reference material that must mirror the code exactly. But generated documentation is rarely good documentation on its own. It SHOULD be supplemented with hand-written narrative that explains why and how, not just what. When choosing this second toolchain, consider the following.
- Authoring experience. Can the people who need to write documentation do so without fighting the tool? A tool that requires a specialist to author content creates both a bottleneck and a silo.
- Review workflow. Does the tool support review through the same pull-request workflow the code uses?
- Build and publishing. Can the documentation be built and published automatically, on every change, without manual intervention?
- Versioning. Can the tool publish and maintain multiple versions of the documentation, for readers who remain on older versions of the software?
- Discoverability. Is documentation easily discoverable in the generated publications, eg. through search?
- Extensibility. Can the tool be extended and customized — eg. custom blocks, includes, plugins — without locking the content into a proprietary format that cannot be migrated?
- Hosting. Where will the published documentation live? A tool that only publishes to a proprietary platform limits portability.
It is RECOMMENDED to prefer a lightweight, plain-text markup language. This allows documentation to be authored in the same editor and reviewed in the same workflow as code. Markdown or AsciiDoc are the RECOMMENDED options. Markdown is more ubiquitous and has a wider ecosystem of tools. AsciiDoc is more tailored to the publication of technical documentation specifically, but it’s much more niche.
Avoid binary formats (word processor documents, proprietary help-authoring formats) for documentation that is maintained alongside code. They resist diffing, reviewing, and automation, and they tie the content to a specific tool.
Ownership
Each type of documentation MUST have clear ownership. Documentation without an owner tends to rot silently, because no one notices when it drifts from the system it describes, and no one is accountable for fixing it.
Reference documentation — eg. READMEs, API docs, and architecture docs — SHOULD be owned by whoever owns the code it describes, and updated in the same revisions that change the described behavior. Thus, a pull request that changes an API’s behavior without also updating the API documentation SHOULD be treated as incomplete.
Where documentation cannot reasonably be kept current — eg. because it describes a system no one maintains, or a decision that predates the current team — mark it explicitly as historical.
Content
Content is the conceptual information within documentation. The following principles help to create good content.
- Current. Incorrect documentation is worse than missing documentation, because it misleads with authority. Keep documentation up to date with the software it describes. Prefer version-agnostic content that needs less maintenance, and accommodate readers who remain on older versions of the software. A small set of fresh, accurate documentation is better than a sprawling, loose assembly in various states of disrepair, so outdated content SHOULD be deleted in small, regular increments rather than left to accumulate into a single overdue cleanup that never happens.
- Comprehensive. Together, all the documentation, in all its forms, should answer all the important questions that all stakeholders are likely to have about a system. Satisfying every obscure question is unattainable, so the objective is to cover the most important views and questions in full. Within each publication, the concepts it covers SHOULD be covered in full, or not at all. A document that describes fifty of one hundred configuration options is worse than one that describes none, because readers will assume the missing fifty do not exist. Completeness is scoped to what the document sets out to cover. A man page for
iconvthat documents every command-line option but points toiconv -lfor the supported encodings is complete, because the encodings are a separate publication. Where partial coverage is unavoidable, the document MUST say so explicitly and up front, so readers do not mistake the absence of information for the absence of the thing itself. - Focused. Each document MUST cover a single, narrow topic. Readers arriving via a search result or cross-reference want the answer to one question, and a document that covers several becomes harder to navigate, keep accurate, and link to. Narrow documents can be referenced precisely and updated in isolation. When a document grows to cover more than one topic, split it: extract each topic into its own file and link between them, rather than nesting ever more specific subsections in one growing page. Narrow does not mean short. A single topic may need considerable depth. The constraint is on breadth of subject matter, not word count.
- Ordered. Within a publication, prerequisite concepts SHOULD come before the concepts that depend on them, so a reader following the documentation linearly never meets a term that has not yet been introduced. Perfect ordering is not always possible, especially in reference documentation that is consulted non-linearly. But it SHOULD be the goal in tutorials, onboarding material, and any content read in sequence. Where tutorials and reference are separated, tutorials come first. The aim is not to force linear reading, since most readers skip around, but to help a reader with partial knowledge narrow their search. A reader who starts at the 25% mark and gets confused should be able to rewind to earlier content, not find that the prerequisites were never written down.
- Skim-able. Structure content so readers can identify and skip concepts they already understand, or that are irrelevant to their immediate question. Use descriptive headings, link text that describes its target (never "click here" or "this page"), and lead paragraphs and list items with the name of their key concept.
- Exemplary. Include examples and tutorials for the most common use cases. Many readers look at examples first. But do not attempt to exemplify everything. There are tradeoffs. Too many examples reduce skim-ability, and reduce their own individual usefulness.
- Consistent. Use consistent language and formatting. Where documentation is maintained by multiple people, adopt a style guide to enforce consistency.
- ARID — Accept (some) Repetition In Documentation. The DRY principle does not transfer cleanly from code to prose. Some business logic described by the code MUST be described again in the documentation. We SHOULD work to minimize that repetition, but accept that some is inevitable.
TS-26 provides guidance on concerns such as tone of voice, tense, headings, terminology, and formatting conventions.
READMEs
The README is usually the first document a maintainer or user will open. It SHOULD orient a new reader quickly, rather than attempt to be a complete reference.
At minimum, a README SHOULD cover the following.
- Purpose. What the project is and what problem it solves. Keep this short. A few sentences are normally ideal.
- Status. Whether the project is actively maintained, experimental, beta, or deprecated or superseded.
- Setup. The minimum steps to get the project running locally.
- Usage. The most common operations a reader will want to perform, with runnable examples. (See TS-26 for conventions on formatting commands and code blocks.)
A project README SHOULD cross-reference more in-depth documentation, located elsewhere, such as architecture docs, API references, and contribution guides. The README MUST NOT duplicate documentation that is maintained in other artifacts — this invites drift.
API documentation
Where practical, API documentation SHOULD be generated from the code itself, rather than hand-written and maintained separately. For example, an OpenAPI specification generated from route and type definitions, or reference docs generated from typed function signatures and docstrings.
Where generation isn’t practical, eg. documenting a third-party API, or a protocol rather than a code interface, hand-written API documentation MUST still describe the interface as it currently behaves, not as it was designed to behave.
At a minimum, API documentation MUST specify the following, per operation.
- Inputs (including required versus optional, and types).
- Outputs.
- Error conditions.
- Side effects.
Diagrams
A diagram is warranted when a structure or flow is easier to understand visually than as prose. Good use cases for diagrams include the documentation of component boundaries, data flows, state machines, and sequences of calls across systems.
A diagram SHOULD NOT be the only place a piece of information is recorded. Diagrams SHOULD be accompanied by prose that conveys the same information the diagram conveys. This is necessary to enable keyword search, and it also supports the use of assistive technologies to consume the documentation.
Prefer diagrams defined as code, eg. Mermaid or PlantUML, over opaque image files generated using a drawing tool. Text-based diagrams have the following advantages.
- They live alongside the code and are versioned in the same commits.
- They can be reviewed as a diff, so changes to the diagram are visible in code review.
- They don’t require proprietary or specialized software to edit.
Where an opaque image format is unavoidable, eg. a screenshot, keep the source file that generated it, if one exists, so the image can be easily regenerated in the future, if things change.
Changelogs
The changelog is the one form of documentation that is explicitly historical rather than descriptive of the current state of a system. It exists to tell a reader what changed between versions, not what the system currently does.
Where a project maintains a changelog, it SHOULD follow the Keep a Changelog conventions, though not necessarily the exact formatting. Follow these guidelines:
- Group changes by release.
- Write releases in reverse-chronological order.
- Further organize changes by what was added, changed, deprecated, removed, or fixed.
- It is further RECOMMENDED to list security updates as a distinct category.
The primary audience for changelogs is users of the software. Record the changes from the perspective of the consumer. For example, "the API now returns 429 on rate limit" is more informative for this audience than "add rate limiting middleware."
Changelogs MAY be generated automatically from other structured information, such as commit histories and pull request titles, as long as the output is optimized for users rather than maintainers.
Process artifacts
Most forms of documentation SHOULD be descriptive rather the prescriptive. They should describe the current state of the software, not future aspirations or intentions for it. For this reason, to stay accurate, most documentation SHOULD follow implementation rather than precede it. Documentation written before implementation quickly drifts from reality as plans, designs, and implementations naturally evolve.
However, there are plenty of valid exceptions to this rule. Indeed, there are whole categories of technical documentation where the docs are written (at least in draft form) ahead of the code. Examples include proposals for changes to software requirements specifications, architectural decision records (ADRs), requests for comments (RFCs), and design docs. What all these types of documentation have in common is they are concerned more with the process of designing, developing, maintaining, and operating software, rather than with describing the software itself.
These process artifacts MUST be clearly separated from the system documentation, preferably in entirely different locations — eg. their own repositories, rather than nested within code repositories.
Embedded documentation
Embedded documentation is user documentation that is built into the user interface of a product, and is presented to the reader at the point of need, while they are using the product, rather than in a separate document they must seek out. It is also known as in-app help or user assistance. It takes many forms, including the following.
- Microcopy. The small strings of text in the interface itself, such as button labels, menu items, field labels, placeholder text, and error messages.
- Tooltips. Short explanations that appear when the reader hovers over or focuses on an element.
- Inline help. Hints and explanations displayed alongside a control, such as the text beneath a form field describing the expected format.
- Contextual help. Help that is tied to the screen, element, or task the reader is working on, such as a "?" icon next to a field, or a help panel whose content changes with the page.
- Empty states. The messages shown where content would normally appear, which explain what the space is for and how to fill it.
- Onboarding flows and product tours. Guided walkthroughs, coach marks, and checklists that introduce new readers to the product.
- Progressive disclosure. Detail that is hidden by default and revealed on demand, such as expandable "Learn more" sections, so that the reader sees only as much help as they ask for.
UX writing is the discipline of writing the words that appear inside a user interface — the button labels, menu items, error messages, empty states, tooltips, onboarding flows, etc.
UI text MUST be clear before it is clever. A witty error message that the reader has to decode is worse than a plain one that tells them what happened and what to do next. Wordplay, brand voice, and personality are welcome where they do not obscure meaning, but clarity always wins when the two conflict.
Button and link labels SHOULD describe the action the control performs, in the reader’s voice rather than the system’s. Prefer "Delete file" to "OK" or "Submit." A reader scanning a dialog SHOULD be able to understand what each button does from its label alone, without reading the surrounding text.
The same action SHOULD be labeled the same way everywhere it appears. If the primary action is "Save" on one screen, it SHOULD NOT be "Submit" on the next. Inconsistent UI text forces the reader to re-parse familiar actions and erodes trust.
Error messages SHOULD:
- State what happened, in plain language. Avoid raw error codes as the only message.
- State what the reader can do about it, if anything. If the error is recoverable, the message SHOULD point to the recovery action.
- Avoid blaming the reader. "Invalid email" is better than "You entered an invalid email". "Email address not recognized" is better still.
See also TS-15, TS-17, and TS-18.
References
- Elhage, N (2020). Computers Can Be Understood.
- Google. Google Documentation Style Guide.
- Write the Docs. Write the Docs.