TS-57: Logging, Monitoring, Observability
This technical standard sets out principles and best practices for logging, monitoring, and observability – collectively known as the visibility quality attribute.
Logging
Personally identifiable information (PII) MUST NOT be sent to log output. The same applies to any other private information, and to protected content, such as credentials, secrets, access tokens, session identifiers, payment details, and the content of users' documents and messages. Code that logs data passed to it from elsewhere, particularly framework and library code, cannot know in advance which of that data is private. Such code SHOULD log identifiers and metadata rather than payloads.
Runtime errors MUST be logged, so that issues can be detected and investigated.
Log levels
Every log entry MUST be assigned a level that reflects the consequence of the event it records. This standard uses the five levels defined below, from most to least severe. Logging libraries name them differently, eg. WARNING for WARN or VERBOSE for TRACE, and some add a level above ERROR, such as FATAL or CRITICAL. Map a library’s levels onto this scale.
- ERROR
Something has failed with consequences that users will see, and the system cannot recover from it without intervention. A request could not be served, data could not be written, or a dependency has stayed unavailable beyond its retry budget. Always logged. An
ERRORentry is a candidate for alerting or error reporting.- WARN
Something serious and unexpected has happened, with consequences that users may see, but the system has recovered, or can recover without data loss, by retrying, falling back, or restarting. Always logged.
- INFO
Something noteworthy has happened that is likely to have a wide impact, though it is not an error, such as a service starting or stopping, a configuration change taking effect, or a failover. Always logged. Only the component that is authoritative for the event SHOULD log it at this level, so that one event is not recorded again by every component it passes through.
- DEBUG
Detail that helps to investigate unexpected behavior. Log only what is needed to understand what a component is doing.
DEBUGlogging MAY be enabled in production, and it MUST be possible to disable it there without a code change.- TRACE
Everything else, including step-by-step detail of normal operation. Disabled in production by default.
Conditions that are expected in normal operation are not errors, and SHOULD NOT be logged above DEBUG. Invalid input from an untrusted source, such as a malformed request or a corrupt file on shared storage, is expected. So is a transient loss of network connectivity that the system retries.
Log each failure once. Within a component, the code that detects a failure and handles it SHOULD log it, and code that only passes the failure on SHOULD NOT. Do not log an error and then throw it, since whichever code catches it will log it again. Where an exception or error value already carries everything a log entry would record, logging it at the point it is raised adds nothing.
Rate limiting
A condition that can recur many times in quick succession, such as a failing dependency that is called on every request, can flood the log with copies of the same message. The flood costs storage and ingestion, buries the messages around it, and in a bounded buffer can push other messages out before anyone reads them.
Log statements on paths where this can happen SHOULD be rate-limited. Log the first occurrence, suppress repeats for an interval, and then log a count of the occurrences that were suppressed, so that the log still records how often the condition arose. Many logging libraries and log shippers provide rate limiting. Prefer these to a hand-written limiter.
The cost of disabled log statements
A log statement below the enabled level writes nothing, but it is not free. In most languages, a function’s arguments are evaluated before it is called, so a message built by string concatenation or formatting is built in full and then discarded. Where the message is cheap to build, this does not matter. On a hot path, or where building the message serializes a large object, it does.
Log statements SHOULD NOT do expensive work that is discarded when their level is disabled. Prefer a logging API that defers formatting, taking a message template and its arguments and formatting them only when the entry will be written. Where the message needs an expensive computation of its own, guard the statement with a check of the enabled level, or pass a function that the logging API calls only when the level is enabled. Any such guard MUST contain only the logging, never logic that the program depends on, because disabling the level removes it.
Language standards set out the idioms for their language. For Java, see TS-33: Java.
Monitoring
Uptime monitoring MUST be implemented for software-as-a-service applications. Individual services within a distributed system MUST be monitored separately to the user interface of the application as a whole.
Alerting
Production alerting – notifications of errors or other issues – MUST be implemented in production environments, and SHOULD be implemented in pre-production environments.
Oncall health
Alerting exists to reach a human, and the oncall rotation that receives those alerts is a system with its own failure modes — alert fatigue, excessive page volume, and burnout — that degrades the effectiveness of every other observability investment if left unmanaged. Oncall health SHOULD be measured and tracked with the same rigor as any other production metric, using signals such as:
- Page volume per rotation and per engineer, including a breakdown of actionable versus non-actionable pages.
- Time-of-day distribution of pages, since a page load that is otherwise acceptable becomes a wellbeing problem when it concentrates overnight.
- Time to resolution, and the proportion of pages that resolve without human intervention once acknowledged.
Where these signals show an oncall rotation is unhealthy, fixing it SHOULD be prioritized over product work. An engineering organization that lets an unhealthy oncall rotation persist while continuing to ship features is trading a visible, short-term output for an invisible, compounding cost: alert fatigue causes real alerts to be missed, and burnout drives the engineers who best understand the system to leave. Remediation includes tuning alert thresholds to cut non-actionable pages, fixing the underlying causes of recurring pages rather than re-paging on them indefinitely, and adjusting rotation size or shift length where the volume of a single rotation is unsustainable.
Metrics
Metrics MUST be gathered to inform and evaluate:
- Success criteria for changes made.
- Future development.
- Stability and reliability monitoring.
- Customer usage / activity patterns.
Key metrics – eg. API response times, batch processing times, speed of transactions etc. – MUST be tracked and monitored.
Customer-facing applications MUST have front-end analytics integrated, to provide insights into how customers use the application.
Observability strategy
Observability is the degree to which a system’s internal state can be inferred from its external outputs. Logging, monitoring, alerting, and metrics are the mechanisms; observability is the outcome — the ability to ask arbitrary questions of a running system without having to instrument it anew for each question.
A system with extensive telemetry but poor structure is not observable. The value of observability comes from treating it as a product: designed for the engineers who use it, consistent across systems, and exercised before it is needed.
Treat observability as a product
Observability tooling is used under pressure — during incidents, when the cost of confusion is highest. It SHOULD be designed with the same care as a user-facing product:
- Usability. Dashboards and logs should be easy to find, read, and navigate. An engineer who has never seen a particular dashboard SHOULD be able to orient themselves within seconds.
- Consistency. The structure of dashboards and logs SHOULD be consistent across all systems, so familiarity with one transfers to all others.
- Ownership. Observability is an engineering-team-wide responsibility, not the specialty of a few infrastructure engineers. Every team that owns a service owns its dashboards, logs, and alerts.
Ownership requires self-service tooling to be meaningful. Where changing a dashboard, adding a metric, or configuring an alert requires a ticket to a centralized observability or platform team, ownership sits with that team in practice, whatever the stated policy says — and a request routed through a queue turns a same-day investigation into a week-long wait. Development teams MUST be able to configure their own dashboards, alerts, and metrics directly, without going through a centralized team. A centralized team MAY still own the underlying platform — provisioning, retention policy, cost control — but MUST NOT be the sole route through which a development team’s own observability configuration changes.
Structure dashboards as a hierarchy
Dashboards SHOULD be organized as a drill-down hierarchy that lets an engineer move from "something is wrong" to "this specific request failed" in a few clicks. A RECOMMENDED hierarchy has four levels:
- Overview dashboard — A single, bird’s-eye view of every subsystem, acting as a traffic light. Each subsystem’s health is summarized by a small number of high-level signals that go red when anything is wrong. This is the first stop when an alert fires: it points to which subsystem is unhappy. It SHOULD be the default landing page for the dashboard tool and SHOULD be linked from team channels.
- System dashboards — One per subsystem, giving an exhaustive picture of that subsystem’s health. A system dashboard SHOULD answer: where are requests being rejected, where are the bottlenecks, what are the outcomes, and how much strain is the system under? Each system dashboard SHOULD link out to the logs for that subsystem.
- Logs — The individual events. See Logging for log design.
- Traces — The most zoomed-in view, showing where a single request spent its time. Traces are the tool of last resort for slow requests.
Each level SHOULD link to the next, so an engineer moves down the hierarchy without leaving the dashboard. Moving back up — from a trace to the broader context — SHOULD be equally easy.
Dashboard design
- Optimize for a glance. An engineer SHOULD be able to tell whether a dashboard is relevant to their investigation at a glance. Each dashboard SHOULD have a clear focus and direct the reader to where they need to go next.
- Avoid false negatives. An overview dashboard SHOULD go red when anything is wrong, no matter how minor. It is better to investigate a false alarm than to miss a real problem because the dashboard looked green.
- Use a consistent design system. Dashboards created by different teams SHOULD follow the same layout, color, and panel conventions, so an engineer comfortable with one dashboard can transfer that familiarity to any other.
- Split metrics by outcome. Where a request can have several outcomes (eg.
success,rate_limited,error), metrics SHOULD be split by outcome so the dashboard shows not just how many requests there were, but how they were handled.
Event logs
Each significant unit of work SHOULD emit a single, consistently formatted event log on completion, regardless of outcome (success, error, or panic). An event log records the event name, duration, outcome, and the IDs of the resources involved — the high-cardinality data that cannot be captured in metrics.
Event logs bridge the gap between metrics and logs: a metric tells you that something is wrong; the event log gives you a specific example of what went wrong. Event logs SHOULD share labels with the corresponding metrics so an engineer can move from a chart to a log line without changing query.
Note
Emit the event log in a defer block (or equivalent) so it is always written, even when the operation panics or returns early.
Exemplars
Where a metric is tracked, an exemplar — a reference to a specific request that contributed to that metric value — MAY be attached. Exemplars let an engineer jump from a point on a chart directly to the trace for the request that produced it, collapsing the metric-to-trace step into a single click.
Tracing
Traces show where a single request spent its time. They are the tool of last resort for investigating slow requests. To make traces useful:
- Trace third-party interactions. Every call to an external API SHOULD be traced, so a slow request is immediately attributable to an upstream dependency. Use a shared base client that applies tracing as standard.
- Trace database queries. Each query SHOULD be traced with its duration and, where practical, the query itself, so an engineer can see both the slow query and an example of it. Time spent waiting for a connection SHOULD be tracked separately from time spent executing the query, to distinguish contention from inefficiency.
Debugging as a discipline
Observability supplies the raw material — logs, traces, metrics, dumps — but turning that material into a root cause is a separate skill: forming a hypothesis, finding the observation that would confirm or refute it, and repeating until the cause is isolated. Treat debugging as a discipline to practice, not an ability engineers either have or lack:
- Reason across layers. The hardest bugs span multiple abstraction layers — application code, runtime, kernel, network, hardware — and root-causing them requires moving between layers rather than staying at the one where the symptom appeared. An engineer with a strong mental model of the full stack can sometimes single-shot a bug from one observation, by reasoning through what system state would produce it; an engineer without that model has to fall back to trial and error.
- Form a hypothesis before gathering more data. Undirected log-scrolling is slow. State what you expect to be true if the hypothesis holds, find the cheapest observation that would confirm or refute it, and only then look.
- Prefer the observation that eliminates the most hypotheses. Where several observations would each move the investigation forward, the one that splits the remaining hypothesis space roughly in half is worth more than one that only confirms the leading theory.
When to stop deducing and start observing
Deductive root-causing does not scale to every system. For a distributed system, a codebase with years of undocumented accretion (a "big ball of mud"), or heterogeneous client-side code running on devices you do not control, building a mental model precise enough to reason from is often impractical — there are too many interacting components, too much unobserved state, or too little control over the runtime environment.
Where deductive understanding is impractical, switch to an empirical strategy instead:
- Treat behavior statistically. Instead of explaining any one failure, characterize the failure rate, its distribution across inputs or instances, and how it moves in response to changes. A statistical answer ("errors correlate with payload size above 1 MB") is often actionable even where a full causal explanation is not.
- Invest in fault tolerance over root-causing every failure. In a distributed system, some component failures are not worth explaining individually — retries, circuit breakers, and graceful degradation absorb them more cheaply than root-causing each one. Root-cause the failures that recur or that the fault-tolerance layer does not absorb; let the fault-tolerance layer handle the rest. See TS-6 for the resilience patterns this relies on.
This is a judgment call, not a rule with a fixed threshold: attempt deductive root-causing first, and switch to empirical methods once the cost of building an accurate mental model clearly exceeds the cost of living with a statistical answer.
Exercise the setup
An observability setup that has never been used under pressure is unproven. Game days — scheduled exercises in which the team deliberately breaks something and uses the dashboards to diagnose it — SHOULD be run regularly. Game days both validate that the dashboards work and build the team’s familiarity with them, so that real incidents are not the first time anyone opens the tooling.
Treat game days like user testing: give participants minimal context, watch what they find easy and what they struggle with, and adjust the dashboards together afterward.
References
- Android Open Source Project. Java Code Style for Contributors.
- Elhage, N (2020). Computers Can Be Understood.
- Kosmulski, M (2024). Ten Years of Microservices at Allegro.
- Orosz, G (2021). The Pragmatic Engineer Test.
- Wiggins, A (2017). The Twelve-Factor App: XI. Logs.