Symptom
What the user or on-call sees first.
- "Checkout is spinning"
- "API is timing out"
- Black-box checks fail
A vendor-neutral reference for deciding what to collect, how to name it, when to page on it, and where OpenTelemetry, Prometheus, Datadog, and New Relic fit without becoming the architecture.
Start with the thing the user feels, then move through the signals that narrow the search. The page is built around this flow because that is how incidents actually get solved.
What the user or on-call sees first.
Trend, rate, and saturation at low cost.
Follow the request path across service boundaries.
Exact facts, stack traces, and event payloads.
The reliability contract that decides paging.
The next move once the signal points at a cause.
| Signal | Use it when | Do not use it when | Gotcha |
|---|---|---|---|
| Logs | You need the exact event payload, stack trace, or audit trail. | You are trying to trend latency or page on aggregate health. | Unstructured strings are hard to query; make them JSON with stable keys. |
| Metrics | You need a cheap time series for rates, errors, latency, or saturation. | You need the full story of one request or a precise per-user payload. | Every label combination is a new time series; cardinality is a budget. |
| Traces | You need the causal path across services, queues, or async hops. | You want a permanent archive of every request without sampling. | Broken context propagation turns a trace into disconnected spans. |
| Profiles | You need to find hot code paths, CPU burn, or memory pressure. | You need business semantics or exact request sequencing. | Profiles are usually sampled, so they are best for hotspots, not proof. |
| Events | You need discrete state changes: deploys, failovers, feature flag flips. | You want a continuously aggregated trend. | Events need shared IDs and timestamps to line up with the other signals. |
| Synthetics | You want a scripted, repeatable check of a user journey or endpoint. | You need to debug a server-side code path directly. | Synthetics only prove what the scripted path can see. |
| RUM | You need frontend latency, browser errors, or user experience. | You need backend-only internals. | Frontend sessions are useful only if you can tie them back to backend IDs. |
| SLOs | You need to decide when to page and when to hold the line. | You are still arguing about whether the service is worth protecting. | Page on symptom burn; route cause details to a dashboard or ticket. |
OpenTelemetry defines observability as the ability to understand behavior from outside the system by using correlated telemetry. The practical rule is simpler: use logs for detail, metrics for trend, traces for causality, and SLOs for the paging contract.
A way to ask new questions about a system without first changing code for every question. Example: "Which deploy made checkout slower?" Gotcha: dashboards alone do not create observability if the underlying signals are uncorrelated.
Monitoring based on internals like counters, logs, and runtime stats. Example: a latency histogram and a queue-depth metric. Gotcha: if you only watch black-box uptime, you miss which dependency degraded first.
External checks that answer "is the service usable?" Example: a synthetic login probe. Gotcha: external success does not tell you where the failure is, only that the symptom exists.
| Signal | Definition | Example | Primary gotcha |
|---|---|---|---|
| Logs | Timestamped records of discrete events, ideally structured. | {"level":"error","service":"checkout","trace_id":"4bf92f..."} |
They get expensive and noisy fast if you log everything at info level. |
| Metrics | Measurements captured over time and aggregated into time series. | http_requests_total, http_request_duration_seconds_bucket |
Averages hide the long tail; cardinality can blow up storage and query cost. |
| Traces | A causality graph showing the path of one request or workflow. | Gateway span -> DB span -> downstream API span | Sampling too aggressively can erase the rare failure you needed to keep. |
| Profiles | Samples of CPU, wall time, or memory usage that show hot paths. | 60-second CPU profile showing JSON encoding at the top | Great for optimization, weak for business context. |
| Events | State transitions or discrete occurrences with a timestamp. | Deploy started, deploy finished, feature flag toggled | Events without stable service and version tags are hard to correlate. |
| Synthetics | Scripted outside-in checks that mimic a user path. | POST /login and then load /account | They only test what the script covers, not the whole production surface. |
| RUM | Browser-side telemetry about real user experience. | Largest Contentful Paint, JavaScript errors, client-side latency | Frontend data is not a substitute for backend tracing. |
| SLOs | A target for a service-level indicator over a time window. | 99.9% successful checkout requests over 30 days | Without an error budget, the target is just a slogan. |
The best observability data is boringly consistent: stable service identity, a shared set of resource fields, and enough context to join logs, metrics, and traces without inventing ad hoc keys in every service.
| Field | Definition | Example | Gotcha |
|---|---|---|---|
service.name | Stable logical service name used across telemetry. | checkout-api | Do not use pod names or hostnames; they change too often. |
deployment.environment | Environment label such as dev, staging, prod. | prod | Keep the vocabulary small or dashboards splinter. |
service.version | The build or release version of the service. | 2026.07.07-4f2c1d1 | Without version, deploy markers cannot explain regressions. |
cloud.region | Where the workload runs. | us-west-2 | Region is a common culprit in partial outages, so make it explicit. |
trace_id | Trace-wide identifier that lets logs and spans join. | 4bf92f3577b34da6a3ce929d0e0e4736 | If logs drop it, correlation breaks at the exact moment you need it. |
span_id | Identifier for one unit of work inside a trace. | 00f067aa0ba902b7 | Use it to pin one log line to one span, not as a business identifier. |
request_id / correlation_id | Request-scoped identifier used by your app and gateway. | req_01J0ZQ6A1B8P2 | Prefer one canonical ID per request, not three competing ones. |
severity | Log urgency such as debug/info/warn/error. | error | Do not make every failed validation an error if users can recover. |
error.type / code | Structured error class or status code. | SQLTimeoutError, HTTP 503 | Human text alone is hard to group or alert on. |
resource.attrs | Shared resource context attached everywhere. | host.name, k8s.pod.name | Attach through the Collector so logs, traces, and metrics match exactly. |
deploy.marker | Event showing when a version rolled out. | checkout-api 2026.07.07-4f2c1d1 deployed | No marker means you will guess at causality from a chart gap. |
user/session | Anonymous identifier for a user or session. | user_7f9a... | Hash or pseudonymize; do not put raw email or account IDs in broad labels. |
Emit JSON fields that machines can query. Example: {"msg":"payment failed","order_id":"ord_1287","trace_id":"..."}. Gotcha: plain text makes correlation and dashboards brittle.
Choose buckets that reflect user experience and SLO thresholds. Example: latency buckets around 50 ms, 100 ms, 250 ms, 500 ms, 1 s, 2 s. Gotcha: buckets that are too coarse hide the tail; buckets that are too dense increase cost.
Attach a trace sample to a metric point so the graph can jump to an example request. Example: a latency spike annotated with a trace ID. Gotcha: exemplars are most useful when trace sampling still keeps enough rare paths.
Metrics are for trend and alerting. The right metric answers "is the system healthy enough right now?" and the right graph tells you whether the answer is stable, worsening, or already out of budget.
| Pattern | Definition | Example | Anti-pattern |
|---|---|---|---|
| RED | Rate, errors, duration for request-driven services. | http_requests_total, http_request_errors_total, http_request_duration_seconds |
Monitoring CPU first and user latency second. |
| USE | Utilization, saturation, errors for resources. | CPU busy, disk queue depth, memory pressure | Assuming utilization alone predicts customer pain. |
| Golden Signals | Latency, traffic, errors, saturation for service health. | p95 latency + RPS + 5xx rate + queue depth | Building dashboards with no owner or decision attached. |
| Business metrics | User-visible or revenue-visible outcomes. | Checkout completion rate, signups, order failures | Confusing a healthy infra chart with a healthy product. |
| Backlog and queue depth | How much asynchronous work is waiting. | queue_depth or oldest message age |
Queues can look fine until age starts increasing faster than drain. |
| Availability ratio | Good events divided by total events. | 99.9% successful requests over 30 days | Counting only 2xx responses if the user still sees a broken page. |
| Latency percentiles | Tail measurements of user experience. | p50, p95, p99 from a histogram | Averages hide a painful tail and make regressions look smaller than they are. |
| Histogram | Bucketed distribution that supports percentiles and SLO math. | histogram_quantile(0.95, sum by (le) (...)) |
Choose bucket boundaries with the SLO in mind or the graph misleads you. |
| Counter rate | Monotonic count turned into a rate with a window. | rate(http_requests_total[5m]) |
Do not store a percentage when a pair of counters would be safer and more composable. |
Prometheus warns that every unique label set is a new time series. Example: `user_id` as a label is a bad idea because the cardinality is unbounded. Gotcha: a "helpful" label can become the most expensive field in the system.
Store seconds, bytes, and counts, then convert in the query or chart. Example: `request_duration_seconds`. Gotcha: encoding milliseconds in the metric name locks you into one unit and makes cross-service comparisons harder.
Monotonic counters are easy to roll up into rates and burn calculations. Example: `errors_total` over a 5-minute window. Gotcha: a gauge of "requests this minute" is less robust than a counter plus `rate()`.
Traces show causality: one request, one path, many spans. The point is not to collect more spans; it is to keep enough structure that a weird latency spike can be followed across services, databases, queues, and external APIs.
| Item | Definition | Example | Gotcha |
|---|---|---|---|
| Root span | The top-level span for the request or job. | GET /checkout | If the root is missing, the whole trace loses shape. |
| Child span | A nested unit of work inside the parent path. | SELECT order by id | Do not create a span for every tiny helper call. |
| Async boundary | A hop across queues, tasks, or callbacks where context must be propagated. | Publish event -> consume event | If context drops here, tracing stops exactly where incidents get interesting. |
| DB span | A span around a datastore operation. | mysql.query SELECT * FROM orders | Include the operation and resource, but avoid dumping full query text if it exposes secrets. |
| External call span | A span for HTTP, RPC, or third-party dependency work. | POST https://api.stripe.com/v1/payment_intents | Keep service and peer attributes consistent or service maps drift. |
| Baggage | Cross-process key/value context carried alongside the trace. | tenant_id=acme | Baggage is for propagation, not a dumping ground for every app attribute. |
| Span attributes | Structured metadata about one span. | http.status_code=503 | Use attributes for facts about the span, not a second log record. |
| Events in spans | Timestamped annotations inside one span. | exception, retry, cache_miss | Do not rely on events if the parent context never arrives. |
| Exemplars | A trace reference attached to a metric sample. | Latency bucket with one trace ID | Without enough trace samples, the jump from graph to trace is hit-or-miss. |
Decide at trace start whether to keep it. Example: keep 10% of requests. Gotcha: head sampling can miss a rare error if the signal only appears later in the request.
Decide after the trace is complete, often based on latency or errors. Example: retain slow traces and failed traces. Gotcha: it costs more to operate, but it is often better for incident diagnosis.
The mechanism that keeps the trace connected across service boundaries. Example: inject and extract trace context on every HTTP hop. Gotcha: one missing propagation point can make logs, metrics, and traces look unrelated when they are not.
Google’s SRE guidance is clear: the error budget is the difference between perfect and the SLO. That budget should steer alerting, release pace, and where the team spends its attention.
| Term | Definition | Example | Gotcha |
|---|---|---|---|
| SLI | A measured signal that reflects user experience. | Successful checkout requests / total checkout requests | If the SLI is not user-facing, the SLO can reward the wrong thing. |
| SLO | The target value for an SLI over a window. | 99.9% success over 30 days | A target without an error budget is not operationally actionable. |
| Error budget | How much unreliability the service can spend before the SLO fails. | At 99.9%, budget is 0.1% of requests over the window. | If the team never measures budget burn, it cannot tell whether it is spending too fast. |
| Burn rate | How fast the error budget is being consumed. | Burn rate 1 means spending budget exactly as fast as the window allows. | Fast-burn and slow-burn alerts serve different operational purposes. |
| Symptom page | An alert that tells you users are hurting now. | 5xx rate or latency SLO burn | Pages that require no action become background noise. |
| Cause ticket | An alert or report that feeds diagnosis, not paging. | Disk nearing capacity, replica lag, noisy neighbors | If you page on every cause, the page becomes untrustworthy. |
| Maintenance window | A period when expected unreliability should not count against the budget. | Planned deploy or migration window | Do not silently hide real outages inside broad windows. |
| Runbook | The next steps once the alert fires. | Link to rollback, feature flag, and owner info | Pages without runbooks create slow, random response. |
Use a numerator and denominator that reflect user success, not an internal proxy.
Pick a window and target that match how much risk the service can absorb.
Alert when the budget is burning fast enough that action is needed now.
Pages should be actionable in minutes; tickets can wait for the next work cycle.
Every page should tell the responder what to check first.
The trap is to let the platform vendor define the architecture. The safer pattern is to standardize on OpenTelemetry and Prometheus-style metrics, then treat Datadog or New Relic as backends and workflow surfaces.
| Tooling layer | What it is | Best use | Lock-in risk / gotcha |
|---|---|---|---|
| OpenTelemetry Collector | Vendor-agnostic receive/process/export pipeline for telemetry. | Normalize context, scrub data, batch, sample, and route to one or more backends. | It is not the system of record; it is the transport and transformation layer. |
| Prometheus + Alertmanager | Time-series model plus alert routing and notification handling. | Metrics, recording rules, and symptom-based alerting at low operational cost. | High-cardinality labels and bad histograms are the fastest way to make it hurt. |
| Datadog | Hosted observability platform with logs, traces, metrics, SLOs, and OpenTelemetry ingestion. | Fast setup, cross-telemetry correlation, and operational dashboards for teams that want a managed surface. | Correlate with DD_ENV, DD_SERVICE, and DD_VERSION; otherwise the data stays fragmented. |
| New Relic | Hosted observability platform with service levels and OpenTelemetry support. | Service-level management, entity views, and an operations workflow centered on SLIs and SLOs. | Do not confuse the UI features with the underlying signal design; the SLI still needs to be user-centered. |
| Cloud-native monitors | Provider-specific metrics, logs, and alerting in AWS, Azure, or GCP. | Baseline infra monitoring and control-plane visibility. | Useful for platform hygiene, but rarely sufficient alone for service-level alerting. |
Datadog docs currently describe logs, traces, metrics, OpenTelemetry export, and SLO/burn-rate workflows. Use it when you want a managed platform with tight cross-linking between telemetry types.
New Relic docs emphasize service levels, entity-centered SLO creation, and OpenTelemetry ingestion. Use it when the workflow should start from service health rather than raw infrastructure graphs.
OTel is the standardization layer: semantic conventions, Collector pipelines, and shared context. Use it to keep backend choice flexible and observability data interoperable.
The fastest teams do not inspect every dashboard. They move through a short sequence: detect the symptom, localize the fault domain, mitigate the user impact, verify the fix, and capture the lesson.
Start from the SLO burn, synthetic failure, or user report that proves a user-visible problem exists.
Use metrics to spot the shape, traces to follow the path, and logs to inspect the exact error.
Roll back, disable the feature, drain the queue, or add capacity if that restores user success fastest.
Confirm the SLO, synthetic, and key user path recover before you declare victory.
Write the postmortem, preserve the evidence, and update the runbook and alert design.
Most observability failures are self-inflicted. The tooling is usually fine; the data model or alert design is not.
| Anti-pattern | Why it hurts | Concrete fix |
|---|---|---|
| Logging secrets or raw personal data | Creates security and privacy risk, and turns logs into liability. | Redact at source or in the Collector; log stable identifiers instead. |
| Unbounded labels or tags | Explodes cardinality, storage, and query latency. | Keep dimensions small and safe; never use user IDs or emails as labels. |
| Alerting on CPU alone | CPU can be high while users are fine, or low while users are failing. | Page on symptoms: error rate, latency, and SLO burn. |
| Averaging latency | Average latency hides the tail that users notice. | Use histograms and p95/p99 percentiles. |
| No deploy markers | You cannot tell whether the regression started with a release. | Emit deploy events and annotate dashboards. |
| Sampling away rare errors | The exact incident you need may be the one that gets dropped. | Use tail sampling or error-biased retention. |
| Dashboards with no owner | Charts accumulate and stop reflecting operational decisions. | Assign ownership and tie each dashboard to a question. |
| Cause pages disguised as symptom pages | Pages arrive before anyone knows there is user pain. | Keep causes in tickets or dashboards; reserve pages for user impact. |
| Signals with no shared context | Logs, metrics, and traces cannot be joined. | Standardize service, version, environment, and trace IDs everywhere. |
| One backend to rule everything | Vendor lock-in becomes the design constraint. | Push standards into the app and Collector; treat the backend as replaceable. |