Observability field guide Software & DevOps · incident prep · vendor-aware
Research mode Global technical content

Observability Field Guide

A vendor-neutral reference for deciding what to collect, how to name it, when to page on it, and where OpenTelemetry, Prometheus, Datadog, and New Relic fit without becoming the architecture.

logs metrics traces SLOs OpenTelemetry
View
Theme control
Definition of working Instrument a service, choose the right signal for the question, and page on an SLO instead of noise.
Primary search hook Logs vs metrics vs traces, SLOs, alert fatigue, OpenTelemetry, Datadog, New Relic.
Signature artifact The symptom-to-cause map above the fold, sized to screenshot cleanly at 1200×630.
Scope Service and platform observability. No vendor pricing, no product endorsement, no cloud lock-in sermon.
Signature map

Symptom to cause

Start with the thing the user feels, then move through the signals that narrow the search. The page is built around this flow because that is how incidents actually get solved.

Symptom

What the user or on-call sees first.

  • "Checkout is spinning"
  • "API is timing out"
  • Black-box checks fail

Metrics

Trend, rate, and saturation at low cost.

  • Request rate, error rate, latency
  • Queue depth, CPU, memory, disk
  • Alert on symptoms, not just causes

Traces

Follow the request path across service boundaries.

  • Root span to child spans
  • Async, DB, and external calls
  • Context propagation is the glue

Logs

Exact facts, stack traces, and event payloads.

  • Structured JSON, not opaque strings
  • Include trace_id and span_id
  • Great for one-off debugging

SLO

The reliability contract that decides paging.

  • SLI + target + window
  • Burn-rate alert when budget burns fast
  • Pages should tell you what broke for users

Runbook

The next move once the signal points at a cause.

  • Mitigate, verify, then write up
  • Use dashboards for diagnosis, not decoration
  • Keep the fix path close to the alert
Quick reference

What to reach for first

Signal Use it when Do not use it when Gotcha
Logs You need the exact event payload, stack trace, or audit trail. You are trying to trend latency or page on aggregate health. Unstructured strings are hard to query; make them JSON with stable keys.
Metrics You need a cheap time series for rates, errors, latency, or saturation. You need the full story of one request or a precise per-user payload. Every label combination is a new time series; cardinality is a budget.
Traces You need the causal path across services, queues, or async hops. You want a permanent archive of every request without sampling. Broken context propagation turns a trace into disconnected spans.
Profiles You need to find hot code paths, CPU burn, or memory pressure. You need business semantics or exact request sequencing. Profiles are usually sampled, so they are best for hotspots, not proof.
Events You need discrete state changes: deploys, failovers, feature flag flips. You want a continuously aggregated trend. Events need shared IDs and timestamps to line up with the other signals.
Synthetics You want a scripted, repeatable check of a user journey or endpoint. You need to debug a server-side code path directly. Synthetics only prove what the scripted path can see.
RUM You need frontend latency, browser errors, or user experience. You need backend-only internals. Frontend sessions are useful only if you can tie them back to backend IDs.
SLOs You need to decide when to page and when to hold the line. You are still arguing about whether the service is worth protecting. Page on symptom burn; route cause details to a dashboard or ticket.
Fundamentals

What observability is, and what it is not

OpenTelemetry defines observability as the ability to understand behavior from outside the system by using correlated telemetry. The practical rule is simpler: use logs for detail, metrics for trend, traces for causality, and SLOs for the paging contract.

Observability

A way to ask new questions about a system without first changing code for every question. Example: "Which deploy made checkout slower?" Gotcha: dashboards alone do not create observability if the underlying signals are uncorrelated.

White-box monitoring

Monitoring based on internals like counters, logs, and runtime stats. Example: a latency histogram and a queue-depth metric. Gotcha: if you only watch black-box uptime, you miss which dependency degraded first.

Black-box monitoring

External checks that answer "is the service usable?" Example: a synthetic login probe. Gotcha: external success does not tell you where the failure is, only that the symptom exists.

Signal taxonomy

Signal Definition Example Primary gotcha
Logs Timestamped records of discrete events, ideally structured. {"level":"error","service":"checkout","trace_id":"4bf92f..."} They get expensive and noisy fast if you log everything at info level.
Metrics Measurements captured over time and aggregated into time series. http_requests_total, http_request_duration_seconds_bucket Averages hide the long tail; cardinality can blow up storage and query cost.
Traces A causality graph showing the path of one request or workflow. Gateway span -> DB span -> downstream API span Sampling too aggressively can erase the rare failure you needed to keep.
Profiles Samples of CPU, wall time, or memory usage that show hot paths. 60-second CPU profile showing JSON encoding at the top Great for optimization, weak for business context.
Events State transitions or discrete occurrences with a timestamp. Deploy started, deploy finished, feature flag toggled Events without stable service and version tags are hard to correlate.
Synthetics Scripted outside-in checks that mimic a user path. POST /login and then load /account They only test what the script covers, not the whole production surface.
RUM Browser-side telemetry about real user experience. Largest Contentful Paint, JavaScript errors, client-side latency Frontend data is not a substitute for backend tracing.
SLOs A target for a service-level indicator over a time window. 99.9% successful checkout requests over 30 days Without an error budget, the target is just a slogan.
Decision rule: choose the signal with the shortest path to truth. If the question is "what happened?" start with logs; if it is "is it getting worse?" start with metrics; if it is "where did the time go?" start with traces; if it is "should we page?" use the SLO.
Working knowledge

Instrumentation checklist

The best observability data is boringly consistent: stable service identity, a shared set of resource fields, and enough context to join logs, metrics, and traces without inventing ad hoc keys in every service.

Identity and correlation fields

Field Definition Example Gotcha
service.nameStable logical service name used across telemetry.checkout-apiDo not use pod names or hostnames; they change too often.
deployment.environmentEnvironment label such as dev, staging, prod.prodKeep the vocabulary small or dashboards splinter.
service.versionThe build or release version of the service.2026.07.07-4f2c1d1Without version, deploy markers cannot explain regressions.
cloud.regionWhere the workload runs.us-west-2Region is a common culprit in partial outages, so make it explicit.
trace_idTrace-wide identifier that lets logs and spans join.4bf92f3577b34da6a3ce929d0e0e4736If logs drop it, correlation breaks at the exact moment you need it.
span_idIdentifier for one unit of work inside a trace.00f067aa0ba902b7Use it to pin one log line to one span, not as a business identifier.
request_id / correlation_idRequest-scoped identifier used by your app and gateway.req_01J0ZQ6A1B8P2Prefer one canonical ID per request, not three competing ones.
severityLog urgency such as debug/info/warn/error.errorDo not make every failed validation an error if users can recover.
error.type / codeStructured error class or status code.SQLTimeoutError, HTTP 503Human text alone is hard to group or alert on.
resource.attrsShared resource context attached everywhere.host.name, k8s.pod.nameAttach through the Collector so logs, traces, and metrics match exactly.
deploy.markerEvent showing when a version rolled out.checkout-api 2026.07.07-4f2c1d1 deployedNo marker means you will guess at causality from a chart gap.
user/sessionAnonymous identifier for a user or session.user_7f9a...Hash or pseudonymize; do not put raw email or account IDs in broad labels.

Context and payload discipline

Structured logs

Emit JSON fields that machines can query. Example: {"msg":"payment failed","order_id":"ord_1287","trace_id":"..."}. Gotcha: plain text makes correlation and dashboards brittle.

Histogram buckets

Choose buckets that reflect user experience and SLO thresholds. Example: latency buckets around 50 ms, 100 ms, 250 ms, 500 ms, 1 s, 2 s. Gotcha: buckets that are too coarse hide the tail; buckets that are too dense increase cost.

Exemplars

Attach a trace sample to a metric point so the graph can jump to an example request. Example: a latency spike annotated with a trace ID. Gotcha: exemplars are most useful when trace sampling still keeps enough rare paths.

Collector discipline: OpenTelemetry Collector pipelines can enrich and scrub telemetry before export. Use that place to normalize resource attributes and remove personal data, not a dozen different app-specific hacks.
Metrics

Metrics that matter

Metrics are for trend and alerting. The right metric answers "is the system healthy enough right now?" and the right graph tells you whether the answer is stable, worsening, or already out of budget.

Service and resource patterns

Pattern Definition Example Anti-pattern
RED Rate, errors, duration for request-driven services. http_requests_total, http_request_errors_total, http_request_duration_seconds Monitoring CPU first and user latency second.
USE Utilization, saturation, errors for resources. CPU busy, disk queue depth, memory pressure Assuming utilization alone predicts customer pain.
Golden Signals Latency, traffic, errors, saturation for service health. p95 latency + RPS + 5xx rate + queue depth Building dashboards with no owner or decision attached.
Business metrics User-visible or revenue-visible outcomes. Checkout completion rate, signups, order failures Confusing a healthy infra chart with a healthy product.
Backlog and queue depth How much asynchronous work is waiting. queue_depth or oldest message age Queues can look fine until age starts increasing faster than drain.
Availability ratio Good events divided by total events. 99.9% successful requests over 30 days Counting only 2xx responses if the user still sees a broken page.
Latency percentiles Tail measurements of user experience. p50, p95, p99 from a histogram Averages hide a painful tail and make regressions look smaller than they are.
Histogram Bucketed distribution that supports percentiles and SLO math. histogram_quantile(0.95, sum by (le) (...)) Choose bucket boundaries with the SLO in mind or the graph misleads you.
Counter rate Monotonic count turned into a rate with a window. rate(http_requests_total[5m]) Do not store a percentage when a pair of counters would be safer and more composable.

What Prometheus-style guidance is really saying

Keep labels small

Prometheus warns that every unique label set is a new time series. Example: `user_id` as a label is a bad idea because the cardinality is unbounded. Gotcha: a "helpful" label can become the most expensive field in the system.

Use base units

Store seconds, bytes, and counts, then convert in the query or chart. Example: `request_duration_seconds`. Gotcha: encoding milliseconds in the metric name locks you into one unit and makes cross-service comparisons harder.

Prefer counters for rates

Monotonic counters are easy to roll up into rates and burn calculations. Example: `errors_total` over a 5-minute window. Gotcha: a gauge of "requests this minute" is less robust than a counter plus `rate()`.

Traces

Tracing and correlation

Traces show causality: one request, one path, many spans. The point is not to collect more spans; it is to keep enough structure that a weird latency spike can be followed across services, databases, queues, and external APIs.

Span design

Item Definition Example Gotcha
Root spanThe top-level span for the request or job.GET /checkoutIf the root is missing, the whole trace loses shape.
Child spanA nested unit of work inside the parent path.SELECT order by idDo not create a span for every tiny helper call.
Async boundaryA hop across queues, tasks, or callbacks where context must be propagated.Publish event -> consume eventIf context drops here, tracing stops exactly where incidents get interesting.
DB spanA span around a datastore operation.mysql.query SELECT * FROM ordersInclude the operation and resource, but avoid dumping full query text if it exposes secrets.
External call spanA span for HTTP, RPC, or third-party dependency work.POST https://api.stripe.com/v1/payment_intentsKeep service and peer attributes consistent or service maps drift.
BaggageCross-process key/value context carried alongside the trace.tenant_id=acmeBaggage is for propagation, not a dumping ground for every app attribute.
Span attributesStructured metadata about one span.http.status_code=503Use attributes for facts about the span, not a second log record.
Events in spansTimestamped annotations inside one span.exception, retry, cache_missDo not rely on events if the parent context never arrives.
ExemplarsA trace reference attached to a metric sample.Latency bucket with one trace IDWithout enough trace samples, the jump from graph to trace is hit-or-miss.

Sampling and propagation

Head sampling

Decide at trace start whether to keep it. Example: keep 10% of requests. Gotcha: head sampling can miss a rare error if the signal only appears later in the request.

Tail sampling

Decide after the trace is complete, often based on latency or errors. Example: retain slow traces and failed traces. Gotcha: it costs more to operate, but it is often better for incident diagnosis.

Context propagation

The mechanism that keeps the trace connected across service boundaries. Example: inject and extract trace context on every HTTP hop. Gotcha: one missing propagation point can make logs, metrics, and traces look unrelated when they are not.

OpenTelemetry Collector note: the Collector is vendor-agnostic and can receive, process, and export traces, metrics, and logs. That makes it the right place to normalize metadata, batch, sample, and transform before data reaches a backend.
SLOs and alerting

SLOs turn noise into a budget

Google’s SRE guidance is clear: the error budget is the difference between perfect and the SLO. That budget should steer alerting, release pace, and where the team spends its attention.

Core terms

Term Definition Example Gotcha
SLI A measured signal that reflects user experience. Successful checkout requests / total checkout requests If the SLI is not user-facing, the SLO can reward the wrong thing.
SLO The target value for an SLI over a window. 99.9% success over 30 days A target without an error budget is not operationally actionable.
Error budget How much unreliability the service can spend before the SLO fails. At 99.9%, budget is 0.1% of requests over the window. If the team never measures budget burn, it cannot tell whether it is spending too fast.
Burn rate How fast the error budget is being consumed. Burn rate 1 means spending budget exactly as fast as the window allows. Fast-burn and slow-burn alerts serve different operational purposes.
Symptom page An alert that tells you users are hurting now. 5xx rate or latency SLO burn Pages that require no action become background noise.
Cause ticket An alert or report that feeds diagnosis, not paging. Disk nearing capacity, replica lag, noisy neighbors If you page on every cause, the page becomes untrustworthy.
Maintenance window A period when expected unreliability should not count against the budget. Planned deploy or migration window Do not silently hide real outages inside broad windows.
Runbook The next steps once the alert fires. Link to rollback, feature flag, and owner info Pages without runbooks create slow, random response.

How to alert on SLOs

1

Measure the SLI

Use a numerator and denominator that reflect user success, not an internal proxy.

  • Example: successful requests / total requests
  • Gotcha: internal-only health checks can miss user pain
2

Set the SLO

Pick a window and target that match how much risk the service can absorb.

  • Example: 99.9% over 30 days
  • Gotcha: too-tight targets create noise; too-loose targets hide drift
3

Watch burn rate

Alert when the budget is burning fast enough that action is needed now.

  • Example: multi-window burn-rate logic
  • Gotcha: one window catches fast outages; another catches smaller but serious leaks
4

Separate page and ticket

Pages should be actionable in minutes; tickets can wait for the next work cycle.

  • Page on symptoms
  • Ticket the likely cause
5

Attach the runbook

Every page should tell the responder what to check first.

  • Rollback? flag? capacity? dependency?
  • Gotcha: no runbook means slower recovery and more paging fatigue
Concrete math: a 99.9% SLO over 30 days means a 0.1% error budget. If the service handles 3,000,000 requests in that window, the budget is 3,000 bad requests. A burn rate of 1 means you are spending exactly at the allowed pace.
Prometheus alerting rule of thumb: keep alerting simple, alert on symptoms, and let dashboards explain causes. Alertmanager handles grouping, silencing, inhibition, and notification routing; the Prometheus server decides whether the alert is firing.
Vendor layer

Where the tools fit

The trap is to let the platform vendor define the architecture. The safer pattern is to standardize on OpenTelemetry and Prometheus-style metrics, then treat Datadog or New Relic as backends and workflow surfaces.

Tooling layer What it is Best use Lock-in risk / gotcha
OpenTelemetry Collector Vendor-agnostic receive/process/export pipeline for telemetry. Normalize context, scrub data, batch, sample, and route to one or more backends. It is not the system of record; it is the transport and transformation layer.
Prometheus + Alertmanager Time-series model plus alert routing and notification handling. Metrics, recording rules, and symptom-based alerting at low operational cost. High-cardinality labels and bad histograms are the fastest way to make it hurt.
Datadog Hosted observability platform with logs, traces, metrics, SLOs, and OpenTelemetry ingestion. Fast setup, cross-telemetry correlation, and operational dashboards for teams that want a managed surface. Correlate with DD_ENV, DD_SERVICE, and DD_VERSION; otherwise the data stays fragmented.
New Relic Hosted observability platform with service levels and OpenTelemetry support. Service-level management, entity views, and an operations workflow centered on SLIs and SLOs. Do not confuse the UI features with the underlying signal design; the SLI still needs to be user-centered.
Cloud-native monitors Provider-specific metrics, logs, and alerting in AWS, Azure, or GCP. Baseline infra monitoring and control-plane visibility. Useful for platform hygiene, but rarely sufficient alone for service-level alerting.

Datadog fit

Datadog docs currently describe logs, traces, metrics, OpenTelemetry export, and SLO/burn-rate workflows. Use it when you want a managed platform with tight cross-linking between telemetry types.

New Relic fit

New Relic docs emphasize service levels, entity-centered SLO creation, and OpenTelemetry ingestion. Use it when the workflow should start from service health rather than raw infrastructure graphs.

OpenTelemetry fit

OTel is the standardization layer: semantic conventions, Collector pipelines, and shared context. Use it to keep backend choice flexible and observability data interoperable.

Operations

Incident workflow

The fastest teams do not inspect every dashboard. They move through a short sequence: detect the symptom, localize the fault domain, mitigate the user impact, verify the fix, and capture the lesson.

1

Detect

Start from the SLO burn, synthetic failure, or user report that proves a user-visible problem exists.

  • Question: who is affected and how bad is it?
2

Localize

Use metrics to spot the shape, traces to follow the path, and logs to inspect the exact error.

  • Question: which dependency or hop is the first clear break?
3

Mitigate

Roll back, disable the feature, drain the queue, or add capacity if that restores user success fastest.

  • Question: what is the smallest safe action?
4

Verify

Confirm the SLO, synthetic, and key user path recover before you declare victory.

  • Question: is the symptom actually gone?
5

Learn

Write the postmortem, preserve the evidence, and update the runbook and alert design.

  • Question: what should never require this much manual effort again?
Trace-to-log-to-metric drill: start from one slow trace, jump to the span’s logs, confirm the error class or request payload, then compare the surrounding metric window to see whether the problem is a single failure or a broad regression.
Edge and advanced

Common mistakes and anti-patterns

Most observability failures are self-inflicted. The tooling is usually fine; the data model or alert design is not.

Anti-pattern Why it hurts Concrete fix
Logging secrets or raw personal dataCreates security and privacy risk, and turns logs into liability.Redact at source or in the Collector; log stable identifiers instead.
Unbounded labels or tagsExplodes cardinality, storage, and query latency.Keep dimensions small and safe; never use user IDs or emails as labels.
Alerting on CPU aloneCPU can be high while users are fine, or low while users are failing.Page on symptoms: error rate, latency, and SLO burn.
Averaging latencyAverage latency hides the tail that users notice.Use histograms and p95/p99 percentiles.
No deploy markersYou cannot tell whether the regression started with a release.Emit deploy events and annotate dashboards.
Sampling away rare errorsThe exact incident you need may be the one that gets dropped.Use tail sampling or error-biased retention.
Dashboards with no ownerCharts accumulate and stop reflecting operational decisions.Assign ownership and tie each dashboard to a question.
Cause pages disguised as symptom pagesPages arrive before anyone knows there is user pain.Keep causes in tickets or dashboards; reserve pages for user impact.
Signals with no shared contextLogs, metrics, and traces cannot be joined.Standardize service, version, environment, and trace IDs everywhere.
One backend to rule everythingVendor lock-in becomes the design constraint.Push standards into the app and Collector; treat the backend as replaceable.
Cardinality budget rule: if you cannot explain why a label is bounded, low-cardinality, and useful for aggregation, do not add it. It will cost more than you think and help less than you want.