Community/Tech articles

Logs, metrics and traces: a practical introduction to observability with OpenTelemetry

What each signal is for, how they fit together through trace ids, and how to instrument a Python service with OpenTelemetry without drowning in data.

When something breaks in production, you need to answer three questions fast: is something wrong, where is it wrong, and why? Metrics, traces and logs each answer one of those best. Used together, linked by a shared trace id, they turn a two-hour hunt into a ten-minute fix.

The three signals

Signal Answers Example
Metrics Is something wrong? p95 latency of /checkout jumped from 180 ms to 2.4 s
Traces Where is the time going? 2.1 s of that is one database query in the payment step
Logs Why did it happen? lock wait timeout on table orders for request abc123

Metrics are cheap numeric time series: counters, gauges and histograms. Keep them aggregated and low-cardinality; they power dashboards and alerts.

Traces follow one request across services. A trace is a tree of spans; each span is a timed operation with attributes (HTTP route, database statement, user tier).

Logs are discrete events with detail. Structured logs (JSON with consistent fields) are far more useful than free text, especially when every log line carries the current trace id.

Why OpenTelemetry

OpenTelemetry (OTel) is a vendor-neutral standard and set of SDKs for producing all three signals. You instrument your code once and send the data to whatever backend you choose: Grafana's stack, Jaeger, Honeycomb, Datadog, or a cloud provider's tracing service. Switching backends becomes a configuration change rather than a rewrite.

Instrumenting a FastAPI service

Install the SDK, an exporter and the instrumentations for your libraries:

pip install opentelemetry-sdk opentelemetry-exporter-otlp \
  opentelemetry-instrumentation-fastapi opentelemetry-instrumentation-httpx

Set up tracing once at startup:

from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.instrumentation.httpx import HTTPXClientInstrumentor

provider = TracerProvider(resource=Resource.create({"service.name": "orders-api"}))
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))  # endpoint from env vars
trace.set_tracer_provider(provider)

FastAPIInstrumentor.instrument_app(app)   # a span for every request
HTTPXClientInstrumentor().instrument()    # spans for outgoing calls, context propagated

Auto-instrumentation gives you request and client spans for free. Add your own spans around the business steps that matter:

tracer = trace.get_tracer(__name__)

async def checkout(cart):
    with tracer.start_as_current_span("checkout.price_cart") as span:
        span.set_attribute("cart.items", len(cart.items))
        total = await price(cart)
    with tracer.start_as_current_span("checkout.charge"):
        return await charge(total)

Connect logs to traces

Put the trace id on every log line so you can jump from a slow trace to its logs and back:

import logging
from opentelemetry import trace

class TraceIdFilter(logging.Filter):
    def filter(self, record):
        ctx = trace.get_current_span().get_span_context()
        record.trace_id = format(ctx.trace_id, "032x") if ctx.is_valid else "-"
        return True

handler = logging.StreamHandler()
handler.addFilter(TraceIdFilter())
handler.setFormatter(logging.Formatter('{"level":"%(levelname)s","msg":"%(message)s","trace_id":"%(trace_id)s"}'))
logging.getLogger().addHandler(handler)

Metrics that matter: start with RED

For each service, track:

  • Rate: requests per second
  • Errors: failed requests per second (or as a percentage)
  • Duration: latency as a histogram, so you can read p50, p95 and p99

Alert on symptoms users feel (error rate, p95 latency against your target), not on every CPU spike.

Keep the cost under control

  • Sample traces. Keeping 100% of traces at high traffic gets expensive. Head sampling (keep, say, 10%) is simple; tail sampling in an OpenTelemetry Collector can keep all errors and slow requests while dropping most of the fast, boring ones.
  • Watch metric cardinality. Never put user ids, emails or full URLs in metric labels. Each unique combination becomes a separate time series.
  • Don't log secrets or personal data. Scrub tokens, passwords and emails before they reach the log pipeline.

A good first week

  1. Add OpenTelemetry auto-instrumentation to one service and send traces to a backend.
  2. Add RED metrics and a dashboard for the top five routes.
  3. Add trace ids to logs.
  4. Write two alerts: error rate above target, and p95 latency above target.

From there, add custom spans wherever an incident showed you a blind spot. Observability grows best one real question at a time.

Written by

RecallRun Editors

Practical guides and independent tool overviews from the RecallRun team. Every post is written to be tested on your own machine.

Website

Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.

More from the community

Write for RecallRun

Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.

Start writing