Logs, metrics and traces: a practical introduction to observability with OpenTelemetry
What each signal is for, how they fit together through trace ids, and how to instrument a Python service with OpenTelemetry without drowning in data.
When something breaks in production, you need to answer three questions fast: is something wrong, where is it wrong, and why? Metrics, traces and logs each answer one of those best. Used together, linked by a shared trace id, they turn a two-hour hunt into a ten-minute fix.
The three signals
| Signal | Answers | Example |
|---|---|---|
| Metrics | Is something wrong? | p95 latency of /checkout jumped from 180 ms to 2.4 s |
| Traces | Where is the time going? | 2.1 s of that is one database query in the payment step |
| Logs | Why did it happen? | lock wait timeout on table orders for request abc123 |
Metrics are cheap numeric time series: counters, gauges and histograms. Keep them aggregated and low-cardinality; they power dashboards and alerts.
Traces follow one request across services. A trace is a tree of spans; each span is a timed operation with attributes (HTTP route, database statement, user tier).
Logs are discrete events with detail. Structured logs (JSON with consistent fields) are far more useful than free text, especially when every log line carries the current trace id.
Why OpenTelemetry
OpenTelemetry (OTel) is a vendor-neutral standard and set of SDKs for producing all three signals. You instrument your code once and send the data to whatever backend you choose: Grafana's stack, Jaeger, Honeycomb, Datadog, or a cloud provider's tracing service. Switching backends becomes a configuration change rather than a rewrite.
Instrumenting a FastAPI service
Install the SDK, an exporter and the instrumentations for your libraries:
pip install opentelemetry-sdk opentelemetry-exporter-otlp \
opentelemetry-instrumentation-fastapi opentelemetry-instrumentation-httpx
Set up tracing once at startup:
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.instrumentation.httpx import HTTPXClientInstrumentor
provider = TracerProvider(resource=Resource.create({"service.name": "orders-api"}))
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter())) # endpoint from env vars
trace.set_tracer_provider(provider)
FastAPIInstrumentor.instrument_app(app) # a span for every request
HTTPXClientInstrumentor().instrument() # spans for outgoing calls, context propagated
Auto-instrumentation gives you request and client spans for free. Add your own spans around the business steps that matter:
tracer = trace.get_tracer(__name__)
async def checkout(cart):
with tracer.start_as_current_span("checkout.price_cart") as span:
span.set_attribute("cart.items", len(cart.items))
total = await price(cart)
with tracer.start_as_current_span("checkout.charge"):
return await charge(total)
Connect logs to traces
Put the trace id on every log line so you can jump from a slow trace to its logs and back:
import logging
from opentelemetry import trace
class TraceIdFilter(logging.Filter):
def filter(self, record):
ctx = trace.get_current_span().get_span_context()
record.trace_id = format(ctx.trace_id, "032x") if ctx.is_valid else "-"
return True
handler = logging.StreamHandler()
handler.addFilter(TraceIdFilter())
handler.setFormatter(logging.Formatter('{"level":"%(levelname)s","msg":"%(message)s","trace_id":"%(trace_id)s"}'))
logging.getLogger().addHandler(handler)
Metrics that matter: start with RED
For each service, track:
- Rate: requests per second
- Errors: failed requests per second (or as a percentage)
- Duration: latency as a histogram, so you can read p50, p95 and p99
Alert on symptoms users feel (error rate, p95 latency against your target), not on every CPU spike.
Keep the cost under control
- Sample traces. Keeping 100% of traces at high traffic gets expensive. Head sampling (keep, say, 10%) is simple; tail sampling in an OpenTelemetry Collector can keep all errors and slow requests while dropping most of the fast, boring ones.
- Watch metric cardinality. Never put user ids, emails or full URLs in metric labels. Each unique combination becomes a separate time series.
- Don't log secrets or personal data. Scrub tokens, passwords and emails before they reach the log pipeline.
A good first week
- Add OpenTelemetry auto-instrumentation to one service and send traces to a backend.
- Add RED metrics and a dashboard for the top five routes.
- Add trace ids to logs.
- Write two alerts: error rate above target, and p95 latency above target.
From there, add custom spans wherever an incident showed you a blind spot. Observability grows best one real question at a time.
Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.
More from the community
- Tech articles
Git workflows for small teams: trunk-based development vs feature branches
How trunk-based development, short-lived feature branches and GitFlow compare for teams of two to twenty, and the habits that make whichever you choose run smoothly.
- Tech articles
asyncio, threads or processes? Choosing Python concurrency by workload
A practical decision guide: which Python concurrency model fits I/O-bound, CPU-bound and mixed workloads, with small runnable examples and the traps to avoid.
- Tools
Pydantic v2: validate data at the edges of your Python application
Pydantic turns type hints into fast runtime validation and serialisation. Here are the core patterns for API payloads, settings and LLM outputs, plus the v2 changes that trip people up.
Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.