Skip to main content

Cloud and DevOps

Observability and application performance

I make your applications diagnosable: correlated traces, useful metrics and alerts that signal a real problem. You move from interpreting symptoms to identifying the cause.

What it covers

An observable system is recognised by how quickly a question becomes an answer: why did this request slow down yesterday at three, which service produced an error on this journey, how long did processing actually take. Without correlated traces, those questions turn into hours of discussion.

I start by instrumenting critical journeys rather than every component. Each request carries an identifier propagated across services, calls to databases and external systems are measured, and errors are enriched with the context needed for diagnosis.

Alerts worth waking up for

Alerts are defined from service objectives rather than arbitrary thresholds. They signal degradation perceived by users, with a false-positive rate tracked and reduced over time. Dashboards are organised by business journey.

  • Distributed traces with an identifier propagated across services.
  • Application and technical metrics tied to user journeys.
  • Alerts based on measurable service objectives.
  • Dashboards per journey rather than per server.
  • Cost analysis for trace and log storage.

Problems addressed

  • Incidents are diagnosed by elimination, for lack of traces.
  • Too many alerts end up being ignored.
  • Slowdowns are noticed but never precisely attributed.
  • Logs are voluminous but hard to use in an emergency.
  • Trace storage costs grow without control.

Expected benefits

  • Incident diagnosis in minutes with an identified cause.
  • Fewer useless alerts draining the team.
  • Localised bottlenecks without opinion-based debate.
  • Service indicators tracked and communicated objectively.
  • Diagnostic capability retained during traffic peaks.
  • Observability storage costs controlled and justified.

Method and steps

  1. 1

    Question framing

    Interviews with operations and development teams to list the questions observability must answer, ordered by real frequency.

  2. 2

    Instrumentation

    Adding traces, metrics and structured logs on critical journeys, with context propagation across services and to external systems.

  3. 3

    Dashboards and alerts

    Building views per business journey and defining alerts tied to service objectives, with false-positive rate measurement.

  4. 4

    Noise reduction

    Removing useless alerts, adjusting thresholds, introducing retention policies and tracking storage costs.

Deliverables

  • List of diagnostic questions and instrumented journeys.

  • Application instrumentation with end-to-end correlation.

  • Dashboards per business journey.

  • Alert rules tied to service objectives.

  • Retention policy and observability cost report.

  • Diagnostic guide for on-call teams.

Technologies used

  • Spring Boot
  • PostgreSQL
  • Kubernetes
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Elastic Stack

Frequently asked questions

Should the whole application be instrumented?

No, and doing so is expensive in volume and complexity. We start with critical journeys and external dependencies, then extend only where a diagnostic need is demonstrated.

How do you avoid alert overload?

By tying every alert to a service objective and an expected action. An alert with no clear action is removed or turned into a dashboard indicator.

What is the performance impact?

Instrumentation has a cost, usually a few percent of processing time. We measure it and tune it, particularly on high-volume processing where sampling is preferable.

Related case studies

Talk about my observability

Describe the situations where your team loses the most time. We will target instrumentation at those specific cases.

Talk about my observability