Free sample question

Why standardize on OpenTelemetry, and when is the instrumentation cost worth it?

Senior · Observability · Standards · Chapter 1: Observability Foundations: Metrics, Logs & Telemetry, question 1 of 7 · from Senior DevOps & SRE Handbook: Observability, Reliability & Security

What the interviewer is really testing

whether you treat telemetry as an architectural decision about lock-in and control, or as a tool checklist. They want to hear when the standardization payoff justifies the cost, and when it does not.

The 30-second answer

OpenTelemetry is the CNCF vendor-neutral standard for emitting traces, metrics, and logs through one set of SDKs and the OTLP wire format. I'd standardize on it for the architectural payoff: instrument once, then route anywhere from a central Collector, so swapping Datadog for Grafana, or sending to two backends at once, is a config change, not a re-instrumentation project. The cost is real, you rewrite instrumentation and run Collector infrastructure, so I justify it where I expect to outlive a vendor contract or run a polyglot fleet. For a single small service on a tool I'm happy with, I'd wait.

Follow-ups the interviewer will probe

What does OTLP actually give you over a vendor protocol?
OTLP is one self-describing wire format for all three signals, so any SDK talks to any Collector and any backend that speaks it. That decoupling is the whole point: producers and consumers evolve independently, and you add or swap a backend by editing an exporter, never the application.
Agent per host, sidecar, or gateway Collector?
I usually run both tiers: a node or sidecar agent for local collection and resource attributes, feeding a gateway Collector pool that does the heavy work (tail-based sampling needs whole traces in one place). The gateway is where I enforce redaction, batching, and routing, and I make that pool HA.
Are OTel logs production-ready, or stick with an existing pipeline?
Logs are stable and GA now, so the bridge is real, and routing logs through the same Collector gets you trace correlation for free. I still migrate gradually: keep the current log pipeline, dual-ship through the Collector, validate parity, then cut over rather than flipping everything at once.

Recall hook

“Instrument once, route anywhere.”

OTel owns the instrumentation contract, OTLP carries every signal, and the Collector is the one place you redact, sample, and fan out to any backend.

What the book adds to this question

In the ebook every question runs three pages. Between the 30-second answer and the follow-ups it adds a deep dive with a diagram, a decision framework, and a pitfalls-and-signals table. Senior DevOps & SRE Handbook: Observability, Reliability & Security has 50 questions across 8 chapters and includes the free Interview-Day Playbook.

Want full three-page questions? Download the free 8-question PDF sample (from Cloud Interview Mastery).

Sample questions from the other books