InterviewDB Question

Metrics Aggregation System: Design a Scalable System to Aggregate and Query Time-Series Metrics

Question Details

Round 1 System Design

Problem

Design a metrics aggregation system for a large distributed application. The system receives millions of metric events per second (e.g., CPU usage, request latency) tagged with service name, host, and region. It must support queries like: average latency for service X in region Y over the last 1 hour, and alert when a metric crosses a threshold for 5 consecutive minutes.

Requirements:
- Ingest: 5M events/sec peak
- Query SLA: p99 < 2s for hour-range queries
- Retention: raw data 24h, 1-min rollups 30 days, 1-hour rollups 1 year
- Alerting latency: < 1 min from breach to notification

Follow-ups

  1. Walk through your ingestion pipeline — how do you handle bursts without data loss?
  2. How do you implement the rollup jobs without impacting query performance?
  3. Cardinality explosion: a tag combination explosion can kill a time-series DB. How do you prevent it?
  4. Describe your alerting architecture — how do you avoid false positives from transient spikes?

Full Details

Round 1 System Design

Problem

Design a metrics aggregation system for a large distributed application. The system receives millions of metric events per second (e.g., CPU usage, request latency) tagged with service name, host, and region. It must support queries like: average latency for service X in region Y over the last 1 hour, and alert when a metric crosses a threshold for 5 consecutive minutes.

Requirements:
- Ingest: 5M events/sec peak
- Query SLA: p99 < 2s for hour-range queries
- Retention: raw data 24h, 1-min rollups 30 days, 1-hour rollups 1 year
- Alerting latency: < 1 min from breach to notification

Follow-ups

  1. Walk through your ingestion pipeline — how do you handle bursts without data loss?
  2. How do you implement the rollup jobs without impacting query performance?
  3. Cardinality explosion: a tag combination explosion can kill a time-series DB. How do you prevent it?
  4. Describe your alerting architecture — how do you avoid false positives from transient spikes?

About This Question

This is a reported interview question from a coursera interview during the onsite round.

It covers the following topics: System Design, System Design, Backtracking, Onsite .