InterviewDB Experience · Los Angeles

Job Monitor: Track Long-Running Jobs and Alert on Failures or Timeouts

Interview Experience

Problem

Design a JobMonitor that tracks the lifecycle of asynchronous jobs. Jobs can be registered, updated with status, and queried. Alert (log to a provided callback) when a job fails or has not completed within a timeout period.

python
from typing import Callable
import time

class JobMonitor:
    def __init__(self, alert_callback: Callable[[str, str], None]):
        """alert_callback(job_id, reason) is called on failure/timeout."""
        pass

    def register(self, job_id: str, timeout_sec: int) -> None:
        """Register a new job with a timeout."""
        pass

    def update_status(self, job_id: str, status: str) -> None:
        """status: 'running', 'completed', 'failed'"""
        pass

    def check_timeouts(self, current_time: float | None = None) -> None:
        """Call periodically; alerts on any job past its timeout."""
        pass

    def get_status(self, job_id: str) -> dict:
        pass
monitor = JobMonitor(alert_callback=lambda jid, reason: print(f"{jid}: {reason}"))
monitor.register("job1", timeout_sec=30)
monitor.update_status("job1", "running")
# ... 31 seconds later ...
monitor.check_timeouts()  # -> prints "job1: timeout"

Follow-ups

  1. How do you efficiently find all timed-out jobs without scanning all registered jobs every check?
  2. How would you persist job state so the monitor survives process restarts?
  3. Extend to support job dependencies: job B should not start until job A completes.
  4. How would you expose this as a REST API with endpoints for registration, status updates, and querying?

Full Details

Problem

Design a JobMonitor that tracks the lifecycle of asynchronous jobs. Jobs can be registered, updated with status, and queried. Alert (log to a provided callback) when a job fails or has not completed within a timeout period.

python
from typing import Callable
import time

class JobMonitor:
    def __init__(self, alert_callback: Callable[[str, str], None]):
        """alert_callback(job_id, reason) is called on failure/timeout."""
        pass

    def register(self, job_id: str, timeout_sec: int) -> None:
        """Register a new job with a timeout."""
        pass

    def update_status(self, job_id: str, status: str) -> None:
        """status: 'running', 'completed', 'failed'"""
        pass

    def check_timeouts(self, current_time: float | None = None) -> None:
        """Call periodically; alerts on any job past its timeout."""
        pass

    def get_status(self, job_id: str) -> dict:
        pass
monitor = JobMonitor(alert_callback=lambda jid, reason: print(f"{jid}: {reason}"))
monitor.register("job1", timeout_sec=30)
monitor.update_status("job1", "running")
# ... 31 seconds later ...
monitor.check_timeouts()  # -> prints "job1: timeout"

Follow-ups

  1. How do you efficiently find all timed-out jobs without scanning all registered jobs every check?
  2. How would you persist job state so the monitor survives process restarts?
  3. Extend to support job dependencies: job B should not start until job A completes.
  4. How would you expose this as a REST API with endpoints for registration, status updates, and querying?

About This Question

This is a candidate experience report from a applied intuition interview during the onsite round.

It covers the following topics: Coding, Onsite .