InterviewDB Experience

Text Ingestion Pipeline - System Design for Scalable Document Ingestion and Indexing

Interview Experience

Round 1 System Design

Problem

Design a text ingestion pipeline that accepts large volumes of raw documents (PDFs, HTML, plain text) from multiple upstream sources, processes them, and makes them searchable within 60 seconds of arrival. The system must handle 100K documents/day and support full-text search with relevance ranking.

Requirements

  • Ingest from: REST upload, S3 event trigger, webhook.
  • Processing: extract text, detect language, chunk into passages, embed for semantic search.
  • Storage: raw store (S3), metadata (Postgres), vector index (e.g. Pinecone or pgvector).
  • SLA: p95 end-to-end ingest latency < 60s.

Design Sketch

Upload API -> Queue (SQS/Kafka) -> Worker Pool
  Worker: extract text -> chunk -> embed -> write metadata + vectors
  DLQ for failed docs -> alerting

Discussion Points

  • How do you handle retries without double-indexing a document?
  • How do you manage embedding model versioning — re-embedding 10M docs after a model upgrade?
  • Schema for the metadata table: what columns, what indexes?

Follow-ups

  1. How does your design change if documents can be updated or deleted after ingestion?
  2. How do you handle a 5 GB PDF that exceeds Lambda/worker memory limits?
  3. How do you monitor ingestion lag and alert when the queue backs up?
  4. How would you prioritize VIP customer uploads over standard ones?

Full Details

Round 1 System Design

Problem

Design a text ingestion pipeline that accepts large volumes of raw documents (PDFs, HTML, plain text) from multiple upstream sources, processes them, and makes them searchable within 60 seconds of arrival. The system must handle 100K documents/day and support full-text search with relevance ranking.

Requirements

  • Ingest from: REST upload, S3 event trigger, webhook.
  • Processing: extract text, detect language, chunk into passages, embed for semantic search.
  • Storage: raw store (S3), metadata (Postgres), vector index (e.g. Pinecone or pgvector).
  • SLA: p95 end-to-end ingest latency < 60s.

Design Sketch

Upload API -> Queue (SQS/Kafka) -> Worker Pool
  Worker: extract text -> chunk -> embed -> write metadata + vectors
  DLQ for failed docs -> alerting

Discussion Points

  • How do you handle retries without double-indexing a document?
  • How do you manage embedding model versioning — re-embedding 10M docs after a model upgrade?
  • Schema for the metadata table: what columns, what indexes?

Follow-ups

  1. How does your design change if documents can be updated or deleted after ingestion?
  2. How do you handle a 5 GB PDF that exceeds Lambda/worker memory limits?
  3. How do you monitor ingestion lag and alert when the queue backs up?
  4. How would you prioritize VIP customer uploads over standard ones?

About This Question

This is a candidate experience report from a temporal interview during the phone round.

It covers the following topics: Phone, System Design, System Design, Queue, Onsite .