InterviewDB Question

Stopword Filtering: Remove Common Words from Text Efficiently

Question Details

Problem Given a list of text documents and a stopword list, remove all stopword occurrences from each document. Matching is case-insensitive. Preserve original word casing for non-stopwords. Return the filtered documents. Follow-ups What data structure do you use for stopword lookup, and why? How do you handle punctuation attached to words (e.g., "fox," where "fox" is a stopword)? For a corpus of 10 million documents, how would you parallelize this pipeline? Extend to support language-specific s…

Full Details

🔒

Unlock all Perplexity questions

Full insider details, leaked discussions, and candidate experiences.

Get full access — $100 a year, unlimited access

About This Question

This is a reported interview question from a perplexity interview during the phone round.

It covers the following topics: Coding, Phone .