InterviewDB Experience

Unsafe Words: Implement a Content Filter with Contextual Word Blocklist

Interview Experience

Problem

Build a content moderation system that flags messages containing unsafe words. Blocklist entries can be:
- Exact strings: "badword"
- Wildcards: "bad*" (prefix match)
- Phrase patterns: "buy * now" (any word in the middle)

The filter should be case-insensitive and must not flag safe words that contain the blocked substring (e.g., blocking "ass" should not flag "assistant").

python
class ContentFilter:
    def add_rule(self, pattern: str) -> None:
    def is_unsafe(self, text: str) -> bool:
    def get_violations(self, text: str) -> list[str]:

**returns** matched patterns

Example

filter = ContentFilter()
filter.add_rule("spam")
filter.add_rule("buy * now")

filter.is_unsafe("Buy cheap products now!")  -> True   # matches "buy * now"
filter.is_unsafe("This is spam")             -> True
filter.is_unsafe("I am not a spammer")       -> False  # word-boundary safe
filter.get_violations("Buy it now or spam")  -> ["buy * now", "spam"]

Follow-ups

  1. How do you enforce word-boundary matching efficiently at scale?
  2. How would you handle Unicode normalization and leetspeak obfuscation (sp4m, s.p.a.m)?
  3. With 100K rules and 10K messages/second, how do you make this performant (Aho-Corasick, regex compilation)?

Full Details

Problem

Build a content moderation system that flags messages containing unsafe words. Blocklist entries can be:
- Exact strings: "badword"
- Wildcards: "bad*" (prefix match)
- Phrase patterns: "buy * now" (any word in the middle)

The filter should be case-insensitive and must not flag safe words that contain the blocked substring (e.g., blocking "ass" should not flag "assistant").

python
class ContentFilter:
    def add_rule(self, pattern: str) -> None:
    def is_unsafe(self, text: str) -> bool:
    def get_violations(self, text: str) -> list[str]:

**returns** matched patterns

Example

filter = ContentFilter()
filter.add_rule("spam")
filter.add_rule("buy * now")

filter.is_unsafe("Buy cheap products now!")  -> True   # matches "buy * now"
filter.is_unsafe("This is spam")             -> True
filter.is_unsafe("I am not a spammer")       -> False  # word-boundary safe
filter.get_violations("Buy it now or spam")  -> ["buy * now", "spam"]

Follow-ups

  1. How do you enforce word-boundary matching efficiently at scale?
  2. How would you handle Unicode normalization and leetspeak obfuscation (sp4m, s.p.a.m)?
  3. With 100K rules and 10K messages/second, how do you make this performant (Aho-Corasick, regex compilation)?

About This Question

This is a candidate experience report from a whatnot interview during the phone round.

It covers the following topics: Coding, Phone, Onsite, Strings .