Bad Word Checking: Detect and Censor Profanity in User Text Using Exact and Fuzzy Matching
Question Details
Problem
Implement a content moderation function that detects bad words in user-submitted text. Matching must handle:
1. Exact match (case-insensitive).
2. Leet-speak substitutions: 3->e, @->a, 0->o, 1->i.
3. Repeated characters: haaaate matches hate.
Return the sanitized string with each bad word replaced by ***.
python
def censor(text: str, blocklist: list[str]) -> str:
pass
Example:
blocklist = ["hate", "spam"]
text = "I h@t3 sp@@m and haaaate it!"
**output** -> "I *** *** and *** it!"
Follow-ups
1. Leet-speak normalization should happen before deduplication or after? Why does the order matter?
2. How would you build a Trie over the normalized blocklist for fast multi-pattern matching (Aho-Corasick)?
3. False positives: "assassination" contains "ass". How do you reduce them without a huge allowlist?
4. The system processes 100,000 messages per second. What architecture would you use for real-time filtering?
Full Details
Problem
Implement a content moderation function that detects bad words in user-submitted text. Matching must handle:
1. Exact match (case-insensitive).
2. Leet-speak substitutions: 3->e, @->a, 0->o, 1->i.
3. Repeated characters: haaaate matches hate.
Return the sanitized string with each bad word replaced by ***.
python
def censor(text: str, blocklist: list[str]) -> str:
pass
Example:
blocklist = ["hate", "spam"]
text = "I h@t3 sp@@m and haaaate it!"
**output** -> "I *** *** and *** it!"
Follow-ups
1. Leet-speak normalization should happen before deduplication or after? Why does the order matter?
2. How would you build a Trie over the normalized blocklist for fast multi-pattern matching (Aho-Corasick)?
3. False positives: "assassination" contains "ass". How do you reduce them without a huge allowlist?
4. The system processes 100,000 messages per second. What architecture would you use for real-time filtering?
About This Question
This is a reported interview question from a discord interview during the phone round.
It covers the following topics: Strings, Phone, Trie, Coding, Onsite .