InterviewDB Experience

Keyword Highlighting: Highlight Multiple Keywords in Text While Handling Overlapping Spans

Interview Experience

Problem

Given a text string and a list of keywords,

return the text with each keyword occurrence wrapped in <mark> tags. Keywords are case-insensitive. If two keyword matches overlap, merge them into a single <mark> span covering both.

python
def highlight_keywords(text: str, keywords: list[str]) -> str:
    ...
text = "The quick brown fox jumps over the lazy dog"
keywords = ["quick brown", "brown fox"]

Matches:
  "quick brown" -> [4, 14]
  "brown fox"   -> [10, 18]
  Merged span:  [4, 18] -> "quick brown fox"

**Output**: "The <mark>quick brown fox</mark> jumps over the lazy dog"

---
keywords = ["the", "lazy"]

**Output**: "<mark>The</mark> quick brown fox jumps over <mark>the</mark> <mark>lazy</mark> dog"

Follow-ups

  1. How do you handle overlapping matches efficiently using an interval merge step?
  2. What if keywords include regex special characters — how do you escape them before building your search pattern?
  3. Your output uses raw HTML. What XSS risk exists if text comes from user input, and how do you mitigate it?
  4. Extend to support highlighting in a rich-text document (nested HTML). Why is string replacement no longer sufficient?

Full Details

Problem

Given a text string and a list of keywords,

return the text with each keyword occurrence wrapped in <mark> tags. Keywords are case-insensitive. If two keyword matches overlap, merge them into a single <mark> span covering both.

python
def highlight_keywords(text: str, keywords: list[str]) -> str:
    ...
text = "The quick brown fox jumps over the lazy dog"
keywords = ["quick brown", "brown fox"]

Matches:
  "quick brown" -> [4, 14]
  "brown fox"   -> [10, 18]
  Merged span:  [4, 18] -> "quick brown fox"

**Output**: "The <mark>quick brown fox</mark> jumps over the lazy dog"

---
keywords = ["the", "lazy"]

**Output**: "<mark>The</mark> quick brown fox jumps over <mark>the</mark> <mark>lazy</mark> dog"

Follow-ups

  1. How do you handle overlapping matches efficiently using an interval merge step?
  2. What if keywords include regex special characters — how do you escape them before building your search pattern?
  3. Your output uses raw HTML. What XSS risk exists if text comes from user input, and how do you mitigate it?
  4. Extend to support highlighting in a rich-text document (nested HTML). Why is string replacement no longer sufficient?

About This Question

This is a candidate experience report from a figma interview during the phone round.

It covers the following topics: Coding, Phone, Onsite, Strings .