InterviewDB Experience

Split Text - Tokenize and Split Text by Sentence and Paragraph Boundaries

Interview Experience

Problem

Implement a text splitter that segments a long document into chunks suitable for ML model ingestion. Given a document string and a max_chunk_tokens limit, split the text while:

  • Preserving sentence boundaries (do not cut mid-sentence).
  • Keeping paragraphs together when they fit.
  • Adding configurable overlap (last overlap_tokens tokens of previous chunk appear at the start of the next).
python
def split_text(
    text: str,
    max_chunk_tokens: int,
    overlap_tokens: int = 50
) -> List[str]:
    ...

Example:

text = "Hello world. How are you?

This is paragraph two. It has two sentences."
split_text(text, max_chunk_tokens=10, overlap_tokens=2)
# -> ["Hello world. How are you?",
#     "are you? This is paragraph two.",
#     "paragraph two. It has two sentences."]
# (token counts approximate)

Follow-ups

  1. How do you estimate token count without running a full tokenizer? When does the approximation break?
  2. How would you handle code blocks or tables inside the document - should they be split differently?
  3. What is the effect of overlap_tokens on downstream retrieval quality in a RAG pipeline?
  4. How would you parallelize splitting for a 1M-document corpus?

Full Details

Problem

Implement a text splitter that segments a long document into chunks suitable for ML model ingestion. Given a document string and a max_chunk_tokens limit, split the text while:

  • Preserving sentence boundaries (do not cut mid-sentence).
  • Keeping paragraphs together when they fit.
  • Adding configurable overlap (last overlap_tokens tokens of previous chunk appear at the start of the next).
python
def split_text(
    text: str,
    max_chunk_tokens: int,
    overlap_tokens: int = 50
) -> List[str]:
    ...

Example:

text = "Hello world. How are you?

This is paragraph two. It has two sentences."
split_text(text, max_chunk_tokens=10, overlap_tokens=2)
# -> ["Hello world. How are you?",
#     "are you? This is paragraph two.",
#     "paragraph two. It has two sentences."]
# (token counts approximate)

Follow-ups

  1. How do you estimate token count without running a full tokenizer? When does the approximation break?
  2. How would you handle code blocks or tables inside the document - should they be split differently?
  3. What is the effect of overlap_tokens on downstream retrieval quality in a RAG pipeline?
  4. How would you parallelize splitting for a 1M-document corpus?

About This Question

This is a candidate experience report from a grammarly interview during the phone round.

It covers the following topics: Mle, Strings, Phone, Coding, Onsite .