InterviewDB
Experience
Split Text - Tokenize and Split Text by Sentence and Paragraph Boundaries
phone
Interview Experience
Problem
Implement a text splitter that segments a long document into chunks suitable for ML model ingestion. Given a document string and a max_chunk_tokens limit, split the text while:
- Preserving sentence boundaries (do not cut mid-sentence).
- Keeping paragraphs together when they fit.
- Adding configurable overlap (last
overlap_tokenstokens of previous chunk appear at the start of the next).
python
def split_text(
text: str,
max_chunk_tokens: int,
overlap_tokens: int = 50
) -> List[str]:
...
Example:
text = "Hello world. How are you?
This is paragraph two. It has two sentences."
split_text(text, max_chunk_tokens=10, overlap_tokens=2)
# -> ["Hello world. How are you?",
# "are you? This is paragraph two.",
# "paragraph two. It has two sentences."]
# (token counts approximate)
Follow-ups
- How do you estimate token count without running a full tokenizer? When does the approximation break?
- How would you handle code blocks or tables inside the document - should they be split differently?
- What is the effect of
overlap_tokenson downstream retrieval quality in a RAG pipeline? - How would you parallelize splitting for a 1M-document corpus?
Full Details
Problem
Implement a text splitter that segments a long document into chunks suitable for ML model ingestion. Given a document string and a max_chunk_tokens limit, split the text while:
- Preserving sentence boundaries (do not cut mid-sentence).
- Keeping paragraphs together when they fit.
- Adding configurable overlap (last
overlap_tokenstokens of previous chunk appear at the start of the next).
python
def split_text(
text: str,
max_chunk_tokens: int,
overlap_tokens: int = 50
) -> List[str]:
...
Example:
text = "Hello world. How are you?
This is paragraph two. It has two sentences."
split_text(text, max_chunk_tokens=10, overlap_tokens=2)
# -> ["Hello world. How are you?",
# "are you? This is paragraph two.",
# "paragraph two. It has two sentences."]
# (token counts approximate)
Follow-ups
- How do you estimate token count without running a full tokenizer? When does the approximation break?
- How would you handle code blocks or tables inside the document - should they be split differently?
- What is the effect of
overlap_tokenson downstream retrieval quality in a RAG pipeline? - How would you parallelize splitting for a 1M-document corpus?
Free preview. Unlock all Grammarly questions →
About This Question
This is a candidate experience report from a grammarly interview during the phone round.
It covers the following topics: Mle, Strings, Phone, Coding, Onsite .
More Grammarly Interview Questions
1p3a
grammarly software engineer onsite interview experience
InterviewDB
All K Substrings - Generate All Substrings of Exactly Length K
InterviewDB
Grammarly SWE Phone - Duplicate and Missing Numbers
InterviewDB
Grammarly SWE Phone - Fibonacci Number
InterviewDB
Grammarly SWE Onsite - Merge Correction