InterviewDB
Experience
·
New York
Pronunciation Annotation - Tag Words in Text with Phonetic Metadata
phone
Interview Experience
Round 1 Coding
Problem
Given a sentence and a dictionary mapping words to their phonetic annotation (e.g. IPA or a simplified pronunciation string),
return the sentence with each recognized word wrapped in an annotation tag. Unrecognized words are left as-is. The matching is case-insensitive; the original casing must be preserved in output.
python
def annotate_pronunciation(
sentence: str,
phoneme_dict: dict[str, str]
) -> str:
pass
Example
**Input**:
sentence = "The quick brown fox"
phoneme_dict = {"the": "dh-ah", "fox": "f-aa-k-s"}
**Output**:
"[The|dh-ah] quick brown [fox|f-aa-k-s]"
Example 2
**Input**:
sentence = "Reading is fun"
phoneme_dict = {"reading": "r-ee-d-ih-ng"}
**Output**:
"[Reading|r-ee-d-ih-ng] is fun"
Follow-ups
- How do you handle multi-word phrases in the dictionary, like
{"New York": "n-y-oo-y-aw-r-k"}? - What if the same word has different pronunciations depending on part of speech (e.g. "read" present vs past)?
- How would you build the phoneme_dict automatically from a corpus using a text-to-speech API?
- How do you handle punctuation attached to words (
fox,should still matchfox)?
Full Details
Round 1 Coding
Problem
Given a sentence and a dictionary mapping words to their phonetic annotation (e.g. IPA or a simplified pronunciation string),
return the sentence with each recognized word wrapped in an annotation tag. Unrecognized words are left as-is. The matching is case-insensitive; the original casing must be preserved in output.
python
def annotate_pronunciation(
sentence: str,
phoneme_dict: dict[str, str]
) -> str:
pass
Example
**Input**:
sentence = "The quick brown fox"
phoneme_dict = {"the": "dh-ah", "fox": "f-aa-k-s"}
**Output**:
"[The|dh-ah] quick brown [fox|f-aa-k-s]"
Example 2
**Input**:
sentence = "Reading is fun"
phoneme_dict = {"reading": "r-ee-d-ih-ng"}
**Output**:
"[Reading|r-ee-d-ih-ng] is fun"
Follow-ups
- How do you handle multi-word phrases in the dictionary, like
{"New York": "n-y-oo-y-aw-r-k"}? - What if the same word has different pronunciations depending on part of speech (e.g. "read" present vs past)?
- How would you build the phoneme_dict automatically from a corpus using a text-to-speech API?
- How do you handle punctuation attached to words (
fox,should still matchfox)?
Free preview. Unlock all Heygen questions →