InterviewDB
Experience
·
USA
Find Tickers: Extract Stock Ticker Symbols from Unstructured Text
Onsite
Interview Experience
Problem
Given a block of unstructured text (e.g., a financial news article), extract all valid US stock ticker symbols. A ticker is 1-5 uppercase letters. However, not every sequence of uppercase letters is a ticker — you must filter using a provided set of known valid tickers.
python
def find_tickers(
text: str,
valid_tickers: set[str]
) -> list[str]:
"""Return list of tickers found, in order of first appearance, deduplicated."""
...
text = "Investors are watching AAPL and GOOGL closely. The FDA approved MRNA."
valid_tickers = {"AAPL", "GOOGL", "MRNA", "TSLA"}
**Output**: ["AAPL", "GOOGL", "MRNA"]
Follow-ups
- How do you distinguish the ticker
I(Intelsat) from the word "I" in normal English text? What heuristics help? - Tickers can appear with punctuation around them (e.g., "(AAPL)", "AAPL,"). How does your regex handle this?
- The valid_tickers set has 10,000 entries. Does set lookup remain O(1) in Python? How does this affect overall complexity?
- A new ticker was listed today and is not yet in
valid_tickers. How would you keep your ticker list fresh without manual updates?
Full Details
Problem
Given a block of unstructured text (e.g., a financial news article), extract all valid US stock ticker symbols. A ticker is 1-5 uppercase letters. However, not every sequence of uppercase letters is a ticker — you must filter using a provided set of known valid tickers.
python
def find_tickers(
text: str,
valid_tickers: set[str]
) -> list[str]:
"""Return list of tickers found, in order of first appearance, deduplicated."""
...
text = "Investors are watching AAPL and GOOGL closely. The FDA approved MRNA."
valid_tickers = {"AAPL", "GOOGL", "MRNA", "TSLA"}
**Output**: ["AAPL", "GOOGL", "MRNA"]
Follow-ups
- How do you distinguish the ticker
I(Intelsat) from the word "I" in normal English text? What heuristics help? - Tickers can appear with punctuation around them (e.g., "(AAPL)", "AAPL,"). How does your regex handle this?
- The valid_tickers set has 10,000 entries. Does set lookup remain O(1) in Python? How does this affect overall complexity?
- A new ticker was listed today and is not yet in
valid_tickers. How would you keep your ticker list fresh without manual updates?
Free preview. Unlock all SIG questions →
About This Question
This is a candidate experience report from a sig interview during the onsite round.