InterviewDB Experience · USA

Find Tickers: Extract Stock Ticker Symbols from Unstructured Text

Interview Experience

Problem

Given a block of unstructured text (e.g., a financial news article), extract all valid US stock ticker symbols. A ticker is 1-5 uppercase letters. However, not every sequence of uppercase letters is a ticker — you must filter using a provided set of known valid tickers.

python
def find_tickers(
    text: str,
    valid_tickers: set[str]
) -> list[str]:
    """Return list of tickers found, in order of first appearance, deduplicated."""
    ...
text = "Investors are watching AAPL and GOOGL closely. The FDA approved MRNA."
valid_tickers = {"AAPL", "GOOGL", "MRNA", "TSLA"}

**Output**: ["AAPL", "GOOGL", "MRNA"]

Follow-ups

  1. How do you distinguish the ticker I (Intelsat) from the word "I" in normal English text? What heuristics help?
  2. Tickers can appear with punctuation around them (e.g., "(AAPL)", "AAPL,"). How does your regex handle this?
  3. The valid_tickers set has 10,000 entries. Does set lookup remain O(1) in Python? How does this affect overall complexity?
  4. A new ticker was listed today and is not yet in valid_tickers. How would you keep your ticker list fresh without manual updates?

Full Details

Problem

Given a block of unstructured text (e.g., a financial news article), extract all valid US stock ticker symbols. A ticker is 1-5 uppercase letters. However, not every sequence of uppercase letters is a ticker — you must filter using a provided set of known valid tickers.

python
def find_tickers(
    text: str,
    valid_tickers: set[str]
) -> list[str]:
    """Return list of tickers found, in order of first appearance, deduplicated."""
    ...
text = "Investors are watching AAPL and GOOGL closely. The FDA approved MRNA."
valid_tickers = {"AAPL", "GOOGL", "MRNA", "TSLA"}

**Output**: ["AAPL", "GOOGL", "MRNA"]

Follow-ups

  1. How do you distinguish the ticker I (Intelsat) from the word "I" in normal English text? What heuristics help?
  2. Tickers can appear with punctuation around them (e.g., "(AAPL)", "AAPL,"). How does your regex handle this?
  3. The valid_tickers set has 10,000 entries. Does set lookup remain O(1) in Python? How does this affect overall complexity?
  4. A new ticker was listed today and is not yet in valid_tickers. How would you keep your ticker list fresh without manual updates?

About This Question

This is a candidate experience report from a sig interview during the onsite round.

It covers the following topics: Coding, Onsite .