InterviewDB Question · Paris

Deduplication Pipeline: Remove Near-Duplicate Records Using Hashing

Question Details

Problem You are building a data deduplication pipeline. Given a list of text records, identify and remove near-duplicates. Two records are near-duplicates if their normalized Jaccard similarity on character 3-grams exceeds a threshold t. Return the deduplicated list, keeping the first occurrence of each cluster. Follow-ups Naive O(n^2) comparison is too slow for large datasets. How does MinHash + LSH reduce this to near-linear? How do you choose the 3-gram size? What are the tradeoffs of 2-grams…

Full Details

🔒

Unlock all Perplexity questions

Full insider details, leaked discussions, and candidate experiences.

Get full access — $100 a year, unlimited access

About This Question

This is a reported interview question from a perplexity interview during the phone round.

It covers the following topics: Coding, Mle, Phone .