InterviewDB
Question
·
Paris
Deduplication Pipeline: Remove Near-Duplicate Records Using Hashing
phone
Question Details
Problem You are building a data deduplication pipeline. Given a list of text records, identify and remove near-duplicates. Two records are near-duplicates if their normalized Jaccard similarity on character 3-grams exceeds a threshold t. Return the deduplicated list, keeping the first occurrence of each cluster. Follow-ups Naive O(n^2) comparison is too slow for large datasets. How does MinHash + LSH reduce this to near-linear? How do you choose the 3-gram size? What are the tradeoffs of 2-grams…
Full Details
🔒
Unlock all Perplexity questions
Full insider details, leaked discussions, and candidate experiences.
Get full access — $100 a year, unlimited accessAbout This Question
This is a reported interview question from a perplexity interview during the phone round.
More Perplexity Interview Questions
1p3a
perplexity software engineer tech phone screen insights
InterviewDB
Check Distribution: Verify Load Is Evenly Distributed Across Servers
1p3a
perplexity software engineer tech phone screen: implement ai todo list in python
InterviewDB
Perplexity SWE Phone - Dependency Validation (Graph/Topological Sort)
InterviewDB
Stopword Filtering: Remove Common Words from Text Efficiently