InterviewDB Experience · Paris

LLM Output Annotator: Build a Tool That Labels LLM Responses with Quality Tags

Interview Experience

Problem You are building an annotation tool that takes an LLM-generated response and a reference answer, and labels the response with quality tags from a predefined set: ["correct", "partially_correct", "incorrect", "hallucinated", "refused", "off_topic"]. Implement the annotation logic using simple heuristics (exact match, keyword overlap, refusal detection). Example: Follow-ups What NLP metric (BLEU, ROUGE, BERTScore) gives a better signal than keyword overlap for partial correctness? How woul…

Full Details

🔒

Unlock all Harvey questions

Full insider details, leaked discussions, and candidate experiences.

Get full access — $100 a year, unlimited access

About This Question

This is a candidate experience report from a harvey interview during the phone round.

It covers the following topics: Coding, Phone .