10How do you evaluate an LLM, and why are benchmarks and LLM-as-judge both unreliable?▼hard★ EssentialOpenAIAnthropicGoogle2 repliesunlockedEvaluation is the hardest, most underrated part of shipping LLMs. The signal is knowing why public benchmarks mislead, why LLM-as-judge is biased, and how to build a task-specific eval you actually trust.Open full answer →
26What are MMLU, HumanEval, and GSM8K, and how do you interpret LLM benchmark scores?▼mediumOpenAIGoogleAnthropic1 replies◆ premiumEveryone quotes benchmark scores; few read them honestly. The signal is naming what each probes and why leaderboard numbers run ahead of real-world ability. The follow-up the interviewer is holding back is how you would detect contamination.Open full answer →
29How do you evaluate the safety of an LLM (safety benchmarks and beyond)?▼mediumAnthropicOpenAIGoogle2 replies◆ premiumSafety is not one number. The candidates who pass name the axes, run benchmarks as a gate, and then explain why benchmarks alone certify nothing. Here is the framing interviewers score highest.Open full answer →