Blind-Spots-Bench introduces a diagnostic benchmark of 235 samples targeting tasks that humans often find trivial but modern AI systems frequently mishandle, including string manipulation and unusual image-generation requests such as drawing a dog with five legs. The authors collect questions from students in an AI course, clean and annotate them with structured reference solutions, define a task taxonomy, and build automated grading for language, vision-language, and image-generation models. Their analysis reports that closed-source frontier models outperform open-weight models by approximately 10%, despite similar results on established benchmarks. No model dominates every task category, and some tasks remain difficult for all evaluated systems.
No heat snapshots are available in the last 24 hours.