WifeBench is a Hacker News Show HN project that presents rankings of large language models based on the creator’s wife’s subjective “vibes.” The supplied discussion metadata reports a score of 2 and 0 comments at publication time. The available information does not specify the evaluated models, prompts, sample size, scoring rubric, annotator process, or reproducibility details, so the project is better understood as an informal opinion-based benchmark than as a validated evaluation methodology.
No heat snapshots are available in the last 24 hours.