量子位henry
“Show Me The Lord of the Rings”: Karpathy Highlights a New LLM Benchmark
Original title:Show me《指环王》!卡帕西强推大模型评测新基准
Industry46
The Lord of the Rings Becomes a New Benchmark for Large Language Models
Why it's worth reading
Narrative tasks may expose capabilities missed by static benchmarks, but the original methodology and results are needed before treating this as meaningful model evidence.
Tags
LLM评测BenchmarkKarpathy指环王QbitAI