SocietyBench evaluates whether LLMs can forecast how social events and public opinion evolve, rather than merely complete operational tasks. It builds date-indexed timelines from Web news and social-media posts across five platforms, separating factual developments from public opinion. Named entities are replaced and dates shifted to reduce matching against pretrained memories. Across five events and 125 bilingual prediction points, the best of six frontier LLMs scores 75.0/100 versus a trivial anchor of 50. Three agent frameworks do not improve over their shared base model, while calibration and temporal accuracy show distinct failure patterns.
No heat snapshots are available in the last 24 hours.