This preprint compares Claude, Grok, GPT, and Gemini across four snapshots from October 2025 to February 2026, using both APIs and web interfaces to evaluate ethnonationalist pseudo-science derived from Frank Salter’s biosocial framework. It reports that Grok Fast versions repeatedly assigned credibility scores of 70–75, while other systems scored 15–40. The same model identifier also produced sharply different API and web results, and behavior changed after an undocumented patch. The authors argue that an LLM’s epistemic stance is a property of its deployment configuration, not simply its model weights.
No heat snapshots are available in the last 24 hours.