Canonical Joint Energy-Based Model on CIFAR-10: Failure Modes and Practical Indistinguishability of Predictor-Corrector and SGLD Samplers
Original title:Canonical Joint Energy-Based Model on CIFAR-10: failure modes and practical indistinguishability of Predictor-Corrector and SGLD samplers
This paper reproduces the canonical Joint Energy-Based Model (JEM) with a WideResNet-28-10 without normalization layers and systematically compares Predictor-Corrector (PC) sampling with SGLD. Across two independent runs, the reconstruction reaches 92.88% test accuracy and a buffer-FID of 44.46, versus the canonical FID of 38.40. Both samplers exhibit catastrophic late-training divergence associated with the canonical outlier-buffer mechanism. Across ten checkpoint-OOD pairs, the absolute AUROC difference remains below 0.007, while the FID difference stays below 0.5. The experiments find no practical advantage for PC under fixed-noise canonical JEM training.
Why it's worth reading
JEM results depend on sampler behavior and buffer stability; this replication tests a widely assumed PC advantage and shows where fixed-noise theory fails to predict practical gains.