The paper introduces the quadrilateral loss, a differentiable second-order mixed-difference penalty computed by swapping one coordinate between pairs of training points. It measures interaction as model behavior rather than enforcing additivity architecturally, and is claimed to remain informative for piecewise-linear networks. The authors relate its expectation to per-feature interaction mass in an interventional Shapley-GAM. Experiments compare structural masks, behavioral regularization, weight decay, backfitting, shared-section models, and bagged boosted stumps, reporting that behavioral constraints can improve accuracy and additivity on small datasets and that pre-regularization interaction rankings poorly predict retained interactions.
No heat snapshots are available in the last 24 hours.