ReBRAC-v2 revisits a conventional behavior-regularized actor-critic instead of relying on increasingly specialized generative policies and value guidance. It uses an exact-likelihood normalizing-flow actor, mixed likelihood/MSE/MAE behavior regularization, a classification-based residual critic, staged optimization, and multi-sample action selection. A shared recipe tuned through roughly 600 Bayesian proposals on six OGBench tasks is then evaluated across ten state-based categories, reaching a 74.8 average versus 52.3 for the next-best aggregate result and ranking first in eight categories. The same recipe reports averages of 90.2 on D4RL AntMaze and 33.6 on Adroit.
No heat snapshots are available in the last 24 hours.