This paper presents a value-learning method for generative AI that jointly infers two components from pairwise prompt-response preferences: implementations of individual value groundings through a multi-objective reward model, and a value-system representation expressed as a weighted linear scalarization of that grounding model. The algorithm dynamically prioritizes grounding learning to promote coherent value representations. Evaluations on prompt-response preference datasets reportedly show competitive performance and limited trade-offs relative to baselines and a contemporary method, while improving explainability.
No heat snapshots are available in the last 24 hours.