LEMUR combines multi-objective reinforcement learning with preference-based reward learning. Instead of assuming a predefined reward for every objective, it interactively collects preferences from multiple humans and jointly learns objective-specific reward models and multi-objective policies. This formulation targets settings with competing goals, such as performance and efficiency, whose rewards are difficult to specify directly. The abstract reports superior performance over baselines across several benchmark tasks, but provides no quantitative results, preference-budget comparisons, statistical significance, or details about how conflicting annotator preferences are handled.
No heat snapshots are available in the last 24 hours.