This paper introduces Knowledge- and Gradient-Guided Reinforcement Learning (KGRL) for Parametrized Action Markov Decision Processes (PAMDPs), where each decision combines a symbolic action with numerical parameters. KGRL uses a Datalog knowledge base to derive applicable actions and feasible parameter ranges, pruning invalid choices before learning. A gradient-based refinement loop then improves parameter estimates during training and deployment. The method also records activated rules along trajectories, producing local procedural explanations for action pruning and parameter constraints. The authors report that KGRL outperforms state-of-the-art reinforcement learning baselines in both sample efficiency and episodic return, although the provided abstract does not include numerical results or benchmark details.
No heat snapshots are available in the last 24 hours.