The paper proposes a reward-driven LLM-agent workflow that combines partially observable Markov decision process (POMDP) routing, an internal reward model, self-critique, multimodal inputs, and graph-based memory. The authors report a 24.5 percentage-point improvement in task success rate and trajectory efficiency over standard ReAct baselines on ALFWorld and WebShop. They also state that ablations show the reward-driven critique module reduces hallucinations. The abstract does not provide detailed model configurations, statistical uncertainty, or benchmark-by-benchmark results.
No heat snapshots are available in the last 24 hours.