Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

First seen · 8/5/2026, 12:12 AMLatest activity · 8/5/2026, 12:12 AM

EvoHIL is a human-in-the-loop reinforcement learning framework for contact-rich robotic manipulation. It combines a self-evolving success classifier, flow-matched generation of temporally coherent action chunks, and retention-aware offline fine-tuning with relit interaction data. The abstract reports experiments on six tasks using Franka FR3 and SO-101 arms under a controlled lighting shift, claiming improvements in success, agreement with human-confirmed labels, motion smoothness, and completion time. However, the supplied abstract provides no quantitative results, and the listed August 2026 publication date is future-dated and cannot currently be verified.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv8/5, 12:12 AMnot independentRepresentative
    EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning