Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

First seen · 8/3/2026, 04:00 AMLatest activity · 8/3/2026, 04:00 AM

PCSD addresses sparse-reward training for language-model agents by deriving token-level self-distillation weights from the local persistence of signals favoring a privileged teacher. It combines adaptive windows, exponentially decayed aggregation, trend-aware attenuation, and sigmoid gating, then jointly optimizes the resulting objective with GRPO. According to the supplied abstract, PCSD leads ALFWorld Overall on two backbones, improving over GRPO by 15.6 and 13.3 points and over SDAR by 6.2 and 5.5 points, while adding 15.8 points over GRPO on an unseen ALFWorld split.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/3, 04:00 AMnot independentRepresentative
    PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning