Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

First seen · 8/6/2026, 01:57 AMLatest activity · 8/6/2026, 01:57 AM

This paper introduces Skill Entropy, a measure of how difficult it is for an LLM to switch between reasoning skills during long-horizon tasks. It presents Skill²-Bench, covering 558 skills across nine verifiable and open-ended domains, and evaluates eight frontier and four open-source models. Accuracy declines as task-level skill entropy rises. The proposed Skill-Entropy RL trains models to predict both intermediate answers and the skills used, reportedly raising scores for Qwen3-4B-Instruct from 34.4% to 68.4% and Qwen3-1.7B from 14.6% to 40.1%.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/5, 04:00 AMnot independent
    Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
  2. AggregatorarXiv8/6, 01:57 AMnot independentRepresentative
    Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning