Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain

First seen · 7/28/2026, 04:10 PMLatest activity · 7/28/2026, 04:10 PM

This paper presents an auditable controller for frozen LLM agents. Instead of allowing costly code search or unconstrained self-modification, it defines a small, human-legible action space over prompt templates, tools, memory and retrieval, planning, and verification policies. The controller learns online with an ε-greedy contextual bandit and REINFORCE, using a multi-objective reward covering task success, verifier score, policy compliance, cost, latency, and unsupported-claim penalties. The authors instantiate it with DSPy and evaluate it on tool-use workflows, HumanEval code generation, and HotpotQA multi-hop QA using a local Ollama model and AWS Bedrock. Code, datasets, logs, and a deployment recipe are released.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/28, 04:10 PMnot independentRepresentative
    A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain