Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

First seen · 7/3/2026, 02:53 PMLatest activity · 7/3/2026, 02:53 PM

MiniCache turns Program-of-Thought programs into parameterized cache objects that can be reused across structurally similar requests. On cache hits, a small model extracts semantic variables; during target-model generation, it also performs speculative drafting. The authors report evaluations on shopping-style request datasets, WebShop, Formula, and CodeTAT-QA, with up to 3.1x lower latency and 2.8x higher throughput under parallel serving, while preserving task quality. The paper argues that small models are most useful as interface models around large models, rather than as direct replacements.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/3, 02:53 PMnot independentRepresentative
    MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference