MiniCache turns Program-of-Thought programs into parameterized cache objects that can be reused across structurally similar requests. On cache hits, a small model extracts semantic variables; during target-model generation, it also performs speculative drafting. The authors report evaluations on shopping-style request datasets, WebShop, Formula, and CodeTAT-QA, with up to 3.1x lower latency and 2.8x higher throughput under parallel serving, while preserving task quality. The paper argues that small models are most useful as interface models around large models, rather than as direct replacements.
No heat snapshots are available in the last 24 hours.