Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

ATSInfer: Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

First seen · 7/11/2026, 04:00 PMLatest activity · 7/11/2026, 04:00 PM

This paper introduces ATSInfer, a hybrid CPU-GPU inference system for consumer devices that schedules tensors rather than entire layers or experts. It combines static tensor placement, load-aware dynamic transfers, and asynchronous coordination of storage, data movement, and computation. The authors evaluate the implementation on representative consumer platforms with both dense and mixture-of-experts models. According to the abstract, ATSInfer improves prefill throughput by up to 1.94x and decode throughput by up to 3.29x over existing systems, while also increasing GPU utilization and improving PCIe bandwidth usage.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/11, 04:00 PMnot independentRepresentative
    ATSInfer: Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices