This paper presents an Integer Linear Programming (ILP) framework for partitioning ML workloads across CPUs and Computing-in-Memory (CIM) accelerators. Unlike prior approaches described by the authors, it incorporates RRAM capacity, write latency, and endurance constraints, while modeling parallelism, low-level architectural effects, and CPU execution as a complementary resource. The objective is to minimize end-to-end inference latency. The abstract reports speedups of up to 30.9x over an edge CPU and 7.3x over a high-performance CPU for heterogeneous CPU-CIM execution. A design-space exploration further examines implications for future CIM accelerator architectures.
No heat snapshots are available in the last 24 hours.