The paper presents Cost-Governed RAG, combining the TurboVec vector index with a multi-tenant LLM governance gateway to attribute embedding, vector-memory, similarity-computation, and generation costs to individual tenants. In a simulation with 100 tenants and 10 million vectors distributed according to a log-normal size distribution, the system reports 99.96% end-to-end attribution accuracy and telemetry overhead below 0.04% of query latency. Under the pricing assumptions described in Section IV, it reports 3.1x to 9.0x lower retrieval infrastructure cost than managed vector database services. The paper also explores whether codebook-oblivious quantization reduces shared-codebook leakage exposure.
No heat snapshots are available in the last 24 hours.