This report deploys MiniCPM-V-4.6 entirely on a 6 GB NVIDIA Tesla C2075 from 2011, using Fermi sm_20 CUDA support. The system includes a SigLIP2 vision encoder, a window-attention merger with 16x visual-token compression, and a compact hybrid gated-delta-net backbone. Reported techniques include vendor SGEMM for 8-bit weights, a chunked delta-rule implementation, zero-extra-memory attention rewrites for long context, and stage-by-stage validation against locally generated reference outputs. Prefill remained comparatively stable from 2k to 10k tokens after optimization, and end-to-end image question answering took 1.7 seconds.
No heat snapshots are available in the last 24 hours.