Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

First seen · 7/16/2026, 12:48 PMLatest activity · 7/16/2026, 12:48 PM

This report deploys MiniCPM-V-4.6 entirely on a 6 GB NVIDIA Tesla C2075 from 2011, using Fermi sm_20 CUDA support. The system includes a SigLIP2 vision encoder, a window-attention merger with 16x visual-token compression, and a compact hybrid gated-delta-net backbone. Reported techniques include vendor SGEMM for 8-bit weights, a chunked delta-rule implementation, zero-extra-memory attention rewrites for long context, and stage-by-stage validation against locally generated reference outputs. Prefill remained comparatively stable from 2k to 10k tokens after optimization, and end-to-end image question answering took 1.7 seconds.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/16, 12:48 PMnot independentRepresentative
    A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi