The paper introduces VLT, a multimodal foundation model for industrial intelligence that jointly models time-series signals, frequency-spectrum visual representations, and textual knowledge. Its design combines a Time-aware Mixture-of-Experts (Time-MoE), a Frequency-Text Augmented Learner, and a time-centric gradient alignment mechanism. The authors position the frequency spectrum as a visual bridge between continuous signals and discrete language semantics. According to the abstract, experiments across multiple industrial datasets show improved robustness and generalization in few-shot, noisy, and incomplete-modality settings, although detailed datasets, baselines, and numerical results require inspection of the full paper.
No heat snapshots are available in the last 24 hours.