Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Sixun, Fan, Wei, Wu, Teresa, Fu, Yanjie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Textualize Visual Prompt for Image Editing via Diffusion Bridge
by: Xu, Pengcheng, et al.
Published: (2025)
by: Xu, Pengcheng, et al.
Published: (2025)
TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting
by: Jiang, Lingyu, et al.
Published: (2025)
by: Jiang, Lingyu, et al.
Published: (2025)
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
by: Góral, Gracjan, et al.
Published: (2024)
by: Góral, Gracjan, et al.
Published: (2024)
Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
by: Zhong, Siru, et al.
Published: (2025)
by: Zhong, Siru, et al.
Published: (2025)
ViTime: Foundation Model for Time Series Forecasting Powered by Vision Intelligence
by: Yang, Luoxiao, et al.
Published: (2024)
by: Yang, Luoxiao, et al.
Published: (2024)
Learning Temporal Saliency for Time Series Forecasting with Cross-Scale Attention
by: Delibasoglu, Ibrahim, et al.
Published: (2025)
by: Delibasoglu, Ibrahim, et al.
Published: (2025)
Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution Detection
by: Chen, Jiankang, et al.
Published: (2024)
by: Chen, Jiankang, et al.
Published: (2024)
VIFO: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion
by: Wang, Yanlong, et al.
Published: (2025)
by: Wang, Yanlong, et al.
Published: (2025)
OccamVTS: Distilling Vision Models to 1% Parameters for Time Series Forecasting
by: Lyu, Sisuo, et al.
Published: (2025)
by: Lyu, Sisuo, et al.
Published: (2025)
Probabilistic NDVI Forecasting from Sparse Satellite Time Series and Weather Covariates
by: Iele, Irene, et al.
Published: (2026)
by: Iele, Irene, et al.
Published: (2026)
VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters
by: Chen, Mouxiang, et al.
Published: (2024)
by: Chen, Mouxiang, et al.
Published: (2024)
EEO-TFV: Escape-Explore Optimizer for Web-Scale Time-Series Forecasting and Vision Analysis
by: Wang, Hua, et al.
Published: (2026)
by: Wang, Hua, et al.
Published: (2026)
ZOTTA: Test-Time Adaptation with Gradient-Free Zeroth-Order Optimization
by: Zhang, Ronghao, et al.
Published: (2026)
by: Zhang, Ronghao, et al.
Published: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
by: Sun, Zhengyang, et al.
Published: (2026)
by: Sun, Zhengyang, et al.
Published: (2026)
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
Mitigating Data Redundancy to Revitalize Transformer-based Long-Term Time Series Forecasting System
by: Li, Mingjie, et al.
Published: (2022)
by: Li, Mingjie, et al.
Published: (2022)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
by: Wang, Chuhan, et al.
Published: (2026)
by: Wang, Chuhan, et al.
Published: (2026)
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Prompting Forgetting: Unlearning in GANs via Textual Guidance
by: Nagasubramaniam, Piyush, et al.
Published: (2025)
by: Nagasubramaniam, Piyush, et al.
Published: (2025)
Forecasting as Rendering: A 2D Gaussian Splatting Framework for Time Series Forecasting
by: Wang, Yixin, et al.
Published: (2026)
by: Wang, Yixin, et al.
Published: (2026)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
by: Han, Junlin, et al.
Published: (2025)
by: Han, Junlin, et al.
Published: (2025)
Differential-Integral Neural Operator for Long-Term Turbulence Forecasting
by: Wu, Hao, et al.
Published: (2025)
by: Wu, Hao, et al.
Published: (2025)
Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning
by: Zhou, Hongkuan, et al.
Published: (2025)
by: Zhou, Hongkuan, et al.
Published: (2025)
Aligning Visual Contrastive learning models via Preference Optimization
by: Afzali, Amirabbas, et al.
Published: (2024)
by: Afzali, Amirabbas, et al.
Published: (2024)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
by: Dong, Sixun, et al.
Published: (2025)
by: Dong, Sixun, et al.
Published: (2025)
Human-Aligned Image Models Improve Visual Decoding from the Brain
by: Rajabi, Nona, et al.
Published: (2025)
by: Rajabi, Nona, et al.
Published: (2025)
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
Improving Position Encoding of Transformers for Multivariate Time Series Classification
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
by: Gao, Yunhe, et al.
Published: (2024)
by: Gao, Yunhe, et al.
Published: (2024)
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
by: Peng, Yibo, et al.
Published: (2026)
by: Peng, Yibo, et al.
Published: (2026)
EmbodiTTA: Resource-Efficient Test-Time Adaptation for Embodied Visual Systems
by: Ma, Xiao, et al.
Published: (2025)
by: Ma, Xiao, et al.
Published: (2025)
Don't Get Me Wrong: How to Apply Deep Visual Interpretations to Time Series
by: Loeffler, Christoffer, et al.
Published: (2022)
by: Loeffler, Christoffer, et al.
Published: (2022)
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
by: Sabour, Amirmojtaba, et al.
Published: (2025)
by: Sabour, Amirmojtaba, et al.
Published: (2025)
IMTS is Worth Time $\times$ Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series Prediction
by: Hu, Zhangyi, et al.
Published: (2025)
by: Hu, Zhangyi, et al.
Published: (2025)
Highlight Every Step: Knowledge Distillation via Collaborative Teaching
by: Zhao, Haoran, et al.
Published: (2019)
by: Zhao, Haoran, et al.
Published: (2019)
Similar Items
-
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025) -
Textualize Visual Prompt for Image Editing via Diffusion Bridge
by: Xu, Pengcheng, et al.
Published: (2025) -
TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting
by: Jiang, Lingyu, et al.
Published: (2025) -
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
by: Góral, Gracjan, et al.
Published: (2024) -
Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
by: Zhong, Siru, et al.
Published: (2025)