From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, Hang, Zhang, Weijia, Wang, Fei, Shao, Zezhi, Liu, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
by: Zhang, Ming, et al.
Published: (2024)
by: Zhang, Ming, et al.
Published: (2024)
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability
by: Zhao, Fei, et al.
Published: (2024)
by: Zhao, Fei, et al.
Published: (2024)
Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
by: Sun, Hao, et al.
Published: (2026)
by: Sun, Hao, et al.
Published: (2026)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
by: Liu, Zixuan, et al.
Published: (2025)
by: Liu, Zixuan, et al.
Published: (2025)
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
by: Cai, Rui, et al.
Published: (2024)
by: Cai, Rui, et al.
Published: (2024)
GeoMM: On Geodesic Perspective for Multi-modal Learning
by: Mei, Shibin, et al.
Published: (2025)
by: Mei, Shibin, et al.
Published: (2025)
ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM Agents
by: Li, Zechen, et al.
Published: (2025)
by: Li, Zechen, et al.
Published: (2025)
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
by: Feng, Jie, et al.
Published: (2025)
by: Feng, Jie, et al.
Published: (2025)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
by: Zhang, Kaichen, et al.
Published: (2024)
by: Zhang, Kaichen, et al.
Published: (2024)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
by: Uehara, Kohei, et al.
Published: (2024)
by: Uehara, Kohei, et al.
Published: (2024)
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
by: Liu, Zhaochen, et al.
Published: (2026)
by: Liu, Zhaochen, et al.
Published: (2026)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
by: Li, Zhaowei, et al.
Published: (2024)
by: Li, Zhaowei, et al.
Published: (2024)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
by: Bhattacharya, Amartya
Published: (2026)
by: Bhattacharya, Amartya
Published: (2026)
FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues
by: Li, Shuang, et al.
Published: (2024)
by: Li, Shuang, et al.
Published: (2024)
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
by: Zhang, Lu, et al.
Published: (2024)
by: Zhang, Lu, et al.
Published: (2024)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
by: Wei, Yanbin, et al.
Published: (2026)
by: Wei, Yanbin, et al.
Published: (2026)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
by: Chen, Junjie, et al.
Published: (2024)
by: Chen, Junjie, et al.
Published: (2024)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
MMBench: Is Your Multi-modal Model an All-around Player?
by: Liu, Yuan, et al.
Published: (2023)
by: Liu, Yuan, et al.
Published: (2023)
Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
by: Zhong, Xinliu, et al.
Published: (2025)
by: Zhong, Xinliu, et al.
Published: (2025)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
CycleChart: A Unified Consistency-Based Learning Framework for Bidirectional Chart Understanding and Generation
by: Deng, Dazhen, et al.
Published: (2025)
by: Deng, Dazhen, et al.
Published: (2025)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
by: Lin, Bingqian, et al.
Published: (2024)
by: Lin, Bingqian, et al.
Published: (2024)
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
by: Du, Yong, et al.
Published: (2025)
by: Du, Yong, et al.
Published: (2025)
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
Can Multi-modal (reasoning) LLMs detect document manipulation?
by: Liang, Zisheng, et al.
Published: (2025)
by: Liang, Zisheng, et al.
Published: (2025)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
by: Lin, Jingyang, et al.
Published: (2023)
by: Lin, Jingyang, et al.
Published: (2023)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
by: Qin, Lixiong, et al.
Published: (2025)
by: Qin, Lixiong, et al.
Published: (2025)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
by: Madaan, Divyam, et al.
Published: (2025)
by: Madaan, Divyam, et al.
Published: (2025)
Similar Items
-
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
by: Zhang, Ming, et al.
Published: (2024) -
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability
by: Zhao, Fei, et al.
Published: (2024) -
Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering
by: Li, Bo, et al.
Published: (2025) -
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
by: Sun, Hao, et al.
Published: (2026) -
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
by: Zhang, Yichi, et al.
Published: (2025)