Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Dahun, Mallya, Ganesh Satish, Angelova, Anelia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Time-Scaling State-Space Models for Dense Video Captioning
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2025)
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2025)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications
von: Mallya, Ganesh, et al.
Veröffentlicht: (2025)
von: Mallya, Ganesh, et al.
Veröffentlicht: (2025)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
von: Kim, Dahun, et al.
Veröffentlicht: (2023)
von: Kim, Dahun, et al.
Veröffentlicht: (2023)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024)
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024)
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2023)
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2023)
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
von: Shao, Hao, et al.
Veröffentlicht: (2024)
von: Shao, Hao, et al.
Veröffentlicht: (2024)
OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
von: Zhang, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhang, Guanghao, et al.
Veröffentlicht: (2025)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
von: Singh, Darshan, et al.
Veröffentlicht: (2026)
von: Singh, Darshan, et al.
Veröffentlicht: (2026)
MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval
von: Ge, Xuri, et al.
Veröffentlicht: (2026)
von: Ge, Xuri, et al.
Veröffentlicht: (2026)
SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification
von: Zhang, Yunkai, et al.
Veröffentlicht: (2025)
von: Zhang, Yunkai, et al.
Veröffentlicht: (2025)
Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection
von: Lee, Sangin, et al.
Veröffentlicht: (2026)
von: Lee, Sangin, et al.
Veröffentlicht: (2026)
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
von: Wu, Jiahe, et al.
Veröffentlicht: (2026)
von: Wu, Jiahe, et al.
Veröffentlicht: (2026)
ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation
von: Tan, Jianwen, et al.
Veröffentlicht: (2025)
von: Tan, Jianwen, et al.
Veröffentlicht: (2025)
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
von: Kim, Hyeonyu, et al.
Veröffentlicht: (2025)
von: Kim, Hyeonyu, et al.
Veröffentlicht: (2025)
DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
von: Zhang, Daoan, et al.
Veröffentlicht: (2024)
von: Zhang, Daoan, et al.
Veröffentlicht: (2024)
Beyond Symbolic Solving: Multi Chain-of-Thought Voting for Geometric Reasoning in Large Language Models
von: Siddique, Md. Abu Bakor, et al.
Veröffentlicht: (2026)
von: Siddique, Md. Abu Bakor, et al.
Veröffentlicht: (2026)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
von: Liu, Xu, et al.
Veröffentlicht: (2026)
von: Liu, Xu, et al.
Veröffentlicht: (2026)
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
von: Ma, Qihang, et al.
Veröffentlicht: (2025)
von: Ma, Qihang, et al.
Veröffentlicht: (2025)
MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder
von: Erdur, Ayhan Can, et al.
Veröffentlicht: (2025)
von: Erdur, Ayhan Can, et al.
Veröffentlicht: (2025)
CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
Learning Visual Grounding from Generative Vision and Language Model
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
Interleaved-Modal Chain-of-Thought
von: Gao, Jun, et al.
Veröffentlicht: (2024)
von: Gao, Jun, et al.
Veröffentlicht: (2024)
Multimodal Chain-of-Thought Reasoning in Language Models
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
von: Chen, Tieyuan, et al.
Veröffentlicht: (2024)
von: Chen, Tieyuan, et al.
Veröffentlicht: (2024)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
von: Chen, Wei, et al.
Veröffentlicht: (2026)
von: Chen, Wei, et al.
Veröffentlicht: (2026)
Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
von: Wen, Junwei, et al.
Veröffentlicht: (2026)
von: Wen, Junwei, et al.
Veröffentlicht: (2026)
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
TumorChain: Interleaved Multimodal Chain-of-Thought Reasoning for Traceable Clinical Tumor Analysis
von: Li, Sijing, et al.
Veröffentlicht: (2026)
von: Li, Sijing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Time-Scaling State-Space Models for Dense Video Captioning
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2025) -
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025) -
Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications
von: Mallya, Ganesh, et al.
Veröffentlicht: (2025) -
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
von: Kim, Dahun, et al.
Veröffentlicht: (2023) -
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024)