FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shuang, Wang, Jiahua, Wen, Lijie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
von: Sun, Hao, et al.
Veröffentlicht: (2026)
von: Sun, Hao, et al.
Veröffentlicht: (2026)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
von: zhang, Kaixin, et al.
Veröffentlicht: (2026)
von: zhang, Kaixin, et al.
Veröffentlicht: (2026)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
Learning Speaker-Invariant Visual Features for Lipreading
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
Beyond the Textual: Generating Coherent Visual Options for MCQs
von: Wang, Wanqiang, et al.
Veröffentlicht: (2025)
von: Wang, Wanqiang, et al.
Veröffentlicht: (2025)
Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing
von: Wijegunarathna, Kalana, et al.
Veröffentlicht: (2025)
von: Wijegunarathna, Kalana, et al.
Veröffentlicht: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
CLEAR: Character Unlearning in Textual and Visual Modalities
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models
von: Zhang, Letian, et al.
Veröffentlicht: (2023)
von: Zhang, Letian, et al.
Veröffentlicht: (2023)
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions
von: Qin, Bowen, et al.
Veröffentlicht: (2025)
von: Qin, Bowen, et al.
Veröffentlicht: (2025)
Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning
von: Madaan, Divyam, et al.
Veröffentlicht: (2024)
von: Madaan, Divyam, et al.
Veröffentlicht: (2024)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
von: Feng, Jie, et al.
Veröffentlicht: (2025)
von: Feng, Jie, et al.
Veröffentlicht: (2025)
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
von: Burapacheep, Jirayu, et al.
Veröffentlicht: (2024)
von: Burapacheep, Jirayu, et al.
Veröffentlicht: (2024)
MMBench: Is Your Multi-modal Model an All-around Player?
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
von: Sun, Hao, et al.
Veröffentlicht: (2026) -
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
von: Uehara, Kohei, et al.
Veröffentlicht: (2024) -
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
von: Wang, Yifan, et al.
Veröffentlicht: (2026) -
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024) -
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)