Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Tao, Shou, Linjun, Liu, Xuejun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning is a Modality
by: Liu, Zhiguang, et al.
Published: (2026)
by: Liu, Zhiguang, et al.
Published: (2026)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
by: Ma, Jie, et al.
Published: (2023)
by: Ma, Jie, et al.
Published: (2023)
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
by: Fan, Lin, et al.
Published: (2024)
by: Fan, Lin, et al.
Published: (2024)
A Multi-Modal Deep Learning Based Approach for House Price Prediction
by: Hasan, Md Hasebul, et al.
Published: (2024)
by: Hasan, Md Hasebul, et al.
Published: (2024)
Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities
by: Mobbs, Rebecca, et al.
Published: (2025)
by: Mobbs, Rebecca, et al.
Published: (2025)
E = T*H/(O+B): A Dimensionless Control Parameter for Mixture-of-Experts Ecology
by: Zhang, Qingjun
Published: (2026)
by: Zhang, Qingjun
Published: (2026)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
by: Chen, Zhangquan, et al.
Published: (2026)
by: Chen, Zhangquan, et al.
Published: (2026)
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
by: Tangtartharakul, Gene, et al.
Published: (2026)
by: Tangtartharakul, Gene, et al.
Published: (2026)
Visual Language Models show widespread visual deficits on neuropsychological tests
by: Tangtartharakul, Gene, et al.
Published: (2025)
by: Tangtartharakul, Gene, et al.
Published: (2025)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
by: Ma, Jie, et al.
Published: (2024)
by: Ma, Jie, et al.
Published: (2024)
Rotation-Adaptive Point Cloud Domain Generalization via Intricate Orientation Learning
by: Liu, Bangzhen, et al.
Published: (2025)
by: Liu, Bangzhen, et al.
Published: (2025)
Deep Learning for automated multi-scale functional field boundaries extraction using multi-date Sentinel-2 and PlanetScope imagery: Case Study of Netherlands and Pakistan
by: Zahid, Saba, et al.
Published: (2024)
by: Zahid, Saba, et al.
Published: (2024)
Region Mixup
by: Saha, Saptarshi, et al.
Published: (2024)
by: Saha, Saptarshi, et al.
Published: (2024)
Neuron-based explanations of neural networks sacrifice completeness and interpretability
by: Dey, Nolan, et al.
Published: (2020)
by: Dey, Nolan, et al.
Published: (2020)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
by: Fan, Lin, et al.
Published: (2026)
by: Fan, Lin, et al.
Published: (2026)
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
by: Agarwal, Rachit, et al.
Published: (2026)
by: Agarwal, Rachit, et al.
Published: (2026)
AMANet: Advancing SAR Ship Detection with Adaptive Multi-Hierarchical Attention Network
by: Ma, Xiaolin, et al.
Published: (2024)
by: Ma, Xiaolin, et al.
Published: (2024)
Representation Learning via Non-Contrastive Mutual Information
by: Guo, Zhaohan Daniel, et al.
Published: (2025)
by: Guo, Zhaohan Daniel, et al.
Published: (2025)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
Perceptual Flow Network for Visually Grounded Reasoning
by: Li, Yangfu, et al.
Published: (2026)
by: Li, Yangfu, et al.
Published: (2026)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
by: He, Jianxiang, et al.
Published: (2025)
by: He, Jianxiang, et al.
Published: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
by: Chen, Yiping, et al.
Published: (2026)
by: Chen, Yiping, et al.
Published: (2026)
Instruction-based Image Editing with Planning, Reasoning, and Generation
by: Ji, Liya, et al.
Published: (2026)
by: Ji, Liya, et al.
Published: (2026)
CausAdv: A Causal-based Framework for Detecting Adversarial Examples
by: Debbi, Hichem
Published: (2024)
by: Debbi, Hichem
Published: (2024)
Tell me why: Visual foundation models as self-explainable classifiers
by: Turbé, Hugues, et al.
Published: (2025)
by: Turbé, Hugues, et al.
Published: (2025)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
by: Huang, Yuhang, et al.
Published: (2025)
by: Huang, Yuhang, et al.
Published: (2025)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Data Efficiency and Transfer Robustness in Biomedical Image Segmentation: A Study of Redundancy and Forgetting with Cellpose
by: Zhao, Shuo, et al.
Published: (2025)
by: Zhao, Shuo, et al.
Published: (2025)
Leveraging Color Channel Independence for Improved Unsupervised Object Detection
by: Jäckl, Bastian, et al.
Published: (2024)
by: Jäckl, Bastian, et al.
Published: (2024)
Data Augmentation for Image Classification using Generative AI
by: Rahat, Fazle, et al.
Published: (2024)
by: Rahat, Fazle, et al.
Published: (2024)
Hybrid Knowledge Transfer through Attention and Logit Distillation for On-Device Vision Systems in Agricultural IoT
by: Mugisha, Stanley, et al.
Published: (2025)
by: Mugisha, Stanley, et al.
Published: (2025)
Enhancing Long-Term Re-Identification Robustness Using Synthetic Data: A Comparative Analysis
by: Pionzewski, Christian, et al.
Published: (2025)
by: Pionzewski, Christian, et al.
Published: (2025)
Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration
by: Bian, Wentao, et al.
Published: (2026)
by: Bian, Wentao, et al.
Published: (2026)
Rethinking Uncertainty in Segmentation: From Estimation to Decision
by: Maganti, Saket
Published: (2026)
by: Maganti, Saket
Published: (2026)
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
by: Wu, Qingyu, et al.
Published: (2026)
by: Wu, Qingyu, et al.
Published: (2026)
Visual Graph Question Answering with ASP and LLMs for Language Parsing
by: Bauer, Jakob Johannes, et al.
Published: (2025)
by: Bauer, Jakob Johannes, et al.
Published: (2025)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
by: Oliveira, Daniel A. P., et al.
Published: (2024)
by: Oliveira, Daniel A. P., et al.
Published: (2024)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Similar Items
-
Reasoning is a Modality
by: Liu, Zhiguang, et al.
Published: (2026) -
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
by: Ma, Jie, et al.
Published: (2023) -
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
by: Fan, Lin, et al.
Published: (2024) -
A Multi-Modal Deep Learning Based Approach for House Price Prediction
by: Hasan, Md Hasebul, et al.
Published: (2024) -
Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities
by: Mobbs, Rebecca, et al.
Published: (2025)