Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yamaguchi, Shin'ya, Nishida, Kosuke, Chijiwa, Daiki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-shot Concept Bottleneck Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
Explanation Bottleneck Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024)
Adaptive Random Feature Regularization on Fine-tuning Deep Neural Networks
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024)
Parallel In-context Learning for Large Vision Language Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2026)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2026)
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)
Learning Robust Convolutional Neural Networks with Relevant Feature Focusing via Explanations
von: Adachi, Kazuki, et al.
Veröffentlicht: (2022)
von: Adachi, Kazuki, et al.
Veröffentlicht: (2022)
Machine Learning Modeling for Multi-order Human Visual Motion Processing
von: Sun, Zitang, et al.
Veröffentlicht: (2025)
von: Sun, Zitang, et al.
Veröffentlicht: (2025)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
von: Li, Yuan, et al.
Veröffentlicht: (2026)
von: Li, Yuan, et al.
Veröffentlicht: (2026)
Toward Data Efficient Model Merging between Different Datasets without Performance Degradation
von: Yamada, Masanori, et al.
Veröffentlicht: (2023)
von: Yamada, Masanori, et al.
Veröffentlicht: (2023)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
HAPI: A Model for Learning Robot Facial Expressions from Human Preferences
von: Yang, Dongsheng, et al.
Veröffentlicht: (2025)
von: Yang, Dongsheng, et al.
Veröffentlicht: (2025)
Transfer Learning with Pre-trained Conditional Generative Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2022)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2022)
Interleaved-Modal Chain-of-Thought
von: Gao, Jun, et al.
Veröffentlicht: (2024)
von: Gao, Jun, et al.
Veröffentlicht: (2024)
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
von: Yerramilli, Sahiti, et al.
Veröffentlicht: (2025)
von: Yerramilli, Sahiti, et al.
Veröffentlicht: (2025)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
von: Mena, Francisco, et al.
Veröffentlicht: (2025)
von: Mena, Francisco, et al.
Veröffentlicht: (2025)
MultiModal Fine-tuning with Synthetic Captions
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
von: Chen, Leon Liangyu, et al.
Veröffentlicht: (2026)
von: Chen, Leon Liangyu, et al.
Veröffentlicht: (2026)
FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models
von: Sahili, Zahraa Al, et al.
Veröffentlicht: (2024)
von: Sahili, Zahraa Al, et al.
Veröffentlicht: (2024)
Rethinking Chain-of-Thought Reasoning for Videos
von: Zhong, Yiwu, et al.
Veröffentlicht: (2025)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2025)
Guiding Perception-Reasoning Closer to Human in Blind Image Quality Assessment
von: Li, Yuan, et al.
Veröffentlicht: (2025)
von: Li, Yuan, et al.
Veröffentlicht: (2025)
Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
von: Cao, Xusheng, et al.
Veröffentlicht: (2025)
von: Cao, Xusheng, et al.
Veröffentlicht: (2025)
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Ensuring Correct Predictions With Correct Rationales
von: Li, Tang, et al.
Veröffentlicht: (2024)
von: Li, Tang, et al.
Veröffentlicht: (2024)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
von: Qin, Yiming, et al.
Veröffentlicht: (2025)
von: Qin, Yiming, et al.
Veröffentlicht: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
CPA-Enhancer: Chain-of-Thought Prompted Adaptive Enhancer for Object Detection under Unknown Degradations
von: Zhang, Yuwei, et al.
Veröffentlicht: (2024)
von: Zhang, Yuwei, et al.
Veröffentlicht: (2024)
Visual Hallucinations of Multi-modal Large Language Models
von: Huang, Wen, et al.
Veröffentlicht: (2024)
von: Huang, Wen, et al.
Veröffentlicht: (2024)
PRIMEDrive-CoT: A Precognitive Chain-of-Thought Framework for Uncertainty-Aware Object Interaction in Driving Scene Scenario
von: Mandalika, Sriram, et al.
Veröffentlicht: (2025)
von: Mandalika, Sriram, et al.
Veröffentlicht: (2025)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Generative Multi-modal Models are Good Class-Incremental Learners
von: Cao, Xusheng, et al.
Veröffentlicht: (2024)
von: Cao, Xusheng, et al.
Veröffentlicht: (2024)
Multi-modal Vision Pre-training for Medical Image Analysis
von: Rui, Shaohao, et al.
Veröffentlicht: (2024)
von: Rui, Shaohao, et al.
Veröffentlicht: (2024)
Skin Lesion Phenotyping via Nested Multi-modal Contrastive Learning
von: Christopoulos, Dionysis, et al.
Veröffentlicht: (2025)
von: Christopoulos, Dionysis, et al.
Veröffentlicht: (2025)
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
von: Camuffo, Elena, et al.
Veröffentlicht: (2025)
von: Camuffo, Elena, et al.
Veröffentlicht: (2025)
Test-time Adaptation Meets Image Enhancement: Improving Accuracy via Uncertainty-aware Logit Switching
von: Enomoto, Shohei, et al.
Veröffentlicht: (2024)
von: Enomoto, Shohei, et al.
Veröffentlicht: (2024)
Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
von: Karimi, Ehsan, et al.
Veröffentlicht: (2025)
von: Karimi, Ehsan, et al.
Veröffentlicht: (2025)
VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments
von: Yang, Bufang, et al.
Veröffentlicht: (2024)
von: Yang, Bufang, et al.
Veröffentlicht: (2024)
CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly Detection
von: Wang, Xiaolei, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zero-shot Concept Bottleneck Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025) -
Explanation Bottleneck Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024) -
Adaptive Random Feature Regularization on Fine-tuning Deep Neural Networks
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2024) -
Parallel In-context Learning for Large Vision Language Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2026) -
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2025)