GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Fenghua, Wang, Jinxiang, Wang, Sen, Huang, Zi, Li, Xue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretable Multimodal Misinformation Detection with Logic Reasoning
von: Liu, Hui, et al.
Veröffentlicht: (2023)
von: Liu, Hui, et al.
Veröffentlicht: (2023)
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
von: Luo, Wen, et al.
Veröffentlicht: (2024)
von: Luo, Wen, et al.
Veröffentlicht: (2024)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
History-Guided Iterative Visual Reasoning with Self-Correction
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
PTA: Enhancing Multimodal Sentiment Analysis through Pipelined Prediction and Translation-based Alignment
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
Towards Robust Multimodal Sentiment Analysis with Incomplete Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
von: Zhang, Hanlei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2025)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
SemEval-2024 Task 3: Multimodal Emotion Cause Analysis in Conversations
von: Wang, Fanfan, et al.
Veröffentlicht: (2024)
von: Wang, Fanfan, et al.
Veröffentlicht: (2024)
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Cardiverse: Harnessing LLMs for Novel Card Game Prototyping
von: Li, Danrui, et al.
Veröffentlicht: (2025)
von: Li, Danrui, et al.
Veröffentlicht: (2025)
Knowledge-Guided Dynamic Modality Attention Fusion Framework for Multimodal Sentiment Analysis
von: Feng, Xinyu, et al.
Veröffentlicht: (2024)
von: Feng, Xinyu, et al.
Veröffentlicht: (2024)
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
von: Chen, Yanxi, et al.
Veröffentlicht: (2025)
von: Chen, Yanxi, et al.
Veröffentlicht: (2025)
A Survey on Image-text Multimodal Models
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning
von: Zheng, Changmeng, et al.
Veröffentlicht: (2024)
von: Zheng, Changmeng, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Affective Analysis with Learned Live Comment Features
von: Deng, Zhaoyuan, et al.
Veröffentlicht: (2024)
von: Deng, Zhaoyuan, et al.
Veröffentlicht: (2024)
A Survey of Generative Categories and Techniques in Multimodal Generative Models
von: Han, Longzhen, et al.
Veröffentlicht: (2025)
von: Han, Longzhen, et al.
Veröffentlicht: (2025)
MultimodalHugs: Enabling Sign Language Processing in Hugging Face
von: Sant, Gerard, et al.
Veröffentlicht: (2025)
von: Sant, Gerard, et al.
Veröffentlicht: (2025)
Towards Better Text-to-Image Generation Alignment via Attention Modulation
von: Wu, Yihang, et al.
Veröffentlicht: (2024)
von: Wu, Yihang, et al.
Veröffentlicht: (2024)
A Multimodal Framework for Explainable Evaluation of Soft Skills in Educational Environments
von: Guerrero-Sosa, Jared D. T., et al.
Veröffentlicht: (2025)
von: Guerrero-Sosa, Jared D. T., et al.
Veröffentlicht: (2025)
Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI
von: Xue, Dong, et al.
Veröffentlicht: (2025)
von: Xue, Dong, et al.
Veröffentlicht: (2025)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
von: Saha, Anisha, et al.
Veröffentlicht: (2026)
von: Saha, Anisha, et al.
Veröffentlicht: (2026)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models
von: Wang, Qingni, et al.
Veröffentlicht: (2024)
von: Wang, Qingni, et al.
Veröffentlicht: (2024)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis
von: Wang, Pan, et al.
Veröffentlicht: (2024)
von: Wang, Pan, et al.
Veröffentlicht: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2025)
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpretable Multimodal Misinformation Detection with Logic Reasoning
von: Liu, Hui, et al.
Veröffentlicht: (2023) -
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
von: Luo, Wen, et al.
Veröffentlicht: (2024) -
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025) -
History-Guided Iterative Visual Reasoning with Self-Correction
von: Yang, Xinglong, et al.
Veröffentlicht: (2026) -
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)