PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Dongxu, Sun, Yiding, Li, Pengcheng, Liu, Yumou, Lin, Hongqiang, Xu, Haoran, Mu, Xiaoxuan, Lin, Liang, Yan, Wenbiao, Yang, Ning, Fang, Chaowei, Zhao, Juanjuan, Zhu, Jihua, He, Conghui, Tan, Cheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024)
von: Wang, Sen, et al.
Veröffentlicht: (2024)
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment
von: Hao, Fangyu, et al.
Veröffentlicht: (2026)
von: Hao, Fangyu, et al.
Veröffentlicht: (2026)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
Not All Queries Need Deep Thought: CoFiCot for Adaptive Coarse-to-fine Stateful Refinement
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark
von: Gao, Changsheng, et al.
Veröffentlicht: (2024)
von: Gao, Changsheng, et al.
Veröffentlicht: (2024)
PC-JND: Subjective Study and Dataset on Just Noticeable Difference for Point Clouds in 6DoF Virtual Reality
von: Fan, Chunling, et al.
Veröffentlicht: (2025)
von: Fan, Chunling, et al.
Veröffentlicht: (2025)
Deep Mamba Multi-modal Learning
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
von: Fang, Pengcheng, et al.
Veröffentlicht: (2026)
von: Fang, Pengcheng, et al.
Veröffentlicht: (2026)
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
von: Yang, An, et al.
Veröffentlicht: (2025)
von: Yang, An, et al.
Veröffentlicht: (2025)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
Efficient and Accurate Image Provenance Analysis: A Scalable Pipeline for Large-scale Images
von: Lai, Jiewei, et al.
Veröffentlicht: (2025)
von: Lai, Jiewei, et al.
Veröffentlicht: (2025)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
Towards Structure-aware Model for Multi-modal Knowledge Graph Completion
von: Li, Linyu, et al.
Veröffentlicht: (2025)
von: Li, Linyu, et al.
Veröffentlicht: (2025)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Benchmarks: In the Era of Large AI Models
von: Li, Lin, et al.
Veröffentlicht: (2024)
von: Li, Lin, et al.
Veröffentlicht: (2024)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
An Emotion Recognition Framework via Cross-modal Alignment of EEG and Eye Movement Data
von: Wang, Jianlu, et al.
Veröffentlicht: (2025)
von: Wang, Jianlu, et al.
Veröffentlicht: (2025)
Characterizing Multimedia Information Environment through Multi-modal Clustering of YouTube Videos
von: Yousefi, Niloofar, et al.
Veröffentlicht: (2024)
von: Yousefi, Niloofar, et al.
Veröffentlicht: (2024)
High-level Codes and Fine-grained Weights for Online Multi-modal Hashing Retrieval
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
PointPCA: Point Cloud Objective Quality Assessment Using PCA-Based Descriptors
von: Alexiou, Evangelos, et al.
Veröffentlicht: (2021)
von: Alexiou, Evangelos, et al.
Veröffentlicht: (2021)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction
von: Jiang, Yueheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yueheng, et al.
Veröffentlicht: (2025)
Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
von: Zhao, Xianbing, et al.
Veröffentlicht: (2025)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2025)
BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind
von: Mao, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Mao, Yuanyuan, et al.
Veröffentlicht: (2024)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
EidetiCom: A Cross-modal Brain-Computer Semantic Communication Paradigm for Decoding Visual Perception
von: Zheng, Linfeng, et al.
Veröffentlicht: (2024)
von: Zheng, Linfeng, et al.
Veröffentlicht: (2024)
Challenging Dataset and Multi-modal Gated Mixture of Experts Model for Remote Sensing Copy-Move Forgery Understanding
von: Zhang, Ze, et al.
Veröffentlicht: (2025)
von: Zhang, Ze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026) -
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025) -
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024) -
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
von: Shu, Dong, et al.
Veröffentlicht: (2025) -
LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment
von: Hao, Fangyu, et al.
Veröffentlicht: (2026)