Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Timin, Chen, Peixian, Zhang, Mengdan, Fu, Chaoyou, Shen, Yunhang, Zhang, Yan, Zhang, Shengchuan, Zheng, Xiawu, Sun, Xing, Cao, Liujuan, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
von: Li, Xudong, et al.
Veröffentlicht: (2025)
von: Li, Xudong, et al.
Veröffentlicht: (2025)
Multi-Modal Prompt Learning on Blind Image Quality Assessment
von: Pan, Wensheng, et al.
Veröffentlicht: (2024)
von: Pan, Wensheng, et al.
Veröffentlicht: (2024)
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
von: Lin, Jianghang, et al.
Veröffentlicht: (2025)
von: Lin, Jianghang, et al.
Veröffentlicht: (2025)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise Sensitivity
von: Li, Xudong, et al.
Veröffentlicht: (2023)
von: Li, Xudong, et al.
Veröffentlicht: (2023)
Pseudo-Label Quality Decoupling and Correction for Semi-Supervised Instance Segmentation
von: Lin, Jianghang, et al.
Veröffentlicht: (2025)
von: Lin, Jianghang, et al.
Veröffentlicht: (2025)
Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models
von: Li, Xudong, et al.
Veröffentlicht: (2024)
von: Li, Xudong, et al.
Veröffentlicht: (2024)
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
von: Shen, You, et al.
Veröffentlicht: (2025)
von: Shen, You, et al.
Veröffentlicht: (2025)
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
von: Shen, Yunhang, et al.
Veröffentlicht: (2025)
von: Shen, Yunhang, et al.
Veröffentlicht: (2025)
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
von: Long, Zuwei, et al.
Veröffentlicht: (2025)
von: Long, Zuwei, et al.
Veröffentlicht: (2025)
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
von: Huang, You, et al.
Veröffentlicht: (2025)
von: Huang, You, et al.
Veröffentlicht: (2025)
HRSAM: Efficient Interactive Segmentation in High-Resolution Images
von: Huang, You, et al.
Veröffentlicht: (2024)
von: Huang, You, et al.
Veröffentlicht: (2024)
S$^2$Teacher: Step-by-step Teacher for Sparsely Annotated Oriented Object Detection
von: Lin, Yu, et al.
Veröffentlicht: (2025)
von: Lin, Yu, et al.
Veröffentlicht: (2025)
Breaking the Bias: Recalibrating the Attention of Industrial Anomaly Detection
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
Contrastive Local Manifold Learning for No-Reference Image Quality Assessment
von: Huang, Zihao, et al.
Veröffentlicht: (2024)
von: Huang, Zihao, et al.
Veröffentlicht: (2024)
Depth-Guided Semi-Supervised Instance Segmentation
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
SCOUT: Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection
von: Yan, Weiqi, et al.
Veröffentlicht: (2025)
von: Yan, Weiqi, et al.
Veröffentlicht: (2025)
FocSAM: Delving Deeply into Focused Objects in Segmenting Anything
von: Huang, You, et al.
Veröffentlicht: (2024)
von: Huang, You, et al.
Veröffentlicht: (2024)
GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane
von: Qu, Yansong, et al.
Veröffentlicht: (2024)
von: Qu, Yansong, et al.
Veröffentlicht: (2024)
Drag Your Gaussian: Effective Drag-Based Editing with Score Distillation for 3D Gaussian Splatting
von: Qu, Yansong, et al.
Veröffentlicht: (2025)
von: Qu, Yansong, et al.
Veröffentlicht: (2025)
Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation
von: Yang, Yilong, et al.
Veröffentlicht: (2026)
von: Yang, Yilong, et al.
Veröffentlicht: (2026)
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
von: Sun, Zhen, et al.
Veröffentlicht: (2025)
von: Sun, Zhen, et al.
Veröffentlicht: (2025)
HUWSOD: Holistic Self-training for Unified Weakly Supervised Object Detection
von: Cao, Liujuan, et al.
Veröffentlicht: (2024)
von: Cao, Liujuan, et al.
Veröffentlicht: (2024)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label Learning
von: Yan, Weiqi, et al.
Veröffentlicht: (2025)
von: Yan, Weiqi, et al.
Veröffentlicht: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
von: Guo, Song, et al.
Veröffentlicht: (2024)
von: Guo, Song, et al.
Veröffentlicht: (2024)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
von: Li, Lijiang, et al.
Veröffentlicht: (2026)
von: Li, Lijiang, et al.
Veröffentlicht: (2026)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text
von: Li, Xinyang, et al.
Veröffentlicht: (2024)
von: Li, Xinyang, et al.
Veröffentlicht: (2024)
Dual3D: Efficient and Consistent Text-to-3D Generation with Dual-mode Multi-view Latent Diffusion
von: Li, Xinyang, et al.
Veröffentlicht: (2024)
von: Li, Xinyang, et al.
Veröffentlicht: (2024)
Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial Inspection
von: Lu, Yilin, et al.
Veröffentlicht: (2025)
von: Lu, Yilin, et al.
Veröffentlicht: (2025)
BUFF: Bayesian Uncertainty Guided Diffusion Probabilistic Model for Single Image Super-Resolution
von: He, Zihao, et al.
Veröffentlicht: (2025)
von: He, Zihao, et al.
Veröffentlicht: (2025)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
CamoTeacher: Dual-Rotation Consistency Learning for Semi-Supervised Camouflaged Object Detection
von: Lai, Xunfa, et al.
Veröffentlicht: (2024)
von: Lai, Xunfa, et al.
Veröffentlicht: (2024)
Streaming Video Instruction Tuning
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2024)
von: Luo, Gen, et al.
Veröffentlicht: (2024)
Active-SAOOD: Active Sparsely Annotated Oriented Object Detection in Remote Sensing Images
von: Lin, Yu, et al.
Veröffentlicht: (2026)
von: Lin, Yu, et al.
Veröffentlicht: (2026)
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
von: Fu, Chaoyou, et al.
Veröffentlicht: (2025)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024) -
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
von: Li, Xudong, et al.
Veröffentlicht: (2025) -
Multi-Modal Prompt Learning on Blind Image Quality Assessment
von: Pan, Wensheng, et al.
Veröffentlicht: (2024) -
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
von: Lin, Jianghang, et al.
Veröffentlicht: (2025) -
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)