Gespeichert in:
| Hauptverfasser: | Yin, Haojie, Feng, Chengcheng, Liu, Tianyi, Zhang, Tianqi, Huang, Kaizhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.26513 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Faithful Reasoning in Comics for Small MLLMs
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026)
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026)
M3: 3D-Spatial MultiModal Memory
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
MultiModal Action Conditioned Video Generation
von: Li, Yichen, et al.
Veröffentlicht: (2025)
von: Li, Yichen, et al.
Veröffentlicht: (2025)
MultiModal Fine-tuning with Synthetic Captions
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)
ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
von: Sastry, Srikumar, et al.
Veröffentlicht: (2025)
von: Sastry, Srikumar, et al.
Veröffentlicht: (2025)
M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention
von: Panta, Sanjeev, et al.
Veröffentlicht: (2026)
von: Panta, Sanjeev, et al.
Veröffentlicht: (2026)
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
ControlEdit: A MultiModal Local Clothing Image Editing Method
von: Cheng, Di, et al.
Veröffentlicht: (2024)
von: Cheng, Di, et al.
Veröffentlicht: (2024)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
von: Wang, Qijie, et al.
Veröffentlicht: (2024)
von: Wang, Qijie, et al.
Veröffentlicht: (2024)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
von: Qiu, Han, et al.
Veröffentlicht: (2024)
von: Qiu, Han, et al.
Veröffentlicht: (2024)
MMA-Diffusion: MultiModal Attack on Diffusion Models
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
M$^2$CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation
von: Liu, Ziyuan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyuan, et al.
Veröffentlicht: (2025)
TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
Mind the Gap: Promoting Missing Modality Brain Tumor Segmentation with Alignment
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation
von: Lan, Libin, et al.
Veröffentlicht: (2025)
von: Lan, Libin, et al.
Veröffentlicht: (2025)
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
von: Li, Zijie, et al.
Veröffentlicht: (2026)
von: Li, Zijie, et al.
Veröffentlicht: (2026)
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
HAMMR: HierArchical MultiModal React agents for generic VQA
von: Castrejon, Lluis, et al.
Veröffentlicht: (2024)
von: Castrejon, Lluis, et al.
Veröffentlicht: (2024)
Rethinking Information Loss in Medical Image Segmentation with Various-sized Targets
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
MedMAP: Promoting Incomplete Multi-modal Brain Tumor Segmentation with Alignment
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
von: Zou, Heqing, et al.
Veröffentlicht: (2024)
von: Zou, Heqing, et al.
Veröffentlicht: (2024)
Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Rebalancing Multi-Label Class-Incremental Learning
von: Du, Kaile, et al.
Veröffentlicht: (2024)
von: Du, Kaile, et al.
Veröffentlicht: (2024)
MMA-DFER: MultiModal Adaptation of unimodal models for Dynamic Facial Expression Recognition in-the-wild
von: Chumachenko, Kateryna, et al.
Veröffentlicht: (2024)
von: Chumachenko, Kateryna, et al.
Veröffentlicht: (2024)
GeoSDF: Plane Geometry Diagram Synthesis via Signed Distance Field
von: Zhang, Chengrui, et al.
Veröffentlicht: (2025)
von: Zhang, Chengrui, et al.
Veröffentlicht: (2025)
TriLoRA: Integrating SVD for Advanced Style Personalization in Text-to-Image Generation
von: Feng, Chengcheng, et al.
Veröffentlicht: (2024)
von: Feng, Chengcheng, et al.
Veröffentlicht: (2024)
Frequency-enhanced Multi-granularity Context Network for Efficient Vertebrae Segmentation
von: Shi, Jian, et al.
Veröffentlicht: (2025)
von: Shi, Jian, et al.
Veröffentlicht: (2025)
MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification
von: Feng, Yingying, et al.
Veröffentlicht: (2025)
von: Feng, Yingying, et al.
Veröffentlicht: (2025)
BFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysis
von: Zhao, Weiguang, et al.
Veröffentlicht: (2025)
von: Zhao, Weiguang, et al.
Veröffentlicht: (2025)
Controlled Data Rebalancing in Multi-Task Learning for Real-World Image Super-Resolution
von: Lin, Shuchen, et al.
Veröffentlicht: (2025)
von: Lin, Shuchen, et al.
Veröffentlicht: (2025)
Vision Transformer based Random Walk for Group Re-Identification
von: Zhang, Guoqing, et al.
Veröffentlicht: (2024)
von: Zhang, Guoqing, et al.
Veröffentlicht: (2024)
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
von: Yang, Jian, et al.
Veröffentlicht: (2024)
von: Yang, Jian, et al.
Veröffentlicht: (2024)
W-Net: One-Shot Arbitrary-Style Chinese Character Generation with Deep Neural Networks
von: Jiang, Haochuan, et al.
Veröffentlicht: (2024)
von: Jiang, Haochuan, et al.
Veröffentlicht: (2024)
Rethinking Multi-domain Generalization with A General Learning Objective
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
DrFuse: Learning Disentangled Representation for Clinical Multi-Modal Fusion with Missing Modality and Modal Inconsistency
von: Yao, Wenfang, et al.
Veröffentlicht: (2024)
von: Yao, Wenfang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Faithful Reasoning in Comics for Small MLLMs
von: Feng, Chengcheng, et al.
Veröffentlicht: (2026) -
M3: 3D-Spatial MultiModal Memory
von: Zou, Xueyan, et al.
Veröffentlicht: (2025) -
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025) -
MultiModal Action Conditioned Video Generation
von: Li, Yichen, et al.
Veröffentlicht: (2025) -
MultiModal Fine-tuning with Synthetic Captions
von: Enomoto, Shohei, et al.
Veröffentlicht: (2026)