Asymmetric Idiosyncrasies in Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tao, Muzi, Shi, Chufan, Wang, Huijuan, Tong, Shengbang, Ma, Xuezhe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
by: Fang, Irving, et al.
Published: (2025)
by: Fang, Irving, et al.
Published: (2025)
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)
by: Zheng, Boyang, et al.
Published: (2025)
Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models
by: Han, Xu, et al.
Published: (2024)
by: Han, Xu, et al.
Published: (2024)
Ctrl123: Consistent Novel View Synthesis via Closed-Loop Transcription
by: Zhao, Hongxiang, et al.
Published: (2024)
by: Zhao, Hongxiang, et al.
Published: (2024)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models
by: Chiu, Ming-Chang, et al.
Published: (2024)
by: Chiu, Ming-Chang, et al.
Published: (2024)
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
by: Chu, Tianzhe, et al.
Published: (2023)
by: Chu, Tianzhe, et al.
Published: (2023)
HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection
by: Zhou, Han, et al.
Published: (2026)
by: Zhou, Han, et al.
Published: (2026)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
by: Gao, Xin, et al.
Published: (2026)
by: Gao, Xin, et al.
Published: (2026)
ColorSense: A Study on Color Vision in Machine Visual Recognition
by: Chiu, Ming-Chang, et al.
Published: (2022)
by: Chiu, Ming-Chang, et al.
Published: (2022)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models
by: Tong, Yujun, et al.
Published: (2026)
by: Tong, Yujun, et al.
Published: (2026)
Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection
by: Shao, YiKang, et al.
Published: (2025)
by: Shao, YiKang, et al.
Published: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
by: Lei, Lei, et al.
Published: (2025)
by: Lei, Lei, et al.
Published: (2025)
CSHNet: A Novel Information Asymmetric Image Translation Method
by: Yang, Xi, et al.
Published: (2025)
by: Yang, Xi, et al.
Published: (2025)
Asymmetric Flow Models
by: Chen, Hansheng, et al.
Published: (2026)
by: Chen, Hansheng, et al.
Published: (2026)
LLaVA-RE: Binary Image-Text Relevancy Evaluation with Multimodal Large Language Model
by: Sun, Tao, et al.
Published: (2025)
by: Sun, Tao, et al.
Published: (2025)
LLMGA: Multimodal Large Language Model based Generation Assistant
by: Xia, Bin, et al.
Published: (2023)
by: Xia, Bin, et al.
Published: (2023)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Measuring Epistemic Humility in Multimodal Large Language Models
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
MMaDA: Multimodal Large Diffusion Language Models
by: Yang, Ling, et al.
Published: (2025)
by: Yang, Ling, et al.
Published: (2025)
AIDE: Agentically Improve Visual Language Model with Domain Experts
by: Chiu, Ming-Chang, et al.
Published: (2025)
by: Chiu, Ming-Chang, et al.
Published: (2025)
Perception-Oriented Video Frame Interpolation via Asymmetric Blending
by: Wu, Guangyang, et al.
Published: (2024)
by: Wu, Guangyang, et al.
Published: (2024)
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
by: Ma, Ji, et al.
Published: (2026)
by: Ma, Ji, et al.
Published: (2026)
Hallucination of Multimodal Large Language Models: A Survey
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models
by: Shuai, Xincheng, et al.
Published: (2024)
by: Shuai, Xincheng, et al.
Published: (2024)
Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
by: Yu, Runpeng, et al.
Published: (2025)
by: Yu, Runpeng, et al.
Published: (2025)
Docopilot: Improving Multimodal Models for Document-Level Understanding
by: Duan, Yuchen, et al.
Published: (2025)
by: Duan, Yuchen, et al.
Published: (2025)
FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models
by: Wang, Hongyang, et al.
Published: (2025)
by: Wang, Hongyang, et al.
Published: (2025)
RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion
by: Wang, Ruofan, et al.
Published: (2025)
by: Wang, Ruofan, et al.
Published: (2025)
Similar Items
-
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
by: Yang, Cheng, et al.
Published: (2026) -
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026) -
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024) -
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
by: Fang, Irving, et al.
Published: (2025) -
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)