The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bozhou, Xue, Xinda, Yang, Sihan, Shi, Yang, Chen, Xinlong, Guan, Yushuo, Zhang, Yuanxing, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
by: Li, Bozhou, et al.
Published: (2026)
by: Li, Bozhou, et al.
Published: (2026)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
NormAUG: Normalization-guided Augmentation for Domain Generalization
by: Qi, Lei, et al.
Published: (2023)
by: Qi, Lei, et al.
Published: (2023)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
InfoNorm: Mutual Information Shaping of Normals for Sparse-View Reconstruction
by: Wang, Xulong, et al.
Published: (2024)
by: Wang, Xulong, et al.
Published: (2024)
Growing Visual Generative Capacity for Pre-Trained MLLMs
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Parametric $ρ$-Norm Scaling Calibration
by: Zhang, Siyuan, et al.
Published: (2024)
by: Zhang, Siyuan, et al.
Published: (2024)
Quaternion Nuclear Norm minus Frobenius Norm Minimization for color image reconstruction
by: Guo, Yu, et al.
Published: (2024)
by: Guo, Yu, et al.
Published: (2024)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows
by: Han, Yuxuan, et al.
Published: (2026)
by: Han, Yuxuan, et al.
Published: (2026)
Are Bigger Encoders Always Better in Vision Large Models?
by: Li, Bozhou, et al.
Published: (2024)
by: Li, Bozhou, et al.
Published: (2024)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
Matrix Completion Via Reweighted Logarithmic Norm Minimization
by: Wang, Zhijie, et al.
Published: (2025)
by: Wang, Zhijie, et al.
Published: (2025)
EgoNormia: Benchmarking Physical Social Norm Understanding
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
by: Liu, Tengfei, et al.
Published: (2026)
by: Liu, Tengfei, et al.
Published: (2026)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
by: Wang, Yuchi, et al.
Published: (2025)
by: Wang, Yuchi, et al.
Published: (2025)
DiffusionAD: Norm-guided One-step Denoising Diffusion for Anomaly Detection
by: Zhang, Hui, et al.
Published: (2023)
by: Zhang, Hui, et al.
Published: (2023)
Explore How to Inject Beneficial Noise in MLLMs
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
by: Liu, Xiaolin, et al.
Published: (2026)
by: Liu, Xiaolin, et al.
Published: (2026)
Low-Rank Tensor Recovery via Variational Schatten-p Quasi-Norm and Jacobian Regularization
by: Cheng, Zhengyun, et al.
Published: (2025)
by: Cheng, Zhengyun, et al.
Published: (2025)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
by: Zhu, Xuanyu, et al.
Published: (2026)
by: Zhu, Xuanyu, et al.
Published: (2026)
DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic States
by: Zhang, Bozhou, et al.
Published: (2024)
by: Zhang, Bozhou, et al.
Published: (2024)
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
by: Heng, Yongrui, et al.
Published: (2026)
by: Heng, Yongrui, et al.
Published: (2026)
Unseen Visual Anomaly Generation
by: Sun, Han, et al.
Published: (2024)
by: Sun, Han, et al.
Published: (2024)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Imagine the Unseen: Occluded Pedestrian Detection via Adversarial Feature Completion
by: Zhang, Shanshan, et al.
Published: (2024)
by: Zhang, Shanshan, et al.
Published: (2024)
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors
by: Guan, Tongkun, et al.
Published: (2023)
by: Guan, Tongkun, et al.
Published: (2023)
Similar Items
-
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025) -
Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
by: Li, Bozhou, et al.
Published: (2026) -
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025) -
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
by: Yin, Yuanyang, et al.
Published: (2024) -
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
by: Li, Bozhou, et al.
Published: (2025)