SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Yuanyang, Zhao, Yaqi, Zhang, Yajie, Zhang, Yuanxing, Lin, Ke, Wang, Jiahao, Tao, Xin, Wan, Pengfei, Zhang, Wentao, Zhao, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
by: Zhao, Yaqi, et al.
Published: (2024)
by: Zhao, Yaqi, et al.
Published: (2024)
Towards Precise Scaling Laws for Video Diffusion Transformers
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
Self-Supervised Visual Preference Alignment
by: Zhu, Ke, et al.
Published: (2024)
by: Zhu, Ke, et al.
Published: (2024)
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
by: Jiang, Pengfei, et al.
Published: (2025)
by: Jiang, Pengfei, et al.
Published: (2025)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
T-REG: Preference Optimization with Token-Level Reward Regularization
by: Zhou, Wenxuan, et al.
Published: (2024)
by: Zhou, Wenxuan, et al.
Published: (2024)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025)
by: Zou, Xin, et al.
Published: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Beyond Textual Context: Structural Graph Encoding with Adaptive Space Alignment to alleviate the hallucination of LLMs
by: Zhang, Yifang, et al.
Published: (2025)
by: Zhang, Yifang, et al.
Published: (2025)
REAL: Response Embedding-based Alignment for LLMs
by: Zhang, Honggen, et al.
Published: (2024)
by: Zhang, Honggen, et al.
Published: (2024)
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Token-Level Inference-Time Alignment for Vision-Language Models
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models
by: Yin, Yuanyang, et al.
Published: (2026)
by: Yin, Yuanyang, et al.
Published: (2026)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
by: He, Wei, et al.
Published: (2024)
by: He, Wei, et al.
Published: (2024)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
by: Zhu, Xuanyu, et al.
Published: (2026)
by: Zhu, Xuanyu, et al.
Published: (2026)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
by: Zhang, Christine, et al.
Published: (2026)
by: Zhang, Christine, et al.
Published: (2026)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Enhancing Visual Continual Learning with Language-Guided Supervision
by: Ni, Bolin, et al.
Published: (2024)
by: Ni, Bolin, et al.
Published: (2024)
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
by: Zhang, Songming, et al.
Published: (2025)
by: Zhang, Songming, et al.
Published: (2025)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
by: Zhong, Tianxiong, et al.
Published: (2025)
by: Zhong, Tianxiong, et al.
Published: (2025)
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
by: Zhao, Qiyan, et al.
Published: (2026)
by: Zhao, Qiyan, et al.
Published: (2026)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
by: Wei, Yanbin, et al.
Published: (2024)
by: Wei, Yanbin, et al.
Published: (2024)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
by: Zhang, Qizhe, et al.
Published: (2025)
by: Zhang, Qizhe, et al.
Published: (2025)
Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
by: Zhao, Yaqi, et al.
Published: (2025)
by: Zhao, Yaqi, et al.
Published: (2025)
Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning
by: Wang, Yating, et al.
Published: (2026)
by: Wang, Yating, et al.
Published: (2026)
Supervised Gromov-Wasserstein Optimal Transport
by: Cang, Zixuan, et al.
Published: (2024)
by: Cang, Zixuan, et al.
Published: (2024)
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
by: Tan, Yifan, et al.
Published: (2026)
by: Tan, Yifan, et al.
Published: (2026)
Similar Items
-
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
by: Zhao, Yaqi, et al.
Published: (2024) -
Towards Precise Scaling Laws for Video Diffusion Transformers
by: Yin, Yuanyang, et al.
Published: (2024) -
Self-Supervised Visual Preference Alignment
by: Zhu, Ke, et al.
Published: (2024) -
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025) -
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
by: Li, Qi, et al.
Published: (2026)