SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Yin, Yuanyang, Zhao, Yaqi, Zhang, Yajie, Zhang, Yuanxing, Lin, Ke, Wang, Jiahao, Tao, Xin, Wan, Pengfei, Zhang, Wentao, Zhao, Feng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
di: Zhao, Yaqi, et al.
Pubblicazione: (2024)
di: Zhao, Yaqi, et al.
Pubblicazione: (2024)
Towards Precise Scaling Laws for Video Diffusion Transformers
di: Yin, Yuanyang, et al.
Pubblicazione: (2024)
di: Yin, Yuanyang, et al.
Pubblicazione: (2024)
Self-Supervised Visual Preference Alignment
di: Zhu, Ke, et al.
Pubblicazione: (2024)
di: Zhu, Ke, et al.
Pubblicazione: (2024)
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
di: Li, Bozhou, et al.
Pubblicazione: (2025)
di: Li, Bozhou, et al.
Pubblicazione: (2025)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
di: Li, Qi, et al.
Pubblicazione: (2026)
di: Li, Qi, et al.
Pubblicazione: (2026)
VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
di: Jiang, Pengfei, et al.
Pubblicazione: (2025)
di: Jiang, Pengfei, et al.
Pubblicazione: (2025)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
di: Li, Bozhou, et al.
Pubblicazione: (2025)
di: Li, Bozhou, et al.
Pubblicazione: (2025)
FG-CLIP: Fine-Grained Visual and Textual Alignment
di: Xie, Chunyu, et al.
Pubblicazione: (2025)
di: Xie, Chunyu, et al.
Pubblicazione: (2025)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
di: Tong, Chengzhuo, et al.
Pubblicazione: (2026)
di: Tong, Chengzhuo, et al.
Pubblicazione: (2026)
T-REG: Preference Optimization with Token-Level Reward Regularization
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
di: Zou, Xin, et al.
Pubblicazione: (2025)
di: Zou, Xin, et al.
Pubblicazione: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
di: Ye, Zixuan, et al.
Pubblicazione: (2025)
di: Ye, Zixuan, et al.
Pubblicazione: (2025)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
di: Chen, Yuhao, et al.
Pubblicazione: (2026)
di: Chen, Yuhao, et al.
Pubblicazione: (2026)
Beyond Textual Context: Structural Graph Encoding with Adaptive Space Alignment to alleviate the hallucination of LLMs
di: Zhang, Yifang, et al.
Pubblicazione: (2025)
di: Zhang, Yifang, et al.
Pubblicazione: (2025)
REAL: Response Embedding-based Alignment for LLMs
di: Zhang, Honggen, et al.
Pubblicazione: (2024)
di: Zhang, Honggen, et al.
Pubblicazione: (2024)
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
di: Zhang, Xin, et al.
Pubblicazione: (2026)
di: Zhang, Xin, et al.
Pubblicazione: (2026)
Token-Level Inference-Time Alignment for Vision-Language Models
di: Chen, Kejia, et al.
Pubblicazione: (2025)
di: Chen, Kejia, et al.
Pubblicazione: (2025)
Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models
di: Yin, Yuanyang, et al.
Pubblicazione: (2026)
di: Yin, Yuanyang, et al.
Pubblicazione: (2026)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
di: He, Wei, et al.
Pubblicazione: (2024)
di: He, Wei, et al.
Pubblicazione: (2024)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
di: Wang, Qixun, et al.
Pubblicazione: (2025)
di: Wang, Qixun, et al.
Pubblicazione: (2025)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
di: Zhang, Yang, et al.
Pubblicazione: (2026)
di: Zhang, Yang, et al.
Pubblicazione: (2026)
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
di: Zhang, Christine, et al.
Pubblicazione: (2026)
di: Zhang, Christine, et al.
Pubblicazione: (2026)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
di: Ou, Siqu, et al.
Pubblicazione: (2026)
di: Ou, Siqu, et al.
Pubblicazione: (2026)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
di: Zhu, Jie, et al.
Pubblicazione: (2026)
di: Zhu, Jie, et al.
Pubblicazione: (2026)
Enhancing Visual Continual Learning with Language-Guided Supervision
di: Ni, Bolin, et al.
Pubblicazione: (2024)
di: Ni, Bolin, et al.
Pubblicazione: (2024)
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
di: Jiang, Yankai, et al.
Pubblicazione: (2026)
di: Jiang, Yankai, et al.
Pubblicazione: (2026)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025)
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025)
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
di: Zhao, Qiyan, et al.
Pubblicazione: (2026)
di: Zhao, Qiyan, et al.
Pubblicazione: (2026)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
di: Zhang, Kejia, et al.
Pubblicazione: (2025)
di: Zhang, Kejia, et al.
Pubblicazione: (2025)
GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
di: Wei, Yanbin, et al.
Pubblicazione: (2024)
di: Wei, Yanbin, et al.
Pubblicazione: (2024)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
di: Li, Hongliang, et al.
Pubblicazione: (2025)
di: Li, Hongliang, et al.
Pubblicazione: (2025)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
di: Zhang, Qizhe, et al.
Pubblicazione: (2025)
di: Zhang, Qizhe, et al.
Pubblicazione: (2025)
Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
di: Zhao, Yaqi, et al.
Pubblicazione: (2025)
di: Zhao, Yaqi, et al.
Pubblicazione: (2025)
Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning
di: Wang, Yating, et al.
Pubblicazione: (2026)
di: Wang, Yating, et al.
Pubblicazione: (2026)
Supervised Gromov-Wasserstein Optimal Transport
di: Cang, Zixuan, et al.
Pubblicazione: (2024)
di: Cang, Zixuan, et al.
Pubblicazione: (2024)
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
di: Tan, Yifan, et al.
Pubblicazione: (2026)
di: Tan, Yifan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
di: Zhao, Yaqi, et al.
Pubblicazione: (2024) -
Towards Precise Scaling Laws for Video Diffusion Transformers
di: Yin, Yuanyang, et al.
Pubblicazione: (2024) -
Self-Supervised Visual Preference Alignment
di: Zhu, Ke, et al.
Pubblicazione: (2024) -
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
di: Li, Bozhou, et al.
Pubblicazione: (2025) -
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
di: Li, Qi, et al.
Pubblicazione: (2026)