Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zihang, Zhang, Siyue, Zhao, Yilun, Yang, Jingyi, Song, Tingyu, Luu, Anh Tuan, Zhao, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
von: Zhang, Siyue, et al.
Veröffentlicht: (2025)
von: Zhang, Siyue, et al.
Veröffentlicht: (2025)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
von: Zhang, Zilun, et al.
Veröffentlicht: (2023)
von: Zhang, Zilun, et al.
Veröffentlicht: (2023)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
von: Stamatakis, Markos, et al.
Veröffentlicht: (2025)
von: Stamatakis, Markos, et al.
Veröffentlicht: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Language Models as Black-Box Optimizers for Vision-Language Models
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
von: Chen, Yijing, et al.
Veröffentlicht: (2025)
von: Chen, Yijing, et al.
Veröffentlicht: (2025)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
von: He, Zongtao, et al.
Veröffentlicht: (2023)
von: He, Zongtao, et al.
Veröffentlicht: (2023)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice
von: Westermann, Hannes, et al.
Veröffentlicht: (2024)
von: Westermann, Hannes, et al.
Veröffentlicht: (2024)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
von: Nishimoto, Tomohiro, et al.
Veröffentlicht: (2024)
von: Nishimoto, Tomohiro, et al.
Veröffentlicht: (2024)
Hyperbolic Safety-Aware Vision-Language Models
von: Poppi, Tobia, et al.
Veröffentlicht: (2025)
von: Poppi, Tobia, et al.
Veröffentlicht: (2025)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
von: Masumura, Ryo, et al.
Veröffentlicht: (2025)
von: Masumura, Ryo, et al.
Veröffentlicht: (2025)
The Revolution of Multimodal Large Language Models: A Survey
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
Causal Graphical Models for Vision-Language Compositional Understanding
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
von: Song, Jiale, et al.
Veröffentlicht: (2026)
von: Song, Jiale, et al.
Veröffentlicht: (2026)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
von: Hayashi, Kazuki, et al.
Veröffentlicht: (2024)
von: Hayashi, Kazuki, et al.
Veröffentlicht: (2024)
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
von: Wu, Jiaying, et al.
Veröffentlicht: (2025) -
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024) -
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
von: Zhang, Siyue, et al.
Veröffentlicht: (2025) -
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023) -
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)