Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yue, Xinli, Sun, JianHui, Lu, Junda, Yao, Liangchao, Xia, Fan, Wang, Tianyi, Rao, Fengyun, Lyu, Jing, Deng, Yuetang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
iDiff: Interpretable Difference-aware Framework for Pairwise Image Quality Assessment
von: Yue, Xinli, et al.
Veröffentlicht: (2026)
von: Yue, Xinli, et al.
Veröffentlicht: (2026)
Revisiting Video Quality Assessment from the Perspective of Generalization
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
Advancing Video Quality Assessment for AIGC
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2025)
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2025)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
von: Tang, Changli, et al.
Veröffentlicht: (2026)
von: Tang, Changli, et al.
Veröffentlicht: (2026)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
von: Suo, Yucheng, et al.
Veröffentlicht: (2025)
von: Suo, Yucheng, et al.
Veröffentlicht: (2025)
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
von: Zhou, Zitang, et al.
Veröffentlicht: (2025)
von: Zhou, Zitang, et al.
Veröffentlicht: (2025)
ObjEmbed: Towards Universal Multimodal Object Embeddings
von: Fu, Shenghao, et al.
Veröffentlicht: (2026)
von: Fu, Shenghao, et al.
Veröffentlicht: (2026)
REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization
von: Li, Yong, et al.
Veröffentlicht: (2026)
von: Li, Yong, et al.
Veröffentlicht: (2026)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
InstructEngine: Instruction-driven Text-to-Image Alignment
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching
von: Yang, Meng, et al.
Veröffentlicht: (2025)
von: Yang, Meng, et al.
Veröffentlicht: (2025)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2026)
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2026)
Instant Preference Alignment for Text-to-Image Diffusion Models
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
InstructGraph: Boosting Large Language Models via Graph-centric Instruction Tuning and Preference Alignment
von: Wang, Jianing, et al.
Veröffentlicht: (2024)
von: Wang, Jianing, et al.
Veröffentlicht: (2024)
Can the Establishment of Bankruptcy Courts Improve Firm Investment Efficiency? Evidence from China
von: Xiangjun Fan, et al.
Veröffentlicht: (2026)
von: Xiangjun Fan, et al.
Veröffentlicht: (2026)
CoMMIT: Coordinated Multimodal Instruction Tuning
von: Li, Xintong, et al.
Veröffentlicht: (2024)
von: Li, Xintong, et al.
Veröffentlicht: (2024)
Alignment of Diffusion Model and Flow Matching for Text-to-Image Generation
von: Ouyang, Yidong, et al.
Veröffentlicht: (2026)
von: Ouyang, Yidong, et al.
Veröffentlicht: (2026)
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization
von: Luo, Yucong, et al.
Veröffentlicht: (2024)
von: Luo, Yucong, et al.
Veröffentlicht: (2024)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
MRI Image Generation Based on Text Prompts
von: Fan, Xinxian, et al.
Veröffentlicht: (2025)
von: Fan, Xinxian, et al.
Veröffentlicht: (2025)
Spatial-Semantic Collaborative Cropping for User Generated Content
von: Su, Yukun, et al.
Veröffentlicht: (2024)
von: Su, Yukun, et al.
Veröffentlicht: (2024)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
Language-Image Alignment with Fixed Text Encoders
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
The Quadratic Geometry of Flow Matching: Semantic Granularity Alignment for Text-to-Image Synthesis
von: Xiong, Zhinan, et al.
Veröffentlicht: (2026)
von: Xiong, Zhinan, et al.
Veröffentlicht: (2026)
REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
von: Shi, Fulin, et al.
Veröffentlicht: (2025)
von: Shi, Fulin, et al.
Veröffentlicht: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
Semantic-Enriched Latent Visual Reasoning
von: Xu, Tianrun, et al.
Veröffentlicht: (2026)
von: Xu, Tianrun, et al.
Veröffentlicht: (2026)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
A Solution of Ultra Wideband Based High-resolution and Lossless Audio Transmission
von: Zhang, Fengyun
Veröffentlicht: (2025)
von: Zhang, Fengyun
Veröffentlicht: (2025)
ZeroStereo: Zero-shot Stereo Matching from Single Images
von: Wang, Xianqi, et al.
Veröffentlicht: (2025)
von: Wang, Xianqi, et al.
Veröffentlicht: (2025)
Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment
von: Huang, Huayang, et al.
Veröffentlicht: (2026)
von: Huang, Huayang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
iDiff: Interpretable Difference-aware Framework for Pairwise Image Quality Assessment
von: Yue, Xinli, et al.
Veröffentlicht: (2026) -
Revisiting Video Quality Assessment from the Perspective of Generalization
von: Yue, Xinli, et al.
Veröffentlicht: (2024) -
Advancing Video Quality Assessment for AIGC
von: Yue, Xinli, et al.
Veröffentlicht: (2024) -
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2025) -
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
von: Tang, Changli, et al.
Veröffentlicht: (2026)