Text-centric Alignment for Multi-Modality Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tsai, Yun-Da, Yen, Ting-Yu, Guo, Pei-Fu, Li, Zhe-Yan, Lin, Shou-De |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
Enhance the Robustness of Text-Centric Multimodal Alignments
von: Yen, Ting-Yu, et al.
Veröffentlicht: (2024)
von: Yen, Ting-Yu, et al.
Veröffentlicht: (2024)
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
von: Lai, Songning, et al.
Veröffentlicht: (2023)
von: Lai, Songning, et al.
Veröffentlicht: (2023)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
von: Huang, Qidong, et al.
Veröffentlicht: (2024)
von: Huang, Qidong, et al.
Veröffentlicht: (2024)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
von: Li, Qian, et al.
Veröffentlicht: (2024)
von: Li, Qian, et al.
Veröffentlicht: (2024)
Multi-Modal 3D Mesh Reconstruction from Images and Text
von: Reka, Melvin, et al.
Veröffentlicht: (2025)
von: Reka, Melvin, et al.
Veröffentlicht: (2025)
Vision-centric Token Compression in Large Language Model
von: Xing, Ling, et al.
Veröffentlicht: (2025)
von: Xing, Ling, et al.
Veröffentlicht: (2025)
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
Cross-Modal Retrieval for Motion and Text via DropTriple Loss
von: Yan, Sheng, et al.
Veröffentlicht: (2023)
von: Yan, Sheng, et al.
Veröffentlicht: (2023)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
von: Hu, Zhe, et al.
Veröffentlicht: (2025)
von: Hu, Zhe, et al.
Veröffentlicht: (2025)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
von: Shi, Yi, et al.
Veröffentlicht: (2025)
von: Shi, Yi, et al.
Veröffentlicht: (2025)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
A Multi-view Mask Contrastive Learning Graph Convolutional Neural Network for Age Estimation
von: Zhang, Yiping, et al.
Veröffentlicht: (2024)
von: Zhang, Yiping, et al.
Veröffentlicht: (2024)
X-VILA: Cross-Modality Alignment for Large Language Model
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
Fast Prompt Alignment for Text-to-Image Generation
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slip Scripts
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Multi-Modal Language Models as Text-to-Image Model Evaluators
von: Chen, Jiahui, et al.
Veröffentlicht: (2025)
von: Chen, Jiahui, et al.
Veröffentlicht: (2025)
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
HARE: an entity and relation centric evaluation framework for histopathology reports
von: Kim, Yunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
von: Li, Shengzhi, et al.
Veröffentlicht: (2024)
von: Li, Shengzhi, et al.
Veröffentlicht: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Is Extending Modality The Right Path Towards Omni-Modality?
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
von: Duan, Chen, et al.
Veröffentlicht: (2024)
von: Duan, Chen, et al.
Veröffentlicht: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
Ego-centric Predictive Model Conditioned on Hand Trajectories
von: Zhang, Binjie, et al.
Veröffentlicht: (2025)
von: Zhang, Binjie, et al.
Veröffentlicht: (2025)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
von: Phukan, Arpan, et al.
Veröffentlicht: (2024)
von: Phukan, Arpan, et al.
Veröffentlicht: (2024)
Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality Learning
von: Yan, Yang, et al.
Veröffentlicht: (2025)
von: Yan, Yang, et al.
Veröffentlicht: (2025)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025) -
Enhance the Robustness of Text-Centric Multimodal Alignments
von: Yen, Ting-Yu, et al.
Veröffentlicht: (2024) -
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
von: Lai, Songning, et al.
Veröffentlicht: (2023) -
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
von: Huang, Qidong, et al.
Veröffentlicht: (2024) -
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
von: Li, Qian, et al.
Veröffentlicht: (2024)