ComAlign: Compositional Alignment in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdollah, Ali, Izadi, Amirmohammad, Saghafian, Armin, Vahidimajd, Reza, Mozafari, Mohammad, Mirzaei, Amirreza, Samiei, Mohammadmahdi, Baghshah, Mahdieh Soleymani |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
von: Saghafian, Armin, et al.
Veröffentlicht: (2024)
von: Saghafian, Armin, et al.
Veröffentlicht: (2024)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
von: Mozafari, Mohammad, et al.
Veröffentlicht: (2024)
von: Mozafari, Mohammad, et al.
Veröffentlicht: (2024)
T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2025)
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2025)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
von: Izadi, Amir Mohammad, et al.
Veröffentlicht: (2025)
von: Izadi, Amir Mohammad, et al.
Veröffentlicht: (2025)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
von: Ahmadpour, Mohammadjavad, et al.
Veröffentlicht: (2025)
von: Ahmadpour, Mohammadjavad, et al.
Veröffentlicht: (2025)
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
von: Paqaleh, Mohammad Mahdi Samiei, et al.
Veröffentlicht: (2025)
von: Paqaleh, Mohammad Mahdi Samiei, et al.
Veröffentlicht: (2025)
The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
von: Samiei, Mahdi, et al.
Veröffentlicht: (2025)
von: Samiei, Mahdi, et al.
Veröffentlicht: (2025)
Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
von: Izadi, Amirmohammad, et al.
Veröffentlicht: (2025)
von: Izadi, Amirmohammad, et al.
Veröffentlicht: (2025)
To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
von: Fang, Wanlong, et al.
Veröffentlicht: (2025)
von: Fang, Wanlong, et al.
Veröffentlicht: (2025)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
EidetiCom: A Cross-modal Brain-Computer Semantic Communication Paradigm for Decoding Visual Perception
von: Zheng, Linfeng, et al.
Veröffentlicht: (2024)
von: Zheng, Linfeng, et al.
Veröffentlicht: (2024)
Investigating the Generalizability of Physiological Characteristics of Anxiety
von: Zhou, Emily, et al.
Veröffentlicht: (2024)
von: Zhou, Emily, et al.
Veröffentlicht: (2024)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
von: Mehri, Faridoun, et al.
Veröffentlicht: (2024)
von: Mehri, Faridoun, et al.
Veröffentlicht: (2024)
Reversing the Damage: A QP-Aware Transformer-Diffusion Approach for 8K Video Restoration under Codec Compression
von: Dehaghi, Ali Mollaahmadi, et al.
Veröffentlicht: (2024)
von: Dehaghi, Ali Mollaahmadi, et al.
Veröffentlicht: (2024)
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
Causal Attribution via Activation Patching
von: Izadi, Amirmohammad, et al.
Veröffentlicht: (2026)
von: Izadi, Amirmohammad, et al.
Veröffentlicht: (2026)
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
von: Liu, Yating, et al.
Veröffentlicht: (2025)
von: Liu, Yating, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
von: Hasani, Hosein, et al.
Veröffentlicht: (2026)
von: Hasani, Hosein, et al.
Veröffentlicht: (2026)
Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
von: Shahabadi, Hossein, et al.
Veröffentlicht: (2025)
von: Shahabadi, Hossein, et al.
Veröffentlicht: (2025)
EMID: An Emotional Aligned Dataset in Audio-Visual Modality
von: Zou, Jialing, et al.
Veröffentlicht: (2023)
von: Zou, Jialing, et al.
Veröffentlicht: (2023)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
Revisiting Vision-Language Features Adaptation and Inconsistency for Social Media Popularity Prediction
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2024)
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2024)
BOLA360: Near-optimal View and Bitrate Adaptation for 360-degree Video Streaming
von: Zeynali, Ali, et al.
Veröffentlicht: (2023)
von: Zeynali, Ali, et al.
Veröffentlicht: (2023)
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
von: Wang, Youze, et al.
Veröffentlicht: (2023)
von: Wang, Youze, et al.
Veröffentlicht: (2023)
EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
von: Lee, SangEun, et al.
Veröffentlicht: (2025)
von: Lee, SangEun, et al.
Veröffentlicht: (2025)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
ALMol: Aligned Language-Molecule Translation LLMs through Offline Preference Contrastive Optimisation
von: Gkoumas, Dimitris
Veröffentlicht: (2024)
von: Gkoumas, Dimitris
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
SciCom Wiki: A Digital Library to Support the Science Communication Knowledge Infrastructure for Videos and Podcasts
von: Wittenborg, Tim, et al.
Veröffentlicht: (2025)
von: Wittenborg, Tim, et al.
Veröffentlicht: (2025)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
von: Saghafian, Armin, et al.
Veröffentlicht: (2024) -
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
von: Mozafari, Mohammad, et al.
Veröffentlicht: (2024) -
T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2025) -
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
von: Izadi, Amir Mohammad, et al.
Veröffentlicht: (2025) -
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)