Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yang, Liu, Mengyuan, Huang, Shudong, Lv, Jiancheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Semantic-Aware Ship Detection with Vision-Language Integration
by: Li, Jiahao, et al.
Published: (2025)
by: Li, Jiahao, et al.
Published: (2025)
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
by: Tian, Yuxin, et al.
Published: (2024)
by: Tian, Yuxin, et al.
Published: (2024)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
ACPO: Counteracting Likelihood Displacement in Vision-Language Alignment with Asymmetric Constraints
by: Huang, Kaili, et al.
Published: (2026)
by: Huang, Kaili, et al.
Published: (2026)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
Modest-Align: Data-Efficient Alignment for Vision-Language Models
by: Liu, Jiaxiang, et al.
Published: (2025)
by: Liu, Jiaxiang, et al.
Published: (2025)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
by: Wei, Xinyu, et al.
Published: (2025)
by: Wei, Xinyu, et al.
Published: (2025)
PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence Learning
by: Liu, Zhuoyao, et al.
Published: (2025)
by: Liu, Zhuoyao, et al.
Published: (2025)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
by: Dong, Jihao, et al.
Published: (2024)
by: Dong, Jihao, et al.
Published: (2024)
HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection
by: Shi, Senyuan, et al.
Published: (2026)
by: Shi, Senyuan, et al.
Published: (2026)
Enhancing Vision-Language Model with Unmasked Token Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
DAPE: Dynamic Non-uniform Alignment and Progressive Detail Enhancement Techniques for Improving the Performance of Efficient Visual Language Models
by: Tian, Mengyuan, et al.
Published: (2026)
by: Tian, Mengyuan, et al.
Published: (2026)
Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language Models
by: Chen, Mengyuan, et al.
Published: (2024)
by: Chen, Mengyuan, et al.
Published: (2024)
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
by: Huang, Youcheng, et al.
Published: (2024)
by: Huang, Youcheng, et al.
Published: (2024)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
by: Lin, Tao, et al.
Published: (2025)
by: Lin, Tao, et al.
Published: (2025)
EVLM: An Efficient Vision-Language Model for Visual Understanding
by: Chen, Kaibing, et al.
Published: (2024)
by: Chen, Kaibing, et al.
Published: (2024)
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
by: Gong, Shizhan, et al.
Published: (2025)
by: Gong, Shizhan, et al.
Published: (2025)
Attention! Your Vision Language Model Could Be Maliciously Manipulated
by: Wang, Xiaosen, et al.
Published: (2025)
by: Wang, Xiaosen, et al.
Published: (2025)
Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking
by: Ge, Jiawei, et al.
Published: (2023)
by: Ge, Jiawei, et al.
Published: (2023)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
by: Yang, Lijin, et al.
Published: (2026)
by: Yang, Lijin, et al.
Published: (2026)
Diagnosing Urban Street Vitality via a Visual-Semantic and Spatiotemporal Framework for Street-Level Economics
by: Zhuo, Xinxin, et al.
Published: (2026)
by: Zhuo, Xinxin, et al.
Published: (2026)
An Efficient Approach for Muscle Segmentation and 3D Reconstruction Using Keypoint Tracking in MRI Scan
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
by: Wu, Ruijia, et al.
Published: (2025)
by: Wu, Ruijia, et al.
Published: (2025)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models
by: Liao, Zehui, et al.
Published: (2025)
by: Liao, Zehui, et al.
Published: (2025)
Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
by: Jiang, Jerry, et al.
Published: (2026)
by: Jiang, Jerry, et al.
Published: (2026)
Leveraging Contrast Information for Efficient Document Shadow Removal
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Semantic-Space-Intervened Diffusive Alignment for Visual Classification
by: Li, Zixuan, et al.
Published: (2025)
by: Li, Zixuan, et al.
Published: (2025)
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
by: Feng, Ze, et al.
Published: (2025)
by: Feng, Ze, et al.
Published: (2025)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
by: Tao, Chenxin, et al.
Published: (2024)
by: Tao, Chenxin, et al.
Published: (2024)
Visual-Advantage On-Policy Distillation for Vision-Language Models
by: Liu, Ruiqi, et al.
Published: (2026)
by: Liu, Ruiqi, et al.
Published: (2026)
Semantically Grounded QFormer for Efficient Vision Language Understanding
by: Choraria, Moulik, et al.
Published: (2023)
by: Choraria, Moulik, et al.
Published: (2023)
Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment
by: Wu, Lukun, et al.
Published: (2025)
by: Wu, Lukun, et al.
Published: (2025)
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative Pruning
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
Zero-Shot Aerial Object Detection with Visual Description Regularization
by: Zang, Zhengqing, et al.
Published: (2024)
by: Zang, Zhengqing, et al.
Published: (2024)
Similar Items
-
Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching
by: Liu, Yang, et al.
Published: (2025) -
Semantic-Aware Ship Detection with Vision-Language Integration
by: Li, Jiahao, et al.
Published: (2025) -
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
by: Tian, Yuxin, et al.
Published: (2024) -
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
by: Wang, Yifan, et al.
Published: (2025) -
ACPO: Counteracting Likelihood Displacement in Vision-Language Alignment with Asymmetric Constraints
by: Huang, Kaili, et al.
Published: (2026)