Attention Guided Alignment in Efficient Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mahajan, Shweta, Le, Hoang, Park, Hyojin, Farhadzadeh, Farzad, Hayat, Munawar, Porikli, Fatih |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
Resolving the Identity Crisis in Text-to-Image Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
Generalized Contrastive Learning for Universal Multimodal Retrieval
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
PosSAM: Panoptic Open-vocabulary Segment Anything
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
von: Kadambi, Shreya, et al.
Veröffentlicht: (2025)
von: Kadambi, Shreya, et al.
Veröffentlicht: (2025)
Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
Sort-free Gaussian Splatting via Weighted Sum Rendering
von: Hou, Qiqi, et al.
Veröffentlicht: (2024)
von: Hou, Qiqi, et al.
Veröffentlicht: (2024)
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
Neural Graphics Texture Compression Supporting Random Access
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2024)
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2024)
Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories
von: Hedlin, Eric, et al.
Veröffentlicht: (2024)
von: Hedlin, Eric, et al.
Veröffentlicht: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
von: Das, Debasmit, et al.
Veröffentlicht: (2025)
von: Das, Debasmit, et al.
Veröffentlicht: (2025)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware Interpolation
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
von: Guo, Yichen, et al.
Veröffentlicht: (2025)
von: Guo, Yichen, et al.
Veröffentlicht: (2025)
Flashbacks to Harmonize Stability and Plasticity in Continual Learning
von: Mahmoodi, Leila, et al.
Veröffentlicht: (2025)
von: Mahmoodi, Leila, et al.
Veröffentlicht: (2025)
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
von: Porikli, Sedat, et al.
Veröffentlicht: (2025)
von: Porikli, Sedat, et al.
Veröffentlicht: (2025)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
von: Park, Sunghyun, et al.
Veröffentlicht: (2026)
von: Park, Sunghyun, et al.
Veröffentlicht: (2026)
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
von: Ilhan, Fatih, et al.
Veröffentlicht: (2026)
von: Ilhan, Fatih, et al.
Veröffentlicht: (2026)
ToSA: Token Selective Attention for Efficient Vision Transformers
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
von: Bahng, Hyojin, et al.
Veröffentlicht: (2025)
von: Bahng, Hyojin, et al.
Veröffentlicht: (2025)
Low-Latency Neural Stereo Streaming
von: Hou, Qiqi, et al.
Veröffentlicht: (2024)
von: Hou, Qiqi, et al.
Veröffentlicht: (2024)
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
Improved Alignment of Modalities in Large Vision Language Models
von: Jangra, Kartik, et al.
Veröffentlicht: (2025)
von: Jangra, Kartik, et al.
Veröffentlicht: (2025)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
von: Lian, Chenyu, et al.
Veröffentlicht: (2025)
von: Lian, Chenyu, et al.
Veröffentlicht: (2025)
Enhancing Domain Adaptation through Prompt Gradient Alignment
von: Phan, Hoang, et al.
Veröffentlicht: (2024)
von: Phan, Hoang, et al.
Veröffentlicht: (2024)
LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing
von: Yan, Huimin, et al.
Veröffentlicht: (2026)
von: Yan, Huimin, et al.
Veröffentlicht: (2026)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language Models
von: Park, Seonghwan, et al.
Veröffentlicht: (2025)
von: Park, Seonghwan, et al.
Veröffentlicht: (2025)
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
von: Park, Minho, et al.
Veröffentlicht: (2025)
von: Park, Minho, et al.
Veröffentlicht: (2025)
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
von: Corley, Isaac, et al.
Veröffentlicht: (2025)
von: Corley, Isaac, et al.
Veröffentlicht: (2025)
Synthesizer Based Efficient Self-Attention for Vision Tasks
von: Zhu, Guangyang, et al.
Veröffentlicht: (2022)
von: Zhu, Guangyang, et al.
Veröffentlicht: (2022)
Ultrasound Vision-Language Alignment via Contrastive Learning
von: Lyu, Zhuoyang, et al.
Veröffentlicht: (2026)
von: Lyu, Zhuoyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024) -
Resolving the Identity Crisis in Text-to-Image Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025) -
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025) -
Generalized Contrastive Learning for Universal Multimodal Retrieval
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025) -
LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)