VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Bulat, Adrian, Baldrati, Alberto, Metaxas, Ioannis Maniadis, Ouali, Yassine, Tzimiropoulos, Georgios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
by: Das, Anurag, et al.
Published: (2026)
by: Das, Anurag, et al.
Published: (2026)
VladVA: Discriminative Fine-tuning of LVLMs
by: Ouali, Yassine, et al.
Published: (2024)
by: Ouali, Yassine, et al.
Published: (2024)
Aligned Unsupervised Pretraining of Object Detectors with Self-training
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
by: Bulat, Adrian, et al.
Published: (2024)
by: Bulat, Adrian, et al.
Published: (2024)
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
Fwd2Bot: LVLM Visual Token Compression with Double Forward Bottleneck
by: Bulat, Adrian, et al.
Published: (2025)
by: Bulat, Adrian, et al.
Published: (2025)
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
by: Ouali, Yassine, et al.
Published: (2024)
by: Ouali, Yassine, et al.
Published: (2024)
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
by: Ntinou, Ioanna, et al.
Published: (2025)
by: Ntinou, Ioanna, et al.
Published: (2025)
You Only Need One Step: Fast Super-Resolution with Stable Diffusion via Scale Distillation
by: Noroozi, Mehdi, et al.
Published: (2024)
by: Noroozi, Mehdi, et al.
Published: (2024)
Knowledge Distillation Meets Open-Set Semi-Supervised Learning
by: Yang, Jing, et al.
Published: (2022)
by: Yang, Jing, et al.
Published: (2022)
Hierarchical Image Tokenization for Multi-Scale Image Super Resolution
by: Hadji, Isma, et al.
Published: (2026)
by: Hadji, Isma, et al.
Published: (2026)
Multi-scale Image Super Resolution with a Single Auto-Regressive Model
by: Sanchez, Enrique, et al.
Published: (2025)
by: Sanchez, Enrique, et al.
Published: (2025)
CLIPCleaner: Cleaning Noisy Labels with CLIP
by: Feng, Chen, et al.
Published: (2024)
by: Feng, Chen, et al.
Published: (2024)
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
by: Ntinou, Ioanna, et al.
Published: (2023)
by: Ntinou, Ioanna, et al.
Published: (2023)
SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
by: Feng, Chen, et al.
Published: (2021)
by: Feng, Chen, et al.
Published: (2021)
FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion
by: Yang, Haosen, et al.
Published: (2024)
by: Yang, Haosen, et al.
Published: (2024)
CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
by: Meng, Debin, et al.
Published: (2024)
by: Meng, Debin, et al.
Published: (2024)
LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration
by: Chen, I-Hsiang, et al.
Published: (2026)
by: Chen, I-Hsiang, et al.
Published: (2026)
One-shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD
by: Ntinou, Ioanna, et al.
Published: (2024)
by: Ntinou, Ioanna, et al.
Published: (2024)
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
by: Khalil, Ahmad, et al.
Published: (2025)
by: Khalil, Ahmad, et al.
Published: (2025)
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
by: Krestenitis, Marios, et al.
Published: (2026)
by: Krestenitis, Marios, et al.
Published: (2026)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
by: Xenos, Alexandros, et al.
Published: (2024)
by: Xenos, Alexandros, et al.
Published: (2024)
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
by: Agnolucci, Lorenzo, et al.
Published: (2024)
by: Agnolucci, Lorenzo, et al.
Published: (2024)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
by: Xie, Peijin, et al.
Published: (2024)
by: Xie, Peijin, et al.
Published: (2024)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
by: Meng, Debin, et al.
Published: (2025)
by: Meng, Debin, et al.
Published: (2025)
FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing
by: Yuan, Tianshuo, et al.
Published: (2024)
by: Yuan, Tianshuo, et al.
Published: (2024)
Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance
by: Shi, Enyi, et al.
Published: (2026)
by: Shi, Enyi, et al.
Published: (2026)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
by: Guo, Yangyang, et al.
Published: (2024)
by: Guo, Yangyang, et al.
Published: (2024)
Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning
by: Noroozi, Mehdi, et al.
Published: (2024)
by: Noroozi, Mehdi, et al.
Published: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
by: Wang, Jiankang, et al.
Published: (2025)
by: Wang, Jiankang, et al.
Published: (2025)
Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis
by: Feng, Chen, et al.
Published: (2026)
by: Feng, Chen, et al.
Published: (2026)
Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation
by: Mistretta, Marco, et al.
Published: (2024)
by: Mistretta, Marco, et al.
Published: (2024)
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
by: Khalil, Ahmad, et al.
Published: (2025)
by: Khalil, Ahmad, et al.
Published: (2025)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
by: Zhan, Yu-Wei, et al.
Published: (2025)
by: Zhan, Yu-Wei, et al.
Published: (2025)
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
by: Nikandrou, Malvina, et al.
Published: (2024)
by: Nikandrou, Malvina, et al.
Published: (2024)
Similar Items
-
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
by: Das, Anurag, et al.
Published: (2026) -
VladVA: Discriminative Fine-tuning of LVLMs
by: Ouali, Yassine, et al.
Published: (2024) -
Aligned Unsupervised Pretraining of Object Detectors with Self-training
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023) -
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
by: Bulat, Adrian, et al.
Published: (2024) -
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)