VladVA: Discriminative Fine-tuning of LVLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ouali, Yassine, Bulat, Adrian, Xenos, Alexandros, Zaganidis, Anestis, Metaxas, Ioannis Maniadis, Martinez, Brais, Tzimiropoulos, Georgios |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
by: Ouali, Yassine, et al.
Published: (2024)
by: Ouali, Yassine, et al.
Published: (2024)
Aligned Unsupervised Pretraining of Object Detectors with Self-training
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
by: Bulat, Adrian, et al.
Published: (2026)
by: Bulat, Adrian, et al.
Published: (2026)
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
by: Ntinou, Ioanna, et al.
Published: (2025)
by: Ntinou, Ioanna, et al.
Published: (2025)
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
by: Das, Anurag, et al.
Published: (2026)
by: Das, Anurag, et al.
Published: (2026)
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
by: Bulat, Adrian, et al.
Published: (2024)
by: Bulat, Adrian, et al.
Published: (2024)
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
Fwd2Bot: LVLM Visual Token Compression with Double Forward Bottleneck
by: Bulat, Adrian, et al.
Published: (2025)
by: Bulat, Adrian, et al.
Published: (2025)
Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning
by: Noroozi, Mehdi, et al.
Published: (2024)
by: Noroozi, Mehdi, et al.
Published: (2024)
You Only Need One Step: Fast Super-Resolution with Stable Diffusion via Scale Distillation
by: Noroozi, Mehdi, et al.
Published: (2024)
by: Noroozi, Mehdi, et al.
Published: (2024)
Knowledge Distillation Meets Open-Set Semi-Supervised Learning
by: Yang, Jing, et al.
Published: (2022)
by: Yang, Jing, et al.
Published: (2022)
Hierarchical Image Tokenization for Multi-Scale Image Super Resolution
by: Hadji, Isma, et al.
Published: (2026)
by: Hadji, Isma, et al.
Published: (2026)
Multi-scale Image Super Resolution with a Single Auto-Regressive Model
by: Sanchez, Enrique, et al.
Published: (2025)
by: Sanchez, Enrique, et al.
Published: (2025)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
by: Xenos, Alexandros, et al.
Published: (2024)
by: Xenos, Alexandros, et al.
Published: (2024)
FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion
by: Yang, Haosen, et al.
Published: (2024)
by: Yang, Haosen, et al.
Published: (2024)
Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration
by: Chen, I-Hsiang, et al.
Published: (2026)
by: Chen, I-Hsiang, et al.
Published: (2026)
CLIPCleaner: Cleaning Noisy Labels with CLIP
by: Feng, Chen, et al.
Published: (2024)
by: Feng, Chen, et al.
Published: (2024)
SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
by: Feng, Chen, et al.
Published: (2021)
by: Feng, Chen, et al.
Published: (2021)
CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
by: Meng, Debin, et al.
Published: (2024)
by: Meng, Debin, et al.
Published: (2024)
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
by: Krestenitis, Marios, et al.
Published: (2026)
by: Krestenitis, Marios, et al.
Published: (2026)
LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
One-shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
by: Ntinou, Ioanna, et al.
Published: (2023)
by: Ntinou, Ioanna, et al.
Published: (2023)
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD
by: Ntinou, Ioanna, et al.
Published: (2024)
by: Ntinou, Ioanna, et al.
Published: (2024)
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
by: Sui, Xiangjie, et al.
Published: (2025)
by: Sui, Xiangjie, et al.
Published: (2025)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
by: Meng, Debin, et al.
Published: (2025)
by: Meng, Debin, et al.
Published: (2025)
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
by: Yan, Bei, et al.
Published: (2025)
by: Yan, Bei, et al.
Published: (2025)
Multi-Class Anomaly Detection based on Regularized Discriminative Coupled hypersphere-based Feature Adaptation
by: Rafiei, Mehdi, et al.
Published: (2023)
by: Rafiei, Mehdi, et al.
Published: (2023)
Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis
by: Feng, Chen, et al.
Published: (2026)
by: Feng, Chen, et al.
Published: (2026)
Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model
by: Sarafis, Ioannis, et al.
Published: (2025)
by: Sarafis, Ioannis, et al.
Published: (2025)
Coffee: Controllable Diffusion Fine-tuning
by: Zeng, Ziyao, et al.
Published: (2025)
by: Zeng, Ziyao, et al.
Published: (2025)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
Getting to the Point: Pointing Improves LVLMs at Counting
by: Alghisi, Simone, et al.
Published: (2026)
by: Alghisi, Simone, et al.
Published: (2026)
MultiModal Fine-tuning with Synthetic Captions
by: Enomoto, Shohei, et al.
Published: (2026)
by: Enomoto, Shohei, et al.
Published: (2026)
Semantic-aware Adversarial Fine-tuning for CLIP
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
Fine-Grained VLM Fine-tuning via Latent Hierarchical Adapter Learning
by: Zhao, Yumiao, et al.
Published: (2025)
by: Zhao, Yumiao, et al.
Published: (2025)
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
by: He, Zhendong, et al.
Published: (2026)
by: He, Zhendong, et al.
Published: (2026)
Batch Augmentation with Unimodal Fine-tuning for Multimodal Learning
by: Kabir, H M Dipu, et al.
Published: (2025)
by: Kabir, H M Dipu, et al.
Published: (2025)
Similar Items
-
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
by: Ouali, Yassine, et al.
Published: (2024) -
Aligned Unsupervised Pretraining of Object Detectors with Self-training
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023) -
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
by: Bulat, Adrian, et al.
Published: (2026) -
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
by: Ntinou, Ioanna, et al.
Published: (2025) -
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
by: Das, Anurag, et al.
Published: (2026)