Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ntinou, Ioanna, Xenos, Alexandros, Ouali, Yassine, Bulat, Adrian, Tzimiropoulos, Georgios |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
von: Bulat, Adrian, et al.
Veröffentlicht: (2024)
von: Bulat, Adrian, et al.
Veröffentlicht: (2024)
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
Fwd2Bot: LVLM Visual Token Compression with Double Forward Bottleneck
von: Bulat, Adrian, et al.
Veröffentlicht: (2025)
von: Bulat, Adrian, et al.
Veröffentlicht: (2025)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
von: Xenos, Alexandros, et al.
Veröffentlicht: (2024)
von: Xenos, Alexandros, et al.
Veröffentlicht: (2024)
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2023)
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2023)
VladVA: Discriminative Fine-tuning of LVLMs
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2024)
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2024)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
von: Bulat, Adrian, et al.
Veröffentlicht: (2026)
von: Bulat, Adrian, et al.
Veröffentlicht: (2026)
You Only Need One Step: Fast Super-Resolution with Stable Diffusion via Scale Distillation
von: Noroozi, Mehdi, et al.
Veröffentlicht: (2024)
von: Noroozi, Mehdi, et al.
Veröffentlicht: (2024)
Knowledge Distillation Meets Open-Set Semi-Supervised Learning
von: Yang, Jing, et al.
Veröffentlicht: (2022)
von: Yang, Jing, et al.
Veröffentlicht: (2022)
Hierarchical Image Tokenization for Multi-Scale Image Super Resolution
von: Hadji, Isma, et al.
Veröffentlicht: (2026)
von: Hadji, Isma, et al.
Veröffentlicht: (2026)
Aligned Unsupervised Pretraining of Object Detectors with Self-training
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2023)
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2023)
Multi-scale Image Super Resolution with a Single Auto-Regressive Model
von: Sanchez, Enrique, et al.
Veröffentlicht: (2025)
von: Sanchez, Enrique, et al.
Veröffentlicht: (2025)
FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion
von: Yang, Haosen, et al.
Veröffentlicht: (2024)
von: Yang, Haosen, et al.
Veröffentlicht: (2024)
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
von: Das, Anurag, et al.
Veröffentlicht: (2026)
von: Das, Anurag, et al.
Veröffentlicht: (2026)
Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration
von: Chen, I-Hsiang, et al.
Veröffentlicht: (2026)
von: Chen, I-Hsiang, et al.
Veröffentlicht: (2026)
CLIPCleaner: Cleaning Noisy Labels with CLIP
von: Feng, Chen, et al.
Veröffentlicht: (2024)
von: Feng, Chen, et al.
Veröffentlicht: (2024)
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2024)
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2024)
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
TextPSG: Panoptic Scene Graph Generation from Textual Descriptions
von: Zhao, Chengyang, et al.
Veröffentlicht: (2023)
von: Zhao, Chengyang, et al.
Veröffentlicht: (2023)
VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions
von: Moon, Seokha, et al.
Veröffentlicht: (2024)
von: Moon, Seokha, et al.
Veröffentlicht: (2024)
SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
von: Feng, Chen, et al.
Veröffentlicht: (2021)
von: Feng, Chen, et al.
Veröffentlicht: (2021)
CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
von: Meng, Debin, et al.
Veröffentlicht: (2024)
von: Meng, Debin, et al.
Veröffentlicht: (2024)
Aligning Actions and Walking to LLM-Generated Textual Descriptions
von: Chivereanu, Radu, et al.
Veröffentlicht: (2024)
von: Chivereanu, Radu, et al.
Veröffentlicht: (2024)
A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
One-shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
von: Meng, Debin, et al.
Veröffentlicht: (2025)
von: Meng, Debin, et al.
Veröffentlicht: (2025)
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval
von: Demić, Emil, et al.
Veröffentlicht: (2025)
von: Demić, Emil, et al.
Veröffentlicht: (2025)
Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning
von: Noroozi, Mehdi, et al.
Veröffentlicht: (2024)
von: Noroozi, Mehdi, et al.
Veröffentlicht: (2024)
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
von: Berman, Nimrod, et al.
Veröffentlicht: (2025)
von: Berman, Nimrod, et al.
Veröffentlicht: (2025)
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
Zoo3D: Zero-Shot 3D Object Detection at Scene Level
von: Lemeshko, Andrey, et al.
Veröffentlicht: (2025)
von: Lemeshko, Andrey, et al.
Veröffentlicht: (2025)
MoReact: Generating Reactive Motion from Textual Descriptions
von: Xu, Xiyan, et al.
Veröffentlicht: (2025)
von: Xu, Xiyan, et al.
Veröffentlicht: (2025)
The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
von: Jung, Hoin, et al.
Veröffentlicht: (2026)
von: Jung, Hoin, et al.
Veröffentlicht: (2026)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
Vision-by-Language for Training-Free Compositional Image Retrieval
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2023)
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2023)
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes
von: Linok, Sergey, et al.
Veröffentlicht: (2025)
von: Linok, Sergey, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
von: Bulat, Adrian, et al.
Veröffentlicht: (2024) -
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
von: Ouali, Yassine, et al.
Veröffentlicht: (2024) -
Fwd2Bot: LVLM Visual Token Compression with Double Forward Bottleneck
von: Bulat, Adrian, et al.
Veröffentlicht: (2025) -
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
von: Xenos, Alexandros, et al.
Veröffentlicht: (2024) -
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2023)