Listener-Rewarded Thinking in VLMs for Image Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Gambashidze, Alexander, Pengyi, Li, Skripkin, Matvey, Galichin, Andrey, Gusarov, Anton, Sobolev, Konstantin, Kuznetsov, Andrey, Oseledets, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
by: Gambashidze, Alexander, et al.
Published: (2025)
by: Gambashidze, Alexander, et al.
Published: (2025)
CoMa: Contextual Massing Generation with Vision-Language Models
by: Maslov, Evgenii, et al.
Published: (2026)
by: Maslov, Evgenii, et al.
Published: (2026)
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Spread them Apart: Towards Robust Watermarking of Generated Content
by: Pautov, Mikhail, et al.
Published: (2025)
by: Pautov, Mikhail, et al.
Published: (2025)
Simple Vision-Language Math Reasoning via Rendered Text
by: Skripkin, Matvey, et al.
Published: (2025)
by: Skripkin, Matvey, et al.
Published: (2025)
OmniFusion Technical Report
by: Goncharova, Elizaveta, et al.
Published: (2024)
by: Goncharova, Elizaveta, et al.
Published: (2024)
Inverting Black-Box Face Recognition Systems via Zero-Order Optimization in Eigenface Space
by: Razzhigaev, Anton, et al.
Published: (2025)
by: Razzhigaev, Anton, et al.
Published: (2025)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Aligning Diffusion Models with Noise-Conditioned Perception
by: Gambashidze, Alexander, et al.
Published: (2024)
by: Gambashidze, Alexander, et al.
Published: (2024)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
by: Korzh, Dmitrii, et al.
Published: (2025)
by: Korzh, Dmitrii, et al.
Published: (2025)
T-LoRA: Single Image Diffusion Model Customization Without Overfitting
by: Soboleva, Vera, et al.
Published: (2025)
by: Soboleva, Vera, et al.
Published: (2025)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
by: Skripkin, Matvey, et al.
Published: (2025)
by: Skripkin, Matvey, et al.
Published: (2025)
Real-World Transferable Adversarial Attack on Face-Recognition Systems
by: Kaznacheev, Andrey, et al.
Published: (2025)
by: Kaznacheev, Andrey, et al.
Published: (2025)
Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration
by: Tokhchukov, Danil, et al.
Published: (2026)
by: Tokhchukov, Danil, et al.
Published: (2026)
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
by: Arkhipkin, Vladimir, et al.
Published: (2024)
by: Arkhipkin, Vladimir, et al.
Published: (2024)
UniDet3D: Multi-dataset Indoor 3D Object Detection
by: Kolodiazhnyi, Maksim, et al.
Published: (2024)
by: Kolodiazhnyi, Maksim, et al.
Published: (2024)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
by: Li, Yiwei, et al.
Published: (2026)
by: Li, Yiwei, et al.
Published: (2026)
General Lipschitz: Certified Robustness Against Resolvable Semantic Transformations via Transformation-Dependent Randomized Smoothing
by: Korzh, Dmitrii, et al.
Published: (2023)
by: Korzh, Dmitrii, et al.
Published: (2023)
CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Visual Preference Optimization with Rubric Rewards
by: Yu, Ya-Qi, et al.
Published: (2026)
by: Yu, Ya-Qi, et al.
Published: (2026)
Random Direct Preference Optimization for Radiography Report Generation
by: Samokhin, Valentin, et al.
Published: (2025)
by: Samokhin, Valentin, et al.
Published: (2025)
Every Image Listens, Every Image Dances: Music-Driven Image Animation
by: Dong, Zhikang, et al.
Published: (2025)
by: Dong, Zhikang, et al.
Published: (2025)
Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs
by: Gusarov, Anton, et al.
Published: (2025)
by: Gusarov, Anton, et al.
Published: (2025)
RClicks: Realistic Click Simulation for Benchmarking Interactive Segmentation
by: Antonov, Anton, et al.
Published: (2024)
by: Antonov, Anton, et al.
Published: (2024)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
by: Li, Chenjun
Published: (2026)
by: Li, Chenjun
Published: (2026)
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
by: Wang, Haozhe, et al.
Published: (2026)
by: Wang, Haozhe, et al.
Published: (2026)
MGRegBench: A Novel Benchmark Dataset with Anatomical Landmarks for Mammography Image Registration
by: Krasnova, Svetlana, et al.
Published: (2025)
by: Krasnova, Svetlana, et al.
Published: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
Rethinking VLMs and LLMs for Image Classification
by: Cooper, Avi, et al.
Published: (2024)
by: Cooper, Avi, et al.
Published: (2024)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
by: Razzhigaev, Anton, et al.
Published: (2023)
by: Razzhigaev, Anton, et al.
Published: (2023)
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
by: Toibazar, Daulet, et al.
Published: (2025)
by: Toibazar, Daulet, et al.
Published: (2025)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
RewardFlow: Generate Images by Optimizing What You Reward
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
by: Razzhigaev, Anton, et al.
Published: (2025)
by: Razzhigaev, Anton, et al.
Published: (2025)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
by: Korznikov, Anton, et al.
Published: (2025)
by: Korznikov, Anton, et al.
Published: (2025)
Caption This, Reason That: VLMs Caught in the Middle
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
by: Khrulkov, Valentin, et al.
Published: (2025)
by: Khrulkov, Valentin, et al.
Published: (2025)
Similar Items
-
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
by: Gambashidze, Alexander, et al.
Published: (2025) -
CoMa: Contextual Massing Generation with Vision-Language Models
by: Maslov, Evgenii, et al.
Published: (2026) -
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
by: Li, Pengyi, et al.
Published: (2025) -
Spread them Apart: Towards Robust Watermarking of Generated Content
by: Pautov, Mikhail, et al.
Published: (2025) -
Simple Vision-Language Math Reasoning via Rendered Text
by: Skripkin, Matvey, et al.
Published: (2025)