From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gumbsch, Christian, Barcellona, Leonardo, Schünemann, Lennard, Karageorgis, Platon, Zadaianchuk, Andrii, Wang, Zehao, Zakharov, Sergey, Despinoy, Fabien, Aljundi, Rahaf, Gavves, Efstratios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
by: Zadaianchuk, Andrii, et al.
Published: (2026)
by: Zadaianchuk, Andrii, et al.
Published: (2026)
Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
by: Barcellona, Leonardo, et al.
Published: (2024)
by: Barcellona, Leonardo, et al.
Published: (2024)
Annotation Free Semantic Segmentation with Vision Foundation Models
by: Seifi, Soroush, et al.
Published: (2024)
by: Seifi, Soroush, et al.
Published: (2024)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Personalization Toolkit: Training Free Personalization of Large Vision Language Models
by: Seifi, Soroush, et al.
Published: (2025)
by: Seifi, Soroush, et al.
Published: (2025)
SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models
by: Sancaktar, Cansu, et al.
Published: (2025)
by: Sancaktar, Cansu, et al.
Published: (2025)
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation
by: Gkotsi, Polytimi Anna, et al.
Published: (2026)
by: Gkotsi, Polytimi Anna, et al.
Published: (2026)
The Effectiveness of Prosocial Rewards and Cash Rewards in Internal Whistleblower Programs
by: Khim Kelly, et al.
Published: (2026)
by: Khim Kelly, et al.
Published: (2026)
Mechanistic Interpretability for AI Safety -- A Review
by: Bereska, Leonard, et al.
Published: (2024)
by: Bereska, Leonard, et al.
Published: (2024)
Incremental Object-Based Novelty Detection with Feedback Loop
by: Caldarella, Simone, et al.
Published: (2023)
by: Caldarella, Simone, et al.
Published: (2023)
When Data Falls Short: Grokking Below the Critical Threshold
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
by: Baxevanakis, Spiros, et al.
Published: (2026)
by: Baxevanakis, Spiros, et al.
Published: (2026)
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
by: Zadaianchuk, Andrii, et al.
Published: (2023)
by: Zadaianchuk, Andrii, et al.
Published: (2023)
Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
by: He, Lehan, et al.
Published: (2024)
by: He, Lehan, et al.
Published: (2024)
Ai-Sampler: Adversarial Learning of Markov kernels with involutive maps
by: Egorov, Evgenii, et al.
Published: (2024)
by: Egorov, Evgenii, et al.
Published: (2024)
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
by: Kofinas, Miltiadis, et al.
Published: (2024)
by: Kofinas, Miltiadis, et al.
Published: (2024)
Mechanistic Neural Networks for Scientific Machine Learning
by: Pervez, Adeel, et al.
Published: (2024)
by: Pervez, Adeel, et al.
Published: (2024)
Mechanistic PDE Networks for Discovery of Governing Equations
by: Pervez, Adeel, et al.
Published: (2025)
by: Pervez, Adeel, et al.
Published: (2025)
Designing Rewards for Rewarding Designs: Demonstrating the Impact of Rewards on the Creative Design Process
by: Nath, Surabhi S, et al.
Published: (2026)
by: Nath, Surabhi S, et al.
Published: (2026)
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
by: Liu, Gongye, et al.
Published: (2026)
by: Liu, Gongye, et al.
Published: (2026)
The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
by: Huang, Sukai, et al.
Published: (2024)
by: Huang, Sukai, et al.
Published: (2024)
Online In-Context Distillation for Low-Resource Vision Language Models
by: Kang, Zhiqi, et al.
Published: (2025)
by: Kang, Zhiqi, et al.
Published: (2025)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
by: Dorovatas, Vaggelis, et al.
Published: (2025)
by: Dorovatas, Vaggelis, et al.
Published: (2025)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
by: Caldarella, Simone, et al.
Published: (2024)
by: Caldarella, Simone, et al.
Published: (2024)
Overcoming Generic Knowledge Loss with Selective Parameter Update
by: Zhang, Wenxuan, et al.
Published: (2023)
by: Zhang, Wenxuan, et al.
Published: (2023)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
by: Hu, Zhiyuan, et al.
Published: (2025)
by: Hu, Zhiyuan, et al.
Published: (2025)
Physics-Guided Radiotherapy Treatment Planning with Deep Learning
by: Achlatis, Stefanos, et al.
Published: (2025)
by: Achlatis, Stefanos, et al.
Published: (2025)
Roto-translated Local Coordinate Frames For Interacting Dynamical Systems
by: Kofinas, Miltiadis, et al.
Published: (2021)
by: Kofinas, Miltiadis, et al.
Published: (2021)
Cross-Layer Attention Probing for Fine-Grained Hallucination Detection
by: Suresh, Malavika, et al.
Published: (2025)
by: Suresh, Malavika, et al.
Published: (2025)
Any-Resolution AI-Generated Image Detection by Spectral Learning
by: Karageorgiou, Dimitrios, et al.
Published: (2024)
by: Karageorgiou, Dimitrios, et al.
Published: (2024)
Amortized Equation Discovery in Hybrid Dynamical Systems
by: Liu, Yongtuo, et al.
Published: (2024)
by: Liu, Yongtuo, et al.
Published: (2024)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
by: Wang, Haozhe, et al.
Published: (2026)
by: Wang, Haozhe, et al.
Published: (2026)
Towards Uniformity and Alignment for Multimodal Representation Learning
by: Yin, Wenzhe, et al.
Published: (2026)
by: Yin, Wenzhe, et al.
Published: (2026)
Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence
by: Yin, Wenzhe, et al.
Published: (2025)
by: Yin, Wenzhe, et al.
Published: (2025)
Efficient Few-Shot Continual Learning in Vision-Language Models
by: Panos, Aristeidis, et al.
Published: (2025)
by: Panos, Aristeidis, et al.
Published: (2025)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025)
by: Gupta, Gunshi, et al.
Published: (2025)
Similar Items
-
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
by: Zadaianchuk, Andrii, et al.
Published: (2026) -
Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
by: Barcellona, Leonardo, et al.
Published: (2024) -
Annotation Free Semantic Segmentation with Vision Foundation Models
by: Seifi, Soroush, et al.
Published: (2024) -
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025) -
Personalization Toolkit: Training Free Personalization of Large Vision Language Models
by: Seifi, Soroush, et al.
Published: (2025)