TTRV: Test-Time Reinforcement Learning for Vision Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Akshit, Marjit, Shyam, Lin, Wei, Gavrikov, Paul, Yeung-Levy, Serena, Kuehne, Hilde, Feris, Rogerio, Doveh, Sivan, Glass, James, Mirza, M. Jehanzeb |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
von: Gavrikov, Paul, et al.
Veröffentlicht: (2025)
von: Gavrikov, Paul, et al.
Veröffentlicht: (2025)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
von: Mehta, Videet, et al.
Veröffentlicht: (2026)
von: Mehta, Videet, et al.
Veröffentlicht: (2026)
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
von: Selch, Lukas, et al.
Veröffentlicht: (2025)
von: Selch, Lukas, et al.
Veröffentlicht: (2025)
Teaching VLMs to Localize Specific Objects from In-context Examples
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
Towards Audio Token Compression in Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2025)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2026)
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2026)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
von: Mirza, M. Jehanzeb, et al.
Veröffentlicht: (2024)
von: Mirza, M. Jehanzeb, et al.
Veröffentlicht: (2024)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
von: Mirza, M. Jehanzeb, et al.
Veröffentlicht: (2024)
von: Mirza, M. Jehanzeb, et al.
Veröffentlicht: (2024)
State-Space Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
Towards Multimodal In-Context Learning for Vision & Language Models
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
von: Huang, Irene, et al.
Veröffentlicht: (2024)
von: Huang, Irene, et al.
Veröffentlicht: (2024)
Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
von: Spoecklberger, Johannes, et al.
Veröffentlicht: (2025)
von: Spoecklberger, Johannes, et al.
Veröffentlicht: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
von: Chen, Brian, et al.
Veröffentlicht: (2023)
von: Chen, Brian, et al.
Veröffentlicht: (2023)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
von: Bousselham, Walid, et al.
Veröffentlicht: (2026)
von: Bousselham, Walid, et al.
Veröffentlicht: (2026)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
von: Endo, Mark, et al.
Veröffentlicht: (2024)
von: Endo, Mark, et al.
Veröffentlicht: (2024)
Tool Verification for Test-Time Reinforcement Learning
von: Liao, Ruotong, et al.
Veröffentlicht: (2026)
von: Liao, Ruotong, et al.
Veröffentlicht: (2026)
TimeLogic: A Temporal Logic Benchmark for Video QA
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
NegVQA: Can Vision Language Models Understand Negation?
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
Can We Talk Models Into Seeing the World Differently?
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation
von: Yashwanth, M, et al.
Veröffentlicht: (2025)
von: Yashwanth, M, et al.
Veröffentlicht: (2025)
TTT-KD: Test-Time Training for 3D Semantic Segmentation through Knowledge Distillation from Foundation Models
von: Weijler, Lisa, et al.
Veröffentlicht: (2024)
von: Weijler, Lisa, et al.
Veröffentlicht: (2024)
O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
von: Gupta, Rishi, et al.
Veröffentlicht: (2025)
von: Gupta, Rishi, et al.
Veröffentlicht: (2025)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives
von: Mathur, Nityanand, et al.
Veröffentlicht: (2023)
von: Mathur, Nityanand, et al.
Veröffentlicht: (2023)
$\texttt{BATCLIP}$: Bimodal Online Test-Time Adaptation for CLIP
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2024)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2024)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
von: Endo, Mark, et al.
Veröffentlicht: (2025)
von: Endo, Mark, et al.
Veröffentlicht: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
von: Gu, Jeffrey, et al.
Veröffentlicht: (2025)
von: Gu, Jeffrey, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
von: Gavrikov, Paul, et al.
Veröffentlicht: (2025) -
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
von: Mehta, Videet, et al.
Veröffentlicht: (2026) -
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
von: Selch, Lukas, et al.
Veröffentlicht: (2025) -
Teaching VLMs to Localize Specific Objects from In-context Examples
von: Doveh, Sivan, et al.
Veröffentlicht: (2024) -
Towards Audio Token Compression in Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2025)