TTRV: Test-Time Reinforcement Learning for Vision Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Singh, Akshit, Marjit, Shyam, Lin, Wei, Gavrikov, Paul, Yeung-Levy, Serena, Kuehne, Hilde, Feris, Rogerio, Doveh, Sivan, Glass, James, Mirza, M. Jehanzeb |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
di: Gavrikov, Paul, et al.
Pubblicazione: (2025)
di: Gavrikov, Paul, et al.
Pubblicazione: (2025)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
di: Mehta, Videet, et al.
Pubblicazione: (2026)
di: Mehta, Videet, et al.
Pubblicazione: (2026)
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
di: Selch, Lukas, et al.
Pubblicazione: (2025)
di: Selch, Lukas, et al.
Pubblicazione: (2025)
Teaching VLMs to Localize Specific Objects from In-context Examples
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
Towards Audio Token Compression in Large Audio Language Models
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2025)
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
di: Araujo, Edson, et al.
Pubblicazione: (2026)
di: Araujo, Edson, et al.
Pubblicazione: (2026)
Comparison Visual Instruction Tuning
di: Lin, Wei, et al.
Pubblicazione: (2024)
di: Lin, Wei, et al.
Pubblicazione: (2024)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
di: Mirza, M. Jehanzeb, et al.
Pubblicazione: (2024)
State-Space Large Audio Language Models
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2024)
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2024)
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
di: Rouditchenko, Andrew, et al.
Pubblicazione: (2025)
Towards Multimodal In-Context Learning for Vision & Language Models
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
di: Huang, Irene, et al.
Pubblicazione: (2024)
di: Huang, Irene, et al.
Pubblicazione: (2024)
Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
di: Spoecklberger, Johannes, et al.
Pubblicazione: (2025)
di: Spoecklberger, Johannes, et al.
Pubblicazione: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
di: Sui, Elaine, et al.
Pubblicazione: (2024)
di: Sui, Elaine, et al.
Pubblicazione: (2024)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
di: Hansen, Jacob, et al.
Pubblicazione: (2025)
di: Hansen, Jacob, et al.
Pubblicazione: (2025)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
di: Chen, Brian, et al.
Pubblicazione: (2023)
di: Chen, Brian, et al.
Pubblicazione: (2023)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
di: Araujo, Edson, et al.
Pubblicazione: (2025)
di: Araujo, Edson, et al.
Pubblicazione: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
di: Jahagirdar, Soumya, et al.
Pubblicazione: (2026)
di: Jahagirdar, Soumya, et al.
Pubblicazione: (2026)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
di: Bousselham, Walid, et al.
Pubblicazione: (2026)
di: Bousselham, Walid, et al.
Pubblicazione: (2026)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
di: Aklilu, Josiah, et al.
Pubblicazione: (2024)
di: Aklilu, Josiah, et al.
Pubblicazione: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
di: Endo, Mark, et al.
Pubblicazione: (2024)
di: Endo, Mark, et al.
Pubblicazione: (2024)
Tool Verification for Test-Time Reinforcement Learning
di: Liao, Ruotong, et al.
Pubblicazione: (2026)
di: Liao, Ruotong, et al.
Pubblicazione: (2026)
TimeLogic: A Temporal Logic Benchmark for Video QA
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
Can We Talk Models Into Seeing the World Differently?
di: Gavrikov, Paul, et al.
Pubblicazione: (2024)
di: Gavrikov, Paul, et al.
Pubblicazione: (2024)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
di: Shabtay, Nimrod, et al.
Pubblicazione: (2024)
di: Shabtay, Nimrod, et al.
Pubblicazione: (2024)
FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation
di: Yashwanth, M, et al.
Pubblicazione: (2025)
di: Yashwanth, M, et al.
Pubblicazione: (2025)
TTT-KD: Test-Time Training for 3D Semantic Segmentation through Knowledge Distillation from Foundation Models
di: Weijler, Lisa, et al.
Pubblicazione: (2024)
di: Weijler, Lisa, et al.
Pubblicazione: (2024)
O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
di: Gupta, Rishi, et al.
Pubblicazione: (2025)
di: Gupta, Rishi, et al.
Pubblicazione: (2025)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives
di: Mathur, Nityanand, et al.
Pubblicazione: (2023)
di: Mathur, Nityanand, et al.
Pubblicazione: (2023)
$\texttt{BATCLIP}$: Bimodal Online Test-Time Adaptation for CLIP
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2024)
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2024)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
di: Endo, Mark, et al.
Pubblicazione: (2025)
di: Endo, Mark, et al.
Pubblicazione: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
di: Gu, Jeffrey, et al.
Pubblicazione: (2025)
di: Gu, Jeffrey, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
di: Gavrikov, Paul, et al.
Pubblicazione: (2025) -
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
di: Mehta, Videet, et al.
Pubblicazione: (2026) -
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
di: Selch, Lukas, et al.
Pubblicazione: (2025) -
Teaching VLMs to Localize Specific Objects from In-context Examples
di: Doveh, Sivan, et al.
Pubblicazione: (2024) -
Towards Audio Token Compression in Large Audio Language Models
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2025)