RVLM: Recursive Vision-Language Models with Adaptive Depth
Fuente:
arXiv
Salvato in:
| Autori principali: | Mayumu, Nicanor, Khan, Zeenath, Stephens, Melodena, Mukala, Patrick, Oroumchian, Farhad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models
di: Mayumu, Nicanor, et al.
Pubblicazione: (2026)
di: Mayumu, Nicanor, et al.
Pubblicazione: (2026)
F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
di: Bi, Hanbo, et al.
Pubblicazione: (2025)
di: Bi, Hanbo, et al.
Pubblicazione: (2025)
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
di: Chen, Bolei, et al.
Pubblicazione: (2025)
di: Chen, Bolei, et al.
Pubblicazione: (2025)
DepthLM: Metric Depth From Vision Language Models
di: Cai, Zhipeng, et al.
Pubblicazione: (2025)
di: Cai, Zhipeng, et al.
Pubblicazione: (2025)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
di: Yuan, Tianyuan, et al.
Pubblicazione: (2025)
di: Yuan, Tianyuan, et al.
Pubblicazione: (2025)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
di: Hannan, Tanveer, et al.
Pubblicazione: (2024)
di: Hannan, Tanveer, et al.
Pubblicazione: (2024)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
di: Lin, Tao, et al.
Pubblicazione: (2026)
di: Lin, Tao, et al.
Pubblicazione: (2026)
Scaling Down to Scale Up: Towards Operationally-Efficient and Deployable Clinical Models via Cross-Modal Low-Rank Adaptation for Medical Vision-Language Models
di: Alzubaidi, Thuraya, et al.
Pubblicazione: (2025)
di: Alzubaidi, Thuraya, et al.
Pubblicazione: (2025)
Seeing the Unseen: Towards Zero-Shot Inspection for Wind Turbine Blades using Knowledge-Augmented Vision Language Models
di: Zhang, Yang, et al.
Pubblicazione: (2025)
di: Zhang, Yang, et al.
Pubblicazione: (2025)
Vision-Language Embodiment for Monocular Depth Estimation
di: Zhang, Jinchang, et al.
Pubblicazione: (2025)
di: Zhang, Jinchang, et al.
Pubblicazione: (2025)
Vision Tiny Recursion Model (ViTRM): Parameter-Efficient Image Classification via Recursive State Refinement
di: Akazan, Ange-Clément, et al.
Pubblicazione: (2026)
di: Akazan, Ange-Clément, et al.
Pubblicazione: (2026)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
di: Chen, Shiming, et al.
Pubblicazione: (2025)
di: Chen, Shiming, et al.
Pubblicazione: (2025)
Recursive Block-Diagonal Coupling for Resource-Efficient Training of Vision Models
di: Henry, Maxim, et al.
Pubblicazione: (2026)
di: Henry, Maxim, et al.
Pubblicazione: (2026)
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
di: Danier, Duolikun, et al.
Pubblicazione: (2024)
di: Danier, Duolikun, et al.
Pubblicazione: (2024)
Recursive Vision Transformer with Dynamic Depth and Width Adjustment for Resource-Efficient Image Semantic Communication
di: Zhang, Zhilong, et al.
Pubblicazione: (2026)
di: Zhang, Zhilong, et al.
Pubblicazione: (2026)
Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts
di: Nooralahzadeh, Farhad, et al.
Pubblicazione: (2026)
di: Nooralahzadeh, Farhad, et al.
Pubblicazione: (2026)
Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
di: Xu, Zhen, et al.
Pubblicazione: (2025)
di: Xu, Zhen, et al.
Pubblicazione: (2025)
Depth Adaptive Efficient Visual Autoregressive Modeling
di: Li, Chunliang, et al.
Pubblicazione: (2026)
di: Li, Chunliang, et al.
Pubblicazione: (2026)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
di: Ren, Yiming, et al.
Pubblicazione: (2026)
di: Ren, Yiming, et al.
Pubblicazione: (2026)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
di: Maaz, Muhammad, et al.
Pubblicazione: (2023)
di: Maaz, Muhammad, et al.
Pubblicazione: (2023)
Vision-Language Model for Accurate Crater Detection
di: Bauer, Patrick, et al.
Pubblicazione: (2026)
di: Bauer, Patrick, et al.
Pubblicazione: (2026)
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection
di: Khan, Sohail Ahmed, et al.
Pubblicazione: (2024)
di: Khan, Sohail Ahmed, et al.
Pubblicazione: (2024)
DPA: Dual Prototypes Alignment for Unsupervised Adaptation of Vision-Language Models
di: Ali, Eman, et al.
Pubblicazione: (2024)
di: Ali, Eman, et al.
Pubblicazione: (2024)
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
di: Danish, Muhammad Sohail, et al.
Pubblicazione: (2024)
di: Danish, Muhammad Sohail, et al.
Pubblicazione: (2024)
O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models
di: Sharifdeen, Ashshak, et al.
Pubblicazione: (2025)
di: Sharifdeen, Ashshak, et al.
Pubblicazione: (2025)
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
di: Zeng, Weili, et al.
Pubblicazione: (2025)
di: Zeng, Weili, et al.
Pubblicazione: (2025)
Vision Language Model for Interpretable and Fine-grained Detection of Safety Compliance in Diverse Workplaces
di: Chen, Zhiling, et al.
Pubblicazione: (2024)
di: Chen, Zhiling, et al.
Pubblicazione: (2024)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
di: Demidov, Dmitry, et al.
Pubblicazione: (2025)
di: Demidov, Dmitry, et al.
Pubblicazione: (2025)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
di: Choi, In Chong, et al.
Pubblicazione: (2026)
di: Choi, In Chong, et al.
Pubblicazione: (2026)
AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models
di: Li, Zhiwei, et al.
Pubblicazione: (2026)
di: Li, Zhiwei, et al.
Pubblicazione: (2026)
Data Adaptive Traceback for Vision-Language Foundation Models in Image Classification
di: Peng, Wenshuo, et al.
Pubblicazione: (2024)
di: Peng, Wenshuo, et al.
Pubblicazione: (2024)
Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
di: Nguyen, Khanh-Binh, et al.
Pubblicazione: (2025)
di: Nguyen, Khanh-Binh, et al.
Pubblicazione: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
di: Li, YuQian, et al.
Pubblicazione: (2025)
di: Li, YuQian, et al.
Pubblicazione: (2025)
Calibration-Aware Prompt Learning for Medical Vision-Language Models
di: Basu, Abhishek, et al.
Pubblicazione: (2025)
di: Basu, Abhishek, et al.
Pubblicazione: (2025)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
di: Thawakar, Omkar, et al.
Pubblicazione: (2023)
di: Thawakar, Omkar, et al.
Pubblicazione: (2023)
Language as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation
di: Zhan, Mingxia, et al.
Pubblicazione: (2026)
di: Zhan, Mingxia, et al.
Pubblicazione: (2026)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
di: Li, Yixuan, et al.
Pubblicazione: (2025)
di: Li, Yixuan, et al.
Pubblicazione: (2025)
The Spatial Blindspot of Vision-Language Models
di: Alam, Nahid, et al.
Pubblicazione: (2026)
di: Alam, Nahid, et al.
Pubblicazione: (2026)
Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question Answering
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models
di: Mayumu, Nicanor, et al.
Pubblicazione: (2026) -
F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
di: Bi, Hanbo, et al.
Pubblicazione: (2025) -
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
di: Chen, Bolei, et al.
Pubblicazione: (2025) -
DepthLM: Metric Depth From Vision Language Models
di: Cai, Zhipeng, et al.
Pubblicazione: (2025) -
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
di: Yuan, Tianyuan, et al.
Pubblicazione: (2025)