Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Anmin, Zhang, Nan, Tao, Wei, Qu, Xiaoyang, Li, Guokuan, Wan, Jiguang, Wang, Jianzong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
von: Lu, Haocheng, et al.
Veröffentlicht: (2026)
von: Lu, Haocheng, et al.
Veröffentlicht: (2026)
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
von: He, Shenglin, et al.
Veröffentlicht: (2024)
von: He, Shenglin, et al.
Veröffentlicht: (2024)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
von: Tao, Wei, et al.
Veröffentlicht: (2026)
von: Tao, Wei, et al.
Veröffentlicht: (2026)
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
von: Tao, Wei, et al.
Veröffentlicht: (2024)
von: Tao, Wei, et al.
Veröffentlicht: (2024)
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic Segmentation
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
von: Liu, Chuhang, et al.
Veröffentlicht: (2026)
von: Liu, Chuhang, et al.
Veröffentlicht: (2026)
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
Federated Domain Generalization with Domain-specific Soft Prompts Generation
von: Wu, Jianhan, et al.
Veröffentlicht: (2025)
von: Wu, Jianhan, et al.
Veröffentlicht: (2025)
Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-based Planner and Graph-based Policy
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models
von: Guo, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoyang, et al.
Veröffentlicht: (2025)
ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
von: Zhang, Xulong, et al.
Veröffentlicht: (2024)
von: Zhang, Xulong, et al.
Veröffentlicht: (2024)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
von: Yokoi, Shingo, et al.
Veröffentlicht: (2025)
von: Yokoi, Shingo, et al.
Veröffentlicht: (2025)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
von: Sun, Haoyi, et al.
Veröffentlicht: (2026)
von: Sun, Haoyi, et al.
Veröffentlicht: (2026)
HAWAII: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
SafeTriage: Facial Video De-identification for Privacy-Preserving Stroke Triage
von: Cai, Tongan, et al.
Veröffentlicht: (2025)
von: Cai, Tongan, et al.
Veröffentlicht: (2025)
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
EVLM: An Efficient Vision-Language Model for Visual Understanding
von: Chen, Kaibing, et al.
Veröffentlicht: (2024)
von: Chen, Kaibing, et al.
Veröffentlicht: (2024)
Conflict Adaptation in Vision-Language Models
von: Hu, Xiaoyang
Veröffentlicht: (2025)
von: Hu, Xiaoyang
Veröffentlicht: (2025)
PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology
von: Huang, Yating, et al.
Veröffentlicht: (2025)
von: Huang, Yating, et al.
Veröffentlicht: (2025)
Tango: Taming Visual Signals for Efficient Video Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2026)
von: Yin, Shukang, et al.
Veröffentlicht: (2026)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision Language Models
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
Enhancing Visual Programming for Visual Reasoning via Probabilistic Graphs
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training
von: Tao, Yiyi, et al.
Veröffentlicht: (2024)
von: Tao, Yiyi, et al.
Veröffentlicht: (2024)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
von: Zhang, Bin, et al.
Veröffentlicht: (2025) -
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
von: Li, Junjie, et al.
Veröffentlicht: (2025) -
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
von: Lu, Haocheng, et al.
Veröffentlicht: (2026) -
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
von: He, Shenglin, et al.
Veröffentlicht: (2024) -
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
von: Tao, Wei, et al.
Veröffentlicht: (2026)