Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
Fuente:
arXiv
Guardado en:
| Autores principales: | Lu, Haocheng, Zhang, Nan, Tao, Wei, Qu, Xiaoyang, Li, Guokuan, Wan, Jiguang, Wang, Jianzong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
por: Wang, Anmin, et al.
Publicado: (2026)
por: Wang, Anmin, et al.
Publicado: (2026)
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
por: Li, Junjie, et al.
Publicado: (2025)
por: Li, Junjie, et al.
Publicado: (2025)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
por: Zhang, Bin, et al.
Publicado: (2025)
por: Zhang, Bin, et al.
Publicado: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
por: Tao, Wei, et al.
Publicado: (2026)
por: Tao, Wei, et al.
Publicado: (2026)
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
por: Tao, Wei, et al.
Publicado: (2025)
por: Tao, Wei, et al.
Publicado: (2025)
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
por: He, Shenglin, et al.
Publicado: (2024)
por: He, Shenglin, et al.
Publicado: (2024)
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
por: Tao, Wei, et al.
Publicado: (2024)
por: Tao, Wei, et al.
Publicado: (2024)
BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic Segmentation
por: Tao, Wei, et al.
Publicado: (2025)
por: Tao, Wei, et al.
Publicado: (2025)
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
por: Zhang, Bin, et al.
Publicado: (2025)
por: Zhang, Bin, et al.
Publicado: (2025)
MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs
por: Tao, Wei, et al.
Publicado: (2025)
por: Tao, Wei, et al.
Publicado: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
por: Di, Shangzhe, et al.
Publicado: (2025)
por: Di, Shangzhe, et al.
Publicado: (2025)
Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning
por: Jia, Ziqi, et al.
Publicado: (2025)
por: Jia, Ziqi, et al.
Publicado: (2025)
CogStream: Context-guided Streaming Video Question Answering
por: Zhao, Zicheng, et al.
Publicado: (2025)
por: Zhao, Zicheng, et al.
Publicado: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
por: Chen, Yilong, et al.
Publicado: (2025)
por: Chen, Yilong, et al.
Publicado: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2025)
por: Zhou, Sheng, et al.
Publicado: (2025)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
por: Shi, Jiaqi, et al.
Publicado: (2026)
por: Shi, Jiaqi, et al.
Publicado: (2026)
Federated Domain Generalization with Domain-specific Soft Prompts Generation
por: Wu, Jianhan, et al.
Publicado: (2025)
por: Wu, Jianhan, et al.
Publicado: (2025)
Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-based Planner and Graph-based Policy
por: Jia, Ziqi, et al.
Publicado: (2025)
por: Jia, Ziqi, et al.
Publicado: (2025)
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
por: Lu, Renjie, et al.
Publicado: (2026)
por: Lu, Renjie, et al.
Publicado: (2026)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
por: Lu, Renjie, et al.
Publicado: (2026)
por: Lu, Renjie, et al.
Publicado: (2026)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
por: Shi, Jiaqi, et al.
Publicado: (2026)
por: Shi, Jiaqi, et al.
Publicado: (2026)
Scene-Text Grounding for Text-Based Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2024)
por: Zhou, Sheng, et al.
Publicado: (2024)
Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
por: Shi, Yumeng, et al.
Publicado: (2025)
por: Shi, Yumeng, et al.
Publicado: (2025)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
por: Li, Haopeng, et al.
Publicado: (2024)
por: Li, Haopeng, et al.
Publicado: (2024)
Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering
por: Tao, Qian, et al.
Publicado: (2025)
por: Tao, Qian, et al.
Publicado: (2025)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
por: Yang, Xuyi, et al.
Publicado: (2025)
por: Yang, Xuyi, et al.
Publicado: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
por: Drago, Mauro Orazio, et al.
Publicado: (2025)
por: Drago, Mauro Orazio, et al.
Publicado: (2025)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
por: Kwon, Minchan, et al.
Publicado: (2026)
por: Kwon, Minchan, et al.
Publicado: (2026)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
por: Liu, Chuhang, et al.
Publicado: (2026)
por: Liu, Chuhang, et al.
Publicado: (2026)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
por: Jiang, Yuanyuan, et al.
Publicado: (2024)
por: Jiang, Yuanyuan, et al.
Publicado: (2024)
Enhancing Scene Transition Awareness in Video Generation via Post-Training
por: Shen, Hanwen, et al.
Publicado: (2025)
por: Shen, Hanwen, et al.
Publicado: (2025)
3D Question Answering for City Scene Understanding
por: Sun, Penglei, et al.
Publicado: (2024)
por: Sun, Penglei, et al.
Publicado: (2024)
Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference
por: Tao, Wei, et al.
Publicado: (2025)
por: Tao, Wei, et al.
Publicado: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
por: Wang, Xingrui, et al.
Publicado: (2024)
por: Wang, Xingrui, et al.
Publicado: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
por: Kim, Hongyeob, et al.
Publicado: (2025)
por: Kim, Hongyeob, et al.
Publicado: (2025)
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
por: Dahal, Ashim, et al.
Publicado: (2025)
por: Dahal, Ashim, et al.
Publicado: (2025)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
por: Tao, Mingzhe, et al.
Publicado: (2026)
por: Tao, Mingzhe, et al.
Publicado: (2026)
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
por: Wu, Tao, et al.
Publicado: (2024)
por: Wu, Tao, et al.
Publicado: (2024)
Grounded Question-Answering in Long Egocentric Videos
por: Di, Shangzhe, et al.
Publicado: (2023)
por: Di, Shangzhe, et al.
Publicado: (2023)
Agentic Keyframe Search for Video Question Answering
por: Fan, Sunqi, et al.
Publicado: (2025)
por: Fan, Sunqi, et al.
Publicado: (2025)
Ejemplares similares
-
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
por: Wang, Anmin, et al.
Publicado: (2026) -
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
por: Li, Junjie, et al.
Publicado: (2025) -
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
por: Zhang, Bin, et al.
Publicado: (2025) -
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
por: Tao, Wei, et al.
Publicado: (2026) -
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
por: Tao, Wei, et al.
Publicado: (2025)