Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Jongwoo, Ranasinghe, Kanchana, Kahatapitiya, Kumara, Ryu, Wonjeong, Kim, Donghyun, Ryoo, Michael S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
Pixel Motion as Universal Representation for Robot Control
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Pixel Motion Diffusion is What We Need for Robot Control
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2025)
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2025)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
LatentCRF: Continuous CRF for Efficient Latent Diffusion
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
von: Yu, Seungjun, et al.
Veröffentlicht: (2025)
von: Yu, Seungjun, et al.
Veröffentlicht: (2025)
Future Optical Flow Prediction Improves Robot Control & Video Generation
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
MetaHarm: Harmful YouTube Video Dataset Annotated by Domain Experts, GPT-4-Turbo, and Crowdworkers
von: Jo, Wonjeong, et al.
Veröffentlicht: (2025)
von: Jo, Wonjeong, et al.
Veröffentlicht: (2025)
Moment Sampling in Video LLMs for Long-Form Video QA
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
SUPER-AD: Semantic Uncertainty-aware Planning for End-to-End Robust Autonomous Driving
von: Ryu, Wonjeong, et al.
Veröffentlicht: (2025)
von: Ryu, Wonjeong, et al.
Veröffentlicht: (2025)
NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
von: Park, Jongwoo, et al.
Veröffentlicht: (2026)
von: Park, Jongwoo, et al.
Veröffentlicht: (2026)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
Hierarchical Memory for Long Video QA
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
CaptionQA: Is Your Caption as Useful as the Image Itself?
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
ReSpike: Residual Frames-based Hybrid Spiking Neural Networks for Efficient Action Recognition
von: Xiao, Shiting, et al.
Veröffentlicht: (2024)
von: Xiao, Shiting, et al.
Veröffentlicht: (2024)
Think as Needed: Geometry-Driven Adaptive Perception for Autonomous Driving
von: Kim, Donghyun, et al.
Veröffentlicht: (2026)
von: Kim, Donghyun, et al.
Veröffentlicht: (2026)
Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
von: Kang, Inha, et al.
Veröffentlicht: (2025)
von: Kang, Inha, et al.
Veröffentlicht: (2025)
Context-Aware Input Orchestration for Video Inpainting
von: Kim, Hoyoung, et al.
Veröffentlicht: (2024)
von: Kim, Hoyoung, et al.
Veröffentlicht: (2024)
Integrating Meshes and 3D Gaussians for Indoor Scene Reconstruction with SAM Mask Guidance
von: Kim, Jiyeop, et al.
Veröffentlicht: (2024)
von: Kim, Jiyeop, et al.
Veröffentlicht: (2024)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
von: Sun, Guangyu, et al.
Veröffentlicht: (2025)
von: Sun, Guangyu, et al.
Veröffentlicht: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
von: Song, Zhende, et al.
Veröffentlicht: (2024)
von: Song, Zhende, et al.
Veröffentlicht: (2024)
Weakly Supervised Video Scene Graph Generation via Natural Language Supervision
von: Kim, Kibum, et al.
Veröffentlicht: (2025)
von: Kim, Kibum, et al.
Veröffentlicht: (2025)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
von: Jo, Claire Wonjeong, et al.
Veröffentlicht: (2024)
von: Jo, Claire Wonjeong, et al.
Veröffentlicht: (2024)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
YTCommentQA: Video Question Answerability in Instructional Videos
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
LC-Flow: Learning Local Continuous Optical Flow and Confidence from events
von: Jeon, Gunwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Gunwoo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024) -
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024) -
VicTR: Video-conditioned Text Representations for Activity Recognition
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023) -
Pixel Motion as Universal Representation for Robot Control
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025) -
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
von: Li, Xiang, et al.
Veröffentlicht: (2024)