Language Repository for Long Video Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Kahatapitiya, Kumara, Ranasinghe, Kanchana, Park, Jongwoo, Ryoo, Michael S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understanding Long Videos with Multimodal Language Models
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2024)
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2024)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
di: Park, Jongwoo, et al.
Pubblicazione: (2024)
di: Park, Jongwoo, et al.
Pubblicazione: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2023)
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2023)
Pixel Motion as Universal Representation for Robot Control
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2025)
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2025)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
di: Mata, Cristina, et al.
Pubblicazione: (2025)
di: Mata, Cristina, et al.
Pubblicazione: (2025)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
Pixel Motion Diffusion is What We Need for Robot Control
di: Nguyen, E-Ro, et al.
Pubblicazione: (2025)
di: Nguyen, E-Ro, et al.
Pubblicazione: (2025)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2024)
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2024)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
Future Optical Flow Prediction Improves Robot Control & Video Generation
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2026)
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2026)
LatentCRF: Continuous CRF for Efficient Latent Diffusion
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2024)
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2024)
Object-Centric Diffusion for Efficient Video Editing
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
di: Xu, Jiaqi, et al.
Pubblicazione: (2023)
di: Xu, Jiaqi, et al.
Pubblicazione: (2023)
HeightLane: BEV Heightmap guided 3D Lane Detection
di: Park, Chaesong, et al.
Pubblicazione: (2024)
di: Park, Chaesong, et al.
Pubblicazione: (2024)
LC-Flow: Learning Local Continuous Optical Flow and Confidence from events
di: Jeon, Gunwoo, et al.
Pubblicazione: (2026)
di: Jeon, Gunwoo, et al.
Pubblicazione: (2026)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
di: Watawana, Hasindri, et al.
Pubblicazione: (2024)
di: Watawana, Hasindri, et al.
Pubblicazione: (2024)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
di: Shen, Xiaoqian, et al.
Pubblicazione: (2024)
di: Shen, Xiaoqian, et al.
Pubblicazione: (2024)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
Streaming Long Video Understanding with Large Language Models
di: Qian, Rui, et al.
Pubblicazione: (2024)
di: Qian, Rui, et al.
Pubblicazione: (2024)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
di: An, Zhaochong, et al.
Pubblicazione: (2025)
di: An, Zhaochong, et al.
Pubblicazione: (2025)
Predicting Penalty Kick Direction Using Multi-Modal Deep Learning with Pose-Guided Attention
di: Ranasinghe, Pasindu, et al.
Pubblicazione: (2025)
di: Ranasinghe, Pasindu, et al.
Pubblicazione: (2025)
LongVLM: Efficient Long Video Understanding via Large Language Models
di: Weng, Yuetian, et al.
Pubblicazione: (2024)
di: Weng, Yuetian, et al.
Pubblicazione: (2024)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
di: Fang, Yu, et al.
Pubblicazione: (2025)
di: Fang, Yu, et al.
Pubblicazione: (2025)
Context-Aware Input Orchestration for Video Inpainting
di: Kim, Hoyoung, et al.
Pubblicazione: (2024)
di: Kim, Hoyoung, et al.
Pubblicazione: (2024)
Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation
di: De Silva, Ulindu, et al.
Pubblicazione: (2025)
di: De Silva, Ulindu, et al.
Pubblicazione: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
di: Lin, Jingyang, et al.
Pubblicazione: (2025)
di: Lin, Jingyang, et al.
Pubblicazione: (2025)
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
di: Shu, Yan, et al.
Pubblicazione: (2024)
di: Shu, Yan, et al.
Pubblicazione: (2024)
ColFigPhotoAttnNet: Reliable Finger Photo Presentation Attack Detection Leveraging Window-Attention on Color Spaces
di: Vurity, Anudeep, et al.
Pubblicazione: (2025)
di: Vurity, Anudeep, et al.
Pubblicazione: (2025)
SC-Lane: Slope-aware and Consistent Road Height Estimation Framework for 3D Lane Detection
di: Park, Chaesong, et al.
Pubblicazione: (2025)
di: Park, Chaesong, et al.
Pubblicazione: (2025)
Integrating Meshes and 3D Gaussians for Indoor Scene Reconstruction with SAM Mask Guidance
di: Kim, Jiyeop, et al.
Pubblicazione: (2024)
di: Kim, Jiyeop, et al.
Pubblicazione: (2024)
Video Token Merging for Long-form Video Understanding
di: Lee, Seon-Ho, et al.
Pubblicazione: (2024)
di: Lee, Seon-Ho, et al.
Pubblicazione: (2024)
Linear Scaling Video VLMs for Long Video Understanding
di: Eyzaguirre, Cristobal, et al.
Pubblicazione: (2026)
di: Eyzaguirre, Cristobal, et al.
Pubblicazione: (2026)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
di: Tang, Canhui, et al.
Pubblicazione: (2025)
di: Tang, Canhui, et al.
Pubblicazione: (2025)
Image Translation with Kernel Prediction Networks for Semantic Segmentation
di: Mata, Cristina, et al.
Pubblicazione: (2025)
di: Mata, Cristina, et al.
Pubblicazione: (2025)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
di: Ma, Martin Q., et al.
Pubblicazione: (2026)
di: Ma, Martin Q., et al.
Pubblicazione: (2026)
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
di: Ranasinghe, Yasiru, et al.
Pubblicazione: (2025)
di: Ranasinghe, Yasiru, et al.
Pubblicazione: (2025)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
di: Qiu, Jihao, et al.
Pubblicazione: (2026)
di: Qiu, Jihao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Understanding Long Videos with Multimodal Language Models
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2024) -
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
di: Park, Jongwoo, et al.
Pubblicazione: (2024) -
VicTR: Video-conditioned Text Representations for Activity Recognition
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2023) -
Pixel Motion as Universal Representation for Robot Control
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2025) -
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
di: Li, Xiang, et al.
Pubblicazione: (2024)