Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tang, Zitian, Krishnan, Rohan Myer, Yu, Zhiqiu, Sun, Chen |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
How Can Objects Help Video-Language Understanding?
par: Tang, Zitian, et autres
Publié: (2025)
par: Tang, Zitian, et autres
Publié: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
par: Chen, Yuxiao, et autres
Publié: (2026)
par: Chen, Yuxiao, et autres
Publié: (2026)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
par: Yin, Yufei, et autres
Publié: (2026)
par: Yin, Yufei, et autres
Publié: (2026)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
par: Tang, Canhui, et autres
Publié: (2025)
par: Tang, Canhui, et autres
Publié: (2025)
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
par: Guo, Wenliang, et autres
Publié: (2025)
par: Guo, Wenliang, et autres
Publié: (2025)
Video Token Merging for Long-form Video Understanding
par: Lee, Seon-Ho, et autres
Publié: (2024)
par: Lee, Seon-Ho, et autres
Publié: (2024)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
par: Zhang, Hongjie, et autres
Publié: (2023)
par: Zhang, Hongjie, et autres
Publié: (2023)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
par: Wang, Youze, et autres
Publié: (2025)
par: Wang, Youze, et autres
Publié: (2025)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
par: Liu, Ruyang, et autres
Publié: (2025)
par: Liu, Ruyang, et autres
Publié: (2025)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
par: Rasheed, Hanoona, et autres
Publié: (2025)
par: Rasheed, Hanoona, et autres
Publié: (2025)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
par: Gao, Hongcheng, et autres
Publié: (2025)
par: Gao, Hongcheng, et autres
Publié: (2025)
Understanding Long Videos with Multimodal Language Models
par: Ranasinghe, Kanchana, et autres
Publié: (2024)
par: Ranasinghe, Kanchana, et autres
Publié: (2024)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
par: Jiang, Jindong, et autres
Publié: (2025)
par: Jiang, Jindong, et autres
Publié: (2025)
Multimodal Language Models for Domain-Specific Procedural Video Summarization
par: Hussain, Nafisa
Publié: (2024)
par: Hussain, Nafisa
Publié: (2024)
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
par: Chen, Seng Nam, et autres
Publié: (2026)
par: Chen, Seng Nam, et autres
Publié: (2026)
LVBench: An Extreme Long Video Understanding Benchmark
par: Wang, Weihan, et autres
Publié: (2024)
par: Wang, Weihan, et autres
Publié: (2024)
ALLVB: All-in-One Long Video Understanding Benchmark
par: Tan, Xichen, et autres
Publié: (2025)
par: Tan, Xichen, et autres
Publié: (2025)
VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
par: Liu, Pengyiang, et autres
Publié: (2026)
par: Liu, Pengyiang, et autres
Publié: (2026)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
par: Shi, Mengqi, et autres
Publié: (2026)
par: Shi, Mengqi, et autres
Publié: (2026)
Spacewalker: Traversing Representation Spaces for Fast Interactive Exploration and Annotation of Unstructured Data
par: Heine, Lukas, et autres
Publié: (2024)
par: Heine, Lukas, et autres
Publié: (2024)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
par: Liu, Shuming, et autres
Publié: (2025)
par: Liu, Shuming, et autres
Publié: (2025)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
par: Ma, David, et autres
Publié: (2025)
par: Ma, David, et autres
Publié: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
par: Chen, Guo, et autres
Publié: (2024)
par: Chen, Guo, et autres
Publié: (2024)
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding
par: Zhong, Ziqi, et autres
Publié: (2025)
par: Zhong, Ziqi, et autres
Publié: (2025)
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
par: He, Haichen, et autres
Publié: (2026)
par: He, Haichen, et autres
Publié: (2026)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
par: Chen, Dongping, et autres
Publié: (2024)
par: Chen, Dongping, et autres
Publié: (2024)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
par: Wu, Te-Lin, et autres
Publié: (2021)
par: Wu, Te-Lin, et autres
Publié: (2021)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
par: Fang, Xinyu, et autres
Publié: (2024)
par: Fang, Xinyu, et autres
Publié: (2024)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
par: Wu, Haoning, et autres
Publié: (2024)
par: Wu, Haoning, et autres
Publié: (2024)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
par: Song, Seokwon, et autres
Publié: (2025)
par: Song, Seokwon, et autres
Publié: (2025)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
Anticipating Object State Changes in Long Procedural Videos
par: Manousaki, Victoria, et autres
Publié: (2024)
par: Manousaki, Victoria, et autres
Publié: (2024)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
par: Pang, Ziqi, et autres
Publié: (2025)
par: Pang, Ziqi, et autres
Publié: (2025)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
par: Lu, Hao, et autres
Publié: (2025)
par: Lu, Hao, et autres
Publié: (2025)
VUDG: A Dataset for Video Understanding Domain Generalization
par: Wang, Ziyi, et autres
Publié: (2025)
par: Wang, Ziyi, et autres
Publié: (2025)
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
par: Li, Jiaze, et autres
Publié: (2025)
par: Li, Jiaze, et autres
Publié: (2025)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
par: Sasse, Kuleen, et autres
Publié: (2025)
par: Sasse, Kuleen, et autres
Publié: (2025)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
par: Wang, Shaoguang, et autres
Publié: (2026)
par: Wang, Shaoguang, et autres
Publié: (2026)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
par: Gao, Jianxiong, et autres
Publié: (2025)
par: Gao, Jianxiong, et autres
Publié: (2025)
Documents similaires
-
How Can Objects Help Video-Language Understanding?
par: Tang, Zitian, et autres
Publié: (2025) -
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
par: Chen, Yuxiao, et autres
Publié: (2026) -
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
par: Yin, Yufei, et autres
Publié: (2026) -
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
par: Tang, Canhui, et autres
Publié: (2025) -
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
par: Guo, Wenliang, et autres
Publié: (2025)