Time-Scaling State-Space Models for Dense Video Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Piergiovanni, AJ, Mallya, Ganesh Satish, Kim, Dahun, Angelova, Anelia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning
von: Kim, Dahun, et al.
Veröffentlicht: (2026)
von: Kim, Dahun, et al.
Veröffentlicht: (2026)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024)
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024)
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2023)
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2023)
Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications
von: Mallya, Ganesh, et al.
Veröffentlicht: (2025)
von: Mallya, Ganesh, et al.
Veröffentlicht: (2025)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
von: Kim, Dahun, et al.
Veröffentlicht: (2023)
von: Kim, Dahun, et al.
Veröffentlicht: (2023)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
Streaming Dense Video Captioning
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
von: Singh, Darshan, et al.
Veröffentlicht: (2026)
von: Singh, Darshan, et al.
Veröffentlicht: (2026)
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
von: Choi, Seung hee, et al.
Veröffentlicht: (2026)
von: Choi, Seung hee, et al.
Veröffentlicht: (2026)
Dense Video Captioning Using Unsupervised Semantic Information
von: Estevam, Valter, et al.
Veröffentlicht: (2021)
von: Estevam, Valter, et al.
Veröffentlicht: (2021)
Dense Video Object Captioning from Disjoint Supervision
von: Zhou, Xingyi, et al.
Veröffentlicht: (2023)
von: Zhou, Xingyi, et al.
Veröffentlicht: (2023)
Technical Report for Soccernet 2023 -- Dense Video Captioning
von: Ruan, Zheng, et al.
Veröffentlicht: (2024)
von: Ruan, Zheng, et al.
Veröffentlicht: (2024)
OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
von: Kim, Ye-Chan, et al.
Veröffentlicht: (2026)
von: Kim, Ye-Chan, et al.
Veröffentlicht: (2026)
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
Bi-directional Contextual Attention for 3D Dense Captioning
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
von: Lin, Zihan, et al.
Veröffentlicht: (2026)
von: Lin, Zihan, et al.
Veröffentlicht: (2026)
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning
von: Li, Yizhe, et al.
Veröffentlicht: (2025)
von: Li, Yizhe, et al.
Veröffentlicht: (2025)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
von: Xie, Zhuyang, et al.
Veröffentlicht: (2024)
von: Xie, Zhuyang, et al.
Veröffentlicht: (2024)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
von: Jia, Mingda, et al.
Veröffentlicht: (2025)
von: Jia, Mingda, et al.
Veröffentlicht: (2025)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
von: Xu, Lin, et al.
Veröffentlicht: (2024)
von: Xu, Lin, et al.
Veröffentlicht: (2024)
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
von: Shoman, Maged, et al.
Veröffentlicht: (2024)
von: Shoman, Maged, et al.
Veröffentlicht: (2024)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
von: Ding, Ning, et al.
Veröffentlicht: (2025)
von: Ding, Ning, et al.
Veröffentlicht: (2025)
Dense Video Captioning using Graph-based Sentence Summarization
von: Zhang, Zhiwang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiwang, et al.
Veröffentlicht: (2025)
VideoMamba: Spatio-Temporal Selective State Space Model
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
Learning Visual Grounding from Generative Vision and Language Model
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
von: Yao, Linli, et al.
Veröffentlicht: (2026)
von: Yao, Linli, et al.
Veröffentlicht: (2026)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning
von: Baek, Seung Hyup, et al.
Veröffentlicht: (2026)
von: Baek, Seung Hyup, et al.
Veröffentlicht: (2026)
Wolf: Dense Video Captioning with a World Summarization Framework
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
BFMD: A Full-Match Badminton Dense Dataset for Dense Shot Captioning
von: Ding, Ning, et al.
Veröffentlicht: (2026)
von: Ding, Ning, et al.
Veröffentlicht: (2026)
RSCaMa: Remote Sensing Image Change Captioning with State Space Model
von: Liu, Chenyang, et al.
Veröffentlicht: (2024)
von: Liu, Chenyang, et al.
Veröffentlicht: (2024)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025) -
Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning
von: Kim, Dahun, et al.
Veröffentlicht: (2026) -
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2024) -
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities
von: Piergiovanni, AJ, et al.
Veröffentlicht: (2023) -
Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications
von: Mallya, Ganesh, et al.
Veröffentlicht: (2025)