Implicit and Explicit Commonsense for Multi-sentence Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Chou, Shih-Han, Little, James J., Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
Test-Time Consistency in Vision Language Models
by: Chou, Shih-Han, et al.
Published: (2025)
by: Chou, Shih-Han, et al.
Published: (2025)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
Factorized Video Autoencoders for Efficient Generative Modelling
by: Suhail, Mohammed, et al.
Published: (2024)
by: Suhail, Mohammed, et al.
Published: (2024)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
by: Chinchure, Aditya, et al.
Published: (2025)
by: Chinchure, Aditya, et al.
Published: (2025)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
by: Bhatt, Gaurav, et al.
Published: (2024)
by: Bhatt, Gaurav, et al.
Published: (2024)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
by: Tian, Mingkai, et al.
Published: (2025)
by: Tian, Mingkai, et al.
Published: (2025)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
by: Mahdizadeh, Ailar, et al.
Published: (2026)
by: Mahdizadeh, Ailar, et al.
Published: (2026)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
by: Lin, Shih-Yao, et al.
Published: (2025)
by: Lin, Shih-Yao, et al.
Published: (2025)
SPIKE-RL: Video-LLMs meet Bayesian Surprise
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
Live Video Captioning
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
by: Jia, Mingda, et al.
Published: (2025)
by: Jia, Mingda, et al.
Published: (2025)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
by: Ding, Ning, et al.
Published: (2025)
by: Ding, Ning, et al.
Published: (2025)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
Video sentence grounding with temporally global textual knowledge
by: Chen, Cai, et al.
Published: (2024)
by: Chen, Cai, et al.
Published: (2024)
EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking
by: Zhang, Yukuan, et al.
Published: (2025)
by: Zhang, Yukuan, et al.
Published: (2025)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
by: Salamatian, Ali, et al.
Published: (2026)
by: Salamatian, Ali, et al.
Published: (2026)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
by: Wang, Zhen, et al.
Published: (2023)
by: Wang, Zhen, et al.
Published: (2023)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
by: Chu, Sanghyeok, et al.
Published: (2025)
by: Chu, Sanghyeok, et al.
Published: (2025)
Implicit Neural Surface Deformation with Explicit Velocity Fields
by: Sang, Lu, et al.
Published: (2025)
by: Sang, Lu, et al.
Published: (2025)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
by: Mkhallati, Hassan, et al.
Published: (2023)
by: Mkhallati, Hassan, et al.
Published: (2023)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025)
by: Zheng, Guangcong, et al.
Published: (2025)
Addressing the ID-Matching Challenge in Long Video Captioning
by: Yang, Zhantao, et al.
Published: (2025)
by: Yang, Zhantao, et al.
Published: (2025)
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
by: Chinchure, Aditya, et al.
Published: (2024)
by: Chinchure, Aditya, et al.
Published: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Similar Items
-
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
by: Chou, Shih-Han, et al.
Published: (2024) -
Test-Time Consistency in Vision Language Models
by: Chou, Shih-Han, et al.
Published: (2025) -
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024) -
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022) -
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)