Implicit and Explicit Commonsense for Multi-sentence Video Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chou, Shih-Han, Little, James J., Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
Test-Time Consistency in Vision Language Models
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2024)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2024)
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2022)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2022)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2025)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2025)
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
Multi-sentence Video Grounding for Long Video Generation
von: Feng, Wei, et al.
Veröffentlicht: (2024)
von: Feng, Wei, et al.
Veröffentlicht: (2024)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
von: Goyal, Raghav, et al.
Veröffentlicht: (2023)
von: Goyal, Raghav, et al.
Veröffentlicht: (2023)
Factorized Video Autoencoders for Efficient Generative Modelling
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2026)
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2026)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
von: Lin, Shih-Yao, et al.
Veröffentlicht: (2025)
von: Lin, Shih-Yao, et al.
Veröffentlicht: (2025)
SPIKE-RL: Video-LLMs meet Bayesian Surprise
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
Live Video Captioning
von: Blanco-Fernández, Eduardo, et al.
Veröffentlicht: (2024)
von: Blanco-Fernández, Eduardo, et al.
Veröffentlicht: (2024)
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
von: Jia, Mingda, et al.
Veröffentlicht: (2025)
von: Jia, Mingda, et al.
Veröffentlicht: (2025)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
von: Ding, Ning, et al.
Veröffentlicht: (2025)
von: Ding, Ning, et al.
Veröffentlicht: (2025)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
Video sentence grounding with temporally global textual knowledge
von: Chen, Cai, et al.
Veröffentlicht: (2024)
von: Chen, Cai, et al.
Veröffentlicht: (2024)
EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking
von: Zhang, Yukuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yukuan, et al.
Veröffentlicht: (2025)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
von: Salamatian, Ali, et al.
Veröffentlicht: (2026)
von: Salamatian, Ali, et al.
Veröffentlicht: (2026)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
von: Yao, Linli, et al.
Veröffentlicht: (2026)
von: Yao, Linli, et al.
Veröffentlicht: (2026)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
von: Wang, Zhen, et al.
Veröffentlicht: (2023)
von: Wang, Zhen, et al.
Veröffentlicht: (2023)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
Streaming Dense Video Captioning
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2025)
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2025)
Implicit Neural Surface Deformation with Explicit Velocity Fields
von: Sang, Lu, et al.
Veröffentlicht: (2025)
von: Sang, Lu, et al.
Veröffentlicht: (2025)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
von: Hu, Shiyu, et al.
Veröffentlicht: (2024)
von: Hu, Shiyu, et al.
Veröffentlicht: (2024)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
Addressing the ID-Matching Challenge in Long Video Captioning
von: Yang, Zhantao, et al.
Veröffentlicht: (2025)
von: Yang, Zhantao, et al.
Veröffentlicht: (2025)
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
von: Chinchure, Aditya, et al.
Veröffentlicht: (2024)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024) -
Test-Time Consistency in Vision Language Models
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025) -
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2024) -
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2022) -
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2025)