More than a Moment: Towards Coherent Sequences of Audio Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Khandelwal, Eshika, Xie, Junyu, Han, Tengda, Bain, Max, Nagrani, Arsha, Zisserman, Andrew, Varol, Gül, Tapaswi, Makarand |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
by: Kumar, Prajneya, et al.
Published: (2023)
by: Kumar, Prajneya, et al.
Published: (2023)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
The Sound of Water: Inferring Physical Properties from Pouring Liquids
by: Bagad, Piyush, et al.
Published: (2024)
by: Bagad, Piyush, et al.
Published: (2024)
MICap: A Unified Model for Identity-aware Movie Descriptions
by: Raajesh, Haran, et al.
Published: (2024)
by: Raajesh, Haran, et al.
Published: (2024)
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024)
by: Amini-Naieni, Niki, et al.
Published: (2024)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
by: Jang, Youngjoon, et al.
Published: (2025)
by: Jang, Youngjoon, et al.
Published: (2025)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
by: Darur, Balaji, et al.
Published: (2026)
by: Darur, Balaji, et al.
Published: (2026)
MALeR: Improving Compositional Fidelity in Layout-Guided Generation
by: Saxena, Shivank, et al.
Published: (2025)
by: Saxena, Shivank, et al.
Published: (2025)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
by: Singh, Darshan, et al.
Published: (2024)
by: Singh, Darshan, et al.
Published: (2024)
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
Investigating Mechanisms for In-Context Vision Language Binding
by: Saravanan, Darshana, et al.
Published: (2025)
by: Saravanan, Darshana, et al.
Published: (2025)
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
by: Jiang, Zifan, et al.
Published: (2025)
by: Jiang, Zifan, et al.
Published: (2025)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
"Previously on ..." From Recaps to Story Summarization
by: Singh, Aditya Kumar, et al.
Published: (2024)
by: Singh, Aditya Kumar, et al.
Published: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
by: Kahatapitiya, Kumara, et al.
Published: (2023)
by: Kahatapitiya, Kumara, et al.
Published: (2023)
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
by: Jang, Youngjoon, et al.
Published: (2025)
by: Jang, Youngjoon, et al.
Published: (2025)
A Tale of Two Languages: Large-Vocabulary Continuous Sign Language Recognition from Spoken Language Supervision
by: Raude, Charles, et al.
Published: (2024)
by: Raude, Charles, et al.
Published: (2024)
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Moving Object Segmentation: All You Need Is SAM (and Flow)
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoning
by: Sung, Junyoung, et al.
Published: (2026)
by: Sung, Junyoung, et al.
Published: (2026)
Recognising BSL Fingerspelling in Continuous Signing Sequences
by: Chan, Alyssa, et al.
Published: (2026)
by: Chan, Alyssa, et al.
Published: (2026)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
Steerable Visual Representations
by: Ruthardt, Jona, et al.
Published: (2026)
by: Ruthardt, Jona, et al.
Published: (2026)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
STRinGS: Selective Text Refinement in Gaussian Splatting
by: Raundhal, Abhinav, et al.
Published: (2025)
by: Raundhal, Abhinav, et al.
Published: (2025)
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
by: Bensabath, Léore, et al.
Published: (2024)
by: Bensabath, Léore, et al.
Published: (2024)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
by: Korbar, Bruno, et al.
Published: (2024)
by: Korbar, Bruno, et al.
Published: (2024)
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
by: Saravanan, Darshana, et al.
Published: (2024)
by: Saravanan, Darshana, et al.
Published: (2024)
Similar Items
-
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025) -
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024) -
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024) -
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025) -
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
by: Kumar, Prajneya, et al.
Published: (2023)