Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Junyu, Han, Tengda, Bain, Max, Nagrani, Arsha, Khandelwal, Eshika, Varol, Gül, Xie, Weidi, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
von: Xie, Junyu, et al.
Veröffentlicht: (2024)
von: Xie, Junyu, et al.
Veröffentlicht: (2024)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
AutoAD III: The Prequel -- Back to the Pixels
von: Han, Tengda, et al.
Veröffentlicht: (2024)
von: Han, Tengda, et al.
Veröffentlicht: (2024)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
von: Xie, Junyu, et al.
Veröffentlicht: (2026)
von: Xie, Junyu, et al.
Veröffentlicht: (2026)
Character-Centric Understanding of Animated Movies
von: Gui, Zhongrui, et al.
Veröffentlicht: (2025)
von: Gui, Zhongrui, et al.
Veröffentlicht: (2025)
Appearance-Based Refinement for Object-Centric Motion Segmentation
von: Xie, Junyu, et al.
Veröffentlicht: (2023)
von: Xie, Junyu, et al.
Veröffentlicht: (2023)
Moving Object Segmentation: All You Need Is SAM (and Flow)
von: Xie, Junyu, et al.
Veröffentlicht: (2024)
von: Xie, Junyu, et al.
Veröffentlicht: (2024)
What You See is What You Ask: Evaluating Audio Descriptions
von: Kala, Divy, et al.
Veröffentlicht: (2025)
von: Kala, Divy, et al.
Veröffentlicht: (2025)
Made to Order: Discovering monotonic temporal changes via self-supervised video ordering
von: Yang, Charig, et al.
Veröffentlicht: (2024)
von: Yang, Charig, et al.
Veröffentlicht: (2024)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
Amodal Ground Truth and Completion in the Wild
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
CountGD: Multi-Modal Open-World Counting
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
Inferring Dynamic Physical Properties from Video Foundation Models
von: Zhan, Guanqi, et al.
Veröffentlicht: (2025)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2025)
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
von: Zhan, Guanqi, et al.
Veröffentlicht: (2025)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
Synchformer: Efficient Synchronization from Sparse Cues
von: Iashin, Vladimir, et al.
Veröffentlicht: (2024)
von: Iashin, Vladimir, et al.
Veröffentlicht: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
Seeing without Pixels: Perception from Camera Trajectories
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
Multi-Sentence Grounding for Long-term Instructional Video
von: Li, Zeqian, et al.
Veröffentlicht: (2023)
von: Li, Zeqian, et al.
Veröffentlicht: (2023)
Hypergraph-Enhanced Training-Free and Language-Free Few-Shot Anomaly Detection
von: Xie, Guohuan, et al.
Veröffentlicht: (2026)
von: Xie, Guohuan, et al.
Veröffentlicht: (2026)
Scaling Audio-Text Retrieval with Multimodal Large Language Models
von: Xu, Jilan, et al.
Veröffentlicht: (2026)
von: Xu, Jilan, et al.
Veröffentlicht: (2026)
A Training-Free Framework for Video License Plate Tracking and Recognition with Only One-Shot
von: Ding, Haoxuan, et al.
Veröffentlicht: (2024)
von: Ding, Haoxuan, et al.
Veröffentlicht: (2024)
LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
von: Li, Jiachen, et al.
Veröffentlicht: (2025)
von: Li, Jiachen, et al.
Veröffentlicht: (2025)
A Tale of Two Languages: Large-Vocabulary Continuous Sign Language Recognition from Spoken Language Supervision
von: Raude, Charles, et al.
Veröffentlicht: (2024)
von: Raude, Charles, et al.
Veröffentlicht: (2024)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
von: Min, Juhong, et al.
Veröffentlicht: (2024)
von: Min, Juhong, et al.
Veröffentlicht: (2024)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
von: Kumar, Prajneya, et al.
Veröffentlicht: (2023)
von: Kumar, Prajneya, et al.
Veröffentlicht: (2023)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
von: Ge, Mingji, et al.
Veröffentlicht: (2026)
von: Ge, Mingji, et al.
Veröffentlicht: (2026)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
von: Bagad, Piyush, et al.
Veröffentlicht: (2025)
von: Bagad, Piyush, et al.
Veröffentlicht: (2025)
CAViAR: Critic-Augmented Video Agentic Reasoning
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
von: Zhang, Da, et al.
Veröffentlicht: (2026)
von: Zhang, Da, et al.
Veröffentlicht: (2026)
Text-Driven 3D Hand Motion Generation from Sign Language Data
von: Bensabath, Léore, et al.
Veröffentlicht: (2025)
von: Bensabath, Léore, et al.
Veröffentlicht: (2025)
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
von: Bensabath, Léore, et al.
Veröffentlicht: (2024)
von: Bensabath, Léore, et al.
Veröffentlicht: (2024)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
Shot-Aware Frame Sampling for Video Understanding
von: Zhao, Mengyu, et al.
Veröffentlicht: (2026)
von: Zhao, Mengyu, et al.
Veröffentlicht: (2026)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
von: Xie, Junyu, et al.
Veröffentlicht: (2024) -
More than a Moment: Towards Coherent Sequences of Audio Descriptions
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025) -
AutoAD III: The Prequel -- Back to the Pixels
von: Han, Tengda, et al.
Veröffentlicht: (2024) -
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
von: Xie, Junyu, et al.
Veröffentlicht: (2026) -
Character-Centric Understanding of Animated Movies
von: Gui, Zhongrui, et al.
Veröffentlicht: (2025)