It's Just Another Day: Unique Video Captioning by Discriminative Prompting
Fuente:
arXiv
Saved in:
| Main Authors: | Perrett, Toby, Han, Tengda, Damen, Dima, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Unique Lives, Shared World: Learning from Single-Life Videos
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024)
by: Plizzari, Chiara, et al.
Published: (2024)
Learning from Streaming Video with Orthogonal Gradients
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
by: Chalk, Jacob, et al.
Published: (2024)
by: Chalk, Jacob, et al.
Published: (2024)
Beyond Caption-Based Queries for Video Moment Retrieval
by: Pujol-Perich, David, et al.
Published: (2026)
by: Pujol-Perich, David, et al.
Published: (2026)
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024)
by: Amini-Naieni, Niki, et al.
Published: (2024)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
by: Zhu, Zhifan, et al.
Published: (2023)
by: Zhu, Zhifan, et al.
Published: (2023)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025)
by: Flanagan, Kevin, et al.
Published: (2025)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Perception Test 2025: Challenge Summary and a Unified VQA Extension
by: Heyward, Joseph, et al.
Published: (2026)
by: Heyward, Joseph, et al.
Published: (2026)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
by: Sinha, Saptarshi, et al.
Published: (2024)
by: Sinha, Saptarshi, et al.
Published: (2024)
EgoPoints: Advancing Point Tracking for Egocentric Videos
by: Darkhalil, Ahmad, et al.
Published: (2024)
by: Darkhalil, Ahmad, et al.
Published: (2024)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)
by: Bansal, Siddhant, et al.
Published: (2024)
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
by: Souček, Tomáš, et al.
Published: (2023)
by: Souček, Tomáš, et al.
Published: (2023)
Video Editing for Video Retrieval
by: Zhu, Bin, et al.
Published: (2024)
by: Zhu, Bin, et al.
Published: (2024)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
PointSt3R: Point Tracking through 3D Grounded Correspondence
by: Guerrier, Rhodri, et al.
Published: (2025)
by: Guerrier, Rhodri, et al.
Published: (2025)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
AMEGO: Active Memory from long EGOcentric videos
by: Goletto, Gabriele, et al.
Published: (2024)
by: Goletto, Gabriele, et al.
Published: (2024)
The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation
by: Hatano, Masashi, et al.
Published: (2025)
by: Hatano, Masashi, et al.
Published: (2025)
Adapting MLLMs for Nuanced Video Retrieval
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
HD-EPIC: A Highly-Detailed Egocentric Video Dataset
by: Perrett, Toby, et al.
Published: (2025)
by: Perrett, Toby, et al.
Published: (2025)
Segmenting Collision Sound Sources in Egocentric Videos
by: Parida, Kranti Kumar, et al.
Published: (2025)
by: Parida, Kranti Kumar, et al.
Published: (2025)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing
by: Perrett, Toby, et al.
Published: (2026)
by: Perrett, Toby, et al.
Published: (2026)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)
by: Khandelwal, Eshika, et al.
Published: (2025)
EPIC Fields: Marrying 3D Geometry and Video Understanding
by: Tschernezki, Vadim, et al.
Published: (2023)
by: Tschernezki, Vadim, et al.
Published: (2023)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025)
by: Korbar, Bruno, et al.
Published: (2025)
Similar Items
-
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025) -
Unique Lives, Shared World: Learning from Single-Life Videos
by: Han, Tengda, et al.
Published: (2025) -
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024) -
Learning from Streaming Video with Orthogonal Gradients
by: Han, Tengda, et al.
Published: (2025) -
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)