TIM: A Time Interval Machine for Audio-Visual Action Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Chalk, Jacob, Huh, Jaesung, Kazakos, Evangelos, Zisserman, Andrew, Damen, Dima |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Epic-Sounds: A Large-scale Dataset of Actions That Sound
por: Huh, Jaesung, et al.
Publicado: (2023)
por: Huh, Jaesung, et al.
Publicado: (2023)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
por: Korbar, Bruno, et al.
Publicado: (2024)
por: Korbar, Bruno, et al.
Publicado: (2024)
Character-aware audio-visual subtitling in context
por: Huh, Jaesung, et al.
Publicado: (2024)
por: Huh, Jaesung, et al.
Publicado: (2024)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
por: Perrett, Toby, et al.
Publicado: (2024)
por: Perrett, Toby, et al.
Publicado: (2024)
Seeing without Pixels: Perception from Camera Trajectories
por: Xue, Zihui, et al.
Publicado: (2025)
por: Xue, Zihui, et al.
Publicado: (2025)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
por: Plizzari, Chiara, et al.
Publicado: (2024)
por: Plizzari, Chiara, et al.
Publicado: (2024)
Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
por: Hatano, Masashi, et al.
Publicado: (2025)
por: Hatano, Masashi, et al.
Publicado: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
por: Heyward, Joseph, et al.
Publicado: (2024)
por: Heyward, Joseph, et al.
Publicado: (2024)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
por: Zhu, Zhifan, et al.
Publicado: (2023)
por: Zhu, Zhifan, et al.
Publicado: (2023)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
por: Bagad, Piyush, et al.
Publicado: (2025)
por: Bagad, Piyush, et al.
Publicado: (2025)
Perception Test 2025: Challenge Summary and a Unified VQA Extension
por: Heyward, Joseph, et al.
Publicado: (2026)
por: Heyward, Joseph, et al.
Publicado: (2026)
Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2024)
por: Kazakos, Evangelos, et al.
Publicado: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2025)
por: Kazakos, Evangelos, et al.
Publicado: (2025)
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
por: Fragomeni, Adriano, et al.
Publicado: (2025)
por: Fragomeni, Adriano, et al.
Publicado: (2025)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
por: Souček, Tomáš, et al.
Publicado: (2023)
por: Souček, Tomáš, et al.
Publicado: (2023)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
por: Bansal, Siddhant, et al.
Publicado: (2024)
por: Bansal, Siddhant, et al.
Publicado: (2024)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
por: Fragomeni, Adriano, et al.
Publicado: (2025)
por: Fragomeni, Adriano, et al.
Publicado: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
por: Flanagan, Kevin, et al.
Publicado: (2025)
por: Flanagan, Kevin, et al.
Publicado: (2025)
Learning from Streaming Video with Orthogonal Gradients
por: Han, Tengda, et al.
Publicado: (2025)
por: Han, Tengda, et al.
Publicado: (2025)
PointSt3R: Point Tracking through 3D Grounded Correspondence
por: Guerrier, Rhodri, et al.
Publicado: (2025)
por: Guerrier, Rhodri, et al.
Publicado: (2025)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
por: Souček, Tomáš, et al.
Publicado: (2024)
por: Souček, Tomáš, et al.
Publicado: (2024)
AMEGO: Active Memory from long EGOcentric videos
por: Goletto, Gabriele, et al.
Publicado: (2024)
por: Goletto, Gabriele, et al.
Publicado: (2024)
The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation
por: Hatano, Masashi, et al.
Publicado: (2025)
por: Hatano, Masashi, et al.
Publicado: (2025)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
por: Zhu, Zhifan, et al.
Publicado: (2025)
por: Zhu, Zhifan, et al.
Publicado: (2025)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
por: Zhu, Zhifan, et al.
Publicado: (2025)
por: Zhu, Zhifan, et al.
Publicado: (2025)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
por: Sinha, Saptarshi, et al.
Publicado: (2024)
por: Sinha, Saptarshi, et al.
Publicado: (2024)
Noise-Tolerant Learning for Audio-Visual Action Recognition
por: Han, Haochen, et al.
Publicado: (2022)
por: Han, Haochen, et al.
Publicado: (2022)
EgoPoints: Advancing Point Tracking for Egocentric Videos
por: Darkhalil, Ahmad, et al.
Publicado: (2024)
por: Darkhalil, Ahmad, et al.
Publicado: (2024)
Unique Lives, Shared World: Learning from Single-Life Videos
por: Han, Tengda, et al.
Publicado: (2025)
por: Han, Tengda, et al.
Publicado: (2025)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
por: Sachdeva, Ragav, et al.
Publicado: (2024)
por: Sachdeva, Ragav, et al.
Publicado: (2024)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
por: Korbar, Bruno, et al.
Publicado: (2025)
por: Korbar, Bruno, et al.
Publicado: (2025)
From Panels to Prose: Generating Literary Narratives from Comics
por: Sachdeva, Ragav, et al.
Publicado: (2025)
por: Sachdeva, Ragav, et al.
Publicado: (2025)
Beyond Caption-Based Queries for Video Moment Retrieval
por: Pujol-Perich, David, et al.
Publicado: (2026)
por: Pujol-Perich, David, et al.
Publicado: (2026)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
por: Xie, Junyu, et al.
Publicado: (2026)
por: Xie, Junyu, et al.
Publicado: (2026)
CountGD++: Generalized Prompting for Open-World Counting
por: Amini-Naieni, Niki, et al.
Publicado: (2025)
por: Amini-Naieni, Niki, et al.
Publicado: (2025)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
por: Kang, Seokun, et al.
Publicado: (2025)
por: Kang, Seokun, et al.
Publicado: (2025)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
por: Xie, Junyu, et al.
Publicado: (2024)
por: Xie, Junyu, et al.
Publicado: (2024)
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
por: Zhan, Guanqi, et al.
Publicado: (2025)
por: Zhan, Guanqi, et al.
Publicado: (2025)
Learning from One Continuous Video Stream
por: Carreira, João, et al.
Publicado: (2023)
por: Carreira, João, et al.
Publicado: (2023)
WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata
por: Sridhar, Prasanna, et al.
Publicado: (2026)
por: Sridhar, Prasanna, et al.
Publicado: (2026)
Ejemplares similares
-
Epic-Sounds: A Large-scale Dataset of Actions That Sound
por: Huh, Jaesung, et al.
Publicado: (2023) -
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
por: Korbar, Bruno, et al.
Publicado: (2024) -
Character-aware audio-visual subtitling in context
por: Huh, Jaesung, et al.
Publicado: (2024) -
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
por: Perrett, Toby, et al.
Publicado: (2024) -
Seeing without Pixels: Perception from Camera Trajectories
por: Xue, Zihui, et al.
Publicado: (2025)