HD-EPIC: A Highly-Detailed Egocentric Video Dataset
Fuente:
arXiv
Guardado en:
| Autores principales: | Perrett, Toby, Darkhalil, Ahmad, Sinha, Saptarshi, Emara, Omar, Pollard, Sam, Parida, Kranti, Liu, Kaiting, Gatti, Prajwal, Bansal, Siddhant, Flanagan, Kevin, Chalk, Jacob, Zhu, Zhifan, Guerrier, Rhodri, Abdelazim, Fahd, Zhu, Bin, Moltisanti, Davide, Wray, Michael, Doughty, Hazel, Damen, Dima |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Segmenting Collision Sound Sources in Egocentric Videos
por: Parida, Kranti Kumar, et al.
Publicado: (2025)
por: Parida, Kranti Kumar, et al.
Publicado: (2025)
EgoPoints: Advancing Point Tracking for Egocentric Videos
por: Darkhalil, Ahmad, et al.
Publicado: (2024)
por: Darkhalil, Ahmad, et al.
Publicado: (2024)
EPIC Fields: Marrying 3D Geometry and Video Understanding
por: Tschernezki, Vadim, et al.
Publicado: (2023)
por: Tschernezki, Vadim, et al.
Publicado: (2023)
PointSt3R: Point Tracking through 3D Grounded Correspondence
por: Guerrier, Rhodri, et al.
Publicado: (2025)
por: Guerrier, Rhodri, et al.
Publicado: (2025)
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
por: Liu, Kaiting, et al.
Publicado: (2026)
por: Liu, Kaiting, et al.
Publicado: (2026)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
por: Zhu, Zhifan, et al.
Publicado: (2023)
por: Zhu, Zhifan, et al.
Publicado: (2023)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
por: Flanagan, Kevin, et al.
Publicado: (2025)
por: Flanagan, Kevin, et al.
Publicado: (2025)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
por: Zhu, Zhifan, et al.
Publicado: (2025)
por: Zhu, Zhifan, et al.
Publicado: (2025)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
por: Bansal, Siddhant, et al.
Publicado: (2024)
por: Bansal, Siddhant, et al.
Publicado: (2024)
Video Editing for Video Retrieval
por: Zhu, Bin, et al.
Publicado: (2024)
por: Zhu, Bin, et al.
Publicado: (2024)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
por: Plizzari, Chiara, et al.
Publicado: (2024)
por: Plizzari, Chiara, et al.
Publicado: (2024)
The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation
por: Hatano, Masashi, et al.
Publicado: (2025)
por: Hatano, Masashi, et al.
Publicado: (2025)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
por: Zhu, Zhifan, et al.
Publicado: (2025)
por: Zhu, Zhifan, et al.
Publicado: (2025)
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
por: Fragomeni, Adriano, et al.
Publicado: (2025)
por: Fragomeni, Adriano, et al.
Publicado: (2025)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
por: Fragomeni, Adriano, et al.
Publicado: (2025)
por: Fragomeni, Adriano, et al.
Publicado: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
por: Perrett, Toby, et al.
Publicado: (2024)
por: Perrett, Toby, et al.
Publicado: (2024)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
por: Souček, Tomáš, et al.
Publicado: (2024)
por: Souček, Tomáš, et al.
Publicado: (2024)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
por: Sinha, Saptarshi, et al.
Publicado: (2024)
por: Sinha, Saptarshi, et al.
Publicado: (2024)
Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
por: Hatano, Masashi, et al.
Publicado: (2025)
por: Hatano, Masashi, et al.
Publicado: (2025)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
por: Chalk, Jacob, et al.
Publicado: (2024)
por: Chalk, Jacob, et al.
Publicado: (2024)
Epic-Sounds: A Large-scale Dataset of Actions That Sound
por: Huh, Jaesung, et al.
Publicado: (2023)
por: Huh, Jaesung, et al.
Publicado: (2023)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
por: Souček, Tomáš, et al.
Publicado: (2023)
por: Souček, Tomáš, et al.
Publicado: (2023)
Beyond Caption-Based Queries for Video Moment Retrieval
por: Pujol-Perich, David, et al.
Publicado: (2026)
por: Pujol-Perich, David, et al.
Publicado: (2026)
Video, How Do Your Tokens Merge?
por: Pollard, Sam, et al.
Publicado: (2025)
por: Pollard, Sam, et al.
Publicado: (2025)
A Video Is Not Worth a Thousand Words
por: Pollard, Sam, et al.
Publicado: (2025)
por: Pollard, Sam, et al.
Publicado: (2025)
AMEGO: Active Memory from long EGOcentric videos
por: Goletto, Gabriele, et al.
Publicado: (2024)
por: Goletto, Gabriele, et al.
Publicado: (2024)
Low-Resource Vision Challenges for Foundation Models
por: Zhang, Yunhua, et al.
Publicado: (2024)
por: Zhang, Yunhua, et al.
Publicado: (2024)
Learn to Categorize or Categorize to Learn? Self-Coding for Generalized Category Discovery
por: Rastegar, Sarah, et al.
Publicado: (2023)
por: Rastegar, Sarah, et al.
Publicado: (2023)
gperrett/scaffolding-responsible-software-use: scaffolding responsible software use
por: George Perrett
Publicado: (2025)
por: George Perrett
Publicado: (2025)
Emergent Explainability: Adding a causal chain to neural network inference
por: Perrett, Adam
Publicado: (2024)
por: Perrett, Adam
Publicado: (2024)
Graduate Information Literacy Skills: The 2003 ANU Skills Audit
por: Perrett, Valerie
Publicado: (2004)
por: Perrett, Valerie
Publicado: (2004)
Information Literacy Skills Training: A Factor in Student Satisfaction with Access to High Demand Material
por: Perrett, Valerie
Publicado: (2010)
por: Perrett, Valerie
Publicado: (2010)
An Outlook into the Future of Egocentric Vision
por: Plizzari, Chiara, et al.
Publicado: (2023)
por: Plizzari, Chiara, et al.
Publicado: (2023)
A DEMOCRATIZAÇÃO DA EDUCAÇÃO NO HAITI: UMA ANÁLISE HISTÓRICA DOS PROBLEMAS, DESAFIOS E PERSPECTIVAS PARA A RECONSTRUÇÃO NACIONAL
por: Milius Guerrier
Publicado: (2025)
por: Milius Guerrier
Publicado: (2025)
LocoMotion: Learning Motion-Focused Video-Language Representations
por: Doughty, Hazel, et al.
Publicado: (2024)
por: Doughty, Hazel, et al.
Publicado: (2024)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
por: Kranti, Chalamalasetti
Publicado: (2025)
por: Kranti, Chalamalasetti
Publicado: (2025)
ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
por: Seminara, Luigi, et al.
Publicado: (2026)
por: Seminara, Luigi, et al.
Publicado: (2026)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
por: Heyward, Joseph, et al.
Publicado: (2024)
por: Heyward, Joseph, et al.
Publicado: (2024)
Seeing without Pixels: Perception from Camera Trajectories
por: Xue, Zihui, et al.
Publicado: (2025)
por: Xue, Zihui, et al.
Publicado: (2025)
Good intent, or just good content? Assessing MrBeast's philanthropy
por: Rhodri Davies
Publicado: (2024)
por: Rhodri Davies
Publicado: (2024)
Ejemplares similares
-
Segmenting Collision Sound Sources in Egocentric Videos
por: Parida, Kranti Kumar, et al.
Publicado: (2025) -
EgoPoints: Advancing Point Tracking for Egocentric Videos
por: Darkhalil, Ahmad, et al.
Publicado: (2024) -
EPIC Fields: Marrying 3D Geometry and Video Understanding
por: Tschernezki, Vadim, et al.
Publicado: (2023) -
PointSt3R: Point Tracking through 3D Grounded Correspondence
por: Guerrier, Rhodri, et al.
Publicado: (2025) -
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
por: Liu, Kaiting, et al.
Publicado: (2026)