VideoMix: Aggregating How-To Videos for Task-Oriented Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Saelyne, Truong, Anh, Kim, Juho, Li, Dingzeyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
von: Yang, Saelyne, et al.
Veröffentlicht: (2026)
von: Yang, Saelyne, et al.
Veröffentlicht: (2026)
Vid2Coach: Transforming How-To Videos into Task Assistants
von: Huh, Mina, et al.
Veröffentlicht: (2025)
von: Huh, Mina, et al.
Veröffentlicht: (2025)
Rewriting Video: Text-Driven Reauthoring of Video Footage
von: Wang, Sitong, et al.
Veröffentlicht: (2026)
von: Wang, Sitong, et al.
Veröffentlicht: (2026)
Morae: Proactively Pausing UI Agents for User Choices
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
VideoA11y: Method and Dataset for Accessible Video Description
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
SOS: A Shuffle Order Strategy for Data Augmentation in Industrial Human Activity Recognition
von: Ha, Anh Tuan, et al.
Veröffentlicht: (2025)
von: Ha, Anh Tuan, et al.
Veröffentlicht: (2025)
Acoustic Field Video for Multimodal Scene Understanding
von: Kim, Daehwa, et al.
Veröffentlicht: (2026)
von: Kim, Daehwa, et al.
Veröffentlicht: (2026)
ExpressEdit: Video Editing with Natural Language and Sketching
von: Tilekbay, Bekzat, et al.
Veröffentlicht: (2024)
von: Tilekbay, Bekzat, et al.
Veröffentlicht: (2024)
PodReels: Human-AI Co-Creation of Video Podcast Teasers
von: Wang, Sitong, et al.
Veröffentlicht: (2023)
von: Wang, Sitong, et al.
Veröffentlicht: (2023)
EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
von: Leng, Zikang, et al.
Veröffentlicht: (2026)
von: Leng, Zikang, et al.
Veröffentlicht: (2026)
Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery Learning
von: Wang, Jingying, et al.
Veröffentlicht: (2024)
von: Wang, Jingying, et al.
Veröffentlicht: (2024)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
von: Li, Zisu, et al.
Veröffentlicht: (2025)
von: Li, Zisu, et al.
Veröffentlicht: (2025)
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
Analyzing Swimming Performance Using Drone Captured Aerial Videos
von: Tran, Thu, et al.
Veröffentlicht: (2025)
von: Tran, Thu, et al.
Veröffentlicht: (2025)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
von: Nadeem, Asmar, et al.
Veröffentlicht: (2024)
von: Nadeem, Asmar, et al.
Veröffentlicht: (2024)
Reframe Anything: LLM Agent for Open World Video Reframing
von: Cao, Jiawang, et al.
Veröffentlicht: (2024)
von: Cao, Jiawang, et al.
Veröffentlicht: (2024)
Video Joint-Embedding Predictive Architectures for Facial Expression Recognition
von: Eing, Lennart, et al.
Veröffentlicht: (2026)
von: Eing, Lennart, et al.
Veröffentlicht: (2026)
"I Can't Keep Up": Accessibility Barriers in Video-Based Learning for Individuals with Borderline Intellectual Functioning
von: Chu, Hyehyun, et al.
Veröffentlicht: (2026)
von: Chu, Hyehyun, et al.
Veröffentlicht: (2026)
MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition
von: Garg, Mallika, et al.
Veröffentlicht: (2024)
von: Garg, Mallika, et al.
Veröffentlicht: (2024)
Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals
von: Zhou, Puqi, et al.
Veröffentlicht: (2026)
von: Zhou, Puqi, et al.
Veröffentlicht: (2026)
Hybrid 3D Human Pose Estimation with Monocular Video and Sparse IMUs
von: Bao, Yiming, et al.
Veröffentlicht: (2024)
von: Bao, Yiming, et al.
Veröffentlicht: (2024)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
von: Nguyen, Giang, et al.
Veröffentlicht: (2024)
von: Nguyen, Giang, et al.
Veröffentlicht: (2024)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
von: He, Xu, et al.
Veröffentlicht: (2024)
von: He, Xu, et al.
Veröffentlicht: (2024)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
von: Garg, Mallika, et al.
Veröffentlicht: (2025)
von: Garg, Mallika, et al.
Veröffentlicht: (2025)
ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
von: He, Yuchen, et al.
Veröffentlicht: (2025)
von: He, Yuchen, et al.
Veröffentlicht: (2025)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
von: Zheng, Ruizhe, et al.
Veröffentlicht: (2024)
von: Zheng, Ruizhe, et al.
Veröffentlicht: (2024)
CinePreGen: Camera Controllable Video Previsualization via Engine-powered Diffusion
von: Chen, Yiran, et al.
Veröffentlicht: (2024)
von: Chen, Yiran, et al.
Veröffentlicht: (2024)
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
von: Xu, Yutong, et al.
Veröffentlicht: (2024)
von: Xu, Yutong, et al.
Veröffentlicht: (2024)
The Visual Experience Dataset: Over 200 Recorded Hours of Integrated Eye Movement, Odometry, and Egocentric Video
von: Greene, Michelle R., et al.
Veröffentlicht: (2024)
von: Greene, Michelle R., et al.
Veröffentlicht: (2024)
A Comparison of Bounding Box and Landmark Detection Methods for Video-Based Heart Rate Estimation
von: Liang, Laurence
Veröffentlicht: (2023)
von: Liang, Laurence
Veröffentlicht: (2023)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
von: Foteinopoulou, Niki Maria, et al.
Veröffentlicht: (2023)
von: Foteinopoulou, Niki Maria, et al.
Veröffentlicht: (2023)
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Contexts
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Detecting Activities of Daily Living in Egocentric Video to Contextualize Hand Use at Home in Outpatient Neurorehabilitation Settings
von: Kadambi, Adesh, et al.
Veröffentlicht: (2024)
von: Kadambi, Adesh, et al.
Veröffentlicht: (2024)
Editing Physiological Signals in Videos Using Latent Representations
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
SVFAP: Self-supervised Video Facial Affect Perceiver
von: Sun, Licai, et al.
Veröffentlicht: (2023)
von: Sun, Licai, et al.
Veröffentlicht: (2023)
PedaCo-Gen: Scaffolding Pedagogical Agency in Human-AI Collaborative Video Authoring
von: Baek, Injun, et al.
Veröffentlicht: (2026)
von: Baek, Injun, et al.
Veröffentlicht: (2026)
Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
von: Personnic, Alexandre, et al.
Veröffentlicht: (2025)
von: Personnic, Alexandre, et al.
Veröffentlicht: (2025)
Beyond Questionnaires: Video Analysis for Social Anxiety Detection
von: Sahu, Nilesh Kumar, et al.
Veröffentlicht: (2024)
von: Sahu, Nilesh Kumar, et al.
Veröffentlicht: (2024)
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
von: Yang, Saelyne, et al.
Veröffentlicht: (2026) -
Vid2Coach: Transforming How-To Videos into Task Assistants
von: Huh, Mina, et al.
Veröffentlicht: (2025) -
Rewriting Video: Text-Driven Reauthoring of Video Footage
von: Wang, Sitong, et al.
Veröffentlicht: (2026) -
Morae: Proactively Pausing UI Agents for User Choices
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025) -
VideoA11y: Method and Dataset for Accessible Video Description
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)