What to Do Next? Memorizing skills from Egocentric Instructional Video
Fuente:
arXiv
Guardado en:
| Autores principales: | Bi, Jing, Xu, Chenliang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EAGLE: Egocentric AGgregated Language-video Engine
por: Bi, Jing, et al.
Publicado: (2024)
por: Bi, Jing, et al.
Publicado: (2024)
OSCaR: Object State Captioning and State Change Representation
por: Nguyen, Nguyen, et al.
Publicado: (2024)
por: Nguyen, Nguyen, et al.
Publicado: (2024)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
por: Ye, Hanrong, et al.
Publicado: (2024)
por: Ye, Hanrong, et al.
Publicado: (2024)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
por: Bhalgat, Yash, et al.
Publicado: (2024)
por: Bhalgat, Yash, et al.
Publicado: (2024)
EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
por: VanVoorst, Brian, et al.
Publicado: (2026)
por: VanVoorst, Brian, et al.
Publicado: (2026)
On Memorization in Diffusion Models
por: Gu, Xiangming, et al.
Publicado: (2023)
por: Gu, Xiangming, et al.
Publicado: (2023)
EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos
por: Fujii, Ryo, et al.
Publicado: (2024)
por: Fujii, Ryo, et al.
Publicado: (2024)
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
por: Fujii, Ryo, et al.
Publicado: (2024)
por: Fujii, Ryo, et al.
Publicado: (2024)
EgoSurgery-HTS: A Dataset for Egocentric Hand-Tool Segmentation in Open Surgery Videos
por: Darjana, Nathan, et al.
Publicado: (2025)
por: Darjana, Nathan, et al.
Publicado: (2025)
Whole-Body Conditioned Egocentric Video Prediction
por: Bai, Yutong, et al.
Publicado: (2025)
por: Bai, Yutong, et al.
Publicado: (2025)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
por: Patel, Alkesh, et al.
Publicado: (2025)
por: Patel, Alkesh, et al.
Publicado: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
por: Yang, Ruihan, et al.
Publicado: (2025)
por: Yang, Ruihan, et al.
Publicado: (2025)
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation
por: Zhou, Chenliang, et al.
Publicado: (2022)
por: Zhou, Chenliang, et al.
Publicado: (2022)
Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection
por: Han, Boyu, et al.
Publicado: (2026)
por: Han, Boyu, et al.
Publicado: (2026)
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
por: Wang, Yunzhe, et al.
Publicado: (2025)
por: Wang, Yunzhe, et al.
Publicado: (2025)
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
por: Lin, Dongyan, et al.
Publicado: (2026)
por: Lin, Dongyan, et al.
Publicado: (2026)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
por: Qu, Leigang, et al.
Publicado: (2025)
por: Qu, Leigang, et al.
Publicado: (2025)
Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
por: Han, Boyu, et al.
Publicado: (2025)
por: Han, Boyu, et al.
Publicado: (2025)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
por: Wang, Wenhao, et al.
Publicado: (2025)
por: Wang, Wenhao, et al.
Publicado: (2025)
Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
por: Kim, Bumjun, et al.
Publicado: (2026)
por: Kim, Bumjun, et al.
Publicado: (2026)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
por: Chen, Baiyu, et al.
Publicado: (2025)
por: Chen, Baiyu, et al.
Publicado: (2025)
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
por: Karnik, Sathwik, et al.
Publicado: (2026)
por: Karnik, Sathwik, et al.
Publicado: (2026)
Detecting and Mitigating Memorization in Diffusion Models through Anisotropy of the Log-Probability
por: Asthana, Rohan, et al.
Publicado: (2026)
por: Asthana, Rohan, et al.
Publicado: (2026)
Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
por: Jain, Anubhav, et al.
Publicado: (2024)
por: Jain, Anubhav, et al.
Publicado: (2024)
Memorized Images in Diffusion Models share a Subspace that can be Located and Deleted
por: Chavhan, Ruchika, et al.
Publicado: (2024)
por: Chavhan, Ruchika, et al.
Publicado: (2024)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
por: Safaei, Bardia, et al.
Publicado: (2025)
por: Safaei, Bardia, et al.
Publicado: (2025)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
por: Tang, Yolo Yunlong, et al.
Publicado: (2024)
por: Tang, Yolo Yunlong, et al.
Publicado: (2024)
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
por: Chen, Lu, et al.
Publicado: (2025)
por: Chen, Lu, et al.
Publicado: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
por: Chu, Tianzhe, et al.
Publicado: (2025)
por: Chu, Tianzhe, et al.
Publicado: (2025)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
por: Hua, Hang, et al.
Publicado: (2024)
por: Hua, Hang, et al.
Publicado: (2024)
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Towards Streaming LiDAR Object Detection with Point Clouds as Egocentric Sequences
por: Zhang, Mellon M., et al.
Publicado: (2025)
por: Zhang, Mellon M., et al.
Publicado: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
por: Singhal, Rishi, et al.
Publicado: (2025)
por: Singhal, Rishi, et al.
Publicado: (2025)
Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
por: Du, Xiaodan, et al.
Publicado: (2023)
por: Du, Xiaodan, et al.
Publicado: (2023)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024)
por: Zhang, Yiyuan, et al.
Publicado: (2024)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
por: Lou, Siyu, et al.
Publicado: (2024)
por: Lou, Siyu, et al.
Publicado: (2024)
Efficient Pre-training for Localized Instruction Generation of Videos
por: Batra, Anil, et al.
Publicado: (2023)
por: Batra, Anil, et al.
Publicado: (2023)
Next Visual Granularity Generation
por: Wang, Yikai, et al.
Publicado: (2025)
por: Wang, Yikai, et al.
Publicado: (2025)
EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
por: Weerasinghe, Keshara, et al.
Publicado: (2025)
por: Weerasinghe, Keshara, et al.
Publicado: (2025)
Ejemplares similares
-
EAGLE: Egocentric AGgregated Language-video Engine
por: Bi, Jing, et al.
Publicado: (2024) -
OSCaR: Object State Captioning and State Change Representation
por: Nguyen, Nguyen, et al.
Publicado: (2024) -
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
por: Ye, Hanrong, et al.
Publicado: (2024) -
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
por: Bhalgat, Yash, et al.
Publicado: (2024) -
EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
por: VanVoorst, Brian, et al.
Publicado: (2026)