Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Junlong, Xu, Huaiyuan, Cheng, Sijie, Wu, Kejun, Yap, Kim-Hui, Chau, Lap-Pui, Wang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025)
by: Deng, Kunyuan, et al.
Published: (2025)
OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
by: Chen, Junliang, et al.
Published: (2025)
by: Chen, Junliang, et al.
Published: (2025)
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024)
by: Su, Yuejiao, et al.
Published: (2024)
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
by: Liu, Wenyang, et al.
Published: (2024)
by: Liu, Wenyang, et al.
Published: (2024)
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
by: Liu, Wenyang, et al.
Published: (2025)
by: Liu, Wenyang, et al.
Published: (2025)
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
by: Wang, Xiaoqi, et al.
Published: (2025)
by: Wang, Xiaoqi, et al.
Published: (2025)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
by: Liu, Tianyi, et al.
Published: (2025)
by: Liu, Tianyi, et al.
Published: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024
by: Wang, Xiaoqi, et al.
Published: (2024)
by: Wang, Xiaoqi, et al.
Published: (2024)
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
by: Su, Yuejiao, et al.
Published: (2025)
by: Su, Yuejiao, et al.
Published: (2025)
A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
by: Xu, Huaiyuan, et al.
Published: (2024)
by: Xu, Huaiyuan, et al.
Published: (2024)
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
by: Su, Yuejiao, et al.
Published: (2026)
by: Su, Yuejiao, et al.
Published: (2026)
Weakly-supervised Part-Attention and Mentored Networks for Vehicle Re-Identification
by: Tang, Lisha, et al.
Published: (2021)
by: Tang, Lisha, et al.
Published: (2021)
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
by: Su, Yuejiao, et al.
Published: (2026)
by: Su, Yuejiao, et al.
Published: (2026)
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
ProCal: Probability Calibration for Neighborhood-Guided Source-Free Domain Adaptation
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
Evolution-Inspired Sample Competition for Deep Neural Network Optimization
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2024)
by: Yao, Lei, et al.
Published: (2024)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
by: Yan, Jiaqi, et al.
Published: (2025)
by: Yan, Jiaqi, et al.
Published: (2025)
Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
by: Zheng, Ying, et al.
Published: (2025)
by: Zheng, Ying, et al.
Published: (2025)
LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2026)
by: Yao, Lei, et al.
Published: (2026)
MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video
by: Zhu, Haoyu, et al.
Published: (2026)
by: Zhu, Haoyu, et al.
Published: (2026)
PADetBench: Towards Benchmarking Physical Attacks against Object Detection
by: Lian, Jiawei, et al.
Published: (2024)
by: Lian, Jiawei, et al.
Published: (2024)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
by: Yao, Lei, et al.
Published: (2026)
by: Yao, Lei, et al.
Published: (2026)
HSNet: Heterogeneous Subgraph Network for Single Image Super-resolution
by: Hu, Qiongyang, et al.
Published: (2025)
by: Hu, Qiongyang, et al.
Published: (2025)
SignEye: Traffic Sign Interpretation from Vehicle First-Person View
by: Yang, Chuang, et al.
Published: (2024)
by: Yang, Chuang, et al.
Published: (2024)
NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results
by: Zou, Wenbin, et al.
Published: (2026)
by: Zou, Wenbin, et al.
Published: (2026)
Video sentence grounding with temporally global textual knowledge
by: Chen, Cai, et al.
Published: (2024)
by: Chen, Cai, et al.
Published: (2024)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
SP-MoMamba: Superpixel-driven Mixture of State Space Experts for Efficient Image Super-Resolution
by: Zou, Wenbin, et al.
Published: (2026)
by: Zou, Wenbin, et al.
Published: (2026)
F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental Learning
by: Zhuang, Huiping, et al.
Published: (2024)
by: Zhuang, Huiping, et al.
Published: (2024)
EgoLife: Towards Egocentric Life Assistant
by: Yang, Jingkang, et al.
Published: (2025)
by: Yang, Jingkang, et al.
Published: (2025)
HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
by: Chen, Mingjin, et al.
Published: (2026)
by: Chen, Mingjin, et al.
Published: (2026)
Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
by: Fan, Zicong, et al.
Published: (2024)
by: Fan, Zicong, et al.
Published: (2024)
EgoSelf: From Memory to Personalized Egocentric Assistant
by: Wang, Yanshuo, et al.
Published: (2026)
by: Wang, Yanshuo, et al.
Published: (2026)
How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos
by: Loginova, Olga, et al.
Published: (2026)
by: Loginova, Olga, et al.
Published: (2026)
Similar Items
-
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025) -
OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
by: Chen, Junliang, et al.
Published: (2025) -
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024) -
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
by: Liu, Wenyang, et al.
Published: (2024) -
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
by: Liu, Wenyang, et al.
Published: (2025)