EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Yuejiao, Zhang, Xinshen, Ye, Zhen, Yao, Lei, Chau, Lap-Pui, Wang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
by: Su, Yuejiao, et al.
Published: (2025)
by: Su, Yuejiao, et al.
Published: (2025)
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024)
by: Su, Yuejiao, et al.
Published: (2024)
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
by: Su, Yuejiao, et al.
Published: (2026)
by: Su, Yuejiao, et al.
Published: (2026)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
by: Zheng, Ying, et al.
Published: (2024)
by: Zheng, Ying, et al.
Published: (2024)
MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video
by: Zhu, Haoyu, et al.
Published: (2026)
by: Zhu, Haoyu, et al.
Published: (2026)
OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
by: Chen, Junliang, et al.
Published: (2025)
by: Chen, Junliang, et al.
Published: (2025)
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
by: Yao, Lei, et al.
Published: (2026)
by: Yao, Lei, et al.
Published: (2026)
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025)
by: Deng, Kunyuan, et al.
Published: (2025)
A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
by: Xu, Huaiyuan, et al.
Published: (2024)
by: Xu, Huaiyuan, et al.
Published: (2024)
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
by: Wang, Xiaoqi, et al.
Published: (2025)
by: Wang, Xiaoqi, et al.
Published: (2025)
PEM: Perception Error Model for Virtual Testing of Autonomous Vehicles
by: Piazzoni, Andrea, et al.
Published: (2023)
by: Piazzoni, Andrea, et al.
Published: (2023)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation
by: Lin, Sixu, et al.
Published: (2026)
by: Lin, Sixu, et al.
Published: (2026)
Towards Unified Interactive Visual Grounding in The Wild
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
HSNet: Heterogeneous Subgraph Network for Single Image Super-resolution
by: Hu, Qiongyang, et al.
Published: (2025)
by: Hu, Qiongyang, et al.
Published: (2025)
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
by: Zhang, Xinshen, et al.
Published: (2025)
by: Zhang, Xinshen, et al.
Published: (2025)
SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2024)
by: Yao, Lei, et al.
Published: (2024)
UMIGen: A Unified Framework for Egocentric Point Cloud Generation and Cross-Embodiment Robotic Imitation Learning
by: Huang, Yan, et al.
Published: (2025)
by: Huang, Yan, et al.
Published: (2025)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction
by: Luo, Shaqi, et al.
Published: (2026)
by: Luo, Shaqi, et al.
Published: (2026)
MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
by: Deigmoeller, Joerg, et al.
Published: (2026)
by: Deigmoeller, Joerg, et al.
Published: (2026)
ProCal: Probability Calibration for Neighborhood-Guided Source-Free Domain Adaptation
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
Toward Grounded Commonsense Reasoning
by: Kwon, Minae, et al.
Published: (2023)
by: Kwon, Minae, et al.
Published: (2023)
Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning
by: Tirumala, Dhruva, et al.
Published: (2024)
by: Tirumala, Dhruva, et al.
Published: (2024)
RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
by: Zang, Hongzhi, et al.
Published: (2025)
by: Zang, Hongzhi, et al.
Published: (2025)
Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024
by: Wang, Xiaoqi, et al.
Published: (2024)
by: Wang, Xiaoqi, et al.
Published: (2024)
Weakly-supervised Part-Attention and Mentored Networks for Vehicle Re-Identification
by: Tang, Lisha, et al.
Published: (2021)
by: Tang, Lisha, et al.
Published: (2021)
EgoKit: Towards Unified Low-Cost Egocentric Data Collection with Heterogeneous Devices
by: Yu, Liuchuan, et al.
Published: (2026)
by: Yu, Liuchuan, et al.
Published: (2026)
Real-Time Communication-Aware Ride-Sharing Route Planning for Urban Air Mobility: A Multi-Source Hybrid Attention Reinforcement Learning Approach
by: Xie, Yuejiao, et al.
Published: (2025)
by: Xie, Yuejiao, et al.
Published: (2025)
Web-Gewu: A Browser-Based Interactive Playground for Robot Reinforcement Learning
by: Chen, Kaixuan, et al.
Published: (2026)
by: Chen, Kaixuan, et al.
Published: (2026)
A Unified Framework for Robots that Influence Humans over Long-Term Interaction
by: Sagheb, Shahabedin, et al.
Published: (2025)
by: Sagheb, Shahabedin, et al.
Published: (2025)
LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2026)
by: Yao, Lei, et al.
Published: (2026)
A Unified Approach to Multi-task Legged Navigation: Temporal Logic Meets Reinforcement Learning
by: Jiang, Jesse, et al.
Published: (2024)
by: Jiang, Jesse, et al.
Published: (2024)
Dynamic Deep Factor Graph for Multi-Agent Reinforcement Learning
by: Shi, Yuchen, et al.
Published: (2024)
by: Shi, Yuchen, et al.
Published: (2024)
IG-RFT: An Interaction-Guided RL Framework for VLA Models in Long-Horizon Robotic Manipulation
by: Su, Zhian, et al.
Published: (2026)
by: Su, Zhian, et al.
Published: (2026)
TouchAnything: A Dataset and Framework for Bimanual Tactile Estimation from Egocentric Video
by: Zhou, Jianyi, et al.
Published: (2026)
by: Zhou, Jianyi, et al.
Published: (2026)
RLIF: Interactive Imitation Learning as Reinforcement Learning
by: Luo, Jianlan, et al.
Published: (2023)
by: Luo, Jianlan, et al.
Published: (2023)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
by: Bao, Muyi, et al.
Published: (2026)
by: Bao, Muyi, et al.
Published: (2026)
Similar Items
-
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
by: Su, Yuejiao, et al.
Published: (2025) -
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024) -
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
by: Su, Yuejiao, et al.
Published: (2026) -
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
by: Zheng, Ying, et al.
Published: (2024) -
MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video
by: Zhu, Haoyu, et al.
Published: (2026)