Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wenxin, Peng, Kunyu, Wen, Di, Liu, Ruiping, Duan, Mengfei, Luo, Kai, Yang, Kailun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can we Trust Unreliable Voxels? Exploring 3D Semantic Occupancy Prediction under Label Noise
by: Li, Wenxin, et al.
Published: (2026)
by: Li, Wenxin, et al.
Published: (2026)
Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
by: Liu, Ruiping, et al.
Published: (2024)
by: Liu, Ruiping, et al.
Published: (2024)
Skeleton-Based Human Action Recognition with Noisy Labels
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Out-of-Distribution Semantic Occupancy Prediction
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
Referring Atomic Video Action Recognition
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
$M^2$-Occ: Resilient 3D Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera Inputs
by: Lin, Kaixin, et al.
Published: (2026)
by: Lin, Kaixin, et al.
Published: (2026)
TransKD: Transformer Knowledge Distillation for Efficient Semantic Segmentation
by: Liu, Ruiping, et al.
Published: (2022)
by: Liu, Ruiping, et al.
Published: (2022)
Panoramic Out-of-Distribution Segmentation
by: Duan, Mengfei, et al.
Published: (2025)
by: Duan, Mengfei, et al.
Published: (2025)
OAFuser: Towards Omni-Aperture Fusion for Light Field Semantic Segmentation
by: Teng, Fei, et al.
Published: (2023)
by: Teng, Fei, et al.
Published: (2023)
Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentation
by: Zheng, Yuanfan, et al.
Published: (2026)
by: Zheng, Yuanfan, et al.
Published: (2026)
Exploring Video-Based Driver Activity Recognition under Noisy Labels
by: Fan, Linjuan, et al.
Published: (2025)
by: Fan, Linjuan, et al.
Published: (2025)
LFX: Towards Unified Light Field Dense Semantic Segmentation and Salient Object Detection
by: Teng, Fei, et al.
Published: (2025)
by: Teng, Fei, et al.
Published: (2025)
Occlusion-Aware Seamless Segmentation
by: Cao, Yihong, et al.
Published: (2024)
by: Cao, Yihong, et al.
Published: (2024)
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
by: Peng, Kunyu, et al.
Published: (2025)
by: Peng, Kunyu, et al.
Published: (2025)
ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy Prediction
by: Zhang, Yuheng, et al.
Published: (2026)
by: Zhang, Yuheng, et al.
Published: (2026)
Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation
by: Cao, Yihong, et al.
Published: (2025)
by: Cao, Yihong, et al.
Published: (2025)
HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios
by: Peng, Kunyu, et al.
Published: (2025)
by: Peng, Kunyu, et al.
Published: (2025)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
Hallucinating 360°: Panoramic Street-View Generation via Local Scenes Diffusion and Probabilistic Prompting
by: Teng, Fei, et al.
Published: (2025)
by: Teng, Fei, et al.
Published: (2025)
Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language Guidance
by: Zhao, Jiayi, et al.
Published: (2025)
by: Zhao, Jiayi, et al.
Published: (2025)
OccTrack360: 4D Panoptic Occupancy Tracking from Surround-View Fisheye Cameras
by: Lin, Yongzhi, et al.
Published: (2026)
by: Lin, Yongzhi, et al.
Published: (2026)
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
by: Wen, Di, et al.
Published: (2025)
by: Wen, Di, et al.
Published: (2025)
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
OmniTrack++: Omnidirectional Multi-Object Tracking by Learning Large-FoV Trajectory Feedback
by: Luo, Kai, et al.
Published: (2025)
by: Luo, Kai, et al.
Published: (2025)
NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models
by: Li, Siyu, et al.
Published: (2025)
by: Li, Siyu, et al.
Published: (2025)
Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation
by: Zhang, Jiaming, et al.
Published: (2022)
by: Zhang, Jiaming, et al.
Published: (2022)
Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots
by: Zhao, Guoqiang, et al.
Published: (2026)
by: Zhao, Guoqiang, et al.
Published: (2026)
Omnidirectional Multi-Object Tracking
by: Luo, Kai, et al.
Published: (2025)
by: Luo, Kai, et al.
Published: (2025)
Towards Source-free Domain Adaptive Semantic Segmentation via Importance-aware and Prototype-contrast Learning
by: Cao, Yihong, et al.
Published: (2023)
by: Cao, Yihong, et al.
Published: (2023)
InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing
by: Yang, Yebin, et al.
Published: (2026)
by: Yang, Yebin, et al.
Published: (2026)
Towards Activated Muscle Group Estimation in the Wild
by: Peng, Kunyu, et al.
Published: (2023)
by: Peng, Kunyu, et al.
Published: (2023)
O3N: Omnidirectional Open-Vocabulary Occupancy Prediction
by: Duan, Mengfei, et al.
Published: (2026)
by: Duan, Mengfei, et al.
Published: (2026)
LF Tracy: A Unified Single-Pipeline Approach for Salient Object Detection in Light Field Cameras
by: Teng, Fei, et al.
Published: (2024)
by: Teng, Fei, et al.
Published: (2024)
Towards Consistent Object Detection via LiDAR-Camera Synergy
by: Luo, Kai, et al.
Published: (2024)
by: Luo, Kai, et al.
Published: (2024)
Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
Exploring Few-Shot Adaptation for Activity Recognition on Diverse Domains
by: Peng, Kunyu, et al.
Published: (2023)
by: Peng, Kunyu, et al.
Published: (2023)
QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots
by: Wu, Sheng, et al.
Published: (2025)
by: Wu, Sheng, et al.
Published: (2025)
NOVA: Next-step Open-Vocabulary Autoregression for 3D Multi-Object Tracking in Autonomous Driving
by: Luo, Kai, et al.
Published: (2026)
by: Luo, Kai, et al.
Published: (2026)
SELECT: A Submodular Approach for Active LiDAR Semantic Segmentation
by: Mao, Ruiyu, et al.
Published: (2025)
by: Mao, Ruiyu, et al.
Published: (2025)
Similar Items
-
Can we Trust Unreliable Voxels? Exploring 3D Semantic Occupancy Prediction under Label Noise
by: Li, Wenxin, et al.
Published: (2026) -
Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
by: Liu, Ruiping, et al.
Published: (2024) -
Skeleton-Based Human Action Recognition with Noisy Labels
by: Xu, Yi, et al.
Published: (2024) -
Out-of-Distribution Semantic Occupancy Prediction
by: Zhang, Yuheng, et al.
Published: (2025) -
Referring Atomic Video Action Recognition
by: Peng, Kunyu, et al.
Published: (2024)