Saved in:
| Main Authors: | Xu, Mengzhu, Liu, Hanzhi, Peng, Ningkang, Chen, Qianyu, Xiao, Canran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.00694 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Achieve Prototypical Birth and Death for OOD Detection?
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
Holistic Reliability Propagation: Decoupling Annotation and Prediction for Robust Noisy-Label
by: Mao, Jingyang, et al.
Published: (2026)
by: Mao, Jingyang, et al.
Published: (2026)
Is Complex Training Necessary for Long-Tailed OOD Detection? A Re-think from Feature Geometry
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
HamBR: Active Decision Boundary Restoration Based on Hamiltonian Dynamics for Learning with Noisy Labels
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
Radial-Angular Geometry for Reliable Update Diagnosis in Noisy-Label Learning
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object
by: Chen, Hanzhi, et al.
Published: (2024)
by: Chen, Hanzhi, et al.
Published: (2024)
Learning Visual Affordance from Audio
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
RRNet: Configurable Real-Time Video Enhancement with Arbitrary Local Lighting Variations
by: Yang, Wenlong, et al.
Published: (2026)
by: Yang, Wenlong, et al.
Published: (2026)
GAMR: Geometric-Aware Manifold Regularization with Virtual Outlier Synthesis for Learning with Noisy Labels
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model
by: Liu, Shang-Ching, et al.
Published: (2024)
by: Liu, Shang-Ching, et al.
Published: (2024)
TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances
by: Xu, Wenting, et al.
Published: (2024)
by: Xu, Wenting, et al.
Published: (2024)
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2026)
by: Wang, Hanqing, et al.
Published: (2026)
When Accuracy Is Not Enough: Uncertainty Collapse between Noisy Label Learning and Out-of-Distribution Detection
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
Early High-Frequency Injection for Geometry-Sensitive OOD Detection
by: Cheng, Chuanjie, et al.
Published: (2026)
by: Cheng, Chuanjie, et al.
Published: (2026)
Efficient Continual Learning through Frequency Decomposition and Integration
by: Liu, Ruiqi, et al.
Published: (2025)
by: Liu, Ruiqi, et al.
Published: (2025)
Populate-A-Scene: Affordance-Aware Human Video Generation
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
Learning with Adaptive Prototype Manifolds for Out-of-Distribution Detection
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
arg-VU: Affordance Reasoning with Physics-Aware 3D Geometry for Visual Understanding in Robotic Surgery
by: Xiao, Nan, et al.
Published: (2026)
by: Xiao, Nan, et al.
Published: (2026)
Multi-View Synergistic Learning with Vision-Language Adaption for Low-Resource Biomedical Image Classification
by: Luo, Xiaoliu, et al.
Published: (2026)
by: Luo, Xiaoliu, et al.
Published: (2026)
DepthSSC: Monocular 3D Semantic Scene Completion via Depth-Spatial Alignment and Voxel Adaptation
by: Yao, Jiawei, et al.
Published: (2023)
by: Yao, Jiawei, et al.
Published: (2023)
HarmonicNeRF: Geometry-Informed Synthetic View Augmentation for 3D Scene Reconstruction in Driving Scenarios
by: Pan, Xiaochao, et al.
Published: (2023)
by: Pan, Xiaochao, et al.
Published: (2023)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Low-light Image Enhancement with Retinex Decomposition in Latent Space
by: Zheng, Bolun, et al.
Published: (2026)
by: Zheng, Bolun, et al.
Published: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
by: Tan, Yue, et al.
Published: (2025)
by: Tan, Yue, et al.
Published: (2025)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
by: Chen, Hanzhi, et al.
Published: (2025)
by: Chen, Hanzhi, et al.
Published: (2025)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
VMF-GOS: Geometry-guided virtual Outlier Synthesis for Long-Tailed OOD Detection
by: Peng, Ningkang, et al.
Published: (2026)
by: Peng, Ningkang, et al.
Published: (2026)
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
MIDAS: Modeling Ground-Truth Distributions with Dark Knowledge for Domain Generalized Stereo Matching
by: Xu, Peng, et al.
Published: (2025)
by: Xu, Peng, et al.
Published: (2025)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
by: Mao, Aihua, et al.
Published: (2026)
by: Mao, Aihua, et al.
Published: (2026)
Self-Explainable Affordance Learning with Embodied Caption
by: Zhang, Zhipeng, et al.
Published: (2024)
by: Zhang, Zhipeng, et al.
Published: (2024)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
by: Shen, Xiaoqian, et al.
Published: (2024)
by: Shen, Xiaoqian, et al.
Published: (2024)
CL-VISTA: Benchmarking Continual Learning in Video Large Language Models
by: Guo, Haiyang, et al.
Published: (2026)
by: Guo, Haiyang, et al.
Published: (2026)
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
by: Zhou, Dingyi, et al.
Published: (2026)
by: Zhou, Dingyi, et al.
Published: (2026)
Similar Items
-
How to Achieve Prototypical Birth and Death for OOD Detection?
by: Peng, Ningkang, et al.
Published: (2026) -
Holistic Reliability Propagation: Decoupling Annotation and Prediction for Robust Noisy-Label
by: Mao, Jingyang, et al.
Published: (2026) -
Is Complex Training Necessary for Long-Tailed OOD Detection? A Re-think from Feature Geometry
by: Peng, Ningkang, et al.
Published: (2026) -
HamBR: Active Decision Boundary Restoration Based on Hamiltonian Dynamics for Learning with Noisy Labels
by: Peng, Ningkang, et al.
Published: (2026) -
Radial-Angular Geometry for Reliable Update Diagnosis in Noisy-Label Learning
by: Peng, Ningkang, et al.
Published: (2026)