Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Boyu, Xu, Qianqian, Bao, Shilong, Yang, Zhiyong, Cui, Ruochen, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2025)
by: Han, Boyu, et al.
Published: (2025)
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
by: Han, Boyu, et al.
Published: (2026)
by: Han, Boyu, et al.
Published: (2026)
Hybrid Generative Fusion for Efficient and Privacy-Preserving Face Recognition Dataset Generation
by: Li, Feiran, et al.
Published: (2025)
by: Li, Feiran, et al.
Published: (2025)
AUCSeg: AUC-oriented Pixel-level Long-tail Semantic Segmentation
by: Han, Boyu, et al.
Published: (2024)
by: Han, Boyu, et al.
Published: (2024)
LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text Encoders
by: Han, Boyu, et al.
Published: (2025)
by: Han, Boyu, et al.
Published: (2025)
Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
by: Bao, Shilong, et al.
Published: (2025)
by: Bao, Shilong, et al.
Published: (2025)
One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework
by: Li, Feiran, et al.
Published: (2025)
by: Li, Feiran, et al.
Published: (2025)
BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
by: Li, Feiran, et al.
Published: (2026)
by: Li, Feiran, et al.
Published: (2026)
ReconBoost: Boosting Can Achieve Modality Reconcilement
by: Hua, Cong, et al.
Published: (2024)
by: Hua, Cong, et al.
Published: (2024)
Bidirectional Logits Tree: Pursuing Granularity Reconcilement in Fine-Grained Classification
by: Lu, Zhiguang, et al.
Published: (2024)
by: Lu, Zhiguang, et al.
Published: (2024)
Size-invariance Matters: Rethinking Metrics and Losses for Imbalanced Multi-object Salient Object Detection
by: Li, Feiran, et al.
Published: (2024)
by: Li, Feiran, et al.
Published: (2024)
DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations
by: Yang, Zhiyong, et al.
Published: (2024)
by: Yang, Zhiyong, et al.
Published: (2024)
Suppress Content Shift: Better Diffusion Features via Off-the-Shelf Generation Techniques
by: Meng, Benyuan, et al.
Published: (2024)
by: Meng, Benyuan, et al.
Published: (2024)
Closing the Approximation Gap of Partial AUC Optimization: A Tale of Two Formulations
by: Jiang, Yangbangyan, et al.
Published: (2025)
by: Jiang, Yangbangyan, et al.
Published: (2025)
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
by: Haneji, Yuto, et al.
Published: (2024)
by: Haneji, Yuto, et al.
Published: (2024)
Not All Diffusion Model Activations Have Been Evaluated as Discriminative Features
by: Meng, Benyuan, et al.
Published: (2024)
by: Meng, Benyuan, et al.
Published: (2024)
A Unified Perspective for Loss-Oriented Imbalanced Learning via Localization
by: Wang, Zitai, et al.
Published: (2023)
by: Wang, Zitai, et al.
Published: (2023)
Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
by: Patsch, Constantin, et al.
Published: (2025)
by: Patsch, Constantin, et al.
Published: (2025)
Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs
by: Xu, Zhikang, et al.
Published: (2026)
by: Xu, Zhikang, et al.
Published: (2026)
Uncertainty-aware Long-tailed Weights Model the Utility of Pseudo-labels for Semi-supervised Learning
by: Wu, Jiaqi, et al.
Published: (2025)
by: Wu, Jiaqi, et al.
Published: (2025)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos
by: Loginova, Olga, et al.
Published: (2026)
by: Loginova, Olga, et al.
Published: (2026)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
by: Zhang, Deheng, et al.
Published: (2025)
by: Zhang, Deheng, et al.
Published: (2025)
HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
by: Wang, Daming, et al.
Published: (2025)
by: Wang, Daming, et al.
Published: (2025)
Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial Purification
by: Pei, Gaozheng, et al.
Published: (2025)
by: Pei, Gaozheng, et al.
Published: (2025)
Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
by: Jung, Ji Hyeok, et al.
Published: (2024)
by: Jung, Ji Hyeok, et al.
Published: (2024)
Adjusting Logit in Gaussian Form for Long-Tailed Visual Recognition
by: Li, Mengke, et al.
Published: (2023)
by: Li, Mengke, et al.
Published: (2023)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2025)
by: Seth, Ashish, et al.
Published: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
by: Mazzamuto, Michele, et al.
Published: (2024)
by: Mazzamuto, Michele, et al.
Published: (2024)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)
by: Seminara, Luigi, et al.
Published: (2024)
Enhancing Sample Utilization in Noise-Robust Deep Metric Learning With Subgroup-Based Positive-Pair Selection
by: Yu, Zhipeng, et al.
Published: (2025)
by: Yu, Zhipeng, et al.
Published: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
by: Pani, Anupam, et al.
Published: (2025)
by: Pani, Anupam, et al.
Published: (2025)
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)
by: Vinod, Ashwin, et al.
Published: (2025)
Compensating Visual Insufficiency with Stratified Language Guidance for Long-Tail Class Incremental Learning
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
by: Li, Yuan-Ming, et al.
Published: (2024)
by: Li, Yuan-Ming, et al.
Published: (2024)
Mitigating Long-Tail Bias in HOI Detection via Adaptive Diversity Cache
by: Jiang, Yuqiu, et al.
Published: (2025)
by: Jiang, Yuqiu, et al.
Published: (2025)
HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
by: Lu, Zhiguang, et al.
Published: (2025)
by: Lu, Zhiguang, et al.
Published: (2025)
VLMine: Long-Tail Data Mining with Vision Language Models
by: Ye, Mao, et al.
Published: (2024)
by: Ye, Mao, et al.
Published: (2024)
Similar Items
-
Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2025) -
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
by: Han, Boyu, et al.
Published: (2026) -
Hybrid Generative Fusion for Efficient and Privacy-Preserving Face Recognition Dataset Generation
by: Li, Feiran, et al.
Published: (2025) -
AUCSeg: AUC-oriented Pixel-level Long-tail Semantic Segmentation
by: Han, Boyu, et al.
Published: (2024) -
LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text Encoders
by: Han, Boyu, et al.
Published: (2025)