Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Boyu, Xu, Qianqian, Bao, Shilong, Yang, Zhiyong, Li, Sicong, Huang, Qingming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909822443585536
author Han, Boyu
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Li, Sicong
Huang, Qingming
author_facet Han, Boyu
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Li, Sicong
Huang, Qingming
contents In this report, we address the problem of determining whether a user performs an action incorrectly from egocentric video data. To handle the challenges posed by subtle and infrequent mistakes, we propose a Dual-Stage Reweighted Mixture-of-Experts (DR-MoE) framework. In the first stage, features are extracted using a frozen ViViT model and a LoRA-tuned ViViT model, which are combined through a feature-level expert module. In the second stage, three classifiers are trained with different objectives: reweighted cross-entropy to mitigate class imbalance, AUC loss to improve ranking under skewed distributions, and label-aware loss with sharpness-aware minimization to enhance calibration and generalization. Their predictions are fused using a classification-level expert module. The proposed method achieves strong performance, particularly in identifying rare and ambiguous mistake instances. The code is available at https://github.com/boyuh/DR-MoE.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
Han, Boyu
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Li, Sicong
Huang, Qingming
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
In this report, we address the problem of determining whether a user performs an action incorrectly from egocentric video data. To handle the challenges posed by subtle and infrequent mistakes, we propose a Dual-Stage Reweighted Mixture-of-Experts (DR-MoE) framework. In the first stage, features are extracted using a frozen ViViT model and a LoRA-tuned ViViT model, which are combined through a feature-level expert module. In the second stage, three classifiers are trained with different objectives: reweighted cross-entropy to mitigate class imbalance, AUC loss to improve ranking under skewed distributions, and label-aware loss with sharpness-aware minimization to enhance calibration and generalization. Their predictions are fused using a classification-level expert module. The proposed method achieves strong performance, particularly in identifying rare and ambiguous mistake instances. The code is available at https://github.com/boyuh/DR-MoE.
title Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.12990