EvoStruggle: A Dataset Capturing the Evolution of Struggle across Activities and Skill Levels
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Shijia, Wray, Michael, Mayol-Cuevas, Walterio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and Activities
by: Feng, Shijia, et al.
Published: (2025)
by: Feng, Shijia, et al.
Published: (2025)
Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
by: Feng, Shijia, et al.
Published: (2024)
by: Feng, Shijia, et al.
Published: (2024)
Re-localization acceleration with Medoid Silhouette Clustering
by: Zhang, Hongyi, et al.
Published: (2024)
by: Zhang, Hongyi, et al.
Published: (2024)
CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene Rendering
by: Zheng, Xinyi, et al.
Published: (2025)
by: Zheng, Xinyi, et al.
Published: (2025)
Why MLLMs Struggle to Determine Object Orientations
by: Gopinath, Anju, et al.
Published: (2026)
by: Gopinath, Anju, et al.
Published: (2026)
Struggle with Adversarial Defense? Try Diffusion
by: Li, Yujie, et al.
Published: (2024)
by: Li, Yujie, et al.
Published: (2024)
SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything
by: Ji, Ge-Peng, et al.
Published: (2023)
by: Ji, Ge-Peng, et al.
Published: (2023)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
by: Zhou, Wenqi, et al.
Published: (2025)
by: Zhou, Wenqi, et al.
Published: (2025)
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
Lost in Translation: Modern Neural Networks Still Struggle With Small Realistic Image Transformations
by: Shifman, Ofir, et al.
Published: (2024)
by: Shifman, Ofir, et al.
Published: (2024)
Why Do Vision Language Models Struggle To Recognize Human Emotions?
by: Agarwal, Madhav, et al.
Published: (2026)
by: Agarwal, Madhav, et al.
Published: (2026)
Catch-Up Mix: Catch-Up Class for Struggling Filters in CNN
by: Kang, Minsoo, et al.
Published: (2024)
by: Kang, Minsoo, et al.
Published: (2024)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
by: Kargaran, Amir Hossein, et al.
Published: (2026)
by: Kargaran, Amir Hossein, et al.
Published: (2026)
A Video Is Not Worth a Thousand Words
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
Specialized Foundation Models Struggle to Beat Supervised Baselines
by: Xu, Zongzhe, et al.
Published: (2024)
by: Xu, Zongzhe, et al.
Published: (2024)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Multimodal LLMs Struggle with Basic Visual Network Analysis: a VNA Benchmark
by: Williams, Evan M., et al.
Published: (2024)
by: Williams, Evan M., et al.
Published: (2024)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
by: Deng, Ken, et al.
Published: (2026)
by: Deng, Ken, et al.
Published: (2026)
StrengthSense: A Dataset of IMU Signals Capturing Everyday Strength-Demanding Activities
by: Yang, Zeyu, et al.
Published: (2025)
by: Yang, Zeyu, et al.
Published: (2025)
Video, How Do Your Tokens Merge?
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
by: Lucy, Li, et al.
Published: (2026)
by: Lucy, Li, et al.
Published: (2026)
LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
by: Kong, Fei, et al.
Published: (2025)
by: Kong, Fei, et al.
Published: (2025)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models
by: Lu, Hongyu, et al.
Published: (2026)
by: Lu, Hongyu, et al.
Published: (2026)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025)
by: Flanagan, Kevin, et al.
Published: (2025)
SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)
by: Bansal, Siddhant, et al.
Published: (2024)
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains
by: Mucha, Wiktor, et al.
Published: (2026)
by: Mucha, Wiktor, et al.
Published: (2026)
Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document Retrieval
by: Li, Weiqing, et al.
Published: (2026)
by: Li, Weiqing, et al.
Published: (2026)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
CurEvo: Curriculum-Guided Self-Evolution for Video Understanding
by: Zeng, Guiyi, et al.
Published: (2026)
by: Zeng, Guiyi, et al.
Published: (2026)
ProSkill: Segment-Level Skill Assessment in Procedural Videos
by: Mazzamuto, Michele, et al.
Published: (2026)
by: Mazzamuto, Michele, et al.
Published: (2026)
Facial Appearance Capture at Home with Patch-Level Reflectance Prior
by: Han, Yuxuan, et al.
Published: (2025)
by: Han, Yuxuan, et al.
Published: (2025)
EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution
by: Wang, Tianfu, et al.
Published: (2026)
by: Wang, Tianfu, et al.
Published: (2026)
MicroEvoEval: A Systematic Evaluation Framework for Image-Based Microstructure Evolution Prediction
by: Zhang, Qinyi, et al.
Published: (2025)
by: Zhang, Qinyi, et al.
Published: (2025)
Similar Items
-
From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and Activities
by: Feng, Shijia, et al.
Published: (2025) -
Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
by: Feng, Shijia, et al.
Published: (2024) -
Re-localization acceleration with Medoid Silhouette Clustering
by: Zhang, Hongyi, et al.
Published: (2024) -
CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene Rendering
by: Zheng, Xinyi, et al.
Published: (2025) -
Why MLLMs Struggle to Determine Object Orientations
by: Gopinath, Anju, et al.
Published: (2026)