IMPACT: A Dataset for Multi-Granularity Human Procedural Action Understanding in Industrial Assembly
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Di, Zhong, Zeyun, Schneider, David, Zaremski, Manuel, Kunzmann, Linus, Shi, Yitian, Liu, Ruiping, Chen, Yufan, Zheng, Junwei, Li, Jiahang, Hemmerich, Jonas, Tong, Qiyi, Grauberger, Patric, Ajoudani, Arash, Paudel, Danda Pani, Matthiesen, Sven, Deml, Barbara, Beyerer, Jürgen, Van Gool, Luc, Stiefelhagen, Rainer, Peng, Kunyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
von: Wen, Di, et al.
Veröffentlicht: (2025)
von: Wen, Di, et al.
Veröffentlicht: (2025)
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
von: Peng, Kunyu, et al.
Veröffentlicht: (2025)
von: Peng, Kunyu, et al.
Veröffentlicht: (2025)
EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
von: Peng, Kunyu, et al.
Veröffentlicht: (2025)
von: Peng, Kunyu, et al.
Veröffentlicht: (2025)
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
Autonomous Vehicle Controllers From End-to-End Differentiable Simulation
von: Nachkov, Asen, et al.
Veröffentlicht: (2024)
von: Nachkov, Asen, et al.
Veröffentlicht: (2024)
EvenNICER-SLAM: Event-based Neural Implicit Encoding SLAM
von: Chen, Shi, et al.
Veröffentlicht: (2024)
von: Chen, Shi, et al.
Veröffentlicht: (2024)
IMPACT-CYCLE: A Contract-Based Multi-Agent System for Claim-Level Supervisory Correction of Long-Video Semantic Memory
von: Kong, Weitong, et al.
Veröffentlicht: (2026)
von: Kong, Weitong, et al.
Veröffentlicht: (2026)
IMPACT-Scribe: Interactive Temporal Action Segmentation with Boundary Scribbles and Query Planning
von: Yin, Qian, et al.
Veröffentlicht: (2026)
von: Yin, Qian, et al.
Veröffentlicht: (2026)
MICA: Multi-Agent Industrial Coordination Assistant
von: Wen, Di, et al.
Veröffentlicht: (2025)
von: Wen, Di, et al.
Veröffentlicht: (2025)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
von: Chen, Yufan, et al.
Veröffentlicht: (2024)
von: Chen, Yufan, et al.
Veröffentlicht: (2024)
HybriDLA: Hybrid Generation for Document Layout Analysis
von: Chen, Yufan, et al.
Veröffentlicht: (2025)
von: Chen, Yufan, et al.
Veröffentlicht: (2025)
Graph-based Document Structure Analysis
von: Chen, Yufan, et al.
Veröffentlicht: (2025)
von: Chen, Yufan, et al.
Veröffentlicht: (2025)
InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing
von: Yang, Yebin, et al.
Veröffentlicht: (2026)
von: Yang, Yebin, et al.
Veröffentlicht: (2026)
IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction
von: Zhang, Haoshen, et al.
Veröffentlicht: (2026)
von: Zhang, Haoshen, et al.
Veröffentlicht: (2026)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
von: Ma, Qi, et al.
Veröffentlicht: (2024)
von: Ma, Qi, et al.
Veröffentlicht: (2024)
Continuous Pose for Monocular Cameras in Neural Implicit Representation
von: Ma, Qi, et al.
Veröffentlicht: (2023)
von: Ma, Qi, et al.
Veröffentlicht: (2023)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2026)
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2026)
Vision encoders should be image size agnostic and task driven
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2025)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2025)
Self-supervised pretraining for an iterative image size agnostic vision transformer
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2026)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2026)
Open Panoramic Segmentation
von: Zheng, Junwei, et al.
Veröffentlicht: (2024)
von: Zheng, Junwei, et al.
Veröffentlicht: (2024)
Inferring Compositional 4D Scenes without Ever Seeing One
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
A Simple and Generalist Approach for Panoptic Segmentation
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2024)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2024)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
von: Balauca, Ada-Astrid, et al.
Veröffentlicht: (2024)
von: Balauca, Ada-Astrid, et al.
Veröffentlicht: (2024)
RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
Incremental Object Detection with Prompt-based Methods
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
Occam's LGS: An Efficient Approach for Language Gaussian Splatting
von: Cheng, Jiahuan, et al.
Veröffentlicht: (2024)
von: Cheng, Jiahuan, et al.
Veröffentlicht: (2024)
Snap, Segment, Deploy: A Visual Data and Detection Pipeline for Wearable Industrial Assistants
von: Wen, Di, et al.
Veröffentlicht: (2025)
von: Wen, Di, et al.
Veröffentlicht: (2025)
ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy Prediction
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
SeasonScapes: Learning Large-scale Re-lightable 3D Landscapes with Seasonal Variation from Sparse Webcams
von: Kleger, Timo, et al.
Veröffentlicht: (2026)
von: Kleger, Timo, et al.
Veröffentlicht: (2026)
Ternary-Type Opacity and Hybrid Odometry for RGB NeRF-SLAM
von: Lin, Junru, et al.
Veröffentlicht: (2023)
von: Lin, Junru, et al.
Veröffentlicht: (2023)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
von: Liu, Ruiping, et al.
Veröffentlicht: (2025)
von: Liu, Ruiping, et al.
Veröffentlicht: (2025)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
von: Wei, Yiping, et al.
Veröffentlicht: (2023)
von: Wei, Yiping, et al.
Veröffentlicht: (2023)
Towards Multi-Source Domain Generalization for Sleep Staging with Noisy Labels
von: Wang, Kening, et al.
Veröffentlicht: (2026)
von: Wang, Kening, et al.
Veröffentlicht: (2026)
ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models
von: Dey, Sombit, et al.
Veröffentlicht: (2024)
von: Dey, Sombit, et al.
Veröffentlicht: (2024)
Autonomous Vehicle Path Planning by Searching With Differentiable Simulation
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
Unlocking Efficient Vehicle Dynamics Modeling via Analytic World Models
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
Generalist Robot Manipulation beyond Action Labeled Data
von: Spiridonov, Alexander, et al.
Veröffentlicht: (2025)
von: Spiridonov, Alexander, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
von: Wen, Di, et al.
Veröffentlicht: (2025) -
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
von: Peng, Kunyu, et al.
Veröffentlicht: (2025) -
EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
von: Peng, Kunyu, et al.
Veröffentlicht: (2025) -
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
von: Jiang, Xin, et al.
Veröffentlicht: (2024) -
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
von: Motamed, Saman, et al.
Veröffentlicht: (2023)