RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yufan, Zhang, Jiaming, Peng, Kunyu, Zheng, Junwei, Liu, Ruiping, Torr, Philip, Stiefelhagen, Rainer |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HybriDLA: Hybrid Generation for Document Layout Analysis
by: Chen, Yufan, et al.
Published: (2025)
by: Chen, Yufan, et al.
Published: (2025)
Graph-based Document Structure Analysis
by: Chen, Yufan, et al.
Published: (2025)
by: Chen, Yufan, et al.
Published: (2025)
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
by: Wen, Di, et al.
Published: (2025)
by: Wen, Di, et al.
Published: (2025)
SFDLA: Source-Free Document Layout Analysis
by: Tewes, Sebastian, et al.
Published: (2025)
by: Tewes, Sebastian, et al.
Published: (2025)
Open Panoramic Segmentation
by: Zheng, Junwei, et al.
Published: (2024)
by: Zheng, Junwei, et al.
Published: (2024)
Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
by: Liu, Ruiping, et al.
Published: (2025)
by: Liu, Ruiping, et al.
Published: (2025)
RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization
by: Zheng, Junwei, et al.
Published: (2026)
by: Zheng, Junwei, et al.
Published: (2026)
CHAOS: Chart Analysis with Outlier Samples
by: Moured, Omar, et al.
Published: (2025)
by: Moured, Omar, et al.
Published: (2025)
SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
Comb, Prune, Distill: Towards Unified Pruning for Vision Model Compression
by: Schmitt, Jonas, et al.
Published: (2024)
by: Schmitt, Jonas, et al.
Published: (2024)
Scene-agnostic Pose Regression for Visual Localization
by: Zheng, Junwei, et al.
Published: (2025)
by: Zheng, Junwei, et al.
Published: (2025)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
by: Liu, Ruiping, et al.
Published: (2026)
by: Liu, Ruiping, et al.
Published: (2026)
MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments
by: Zheng, Junwei, et al.
Published: (2023)
by: Zheng, Junwei, et al.
Published: (2023)
What if? Emulative Simulation with World Models for Situated Reasoning
by: Liu, Ruiping, et al.
Published: (2026)
by: Liu, Ruiping, et al.
Published: (2026)
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
by: Jiang, Xin, et al.
Published: (2024)
by: Jiang, Xin, et al.
Published: (2024)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
by: Liu, Ruiping, et al.
Published: (2024)
by: Liu, Ruiping, et al.
Published: (2024)
Skeleton-Based Human Action Recognition with Noisy Labels
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
by: Wei, Jiale, et al.
Published: (2024)
by: Wei, Jiale, et al.
Published: (2024)
Towards Multi-Source Domain Generalization for Sleep Staging with Noisy Labels
by: Wang, Kening, et al.
Published: (2026)
by: Wang, Kening, et al.
Published: (2026)
MICA: Multi-Agent Industrial Coordination Assistant
by: Wen, Di, et al.
Published: (2025)
by: Wen, Di, et al.
Published: (2025)
More than the Sum: Panorama-Language Models for Adverse Omni-Scenes
by: Fan, Weijia, et al.
Published: (2026)
by: Fan, Weijia, et al.
Published: (2026)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
by: Tao, Mingzhe, et al.
Published: (2026)
by: Tao, Mingzhe, et al.
Published: (2026)
UnSupDLA: Towards Unsupervised Document Layout Analysis
by: Sheikh, Talha Uddin, et al.
Published: (2024)
by: Sheikh, Talha Uddin, et al.
Published: (2024)
Referring Atomic Video Action Recognition
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
Exploring Video-Based Driver Activity Recognition under Noisy Labels
by: Fan, Linjuan, et al.
Published: (2025)
by: Fan, Linjuan, et al.
Published: (2025)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
by: Luo, Yuanhao, et al.
Published: (2026)
by: Luo, Yuanhao, et al.
Published: (2026)
TransKD: Transformer Knowledge Distillation for Efficient Semantic Segmentation
by: Liu, Ruiping, et al.
Published: (2022)
by: Liu, Ruiping, et al.
Published: (2022)
Deformable Mamba for Wide Field of View Segmentation
by: Hu, Jie, et al.
Published: (2024)
by: Hu, Jie, et al.
Published: (2024)
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
by: Vogel, Alexander, et al.
Published: (2025)
by: Vogel, Alexander, et al.
Published: (2025)
IMPACT-CYCLE: A Contract-Based Multi-Agent System for Claim-Level Supervisory Correction of Long-Video Semantic Memory
by: Kong, Weitong, et al.
Published: (2026)
by: Kong, Weitong, et al.
Published: (2026)
IMPACT-Scribe: Interactive Temporal Action Segmentation with Boundary Scribbles and Query Planning
by: Yin, Qian, et al.
Published: (2026)
by: Yin, Qian, et al.
Published: (2026)
EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
by: Peng, Kunyu, et al.
Published: (2025)
by: Peng, Kunyu, et al.
Published: (2025)
Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments
by: Wen, Di, et al.
Published: (2025)
by: Wen, Di, et al.
Published: (2025)
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
by: Peng, Kunyu, et al.
Published: (2025)
by: Peng, Kunyu, et al.
Published: (2025)
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
by: Liu, Ruiping, et al.
Published: (2024)
by: Liu, Ruiping, et al.
Published: (2024)
OAFuser: Towards Omni-Aperture Fusion for Light Field Semantic Segmentation
by: Teng, Fei, et al.
Published: (2023)
by: Teng, Fei, et al.
Published: (2023)
mmWalk: Towards Multi-modal Multi-view Walking Assistance
by: Ying, Kedi, et al.
Published: (2025)
by: Ying, Kedi, et al.
Published: (2025)
Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation
by: Zhang, Jiaming, et al.
Published: (2022)
by: Zhang, Jiaming, et al.
Published: (2022)
Similar Items
-
HybriDLA: Hybrid Generation for Document Layout Analysis
by: Chen, Yufan, et al.
Published: (2025) -
Graph-based Document Structure Analysis
by: Chen, Yufan, et al.
Published: (2025) -
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
by: Wen, Di, et al.
Published: (2025) -
SFDLA: Source-Free Document Layout Analysis
by: Tewes, Sebastian, et al.
Published: (2025) -
Open Panoramic Segmentation
by: Zheng, Junwei, et al.
Published: (2024)