Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Heyne, Catyana, Frikel, Jürgen, Riccio, Filippo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IMUVIE: Pickup Timeline Action Localization via Motion Movies
by: Clapham, John, et al.
Published: (2024)
by: Clapham, John, et al.
Published: (2024)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
by: Sutton, Matthew, et al.
Published: (2026)
by: Sutton, Matthew, et al.
Published: (2026)
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review
by: Rombach, Alexander Michael, et al.
Published: (2024)
by: Rombach, Alexander Michael, et al.
Published: (2024)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
by: Banthia, Saumya, et al.
Published: (2020)
by: Banthia, Saumya, et al.
Published: (2020)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
by: Chien, Sheng-You, et al.
Published: (2026)
by: Chien, Sheng-You, et al.
Published: (2026)
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
by: Idahl, Maximilian, et al.
Published: (2026)
by: Idahl, Maximilian, et al.
Published: (2026)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
by: Nguyen, Huyen, et al.
Published: (2026)
by: Nguyen, Huyen, et al.
Published: (2026)
A Spatio-Temporal Deep Learning Approach For High-Resolution Gridded Monsoon Prediction
by: Borah, Parashjyoti, et al.
Published: (2026)
by: Borah, Parashjyoti, et al.
Published: (2026)
WaveMix: A Resource-efficient Neural Network for Image Analysis
by: Jeevan, Pranav, et al.
Published: (2022)
by: Jeevan, Pranav, et al.
Published: (2022)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
by: Ismail, Saifelden M.
Published: (2025)
by: Ismail, Saifelden M.
Published: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
by: Drchal, Jan, et al.
Published: (2023)
by: Drchal, Jan, et al.
Published: (2023)
Automating Clinical Information Retrieval from Finnish Electronic Health Records Using Large Language Models
by: Saukkoriipi, Mikko, et al.
Published: (2026)
by: Saukkoriipi, Mikko, et al.
Published: (2026)
FLD+: Data-efficient Evaluation Metric for Generative Models
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Normalizing Flow-Based Metric for Image Generation
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
Evaluation Metric for Quality Control and Generative Models in Histopathology Images
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
by: Gupta, Sunny, et al.
Published: (2024)
by: Gupta, Sunny, et al.
Published: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
by: Kashyap, Pankhi, et al.
Published: (2024)
by: Kashyap, Pankhi, et al.
Published: (2024)
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
by: Hua, Wenjie, et al.
Published: (2025)
by: Hua, Wenjie, et al.
Published: (2025)
Deep Spectral Meshes: Multi-Frequency Facial Mesh Processing with Graph Neural Networks
by: Kosk, Robert, et al.
Published: (2024)
by: Kosk, Robert, et al.
Published: (2024)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
by: Perera, Amal S., et al.
Published: (2025)
by: Perera, Amal S., et al.
Published: (2025)
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
by: Wright, Devin R., et al.
Published: (2026)
by: Wright, Devin R., et al.
Published: (2026)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
by: Martin, Michael R., et al.
Published: (2025)
by: Martin, Michael R., et al.
Published: (2025)
Improving Omics-Based Classification: The Role of Feature Selection and Synthetic Data Generation
by: Perazzolo, Diego, et al.
Published: (2025)
by: Perazzolo, Diego, et al.
Published: (2025)
FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Document Understanding for Healthcare Referrals
by: Mistry, Jimit, et al.
Published: (2023)
by: Mistry, Jimit, et al.
Published: (2023)
Alternative Local Discriminant Bases Using Empirical Expectation and Variance Estimation
by: Fossgaard, Eirik
Published: (1999)
by: Fossgaard, Eirik
Published: (1999)
Pairwise Spatiotemporal Partial Trajectory Matching for Co-movement Analysis
by: Cardei, Maria, et al.
Published: (2024)
by: Cardei, Maria, et al.
Published: (2024)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
by: Bouzid, Hamza, et al.
Published: (2023)
by: Bouzid, Hamza, et al.
Published: (2023)
DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology
by: Dimasaka, Joshua, et al.
Published: (2025)
by: Dimasaka, Joshua, et al.
Published: (2025)
Scaling Laws for State Dynamics in Large Language Models
by: Li, Jacob X, et al.
Published: (2025)
by: Li, Jacob X, et al.
Published: (2025)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
by: Borisov, Vadim
Published: (2026)
by: Borisov, Vadim
Published: (2026)
SYNOSIS: Image synthesis pipeline for machine vision in metal surface inspection
by: Fulir, Juraj, et al.
Published: (2024)
by: Fulir, Juraj, et al.
Published: (2024)
FEDTAIL: Federated Long-Tailed Domain Generalization with Sharpness-Guided Gradient Matching
by: Gupta, Sunny, et al.
Published: (2025)
by: Gupta, Sunny, et al.
Published: (2025)
Robust and Clinically Reliable EEG Biomarkers: A Cross Population Framework for Generalizable Parkinson's Disease Detection
by: Rasmussen, Nicholas R., et al.
Published: (2026)
by: Rasmussen, Nicholas R., et al.
Published: (2026)
A Human-Machine Collaboration Framework for the Development of Schemas
by: Isaak, Nicos
Published: (2024)
by: Isaak, Nicos
Published: (2024)
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
by: Ma, Chong, et al.
Published: (2024)
by: Ma, Chong, et al.
Published: (2024)
Similar Items
-
IMUVIE: Pickup Timeline Action Localization via Motion Movies
by: Clapham, John, et al.
Published: (2024) -
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
by: Sutton, Matthew, et al.
Published: (2026) -
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review
by: Rombach, Alexander Michael, et al.
Published: (2024) -
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
by: Banthia, Saumya, et al.
Published: (2020) -
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
by: Chien, Sheng-You, et al.
Published: (2026)