Gespeichert in:
| Hauptverfasser: | Uthayasooriyar, Benno, Ly, Antoine, Vermet, Franck, Corro, Caio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.08606 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training LayoutLM from Scratch for Efficient Named-Entity Recognition in the Insurance Domain
von: Uthayasooriyar, Benno, et al.
Veröffentlicht: (2024)
von: Uthayasooriyar, Benno, et al.
Veröffentlicht: (2024)
Few-Shot Domain Adaptation for Named-Entity Recognition via Joint Constrained k-Means and Subspace Selection
von: Hammal, Ayoub, et al.
Veröffentlicht: (2024)
von: Hammal, Ayoub, et al.
Veröffentlicht: (2024)
A fast and sound tagging method for discontinuous named-entity recognition
von: Corro, Caio
Veröffentlicht: (2024)
von: Corro, Caio
Veröffentlicht: (2024)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
von: Chaffin, Antoine, et al.
Veröffentlicht: (2026)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2026)
A Study on Building Efficient Zero-Shot Relation Extraction Models
von: Thomas, Hugo, et al.
Veröffentlicht: (2026)
von: Thomas, Hugo, et al.
Veröffentlicht: (2026)
DocMamba: Efficient Document Pre-training with State Space Model
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
Suffix-Constrained Greedy Search Algorithms for Causal Language Models
von: Hammal, Ayoub, et al.
Veröffentlicht: (2026)
von: Hammal, Ayoub, et al.
Veröffentlicht: (2026)
FaBERT: Pre-training BERT on Persian Blogs
von: Masumi, Mostafa, et al.
Veröffentlicht: (2024)
von: Masumi, Mostafa, et al.
Veröffentlicht: (2024)
Kad: A Framework for Proxy-based Test-time Alignment with Knapsack Approximation Deferral
von: Hammal, Ayoub, et al.
Veröffentlicht: (2025)
von: Hammal, Ayoub, et al.
Veröffentlicht: (2025)
On the Rejection Criterion for Proxy-based Test-time Alignment
von: Hammal, Ayoub, et al.
Veröffentlicht: (2026)
von: Hammal, Ayoub, et al.
Veröffentlicht: (2026)
Sparse Logistic Regression with High-order Features for Automatic Grammar Rule Extraction from Treebanks
von: Herrera, Santiago, et al.
Veröffentlicht: (2024)
von: Herrera, Santiago, et al.
Veröffentlicht: (2024)
Pre-training technique to localize medical BERT and enhance biomedical BERT
von: Wada, Shoya, et al.
Veröffentlicht: (2020)
von: Wada, Shoya, et al.
Veröffentlicht: (2020)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
Understanding Data Temporality Impact on Large Language Models Pre-training
von: Pilchen, Hippolyte, et al.
Veröffentlicht: (2026)
von: Pilchen, Hippolyte, et al.
Veröffentlicht: (2026)
Bag of Lies: Robustness in Continuous Pre-training BERT
von: Gevers, Ine, et al.
Veröffentlicht: (2024)
von: Gevers, Ine, et al.
Veröffentlicht: (2024)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
von: Suri, Manan, et al.
Veröffentlicht: (2024)
von: Suri, Manan, et al.
Veröffentlicht: (2024)
Bregman Conditional Random Fields: Sequence Labeling with Parallelizable Inference Algorithms
von: Corro, Caio, et al.
Veröffentlicht: (2025)
von: Corro, Caio, et al.
Veröffentlicht: (2025)
Unveiling the Deficiencies of Pre-trained Text-and-Layout Models in Real-world Visually-rich Document Information Extraction
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
von: Li, Jierui, et al.
Veröffentlicht: (2023)
von: Li, Jierui, et al.
Veröffentlicht: (2023)
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
von: Vicentino, Caio
Veröffentlicht: (2026)
von: Vicentino, Caio
Veröffentlicht: (2026)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
von: Yadav, Sumit, et al.
Veröffentlicht: (2025)
von: Yadav, Sumit, et al.
Veröffentlicht: (2025)
DocReward: A Document Reward Model for Structuring and Stylizing
von: Liu, Junpeng, et al.
Veröffentlicht: (2025)
von: Liu, Junpeng, et al.
Veröffentlicht: (2025)
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
von: Zhang, Ying, et al.
Veröffentlicht: (2024)
von: Zhang, Ying, et al.
Veröffentlicht: (2024)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
von: Kim, Sungnyun, et al.
Veröffentlicht: (2024)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2024)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
von: Liang, Chaoqi, et al.
Veröffentlicht: (2023)
von: Liang, Chaoqi, et al.
Veröffentlicht: (2023)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
von: Nandy, Abhilash, et al.
Veröffentlicht: (2023)
von: Nandy, Abhilash, et al.
Veröffentlicht: (2023)
KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates
von: Li, Yudong, et al.
Veröffentlicht: (2026)
von: Li, Yudong, et al.
Veröffentlicht: (2026)
LipidBERT: A Lipid Language Model Pre-trained on METiS de novo Lipid Library
von: Yu, Tianhao, et al.
Veröffentlicht: (2024)
von: Yu, Tianhao, et al.
Veröffentlicht: (2024)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
Large Language Models Understand Layout
von: Li, Weiming, et al.
Veröffentlicht: (2024)
von: Li, Weiming, et al.
Veröffentlicht: (2024)
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
von: Touchent, Rian, et al.
Veröffentlicht: (2023)
von: Touchent, Rian, et al.
Veröffentlicht: (2023)
TempCharBERT: Keystroke Dynamics for Continuous Access Control Based on Pre-trained Language Models
von: Simão, Matheus, et al.
Veröffentlicht: (2024)
von: Simão, Matheus, et al.
Veröffentlicht: (2024)
DomURLs_BERT: Pre-trained BERT-based Model for Malicious Domains and URLs Detection and Classification
von: Mahdaouy, Abdelkader El, et al.
Veröffentlicht: (2024)
von: Mahdaouy, Abdelkader El, et al.
Veröffentlicht: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Training LayoutLM from Scratch for Efficient Named-Entity Recognition in the Insurance Domain
von: Uthayasooriyar, Benno, et al.
Veröffentlicht: (2024) -
Few-Shot Domain Adaptation for Named-Entity Recognition via Joint Constrained k-Means and Subspace Selection
von: Hammal, Ayoub, et al.
Veröffentlicht: (2024) -
A fast and sound tagging method for discontinuous named-entity recognition
von: Corro, Caio
Veröffentlicht: (2024) -
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024) -
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
von: Chaffin, Antoine, et al.
Veröffentlicht: (2026)