Information Extraction from Heterogeneous Documents without Ground Truth Labels using Synthetic Label Generation and Knowledge Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Bhattacharyya, Aniket, Tripathi, Anurag |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
por: Ruiz, Tomas, et al.
Publicado: (2026)
por: Ruiz, Tomas, et al.
Publicado: (2026)
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
por: Sun, Lin, et al.
Publicado: (2025)
por: Sun, Lin, et al.
Publicado: (2025)
Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments
por: Bhattacharyya, Aniket, et al.
Publicado: (2025)
por: Bhattacharyya, Aniket, et al.
Publicado: (2025)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
por: Chew, Robert, et al.
Publicado: (2026)
por: Chew, Robert, et al.
Publicado: (2026)
Crossing Domains without Labels: Distant Supervision for Term Extraction
por: Senger, Elena, et al.
Publicado: (2025)
por: Senger, Elena, et al.
Publicado: (2025)
KDH-MLTC: Knowledge Distillation for Healthcare Multi-Label Text Classification
por: Sakai, Hajar, et al.
Publicado: (2025)
por: Sakai, Hajar, et al.
Publicado: (2025)
UniVIE: A Unified Label Space Approach to Visual Information Extraction from Form-like Documents
por: Hu, Kai, et al.
Publicado: (2024)
por: Hu, Kai, et al.
Publicado: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
por: Pangakis, Nicholas, et al.
Publicado: (2024)
por: Pangakis, Nicholas, et al.
Publicado: (2024)
An Effective Incorporating Heterogeneous Knowledge Curriculum Learning for Sequence Labeling
por: Tang, Xuemei, et al.
Publicado: (2024)
por: Tang, Xuemei, et al.
Publicado: (2024)
Auxiliary Knowledge-Induced Learning for Automatic Multi-Label Medical Document Classification
por: Wang, Xindi, et al.
Publicado: (2024)
por: Wang, Xindi, et al.
Publicado: (2024)
Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations
por: Lai, Peichao, et al.
Publicado: (2025)
por: Lai, Peichao, et al.
Publicado: (2025)
Ranking Large Language Models without Ground Truth
por: Dhurandhar, Amit, et al.
Publicado: (2024)
por: Dhurandhar, Amit, et al.
Publicado: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
por: Waheed, Abdul, et al.
Publicado: (2024)
por: Waheed, Abdul, et al.
Publicado: (2024)
Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction
por: Yang, Lu, et al.
Publicado: (2025)
por: Yang, Lu, et al.
Publicado: (2025)
Ground Truth Generation for Multilingual Historical NLP using LLMs
por: Gladstone, Clovis, et al.
Publicado: (2025)
por: Gladstone, Clovis, et al.
Publicado: (2025)
Authors Should Label Their Own Documents
por: Ma, Marcus, et al.
Publicado: (2025)
por: Ma, Marcus, et al.
Publicado: (2025)
Generating the Ground Truth: Synthetic Data for Soft Label and Label Noise Research
por: de Vries, Sjoerd, et al.
Publicado: (2023)
por: de Vries, Sjoerd, et al.
Publicado: (2023)
Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
por: Fränken, Jan-Philipp, et al.
Publicado: (2024)
por: Fränken, Jan-Philipp, et al.
Publicado: (2024)
Free Process Rewards without Process Labels
por: Yuan, Lifan, et al.
Publicado: (2024)
por: Yuan, Lifan, et al.
Publicado: (2024)
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
ADVOSYNTH: A Synthetic Multi-Advocate Dataset for Speaker Identification in Courtroom Scenarios
por: Deroy, Aniket
Publicado: (2026)
por: Deroy, Aniket
Publicado: (2026)
A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete Labeling
por: Wang, Ye, et al.
Publicado: (2023)
por: Wang, Ye, et al.
Publicado: (2023)
Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling
por: Hwang, Seonjeong, et al.
Publicado: (2024)
por: Hwang, Seonjeong, et al.
Publicado: (2024)
Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers
por: Xin, Rihui, et al.
Publicado: (2025)
por: Xin, Rihui, et al.
Publicado: (2025)
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
por: Lee, Young-Suk, et al.
Publicado: (2024)
por: Lee, Young-Suk, et al.
Publicado: (2024)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
por: Krumdick, Michael, et al.
Publicado: (2025)
por: Krumdick, Michael, et al.
Publicado: (2025)
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
por: Choubey, Prafulla Kumar, et al.
Publicado: (2024)
por: Choubey, Prafulla Kumar, et al.
Publicado: (2024)
Recon, Answer, Verify: Agents in Search of Truth
por: Shukla, Satyam, et al.
Publicado: (2025)
por: Shukla, Satyam, et al.
Publicado: (2025)
Differentially Private Knowledge Distillation via Synthetic Text Generation
por: Flemings, James, et al.
Publicado: (2024)
por: Flemings, James, et al.
Publicado: (2024)
Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection
por: Kazemi, Arefeh, et al.
Publicado: (2025)
por: Kazemi, Arefeh, et al.
Publicado: (2025)
DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's Disease
por: Mo, Tingyu, et al.
Publicado: (2025)
por: Mo, Tingyu, et al.
Publicado: (2025)
Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data
por: Li, Rumeng, et al.
Publicado: (2023)
por: Li, Rumeng, et al.
Publicado: (2023)
PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling
por: Lin, Ying-Jia, et al.
Publicado: (2026)
por: Lin, Ying-Jia, et al.
Publicado: (2026)
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation
por: Huang, Chenyang, et al.
Publicado: (2025)
por: Huang, Chenyang, et al.
Publicado: (2025)
Schema-Driven Information Extraction from Heterogeneous Tables
por: Bai, Fan, et al.
Publicado: (2023)
por: Bai, Fan, et al.
Publicado: (2023)
GuideX: Guided Synthetic Data Generation for Zero-Shot Information Extraction
por: De La Fuente, Neil, et al.
Publicado: (2025)
por: De La Fuente, Neil, et al.
Publicado: (2025)
Editing Arbitrary Propositions in LLMs without Subject Labels
por: Feigenbaum, Itai, et al.
Publicado: (2024)
por: Feigenbaum, Itai, et al.
Publicado: (2024)
Multi-layer Sequence Labeling-based Joint Biomedical Event Extraction
por: Chen, Gongchi, et al.
Publicado: (2024)
por: Chen, Gongchi, et al.
Publicado: (2024)
Can We Reliably Rank Model Performance across Domains without Labeled Data?
por: Rammouz, Veronica, et al.
Publicado: (2025)
por: Rammouz, Veronica, et al.
Publicado: (2025)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
por: Gautam, Sushant, et al.
Publicado: (2026)
por: Gautam, Sushant, et al.
Publicado: (2026)
Ejemplares similares
-
Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
por: Ruiz, Tomas, et al.
Publicado: (2026) -
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
por: Sun, Lin, et al.
Publicado: (2025) -
Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments
por: Bhattacharyya, Aniket, et al.
Publicado: (2025) -
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
por: Chew, Robert, et al.
Publicado: (2026) -
Crossing Domains without Labels: Distant Supervision for Term Extraction
por: Senger, Elena, et al.
Publicado: (2025)