Data-Constrained Synthesis of Training Data for De-Identification
Fuente:
arXiv
Salvato in:
| Autori principali: | Vakili, Thomas, Henriksson, Aron, Dalianis, Hercules |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Impact of Steering Large Language Models with Persona Vectors in Educational Applications
di: Wu, Yongchao, et al.
Pubblicazione: (2026)
di: Wu, Yongchao, et al.
Pubblicazione: (2026)
Efficient Text Classification with Conformal In-Context Learning
di: Pantelidis, Ippokratis, et al.
Pubblicazione: (2025)
di: Pantelidis, Ippokratis, et al.
Pubblicazione: (2025)
D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training
di: Xu, Yuanjian, et al.
Pubblicazione: (2026)
di: Xu, Yuanjian, et al.
Pubblicazione: (2026)
Rethinking Data Synthesis: A Teacher Model Training Recipe with Interpretation
di: Chen, Yifang, et al.
Pubblicazione: (2024)
di: Chen, Yifang, et al.
Pubblicazione: (2024)
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages
di: Chen, Zui, et al.
Pubblicazione: (2025)
di: Chen, Zui, et al.
Pubblicazione: (2025)
Implementing a Nordic-Baltic Federated Health Data Network: a case report
di: Chomutare, Taridzo, et al.
Pubblicazione: (2024)
di: Chomutare, Taridzo, et al.
Pubblicazione: (2024)
Scaling Data-Constrained Language Models
di: Muennighoff, Niklas, et al.
Pubblicazione: (2023)
di: Muennighoff, Niklas, et al.
Pubblicazione: (2023)
Scaling Parameter-Constrained Language Models with Quality Data
di: Chang, Ernie, et al.
Pubblicazione: (2024)
di: Chang, Ernie, et al.
Pubblicazione: (2024)
FLUX: Data Worth Training On
di: Gowtham, et al.
Pubblicazione: (2026)
di: Gowtham, et al.
Pubblicazione: (2026)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
di: Li, Shujie, et al.
Pubblicazione: (2024)
di: Li, Shujie, et al.
Pubblicazione: (2024)
Clinical Text Mining
di: Dalianis, Hercules
Pubblicazione: (2018)
di: Dalianis, Hercules
Pubblicazione: (2018)
Compute-Constrained Data Selection
di: Yin, Junjie Oscar, et al.
Pubblicazione: (2024)
di: Yin, Junjie Oscar, et al.
Pubblicazione: (2024)
JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models
di: Zhou, Kun, et al.
Pubblicazione: (2024)
di: Zhou, Kun, et al.
Pubblicazione: (2024)
Open Data Synthesis For Deep Research
di: Xia, Ziyi, et al.
Pubblicazione: (2025)
di: Xia, Ziyi, et al.
Pubblicazione: (2025)
Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
di: Roque, Matthew Theodore, et al.
Pubblicazione: (2025)
di: Roque, Matthew Theodore, et al.
Pubblicazione: (2025)
Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications
di: Kanjirangat, Vani, et al.
Pubblicazione: (2025)
di: Kanjirangat, Vani, et al.
Pubblicazione: (2025)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
di: Gao, Xin, et al.
Pubblicazione: (2025)
di: Gao, Xin, et al.
Pubblicazione: (2025)
Natural Language Processing for Electronic Health Records in Scandinavian Languages: Norwegian, Swedish, and Danish
di: Woldaregay, Ashenafi Zebene, et al.
Pubblicazione: (2025)
di: Woldaregay, Ashenafi Zebene, et al.
Pubblicazione: (2025)
Beyond Public Access in LLM Pre-Training Data
di: Rosenblat, Sruly, et al.
Pubblicazione: (2025)
di: Rosenblat, Sruly, et al.
Pubblicazione: (2025)
Unlearning Traces the Influential Training Data of Language Models
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
Balanced Data Sampling for Language Model Training with Clustering
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
SlimPajama-DC: Understanding Data Combinations for LLM Training
di: Shen, Zhiqiang, et al.
Pubblicazione: (2023)
di: Shen, Zhiqiang, et al.
Pubblicazione: (2023)
Detecting RLVR Training Data via Structural Convergence of Reasoning
di: Zhang, Hongbo, et al.
Pubblicazione: (2026)
di: Zhang, Hongbo, et al.
Pubblicazione: (2026)
Data Management For Training Large Language Models: A Survey
di: Wang, Zige, et al.
Pubblicazione: (2023)
di: Wang, Zige, et al.
Pubblicazione: (2023)
Enhancing the De-identification of Personally Identifiable Information in Educational Data
di: Ji, Zilyu, et al.
Pubblicazione: (2025)
di: Ji, Zilyu, et al.
Pubblicazione: (2025)
Synthesis by Design: Controlled Data Generation via Structural Guidance
di: Xu, Lei, et al.
Pubblicazione: (2025)
di: Xu, Lei, et al.
Pubblicazione: (2025)
Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation
di: He, Bowei, et al.
Pubblicazione: (2026)
di: He, Bowei, et al.
Pubblicazione: (2026)
Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
di: Imran, Sohaib, et al.
Pubblicazione: (2025)
di: Imran, Sohaib, et al.
Pubblicazione: (2025)
Syntriever: How to Train Your Retriever with Synthetic Data from LLMs
di: Kim, Minsang, et al.
Pubblicazione: (2025)
di: Kim, Minsang, et al.
Pubblicazione: (2025)
Retracing the Past: LLMs Emit Training Data When They Get Lost
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
Translation of Multifaceted Data without Re-Training of Machine Translation Systems
di: Moon, Hyeonseok, et al.
Pubblicazione: (2024)
di: Moon, Hyeonseok, et al.
Pubblicazione: (2024)
Distilling an End-to-End Voice Assistant Without Instruction Training Data
di: Held, William, et al.
Pubblicazione: (2024)
di: Held, William, et al.
Pubblicazione: (2024)
Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation
di: Tang, Shuo, et al.
Pubblicazione: (2024)
di: Tang, Shuo, et al.
Pubblicazione: (2024)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025)
di: Cao, Maosong, et al.
Pubblicazione: (2025)
From Detection to Diagnosis: Advancing Hallucination Analysis with Automated Data Synthesis
di: Liu, Yanyi, et al.
Pubblicazione: (2025)
di: Liu, Yanyi, et al.
Pubblicazione: (2025)
LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
di: Jia, Junlong, et al.
Pubblicazione: (2025)
di: Jia, Junlong, et al.
Pubblicazione: (2025)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
di: Huang, Yiming, et al.
Pubblicazione: (2024)
di: Huang, Yiming, et al.
Pubblicazione: (2024)
ActionStudio: A Lightweight Framework for Data and Training of Large Action Models
di: Zhang, Jianguo, et al.
Pubblicazione: (2025)
di: Zhang, Jianguo, et al.
Pubblicazione: (2025)
TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
di: Perczel, Janos, et al.
Pubblicazione: (2025)
di: Perczel, Janos, et al.
Pubblicazione: (2025)
Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
di: Niimi, Junichiro
Pubblicazione: (2025)
di: Niimi, Junichiro
Pubblicazione: (2025)
Documenti analoghi
-
The Impact of Steering Large Language Models with Persona Vectors in Educational Applications
di: Wu, Yongchao, et al.
Pubblicazione: (2026) -
Efficient Text Classification with Conformal In-Context Learning
di: Pantelidis, Ippokratis, et al.
Pubblicazione: (2025) -
D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training
di: Xu, Yuanjian, et al.
Pubblicazione: (2026) -
Rethinking Data Synthesis: A Teacher Model Training Recipe with Interpretation
di: Chen, Yifang, et al.
Pubblicazione: (2024) -
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages
di: Chen, Zui, et al.
Pubblicazione: (2025)