Data filtering methods for training language models
Fuente:
arXiv
Salvato in:
| Autori principali: | Shevchenko, Egor, Bruches, Elena |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Clinical information extraction for Low-resource languages with Few-shot learning using Pre-trained language models and Prompting
di: Richter-Pechanski, Phillip, et al.
Pubblicazione: (2024)
di: Richter-Pechanski, Phillip, et al.
Pubblicazione: (2024)
Evaluating Named Entity Recognition: A comparative analysis of mono- and multilingual transformer models on a novel Brazilian corporate earnings call transcripts dataset
di: Abilio, Ramon, et al.
Pubblicazione: (2024)
di: Abilio, Ramon, et al.
Pubblicazione: (2024)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
di: Aksoy, Sinan G., et al.
Pubblicazione: (2026)
di: Aksoy, Sinan G., et al.
Pubblicazione: (2026)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
di: Yazdanjue, Navid, et al.
Pubblicazione: (2025)
di: Yazdanjue, Navid, et al.
Pubblicazione: (2025)
Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
di: Faisal, Faizan, et al.
Pubblicazione: (2024)
di: Faisal, Faizan, et al.
Pubblicazione: (2024)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
di: Yankun, Hong, et al.
Pubblicazione: (2025)
di: Yankun, Hong, et al.
Pubblicazione: (2025)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
di: Lai, Kunfeng, et al.
Pubblicazione: (2025)
di: Lai, Kunfeng, et al.
Pubblicazione: (2025)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
di: Wang, Youkang, et al.
Pubblicazione: (2025)
di: Wang, Youkang, et al.
Pubblicazione: (2025)
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?
di: Kovač, Grgur, et al.
Pubblicazione: (2025)
di: Kovač, Grgur, et al.
Pubblicazione: (2025)
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
di: Tian, Jacob-Junqi, et al.
Pubblicazione: (2023)
di: Tian, Jacob-Junqi, et al.
Pubblicazione: (2023)
Large Language Models versus Classical Machine Learning: Performance in COVID-19 Mortality Prediction Using High-Dimensional Tabular Data
di: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Pubblicazione: (2024)
di: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Pubblicazione: (2024)
Parameter-Efficient Transformer Embeddings
di: Ndubuaku, Henry, et al.
Pubblicazione: (2025)
di: Ndubuaku, Henry, et al.
Pubblicazione: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
di: Nogales, Miguel, et al.
Pubblicazione: (2025)
di: Nogales, Miguel, et al.
Pubblicazione: (2025)
ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data
di: Seneque, Gareth, et al.
Pubblicazione: (2026)
di: Seneque, Gareth, et al.
Pubblicazione: (2026)
MultiLegalPile: A 689GB Multilingual Legal Corpus
di: Niklaus, Joel, et al.
Pubblicazione: (2023)
di: Niklaus, Joel, et al.
Pubblicazione: (2023)
LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
di: Niklaus, Joel, et al.
Pubblicazione: (2023)
di: Niklaus, Joel, et al.
Pubblicazione: (2023)
LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
di: Niklaus, Joel, et al.
Pubblicazione: (2024)
di: Niklaus, Joel, et al.
Pubblicazione: (2024)
Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
di: Stern, Ronja, et al.
Pubblicazione: (2023)
di: Stern, Ronja, et al.
Pubblicazione: (2023)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
di: Niklaus, Joel, et al.
Pubblicazione: (2025)
di: Niklaus, Joel, et al.
Pubblicazione: (2025)
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset
di: S, Santosh T. Y. S., et al.
Pubblicazione: (2024)
di: S, Santosh T. Y. S., et al.
Pubblicazione: (2024)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
di: Fan, Yu, et al.
Pubblicazione: (2025)
di: Fan, Yu, et al.
Pubblicazione: (2025)
Predicting the Geolocation of Tweets Using transformer models on Customized Data
di: Lutsai, Kateryna, et al.
Pubblicazione: (2023)
di: Lutsai, Kateryna, et al.
Pubblicazione: (2023)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
di: Platt, Nolan, et al.
Pubblicazione: (2025)
di: Platt, Nolan, et al.
Pubblicazione: (2025)
Surfing the modeling of PoS taggers in low-resource scenarios
di: Ferro, Manuel Vilares, et al.
Pubblicazione: (2024)
di: Ferro, Manuel Vilares, et al.
Pubblicazione: (2024)
The Meta-Prompting Protocol: Orchestrating LLMs via Adversarial Feedback Loops
di: Fu, Fanzhe
Pubblicazione: (2025)
di: Fu, Fanzhe
Pubblicazione: (2025)
KIT-TIP-NLP at MultiPride: Continual Learning with Multilingual Foundation Model
di: HB, Barathi Ganesh, et al.
Pubblicazione: (2026)
di: HB, Barathi Ganesh, et al.
Pubblicazione: (2026)
TRUE: A Trustworthy Unified Explanation Framework for Large Language Model Reasoning
di: Yang, Yujiao
Pubblicazione: (2026)
di: Yang, Yujiao
Pubblicazione: (2026)
Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions
di: Rouzegar, Hamdireza, et al.
Pubblicazione: (2024)
di: Rouzegar, Hamdireza, et al.
Pubblicazione: (2024)
Investigating Distributions of Telecom Adapted Sentence Embeddings for Document Retrieval
di: Roychowdhury, Sujoy, et al.
Pubblicazione: (2024)
di: Roychowdhury, Sujoy, et al.
Pubblicazione: (2024)
CSTRL: Context-Driven Sequential Transfer Learning for Abstractive Radiology Report Summarization
di: Naznin, Mst. Fahmida Sultana, et al.
Pubblicazione: (2025)
di: Naznin, Mst. Fahmida Sultana, et al.
Pubblicazione: (2025)
Lightweight Transformers for Clinical Natural Language Processing
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
Aligning Black-box Language Models with Human Judgments
di: Burg, Gerrit J. J. van den, et al.
Pubblicazione: (2025)
di: Burg, Gerrit J. J. van den, et al.
Pubblicazione: (2025)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
di: Anshumann, et al.
Pubblicazione: (2025)
di: Anshumann, et al.
Pubblicazione: (2025)
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
di: Kim, Seongho, et al.
Pubblicazione: (2024)
di: Kim, Seongho, et al.
Pubblicazione: (2024)
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
di: Günther, Michael, et al.
Pubblicazione: (2023)
di: Günther, Michael, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Clinical information extraction for Low-resource languages with Few-shot learning using Pre-trained language models and Prompting
di: Richter-Pechanski, Phillip, et al.
Pubblicazione: (2024) -
Evaluating Named Entity Recognition: A comparative analysis of mono- and multilingual transformer models on a novel Brazilian corporate earnings call transcripts dataset
di: Abilio, Ramon, et al.
Pubblicazione: (2024) -
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
di: Aksoy, Sinan G., et al.
Pubblicazione: (2026) -
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
di: Yazdanjue, Navid, et al.
Pubblicazione: (2025) -
Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments
di: Zhang, Li, et al.
Pubblicazione: (2025)