We Need to Measure Data Diversity in NLP -- Better and Broader
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Dong, Ploeger, Esther |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What is "Typological Diversity" in NLP?
von: Ploeger, Esther, et al.
Veröffentlicht: (2024)
von: Ploeger, Esther, et al.
Veröffentlicht: (2024)
Towards Supporting Legal Argumentation with NLP: Is More Data Really All You Need?
von: Santosh, T. Y. S. S, et al.
Veröffentlicht: (2024)
von: Santosh, T. Y. S. S, et al.
Veröffentlicht: (2024)
NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty
von: Zotos, Leonidas, et al.
Veröffentlicht: (2025)
von: Zotos, Leonidas, et al.
Veröffentlicht: (2025)
Facilitating Opinion Diversity through Hybrid NLP Approaches
von: van der Meer, Michiel
Veröffentlicht: (2024)
von: van der Meer, Michiel
Veröffentlicht: (2024)
Undesirable Biases in NLP: Addressing Challenges of Measurement
von: van der Wal, Oskar, et al.
Veröffentlicht: (2022)
von: van der Wal, Oskar, et al.
Veröffentlicht: (2022)
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
Evaluation Metrics for Text Data Augmentation in NLP
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
Could We Have Had Better Multilingual LLMs If English Was Not the Central Language?
von: Diandaru, Ryandito, et al.
Veröffentlicht: (2024)
von: Diandaru, Ryandito, et al.
Veröffentlicht: (2024)
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
von: Goldman, Omer, et al.
Veröffentlicht: (2024)
von: Goldman, Omer, et al.
Veröffentlicht: (2024)
Is there really a Citation Age Bias in NLP?
von: Nguyen, Hoa, et al.
Veröffentlicht: (2024)
von: Nguyen, Hoa, et al.
Veröffentlicht: (2024)
Why Low-Resource NLP Needs More Than Cross-Lingual Transfer: Lessons Learned from Luxembourgish
von: Philippy, Fred, et al.
Veröffentlicht: (2026)
von: Philippy, Fred, et al.
Veröffentlicht: (2026)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
von: Chimoto, Everlyn Asiko, et al.
Veröffentlicht: (2026)
von: Chimoto, Everlyn Asiko, et al.
Veröffentlicht: (2026)
How Far Are We From AGI: Are LLMs All We Need?
von: Feng, Tao, et al.
Veröffentlicht: (2024)
von: Feng, Tao, et al.
Veröffentlicht: (2024)
Measuring Diversity in Synthetic Datasets
von: Zhu, Yuchang, et al.
Veröffentlicht: (2025)
von: Zhu, Yuchang, et al.
Veröffentlicht: (2025)
Logits are All We Need to Adapt Closed Models
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2025)
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2025)
NLP Security and Ethics, in the Wild
von: Lent, Heather, et al.
Veröffentlicht: (2025)
von: Lent, Heather, et al.
Veröffentlicht: (2025)
Know Your Needs Better: Towards Structured Understanding of Marketer Demands with Analogical Reasoning Augmented LLMs
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
Not All Layers Need Tuning: Selective Layer Restoration Recovers Diversity
von: Zhang, Bowen, et al.
Veröffentlicht: (2026)
von: Zhang, Bowen, et al.
Veröffentlicht: (2026)
LLM-Augmented Symptom Analysis for Cardiovascular Disease Risk Prediction: A Clinical NLP
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
An Audit on the Perspectives and Challenges of Hallucinations in NLP
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2024)
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2024)
State of NLP in Kenya: A Survey
von: Amol, Cynthia Jayne, et al.
Veröffentlicht: (2024)
von: Amol, Cynthia Jayne, et al.
Veröffentlicht: (2024)
README: Bridging Medical Jargon and Lay Understanding for Patient Education through Data-Centric NLP
von: Yao, Zonghai, et al.
Veröffentlicht: (2023)
von: Yao, Zonghai, et al.
Veröffentlicht: (2023)
Named Entity Recognition for Payment Data Using NLP
von: Nayak, Srikumar
Veröffentlicht: (2026)
von: Nayak, Srikumar
Veröffentlicht: (2026)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2026)
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2026)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
Abstraction-of-Thought Makes Language Models Better Reasoners
von: Hong, Ruixin, et al.
Veröffentlicht: (2024)
von: Hong, Ruixin, et al.
Veröffentlicht: (2024)
Select, Label, Evaluate: Active Testing in NLP
von: Purificato, Antonio, et al.
Veröffentlicht: (2026)
von: Purificato, Antonio, et al.
Veröffentlicht: (2026)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
von: Manna, Supriya, et al.
Veröffentlicht: (2024)
von: Manna, Supriya, et al.
Veröffentlicht: (2024)
Speaking of Language: Reflections on Metalanguage Research in NLP
von: Schneider, Nathan, et al.
Veröffentlicht: (2026)
von: Schneider, Nathan, et al.
Veröffentlicht: (2026)
Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models
von: Xiang, Bajian, et al.
Veröffentlicht: (2026)
von: Xiang, Bajian, et al.
Veröffentlicht: (2026)
Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts
von: Lee, Sang-Woo, et al.
Veröffentlicht: (2025)
von: Lee, Sang-Woo, et al.
Veröffentlicht: (2025)
Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
Practising responsibility: Ethics in NLP as a hands-on course
von: Nissim, Malvina, et al.
Veröffentlicht: (2025)
von: Nissim, Malvina, et al.
Veröffentlicht: (2025)
Towards Open-Ended Discovery for Low-Resource NLP
von: Dossou, Bonaventure F. P., et al.
Veröffentlicht: (2025)
von: Dossou, Bonaventure F. P., et al.
Veröffentlicht: (2025)
Intertwining CP and NLP: The Generation of Unreasonably Constrained Sentences
von: Bonlarron, Alexandre, et al.
Veröffentlicht: (2024)
von: Bonlarron, Alexandre, et al.
Veröffentlicht: (2024)
Large Language Models Meet NLP: A Survey
von: Qin, Libo, et al.
Veröffentlicht: (2024)
von: Qin, Libo, et al.
Veröffentlicht: (2024)
Advancing NLP Security by Leveraging LLMs as Adversarial Engines
von: Srinivasan, Sudarshan, et al.
Veröffentlicht: (2024)
von: Srinivasan, Sudarshan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
What is "Typological Diversity" in NLP?
von: Ploeger, Esther, et al.
Veröffentlicht: (2024) -
Towards Supporting Legal Argumentation with NLP: Is More Data Really All You Need?
von: Santosh, T. Y. S. S, et al.
Veröffentlicht: (2024) -
NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty
von: Zotos, Leonidas, et al.
Veröffentlicht: (2025) -
Facilitating Opinion Diversity through Hybrid NLP Approaches
von: van der Meer, Michiel
Veröffentlicht: (2024) -
Undesirable Biases in NLP: Addressing Challenges of Measurement
von: van der Wal, Oskar, et al.
Veröffentlicht: (2022)