Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Deshpande, Vijeta, Ghose, Debasmita, Patterson, John D., Beaty, Roger, Rumshisky, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
by: Deshpande, Vijeta, et al.
Published: (2025)
by: Deshpande, Vijeta, et al.
Published: (2025)
Emergent Abilities in Reduced-Scale Generative Language Models
by: Muckatira, Sherin, et al.
Published: (2024)
by: Muckatira, Sherin, et al.
Published: (2024)
Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
by: Lialin, Vladislav, et al.
Published: (2023)
by: Lialin, Vladislav, et al.
Published: (2023)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
by: Shivagunde, Namrata, et al.
Published: (2026)
by: Shivagunde, Namrata, et al.
Published: (2026)
A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
by: Muckatira, Sherin, et al.
Published: (2026)
by: Muckatira, Sherin, et al.
Published: (2026)
Controlled Diversity: Length-optimized Natural Language Generation
by: Schenke, Diana Marie, et al.
Published: (2025)
by: Schenke, Diana Marie, et al.
Published: (2025)
Enhancing Contrastive Demonstration Selection with Semantic Diversity for Robust In-Context Machine Translation
by: Patterson, Owen, et al.
Published: (2025)
by: Patterson, Owen, et al.
Published: (2025)
Asking a Language Model for Diverse Responses
by: Troshin, Sergey, et al.
Published: (2025)
by: Troshin, Sergey, et al.
Published: (2025)
Analysis of student understanding in short‐answer explanations to concept questions using a human‐centered AI approach
by: Harpreet Auby, et al.
Published: (2025)
by: Harpreet Auby, et al.
Published: (2025)
ControlLM: Crafting Diverse Personalities for Language Models
by: Weng, Yixuan, et al.
Published: (2024)
by: Weng, Yixuan, et al.
Published: (2024)
LocalTweets to LocalHealth: A Mental Health Surveillance Framework Based on Twitter Data
by: Deshpande, Vijeta, et al.
Published: (2024)
by: Deshpande, Vijeta, et al.
Published: (2024)
How do Humans and Language Models Reason About Creativity? A Comparative Analysis
by: Laverghetta Jr., Antonio, et al.
Published: (2025)
by: Laverghetta Jr., Antonio, et al.
Published: (2025)
Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models
by: Peeperkorn, Max, et al.
Published: (2025)
by: Peeperkorn, Max, et al.
Published: (2025)
Deconstructing In-Context Learning: Understanding Prompts via Corruption
by: Shivagunde, Namrata, et al.
Published: (2024)
by: Shivagunde, Namrata, et al.
Published: (2024)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
by: Yang, Xianjun, et al.
Published: (2025)
by: Yang, Xianjun, et al.
Published: (2025)
Improving Diversity of Commonsense Generation by Large Language Models via In-Context Learning
by: Zhang, Tianhui, et al.
Published: (2024)
by: Zhang, Tianhui, et al.
Published: (2024)
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Post-training Large Language Models for Diverse High-Quality Responses
by: Chen, Yilei, et al.
Published: (2025)
by: Chen, Yilei, et al.
Published: (2025)
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
by: Verine, Alexandre, et al.
Published: (2025)
by: Verine, Alexandre, et al.
Published: (2025)
Prompt Perturbation Consistency Learning for Robust Language Models
by: Qiang, Yao, et al.
Published: (2024)
by: Qiang, Yao, et al.
Published: (2024)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
by: Negoita, Vlad, et al.
Published: (2025)
by: Negoita, Vlad, et al.
Published: (2025)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
by: Fröhling, Leon, et al.
Published: (2024)
by: Fröhling, Leon, et al.
Published: (2024)
Modeling Data Diversity for Joint Instance and Verbalizer Selection in Cold-Start Scenarios
by: Chakraborty, Mohna, et al.
Published: (2025)
by: Chakraborty, Mohna, et al.
Published: (2025)
Diversity Helps Jailbreak Large Language Models
by: Zhao, Weiliang, et al.
Published: (2024)
by: Zhao, Weiliang, et al.
Published: (2024)
Benchmarking Linguistic Diversity of Large Language Models
by: Guo, Yanzhu, et al.
Published: (2024)
by: Guo, Yanzhu, et al.
Published: (2024)
Diversity-oriented Data Augmentation with Large Language Models
by: Wang, Zaitian, et al.
Published: (2025)
by: Wang, Zaitian, et al.
Published: (2025)
Improving Linguistic Diversity of Large Language Models with Possibility Exploration Fine-Tuning
by: Mai, Long, et al.
Published: (2024)
by: Mai, Long, et al.
Published: (2024)
Improving Bilingual Capabilities of Language Models to Support Diverse Linguistic Practices in Education
by: Syamkumar, Anand, et al.
Published: (2024)
by: Syamkumar, Anand, et al.
Published: (2024)
Considering Length Diversity in Retrieval-Augmented Summarization
by: Juseon-Do, et al.
Published: (2025)
by: Juseon-Do, et al.
Published: (2025)
IFDID: Information Filter upon Diversity-Improved Decoding for Diversity-Faithfulness Tradeoff in NLG
by: Meng, Han, et al.
Published: (2022)
by: Meng, Han, et al.
Published: (2022)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
by: Yu, Simon, et al.
Published: (2024)
by: Yu, Simon, et al.
Published: (2024)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
by: Liu, Fengze, et al.
Published: (2025)
by: Liu, Fengze, et al.
Published: (2025)
NarrativeTime: Dense Temporal Annotation on a Timeline
by: Rogers, Anna, et al.
Published: (2019)
by: Rogers, Anna, et al.
Published: (2019)
A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters
by: Desai, Shasvat, et al.
Published: (2025)
by: Desai, Shasvat, et al.
Published: (2025)
Harmonizing Diverse Models: A Layer-wise Merging Strategy for Consistent Generation
by: Peng, Xujun, et al.
Published: (2025)
by: Peng, Xujun, et al.
Published: (2025)
Embedding-Driven Diversity Sampling to Improve Few-Shot Synthetic Data Generation
by: Lopez, Ivan, et al.
Published: (2025)
by: Lopez, Ivan, et al.
Published: (2025)
Similar Items
-
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
by: Deshpande, Vijeta, et al.
Published: (2026) -
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
by: Deshpande, Vijeta, et al.
Published: (2026) -
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
by: Deshpande, Vijeta, et al.
Published: (2025) -
Emergent Abilities in Reduced-Scale Generative Language Models
by: Muckatira, Sherin, et al.
Published: (2024) -
Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
by: Lialin, Vladislav, et al.
Published: (2023)