Less is more: Not all samples are effective for evaluation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Song, Wentang, Li, Jinqiang, Huang, Kele, Lin, Junhui, Wu, Shengxiang, Xie, Zhongshi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Towards General Deepfake Detection with Dynamic Curriculum
par: Song, Wentang, et autres
Publié: (2024)
par: Song, Wentang, et autres
Publié: (2024)
LIMO: Less is More for Reasoning
par: Ye, Yixin, et autres
Publié: (2025)
par: Ye, Yixin, et autres
Publié: (2025)
Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generation
par: Gao, Shengxiang, et autres
Publié: (2025)
par: Gao, Shengxiang, et autres
Publié: (2025)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
par: Hu, Junhao, et autres
Publié: (2026)
par: Hu, Junhao, et autres
Publié: (2026)
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
par: Banerjee, Mohor, et autres
Publié: (2025)
par: Banerjee, Mohor, et autres
Publié: (2025)
Less is More: Resource-Efficient Low-Rank Adaptation
par: Tian, Chunlin, et autres
Publié: (2025)
par: Tian, Chunlin, et autres
Publié: (2025)
A Survey on Large Language Models from General Purpose to Medical Applications: Datasets, Methodologies, and Evaluations
par: Wang, Jinqiang, et autres
Publié: (2024)
par: Wang, Jinqiang, et autres
Publié: (2024)
An empirical evaluation of using ChatGPT to summarize disputes for recommending similar labor and employment cases in Chinese
par: Wu, Po-Hsien, et autres
Publié: (2024)
par: Wu, Po-Hsien, et autres
Publié: (2024)
CTourLLM: Enhancing LLMs with Chinese Tourism Knowledge
par: Wei, Qikai, et autres
Publié: (2024)
par: Wei, Qikai, et autres
Publié: (2024)
Acting Less is Reasoning More! Teaching Model to Act Efficiently
par: Wang, Hongru, et autres
Publié: (2025)
par: Wang, Hongru, et autres
Publié: (2025)
Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset
par: Zhu, Jie, et autres
Publié: (2024)
par: Zhu, Jie, et autres
Publié: (2024)
LoRA Learns Less and Forgets Less
par: Biderman, Dan, et autres
Publié: (2024)
par: Biderman, Dan, et autres
Publié: (2024)
Less is More for RAG: Information Gain Pruning for Generator-Aligned Reranking and Evidence Selection
par: Song, Zhipeng, et autres
Publié: (2026)
par: Song, Zhipeng, et autres
Publié: (2026)
Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation
par: Zheng, Yilun, et autres
Publié: (2025)
par: Zheng, Yilun, et autres
Publié: (2025)
Less Is More: Elevating RAG via Performance-Driven Context Compression
par: Cui, Ziqiang, et autres
Publié: (2025)
par: Cui, Ziqiang, et autres
Publié: (2025)
Comparing Approaches to Automatic Summarization in Less-Resourced Languages
par: Palen-Michel, Chester, et autres
Publié: (2025)
par: Palen-Michel, Chester, et autres
Publié: (2025)
LIMR: Less is More for RL Scaling
par: Li, Xuefeng, et autres
Publié: (2025)
par: Li, Xuefeng, et autres
Publié: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
par: Zhang, Yuhao, et autres
Publié: (2025)
par: Zhang, Yuhao, et autres
Publié: (2025)
Automated evaluation of LLMs for effective machine translation of Mandarin Chinese to English
par: Zhang, Yue, et autres
Publié: (2026)
par: Zhang, Yue, et autres
Publié: (2026)
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation
par: Seraj, Javad, et autres
Publié: (2024)
par: Seraj, Javad, et autres
Publié: (2024)
M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset
par: Zhu, Jie, et autres
Publié: (2025)
par: Zhu, Jie, et autres
Publié: (2025)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
par: Yang, Zhen, et autres
Publié: (2025)
par: Yang, Zhen, et autres
Publié: (2025)
Soft Prompt Tuning for Cross-Lingual Transfer: When Less is More
par: Philippy, Fred, et autres
Publié: (2024)
par: Philippy, Fred, et autres
Publié: (2024)
Chain or tree? Re-evaluating complex reasoning from the perspective of a matrix of thought
par: Tang, Fengxiao, et autres
Publié: (2025)
par: Tang, Fengxiao, et autres
Publié: (2025)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
par: Guo, Yiju, et autres
Publié: (2026)
par: Guo, Yiju, et autres
Publié: (2026)
Societal AI Research Has Become Less Interdisciplinary
par: Markus, Dror Kris, et autres
Publié: (2025)
par: Markus, Dror Kris, et autres
Publié: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
par: Wu, Yuyang, et autres
Publié: (2025)
par: Wu, Yuyang, et autres
Publié: (2025)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
par: Yang, Lijie, et autres
Publié: (2025)
par: Yang, Lijie, et autres
Publié: (2025)
Doing More with Less: Data Augmentation for Sudanese Dialect Automatic Speech Recognition
par: Mansour, Ayman
Publié: (2026)
par: Mansour, Ayman
Publié: (2026)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
par: He, Junhui, et autres
Publié: (2024)
par: He, Junhui, et autres
Publié: (2024)
Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
par: Wu, Suhang, et autres
Publié: (2025)
par: Wu, Suhang, et autres
Publié: (2025)
Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
par: Liu, Lingyuan, et autres
Publié: (2025)
par: Liu, Lingyuan, et autres
Publié: (2025)
Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models
par: De Bellis, Alessandro, et autres
Publié: (2025)
par: De Bellis, Alessandro, et autres
Publié: (2025)
High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
par: Franzmeyer, Tim, et autres
Publié: (2025)
par: Franzmeyer, Tim, et autres
Publié: (2025)
Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
par: Nguyen, Vy, et autres
Publié: (2025)
par: Nguyen, Vy, et autres
Publié: (2025)
FairytaleQA Translated: Enabling Educational Question and Answer Generation in Less-Resourced Languages
par: Leite, Bernardo, et autres
Publié: (2024)
par: Leite, Bernardo, et autres
Publié: (2024)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
par: Nahin, Shahriar Kabir, et autres
Publié: (2025)
par: Nahin, Shahriar Kabir, et autres
Publié: (2025)
Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency
par: Huang, Rapheal, et autres
Publié: (2025)
par: Huang, Rapheal, et autres
Publié: (2025)
Reliable and diverse evaluation of LLM medical knowledge mastery
par: Zhou, Yuxuan, et autres
Publié: (2024)
par: Zhou, Yuxuan, et autres
Publié: (2024)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
par: Ruppik, Benjamin Matthias, et autres
Publié: (2025)
par: Ruppik, Benjamin Matthias, et autres
Publié: (2025)
Documents similaires
-
Towards General Deepfake Detection with Dynamic Curriculum
par: Song, Wentang, et autres
Publié: (2024) -
LIMO: Less is More for Reasoning
par: Ye, Yixin, et autres
Publié: (2025) -
Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generation
par: Gao, Shengxiang, et autres
Publié: (2025) -
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
par: Hu, Junhao, et autres
Publié: (2026) -
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
par: Banerjee, Mohor, et autres
Publié: (2025)