Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Hanhua, Xiao, Chenghao, Wang, Yang, Liu, Yiqi, Rong, Wenge, Lin, Chenghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023)
by: Lucy, Li, et al.
Published: (2023)
Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
by: Zhang, Rufan, et al.
Published: (2025)
by: Zhang, Rufan, et al.
Published: (2025)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
by: Liu, Yiqi, et al.
Published: (2023)
by: Liu, Yiqi, et al.
Published: (2023)
Effective Distillation of Table-based Reasoning Ability from LLMs
by: Yang, Bohao, et al.
Published: (2023)
by: Yang, Bohao, et al.
Published: (2023)
Beyond One-Size-Fits-All Summarization: Customizing Summaries for Diverse Users
by: Duran, Mehmet Samet, et al.
Published: (2025)
by: Duran, Mehmet Samet, et al.
Published: (2025)
Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models
by: Liu, Shuqi, et al.
Published: (2025)
by: Liu, Shuqi, et al.
Published: (2025)
Theorem Provers: One Size Fits All?
by: Oates, Harrison, et al.
Published: (2025)
by: Oates, Harrison, et al.
Published: (2025)
The Achilles' Heel of Angular Margins: A Chebyshev Polynomial Fix for Speaker Verification
by: Wang, Yang, et al.
Published: (2026)
by: Wang, Yang, et al.
Published: (2026)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
by: Tang, Tianyi, et al.
Published: (2023)
by: Tang, Tianyi, et al.
Published: (2023)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
by: Zhang, Jianfei, et al.
Published: (2025)
by: Zhang, Jianfei, et al.
Published: (2025)
No One Fits All: From Fixed Prompting to Learned Routing in Multilingual LLMs
by: Wu, Wei-Chi, et al.
Published: (2026)
by: Wu, Wei-Chi, et al.
Published: (2026)
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
by: Yu, Jianxiang, et al.
Published: (2026)
by: Yu, Jianxiang, et al.
Published: (2026)
One Size doesn't Fit All: A Personalized Conversational Tutoring Agent for Mathematics Instruction
by: Liu, Ben, et al.
Published: (2025)
by: Liu, Ben, et al.
Published: (2025)
Beyond One-Size-Fits-All: Multi-Domain, Multi-Task Framework for Embedding Model Selection
by: Khetan, Vivek
Published: (2024)
by: Khetan, Vivek
Published: (2024)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026)
by: James, Joseph, et al.
Published: (2026)
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights
by: James, Joseph, et al.
Published: (2024)
by: James, Joseph, et al.
Published: (2024)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
No One Size Fits All: QueryBandits for Hallucination Mitigation
by: Cho, Nicole, et al.
Published: (2026)
by: Cho, Nicole, et al.
Published: (2026)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
by: Lu, Liming, et al.
Published: (2026)
by: Lu, Liming, et al.
Published: (2026)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
by: Luo, Yingfeng, et al.
Published: (2025)
by: Luo, Yingfeng, et al.
Published: (2025)
PFME: A Modular Approach for Fine-grained Hallucination Detection and Editing of Large Language Models
by: Deng, Kunquan, et al.
Published: (2024)
by: Deng, Kunquan, et al.
Published: (2024)
LLM-based NLG Evaluation: Current Status and Challenges
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution
by: Hong, Hanhua, et al.
Published: (2026)
by: Hong, Hanhua, et al.
Published: (2026)
NLG Evaluation: Past, Present, Future
by: Reiter, Ehud
Published: (2026)
by: Reiter, Ehud
Published: (2026)
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
by: Zhang, Jianfei, et al.
Published: (2024)
by: Zhang, Jianfei, et al.
Published: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024)
by: Niwa, Ayana, et al.
Published: (2024)
From Instruction to Output: The Role of Prompting in Modern NLG
by: Zaib, Munazza, et al.
Published: (2026)
by: Zaib, Munazza, et al.
Published: (2026)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
by: Sheng, Shuqian, et al.
Published: (2024)
by: Sheng, Shuqian, et al.
Published: (2024)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
by: Wang, Yang, et al.
Published: (2023)
by: Wang, Yang, et al.
Published: (2023)
Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text Ranking
by: Bai, Jun, et al.
Published: (2024)
by: Bai, Jun, et al.
Published: (2024)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
by: Yuan, Peiwen, et al.
Published: (2025)
by: Yuan, Peiwen, et al.
Published: (2025)
Think Beyond Size: Adaptive Prompting for More Effective Reasoning
by: R, Kamesh
Published: (2024)
by: R, Kamesh
Published: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
by: Chang, Jiayi, et al.
Published: (2025)
by: Chang, Jiayi, et al.
Published: (2025)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
by: Liu, Xiaoou, et al.
Published: (2025)
by: Liu, Xiaoou, et al.
Published: (2025)
Large Language Models Are Active Critics in NLG Evaluation
by: Xu, Shuying, et al.
Published: (2024)
by: Xu, Shuying, et al.
Published: (2024)
Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Similar Items
-
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023) -
Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
by: Zhang, Rufan, et al.
Published: (2025) -
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
by: Li, Yizhi, et al.
Published: (2025) -
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
by: Liu, Yiqi, et al.
Published: (2023) -
Effective Distillation of Table-based Reasoning Ability from LLMs
by: Yang, Bohao, et al.
Published: (2023)