Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Hong, Hanhua, Xiao, Chenghao, Wang, Yang, Liu, Yiqi, Rong, Wenge, Lin, Chenghua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
di: Lucy, Li, et al.
Pubblicazione: (2023)
di: Lucy, Li, et al.
Pubblicazione: (2023)
Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
di: Zhang, Rufan, et al.
Pubblicazione: (2025)
di: Zhang, Rufan, et al.
Pubblicazione: (2025)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
di: Li, Yizhi, et al.
Pubblicazione: (2025)
di: Li, Yizhi, et al.
Pubblicazione: (2025)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
di: Liu, Yiqi, et al.
Pubblicazione: (2023)
di: Liu, Yiqi, et al.
Pubblicazione: (2023)
Effective Distillation of Table-based Reasoning Ability from LLMs
di: Yang, Bohao, et al.
Pubblicazione: (2023)
di: Yang, Bohao, et al.
Pubblicazione: (2023)
Beyond One-Size-Fits-All Summarization: Customizing Summaries for Diverse Users
di: Duran, Mehmet Samet, et al.
Pubblicazione: (2025)
di: Duran, Mehmet Samet, et al.
Pubblicazione: (2025)
Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
Theorem Provers: One Size Fits All?
di: Oates, Harrison, et al.
Pubblicazione: (2025)
di: Oates, Harrison, et al.
Pubblicazione: (2025)
The Achilles' Heel of Angular Margins: A Chebyshev Polynomial Fix for Speaker Verification
di: Wang, Yang, et al.
Pubblicazione: (2026)
di: Wang, Yang, et al.
Pubblicazione: (2026)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
di: Tang, Tianyi, et al.
Pubblicazione: (2023)
di: Tang, Tianyi, et al.
Pubblicazione: (2023)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
di: Zhang, Jianfei, et al.
Pubblicazione: (2025)
di: Zhang, Jianfei, et al.
Pubblicazione: (2025)
No One Fits All: From Fixed Prompting to Learned Routing in Multilingual LLMs
di: Wu, Wei-Chi, et al.
Pubblicazione: (2026)
di: Wu, Wei-Chi, et al.
Pubblicazione: (2026)
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
di: Yu, Jianxiang, et al.
Pubblicazione: (2026)
di: Yu, Jianxiang, et al.
Pubblicazione: (2026)
One Size doesn't Fit All: A Personalized Conversational Tutoring Agent for Mathematics Instruction
di: Liu, Ben, et al.
Pubblicazione: (2025)
di: Liu, Ben, et al.
Pubblicazione: (2025)
Beyond One-Size-Fits-All: Multi-Domain, Multi-Task Framework for Embedding Model Selection
di: Khetan, Vivek
Pubblicazione: (2024)
di: Khetan, Vivek
Pubblicazione: (2024)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
di: James, Joseph, et al.
Pubblicazione: (2026)
di: James, Joseph, et al.
Pubblicazione: (2026)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights
di: James, Joseph, et al.
Pubblicazione: (2024)
di: James, Joseph, et al.
Pubblicazione: (2024)
No One Size Fits All: QueryBandits for Hallucination Mitigation
di: Cho, Nicole, et al.
Pubblicazione: (2026)
di: Cho, Nicole, et al.
Pubblicazione: (2026)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
di: Lu, Liming, et al.
Pubblicazione: (2026)
di: Lu, Liming, et al.
Pubblicazione: (2026)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
di: Luo, Yingfeng, et al.
Pubblicazione: (2025)
di: Luo, Yingfeng, et al.
Pubblicazione: (2025)
LLM-based NLG Evaluation: Current Status and Challenges
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
PFME: A Modular Approach for Fine-grained Hallucination Detection and Editing of Large Language Models
di: Deng, Kunquan, et al.
Pubblicazione: (2024)
di: Deng, Kunquan, et al.
Pubblicazione: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
di: Wang, Yicheng, et al.
Pubblicazione: (2024)
di: Wang, Yicheng, et al.
Pubblicazione: (2024)
HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution
di: Hong, Hanhua, et al.
Pubblicazione: (2026)
di: Hong, Hanhua, et al.
Pubblicazione: (2026)
NLG Evaluation: Past, Present, Future
di: Reiter, Ehud
Pubblicazione: (2026)
di: Reiter, Ehud
Pubblicazione: (2026)
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
di: Zhang, Jianfei, et al.
Pubblicazione: (2024)
di: Zhang, Jianfei, et al.
Pubblicazione: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
From Instruction to Output: The Role of Prompting in Modern NLG
di: Zaib, Munazza, et al.
Pubblicazione: (2026)
di: Zaib, Munazza, et al.
Pubblicazione: (2026)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
di: Hu, Xinyu, et al.
Pubblicazione: (2024)
di: Hu, Xinyu, et al.
Pubblicazione: (2024)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
di: Wang, Yang, et al.
Pubblicazione: (2023)
di: Wang, Yang, et al.
Pubblicazione: (2023)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text Ranking
di: Bai, Jun, et al.
Pubblicazione: (2024)
di: Bai, Jun, et al.
Pubblicazione: (2024)
Think Beyond Size: Adaptive Prompting for More Effective Reasoning
di: R, Kamesh
Pubblicazione: (2024)
di: R, Kamesh
Pubblicazione: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
di: Chang, Jiayi, et al.
Pubblicazione: (2025)
di: Chang, Jiayi, et al.
Pubblicazione: (2025)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
di: Liu, Xiaoou, et al.
Pubblicazione: (2025)
di: Liu, Xiaoou, et al.
Pubblicazione: (2025)
Large Language Models Are Active Critics in NLG Evaluation
di: Xu, Shuying, et al.
Pubblicazione: (2024)
di: Xu, Shuying, et al.
Pubblicazione: (2024)
Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
di: Wang, Yang, et al.
Pubblicazione: (2025)
di: Wang, Yang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
di: Lucy, Li, et al.
Pubblicazione: (2023) -
Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
di: Zhang, Rufan, et al.
Pubblicazione: (2025) -
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
di: Li, Yizhi, et al.
Pubblicazione: (2025) -
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
di: Liu, Yiqi, et al.
Pubblicazione: (2023) -
Effective Distillation of Table-based Reasoning Ability from LLMs
di: Yang, Bohao, et al.
Pubblicazione: (2023)