Saved in:
| Main Authors: | Ruan, Jie, Pu, Xiao, Gao, Mingqi, Wan, Xiaojun, Zhu, Yuesheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.07967 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
by: Ruan, Jie, et al.
Published: (2024)
by: Ruan, Jie, et al.
Published: (2024)
LLM-based NLG Evaluation: Current Status and Challenges
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
by: Hu, Xinyu, et al.
Published: (2025)
by: Hu, Xinyu, et al.
Published: (2025)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
by: Chang, Jiayi, et al.
Published: (2025)
by: Chang, Jiayi, et al.
Published: (2025)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
A Systematic Review of Data-to-Text NLG
by: Osuji, Chinonso Cynthia, et al.
Published: (2024)
by: Osuji, Chinonso Cynthia, et al.
Published: (2024)
ConFu: Contemplate the Future for Better Speculative Sampling
by: Qin, Zongyue, et al.
Published: (2026)
by: Qin, Zongyue, et al.
Published: (2026)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
by: Bai, Xueying, et al.
Published: (2024)
by: Bai, Xueying, et al.
Published: (2024)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
by: Hong, Yuzhong, et al.
Published: (2024)
by: Hong, Yuzhong, et al.
Published: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMs
by: Li, Xueyan, et al.
Published: (2025)
by: Li, Xueyan, et al.
Published: (2025)
Large Language Models Are Active Critics in NLG Evaluation
by: Xu, Shuying, et al.
Published: (2024)
by: Xu, Shuying, et al.
Published: (2024)
Constrained Adaptive Rejection Sampling
by: Parys, Paweł, et al.
Published: (2025)
by: Parys, Paweł, et al.
Published: (2025)
TRA: Better Length Generalisation with Threshold Relative Attention
by: Opper, Mattia, et al.
Published: (2025)
by: Opper, Mattia, et al.
Published: (2025)
Stabilizing Policy Optimization via Logits Convexity
by: Chen, Hongzhan, et al.
Published: (2026)
by: Chen, Hongzhan, et al.
Published: (2026)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
Enhancing Text Generation in Joint NLG/NLU Learning Through Curriculum Learning, Semi-Supervised Training, and Advanced Optimization Techniques
by: Shaik, Rahimanuddin, et al.
Published: (2024)
by: Shaik, Rahimanuddin, et al.
Published: (2024)
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
by: Huo, Yifu, et al.
Published: (2026)
by: Huo, Yifu, et al.
Published: (2026)
Deep Prompt Multi-task Network for Abuse Language Detection
by: Zhu, Jian, et al.
Published: (2024)
by: Zhu, Jian, et al.
Published: (2024)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
by: Song, Haiyue, et al.
Published: (2026)
by: Song, Haiyue, et al.
Published: (2026)
$p1$: Better Prompt Optimization with Fewer Prompts
by: Gao, Zhaolin, et al.
Published: (2026)
by: Gao, Zhaolin, et al.
Published: (2026)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
by: Yao, Kai, et al.
Published: (2024)
by: Yao, Kai, et al.
Published: (2024)
Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
by: Liu, Yiming, et al.
Published: (2025)
by: Liu, Yiming, et al.
Published: (2025)
Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal Alignment
by: Ma, Xiang, et al.
Published: (2026)
by: Ma, Xiang, et al.
Published: (2026)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
by: Qiu, Wenjie, et al.
Published: (2025)
by: Qiu, Wenjie, et al.
Published: (2025)
Reliability Under Randomness: An Empirical Analysis of Sparse and Dense Language Models Across Decoding Temperatures
by: Grover, Kabir
Published: (2026)
by: Grover, Kabir
Published: (2026)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Learning to Learn for Few-shot Continual Active Learning
by: Ho, Stella, et al.
Published: (2023)
by: Ho, Stella, et al.
Published: (2023)
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024)
by: Wan, Xingchen, et al.
Published: (2024)
LemmaHead: RAG Assisted Proof Generation Using Large Language Models
by: Yang, Tianbo, et al.
Published: (2025)
by: Yang, Tianbo, et al.
Published: (2025)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation
by: Yin, Xunjian, et al.
Published: (2024)
by: Yin, Xunjian, et al.
Published: (2024)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
by: Ma, Qiyao, et al.
Published: (2026)
by: Ma, Qiyao, et al.
Published: (2026)
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024)
by: Poiroux, Auguste, et al.
Published: (2024)
Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition
by: Feng, Kehua, et al.
Published: (2024)
by: Feng, Kehua, et al.
Published: (2024)
Similar Items
-
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
by: Ruan, Jie, et al.
Published: (2024) -
LLM-based NLG Evaluation: Current Status and Challenges
by: Gao, Mingqi, et al.
Published: (2024) -
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
by: Hu, Xinyu, et al.
Published: (2025) -
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
by: Gao, Mingqi, et al.
Published: (2024) -
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
by: Chang, Jiayi, et al.
Published: (2025)