BatchEval: Towards Human-like Text Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Yuan, Peiwen, Feng, Shaoxiong, Li, Yiwei, Wang, Xinglin, Pan, Boyuan, Wang, Heda, Li, Kan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Poor-Supervised Evaluation for SuperLLM via Mutual Consistency
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning
di: Li, Yiwei, et al.
Pubblicazione: (2024)
di: Li, Yiwei, et al.
Pubblicazione: (2024)
Generative Dense Retrieval: Memory Can Be a Burden
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
CogLM: Tracking Cognitive Development of Large Language Models
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
Instruction Embedding: Latent Representations of Instructions Towards Task Identification
di: Li, Yiwei, et al.
Pubblicazione: (2024)
di: Li, Yiwei, et al.
Pubblicazione: (2024)
Focused Large Language Models are Stable Many-Shot Learners
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MARKERGEN
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective Optimization
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
Make Every Penny Count: Difficulty-Adaptive Self-Consistency for Cost-Efficient Reasoning
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
Silencer: From Discovery to Mitigation of Self-Bias in LLM-as-Benchmark-Generator
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules
di: Zhang, Yueqi, et al.
Pubblicazione: (2025)
di: Zhang, Yueqi, et al.
Pubblicazione: (2025)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
InsBank: Evolving Instruction Subset for Ongoing Alignment
di: Shi, Jiayi, et al.
Pubblicazione: (2025)
di: Shi, Jiayi, et al.
Pubblicazione: (2025)
Revisiting Self-Consistency from Dynamic Distributional Alignment Perspective on Answer Aggregation
di: Li, Yiwei, et al.
Pubblicazione: (2025)
di: Li, Yiwei, et al.
Pubblicazione: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
di: Tan, Chuyi, et al.
Pubblicazione: (2025)
di: Tan, Chuyi, et al.
Pubblicazione: (2025)
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
Speculative Decoding for Multi-Sample Inference
di: Li, Yiwei, et al.
Pubblicazione: (2025)
di: Li, Yiwei, et al.
Pubblicazione: (2025)
On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
Dynamic Stochastic Decoding Strategy for Open-Domain Dialogue Generation
di: Li, Yiwei, et al.
Pubblicazione: (2024)
di: Li, Yiwei, et al.
Pubblicazione: (2024)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
di: Wang, Xinglin, et al.
Pubblicazione: (2025)
di: Wang, Xinglin, et al.
Pubblicazione: (2025)
PatternKV: Flattening KV Representation Expands Quantization Headroom
di: Zhang, Ji, et al.
Pubblicazione: (2025)
di: Zhang, Ji, et al.
Pubblicazione: (2025)
Towards Robustness and Diversity: Continual Learning in Dialog Generation with Text-Mixup and Batch Nuclear-Norm Maximization
di: Wang, Zihan, et al.
Pubblicazione: (2024)
di: Wang, Zihan, et al.
Pubblicazione: (2024)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
di: Yuan, Xiaohan, et al.
Pubblicazione: (2024)
di: Yuan, Xiaohan, et al.
Pubblicazione: (2024)
RepEval: Effective Text Evaluation with LLM Representation
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models
di: Luo, Hengyu, et al.
Pubblicazione: (2025)
di: Luo, Hengyu, et al.
Pubblicazione: (2025)
Learning More from Less: Unlocking Internal Representations for Benchmark Compression
di: Zhang, Yueqi, et al.
Pubblicazione: (2026)
di: Zhang, Yueqi, et al.
Pubblicazione: (2026)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
di: Zhou, Lingfeng, et al.
Pubblicazione: (2025)
di: Zhou, Lingfeng, et al.
Pubblicazione: (2025)
BotEval: Facilitating Interactive Human Evaluation
di: Cho, Hyundong, et al.
Pubblicazione: (2024)
di: Cho, Hyundong, et al.
Pubblicazione: (2024)
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
di: Jiang, Lai, et al.
Pubblicazione: (2025)
di: Jiang, Lai, et al.
Pubblicazione: (2025)
DnA-Eval: Enhancing Large Language Model Evaluation through Decomposition and Aggregation
di: Li, Minzhi, et al.
Pubblicazione: (2024)
di: Li, Minzhi, et al.
Pubblicazione: (2024)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
di: Gu, Zhouhong, et al.
Pubblicazione: (2024)
di: Gu, Zhouhong, et al.
Pubblicazione: (2024)
AlphaEval: Evaluating Agents in Production
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
di: Zhao, Haiquan, et al.
Pubblicazione: (2024)
di: Zhao, Haiquan, et al.
Pubblicazione: (2024)
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
di: yunhan, Li, et al.
Pubblicazione: (2025)
di: yunhan, Li, et al.
Pubblicazione: (2025)
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
di: Ramesh, Krithika, et al.
Pubblicazione: (2025)
di: Ramesh, Krithika, et al.
Pubblicazione: (2025)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
di: Zhao, Jiahao, et al.
Pubblicazione: (2025)
di: Zhao, Jiahao, et al.
Pubblicazione: (2025)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
di: Zhou, Yixi, et al.
Pubblicazione: (2026)
di: Zhou, Yixi, et al.
Pubblicazione: (2026)
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
di: Gritta, Milan, et al.
Pubblicazione: (2024)
di: Gritta, Milan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Poor-Supervised Evaluation for SuperLLM via Mutual Consistency
di: Yuan, Peiwen, et al.
Pubblicazione: (2024) -
Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning
di: Li, Yiwei, et al.
Pubblicazione: (2024) -
Generative Dense Retrieval: Memory Can Be a Burden
di: Yuan, Peiwen, et al.
Pubblicazione: (2024) -
CogLM: Tracking Cognitive Development of Large Language Models
di: Wang, Xinglin, et al.
Pubblicazione: (2024) -
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation
di: Wang, Xinglin, et al.
Pubblicazione: (2024)