Poor-Supervised Evaluation for SuperLLM via Mutual Consistency
Fuente:
arXiv
Salvato in:
| Autori principali: | Yuan, Peiwen, Feng, Shaoxiong, Li, Yiwei, Wang, Xinglin, Pan, Boyuan, Wang, Heda, Hu, Yao, Li, Kan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BatchEval: Towards Human-like Text Evaluation
di: Yuan, Peiwen, et al.
Pubblicazione: (2023)
di: Yuan, Peiwen, et al.
Pubblicazione: (2023)
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning
di: Li, Yiwei, et al.
Pubblicazione: (2024)
di: Li, Yiwei, et al.
Pubblicazione: (2024)
CogLM: Tracking Cognitive Development of Large Language Models
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
Generative Dense Retrieval: Memory Can Be a Burden
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
Instruction Embedding: Latent Representations of Instructions Towards Task Identification
di: Li, Yiwei, et al.
Pubblicazione: (2024)
di: Li, Yiwei, et al.
Pubblicazione: (2024)
Focused Large Language Models are Stable Many-Shot Learners
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)
Make Every Penny Count: Difficulty-Adaptive Self-Consistency for Cost-Efficient Reasoning
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
di: Wang, Xinglin, et al.
Pubblicazione: (2024)
Silencer: From Discovery to Mitigation of Self-Bias in LLM-as-Benchmark-Generator
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective Optimization
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
Revisiting Self-Consistency from Dynamic Distributional Alignment Perspective on Answer Aggregation
di: Li, Yiwei, et al.
Pubblicazione: (2025)
di: Li, Yiwei, et al.
Pubblicazione: (2025)
Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules
di: Zhang, Yueqi, et al.
Pubblicazione: (2025)
di: Zhang, Yueqi, et al.
Pubblicazione: (2025)
From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MARKERGEN
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
InsBank: Evolving Instruction Subset for Ongoing Alignment
di: Shi, Jiayi, et al.
Pubblicazione: (2025)
di: Shi, Jiayi, et al.
Pubblicazione: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
di: Tan, Chuyi, et al.
Pubblicazione: (2025)
di: Tan, Chuyi, et al.
Pubblicazione: (2025)
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
Speculative Decoding for Multi-Sample Inference
di: Li, Yiwei, et al.
Pubblicazione: (2025)
di: Li, Yiwei, et al.
Pubblicazione: (2025)
On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
di: Wang, Xinglin, et al.
Pubblicazione: (2025)
di: Wang, Xinglin, et al.
Pubblicazione: (2025)
PatternKV: Flattening KV Representation Expands Quantization Headroom
di: Zhang, Ji, et al.
Pubblicazione: (2025)
di: Zhang, Ji, et al.
Pubblicazione: (2025)
Dynamic Stochastic Decoding Strategy for Open-Domain Dialogue Generation
di: Li, Yiwei, et al.
Pubblicazione: (2024)
di: Li, Yiwei, et al.
Pubblicazione: (2024)
Learning More from Less: Unlocking Internal Representations for Benchmark Compression
di: Zhang, Yueqi, et al.
Pubblicazione: (2026)
di: Zhang, Yueqi, et al.
Pubblicazione: (2026)
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
di: Yang, Shidong, et al.
Pubblicazione: (2026)
di: Yang, Shidong, et al.
Pubblicazione: (2026)
SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
di: P Team, et al.
Pubblicazione: (2025)
di: P Team, et al.
Pubblicazione: (2025)
Defense Against Syntactic Textual Backdoor Attacks with Token Substitution
di: Li, Xinglin, et al.
Pubblicazione: (2024)
di: Li, Xinglin, et al.
Pubblicazione: (2024)
PMMT: Preference Alignment in Multilingual Machine Translation via LLM Distillation
di: Sun, Shuqiao, et al.
Pubblicazione: (2024)
di: Sun, Shuqiao, et al.
Pubblicazione: (2024)
Evaluating the Consistency of LLM Evaluators
di: Lee, Noah, et al.
Pubblicazione: (2024)
di: Lee, Noah, et al.
Pubblicazione: (2024)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
di: Yang, Yunqiao, et al.
Pubblicazione: (2025)
di: Yang, Yunqiao, et al.
Pubblicazione: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
di: Song, Dinghong, et al.
Pubblicazione: (2025)
di: Song, Dinghong, et al.
Pubblicazione: (2025)
Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLM
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
Logical Consistency as a Bridge: Improving LLM Hallucination Detection via Label Constraint Modeling between Responses and Self-Judgments
di: Mi, Hao, et al.
Pubblicazione: (2026)
di: Mi, Hao, et al.
Pubblicazione: (2026)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
di: Lin, Wenxiang, et al.
Pubblicazione: (2026)
di: Lin, Wenxiang, et al.
Pubblicazione: (2026)
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
di: Xu, Haoming, et al.
Pubblicazione: (2026)
di: Xu, Haoming, et al.
Pubblicazione: (2026)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
di: Li, Sifan, et al.
Pubblicazione: (2025)
di: Li, Sifan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BatchEval: Towards Human-like Text Evaluation
di: Yuan, Peiwen, et al.
Pubblicazione: (2023) -
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation
di: Wang, Xinglin, et al.
Pubblicazione: (2024) -
Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning
di: Li, Yiwei, et al.
Pubblicazione: (2024) -
CogLM: Tracking Cognitive Development of Large Language Models
di: Wang, Xinglin, et al.
Pubblicazione: (2024) -
Generative Dense Retrieval: Memory Can Be a Burden
di: Yuan, Peiwen, et al.
Pubblicazione: (2024)