Self-Evaluation of Large Language Model based on Glass-box Features
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Hui, Qu, Yingqi, Liu, Jing, Yang, Muyun, Xu, Bing, Zhao, Tiejun, Lu, Wenpeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
Mitigating the Bias of Large Language Model Evaluation
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
di: Liu, Dongxu, et al.
Pubblicazione: (2024)
di: Liu, Dongxu, et al.
Pubblicazione: (2024)
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models
di: Xu, Mufan, et al.
Pubblicazione: (2026)
di: Xu, Mufan, et al.
Pubblicazione: (2026)
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
di: Zhou, Hongli, et al.
Pubblicazione: (2025)
di: Zhou, Hongli, et al.
Pubblicazione: (2025)
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
di: Xu, Mufan, et al.
Pubblicazione: (2024)
di: Xu, Mufan, et al.
Pubblicazione: (2024)
A Survey on Human Preference Learning for Large Language Models
di: Jiang, Ruili, et al.
Pubblicazione: (2024)
di: Jiang, Ruili, et al.
Pubblicazione: (2024)
DUAL-REFLECT: Enhancing Large Language Models for Reflective Translation through Dual Learning Feedback Mechanisms
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
BASES: Large-scale Web Search User Simulation with Large Language Model based Agents
di: Ren, Ruiyang, et al.
Pubblicazione: (2024)
di: Ren, Ruiyang, et al.
Pubblicazione: (2024)
DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
di: Chen, Andong, et al.
Pubblicazione: (2025)
di: Chen, Andong, et al.
Pubblicazione: (2025)
Dual Instruction Tuning with Large Language Models for Mathematical Reasoning
di: Zhou, Yongwei, et al.
Pubblicazione: (2024)
di: Zhou, Yongwei, et al.
Pubblicazione: (2024)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
di: Tang, Yihong, et al.
Pubblicazione: (2025)
di: Tang, Yihong, et al.
Pubblicazione: (2025)
Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning
di: Xu, Mufan, et al.
Pubblicazione: (2025)
di: Xu, Mufan, et al.
Pubblicazione: (2025)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
User-Aware Active Knowledge Acquisition for Emotional Support Dialogue
di: Xu, Mufan, et al.
Pubblicazione: (2026)
di: Xu, Mufan, et al.
Pubblicazione: (2026)
SCM: Enhancing Large Language Model with Self-Controlled Memory Framework
di: Wang, Bing, et al.
Pubblicazione: (2023)
di: Wang, Bing, et al.
Pubblicazione: (2023)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
Ranked Voting based Self-Consistency of Large Language Models
di: Wang, Weiqin, et al.
Pubblicazione: (2025)
di: Wang, Weiqin, et al.
Pubblicazione: (2025)
LLM-based Translation Inference with Iterative Bilingual Understanding
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
di: Zhao, Weixiang, et al.
Pubblicazione: (2026)
di: Zhao, Weixiang, et al.
Pubblicazione: (2026)
Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation
di: Ren, Ruiyang, et al.
Pubblicazione: (2023)
di: Ren, Ruiyang, et al.
Pubblicazione: (2023)
Advancing Large Language Model Attribution through Self-Improving
di: Huang, Lei, et al.
Pubblicazione: (2024)
di: Huang, Lei, et al.
Pubblicazione: (2024)
Large Language Model Instruction Following: A Survey of Progresses and Challenges
di: Lou, Renze, et al.
Pubblicazione: (2023)
di: Lou, Renze, et al.
Pubblicazione: (2023)
Enhancing Large Language Models'Machine Translation via Dynamic Focus Anchoring
di: Ding, Qiuyu, et al.
Pubblicazione: (2025)
di: Ding, Qiuyu, et al.
Pubblicazione: (2025)
StreamingThinker: Large Language Models Can Think While Reading
di: Tong, Junlong, et al.
Pubblicazione: (2025)
di: Tong, Junlong, et al.
Pubblicazione: (2025)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
SRLCG: Self-Rectified Large-Scale Code Generation with Multidimensional Chain-of-Thought and Dynamic Backtracking
di: Ma, Hongru, et al.
Pubblicazione: (2025)
di: Ma, Hongru, et al.
Pubblicazione: (2025)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
di: Zhao, Siyan, et al.
Pubblicazione: (2026)
di: Zhao, Siyan, et al.
Pubblicazione: (2026)
CBP-Tuning: Efficient Local Customization for Black-box Large Language Models
di: Zhao, Jiaxuan, et al.
Pubblicazione: (2025)
di: Zhao, Jiaxuan, et al.
Pubblicazione: (2025)
Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel Collaboration
di: Huang, Yichong, et al.
Pubblicazione: (2024)
di: Huang, Yichong, et al.
Pubblicazione: (2024)
Knowledge Editing on Black-box Large Language Models
di: Song, Xiaoshuai, et al.
Pubblicazione: (2024)
di: Song, Xiaoshuai, et al.
Pubblicazione: (2024)
Large Language Models for Mathematical Reasoning: Progresses and Challenges
di: Ahn, Janice, et al.
Pubblicazione: (2024)
di: Ahn, Janice, et al.
Pubblicazione: (2024)
Documenti analoghi
-
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024) -
Mitigating the Bias of Large Language Model Evaluation
di: Zhou, Hongli, et al.
Pubblicazione: (2024) -
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
di: Liu, Dongxu, et al.
Pubblicazione: (2024) -
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
di: Huang, Hui, et al.
Pubblicazione: (2025) -
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
di: Zhou, Hongli, et al.
Pubblicazione: (2026)