Salvato in:
| Autore principale: | Bradley, William F. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2411.01533 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Teach Old SAEs New Domain Tricks with Boosting
di: Koriagin, Nikita, et al.
Pubblicazione: (2025)
di: Koriagin, Nikita, et al.
Pubblicazione: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025)
di: Collot, Stephane, et al.
Pubblicazione: (2025)
User-LLM: Efficient LLM Contextualization with User Embeddings
di: Ning, Lin, et al.
Pubblicazione: (2024)
di: Ning, Lin, et al.
Pubblicazione: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
di: Wang, Yiming, et al.
Pubblicazione: (2024)
di: Wang, Yiming, et al.
Pubblicazione: (2024)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
di: Roytburg, Dani, et al.
Pubblicazione: (2026)
di: Roytburg, Dani, et al.
Pubblicazione: (2026)
Towards Multilingual LLM Evaluation for European Languages
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
Enhancing Annotated Bibliography Generation with LLM Ensembles
di: Bermejo, Sergio
Pubblicazione: (2024)
di: Bermejo, Sergio
Pubblicazione: (2024)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
di: Lin, Hender
Pubblicazione: (2025)
di: Lin, Hender
Pubblicazione: (2025)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
di: Xu, Wenda, et al.
Pubblicazione: (2024)
di: Xu, Wenda, et al.
Pubblicazione: (2024)
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task
di: Fujisaki, Yuya, et al.
Pubblicazione: (2024)
di: Fujisaki, Yuya, et al.
Pubblicazione: (2024)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
di: Wang, Chao, et al.
Pubblicazione: (2024)
di: Wang, Chao, et al.
Pubblicazione: (2024)
LLM-Enhanced Data Management
di: Zhou, Xuanhe, et al.
Pubblicazione: (2024)
di: Zhou, Xuanhe, et al.
Pubblicazione: (2024)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
Evaluating Very Long-Term Conversational Memory of LLM Agents
di: Maharana, Adyasha, et al.
Pubblicazione: (2024)
di: Maharana, Adyasha, et al.
Pubblicazione: (2024)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
Evaluating LLM Understanding via Structured Tabular Decision Simulations
di: Li, Sichao, et al.
Pubblicazione: (2025)
di: Li, Sichao, et al.
Pubblicazione: (2025)
Evaluating Cooperation in LLM Social Groups through Elected Leadership
di: Faulkner, Ryan, et al.
Pubblicazione: (2026)
di: Faulkner, Ryan, et al.
Pubblicazione: (2026)
Context-Enhanced Contrastive Search for Improved LLM Text Generation
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
di: Cohen, Seffi, et al.
Pubblicazione: (2025)
di: Cohen, Seffi, et al.
Pubblicazione: (2025)
Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
di: Zheng, Hang, et al.
Pubblicazione: (2025)
di: Zheng, Hang, et al.
Pubblicazione: (2025)
Concept Layers: Enhancing Interpretability and Intervenability via LLM Conceptualization
di: Bidusa, Or Raphael, et al.
Pubblicazione: (2025)
di: Bidusa, Or Raphael, et al.
Pubblicazione: (2025)
Enhancing LLM Agent Safety via Causal Influence Prompting
di: Hahm, Dongyoon, et al.
Pubblicazione: (2025)
di: Hahm, Dongyoon, et al.
Pubblicazione: (2025)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
di: Ma, Chang, et al.
Pubblicazione: (2024)
di: Ma, Chang, et al.
Pubblicazione: (2024)
CreditAudit: 2$^\text{nd}$ Dimension for LLM Evaluation and Selection
di: Song, Yiliang, et al.
Pubblicazione: (2026)
di: Song, Yiliang, et al.
Pubblicazione: (2026)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
di: Pandey, Atharva, et al.
Pubblicazione: (2025)
di: Pandey, Atharva, et al.
Pubblicazione: (2025)
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
di: Huang, Xin, et al.
Pubblicazione: (2026)
di: Huang, Xin, et al.
Pubblicazione: (2026)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
di: Roytburg, Dani, et al.
Pubblicazione: (2025)
di: Roytburg, Dani, et al.
Pubblicazione: (2025)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
di: Xie, Guofu, et al.
Pubblicazione: (2025)
di: Xie, Guofu, et al.
Pubblicazione: (2025)
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
di: Mizrahi, Moran, et al.
Pubblicazione: (2025)
di: Mizrahi, Moran, et al.
Pubblicazione: (2025)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Teach Old SAEs New Domain Tricks with Boosting
di: Koriagin, Nikita, et al.
Pubblicazione: (2025) -
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025) -
User-LLM: Efficient LLM Contextualization with User Embeddings
di: Ning, Lin, et al.
Pubblicazione: (2024) -
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024) -
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)