Guardado en:
| Autores principales: | Wang, Liang, Wang, Junpeng, Yeh, Chin-chia Michael, Zheng, Yan, Sun, Jiarui, Fan, Xiran, Dai, Xin, Fan, Yujie, Cai, Yiwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.05110 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
por: Wang, Xiaohua, et al.
Publicado: (2026)
por: Wang, Xiaohua, et al.
Publicado: (2026)
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
por: Fostiropoulos, Iordanis, et al.
Publicado: (2026)
por: Fostiropoulos, Iordanis, et al.
Publicado: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
por: Souza, Débora, et al.
Publicado: (2026)
por: Souza, Débora, et al.
Publicado: (2026)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
por: Karinshak, Elise, et al.
Publicado: (2024)
por: Karinshak, Elise, et al.
Publicado: (2024)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
por: Wu, Dekun, et al.
Publicado: (2023)
por: Wu, Dekun, et al.
Publicado: (2023)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
por: He, Kaifeng, et al.
Publicado: (2025)
por: He, Kaifeng, et al.
Publicado: (2025)
I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
por: Dai, Dasen, et al.
Publicado: (2026)
por: Dai, Dasen, et al.
Publicado: (2026)
SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
por: Wang, Tianyu, et al.
Publicado: (2026)
por: Wang, Tianyu, et al.
Publicado: (2026)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
por: Palit, Sayon, et al.
Publicado: (2025)
por: Palit, Sayon, et al.
Publicado: (2025)
CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening
por: Lorenzoni, Giuliano, et al.
Publicado: (2026)
por: Lorenzoni, Giuliano, et al.
Publicado: (2026)
Evaluating LLM Metrics Through Real-World Capabilities
por: Miller, Justin K, et al.
Publicado: (2025)
por: Miller, Justin K, et al.
Publicado: (2025)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
por: Sandan, Isik Baran, et al.
Publicado: (2025)
por: Sandan, Isik Baran, et al.
Publicado: (2025)
The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework
por: Shah, Aakriti, et al.
Publicado: (2025)
por: Shah, Aakriti, et al.
Publicado: (2025)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
por: Lucas, Tom, et al.
Publicado: (2026)
por: Lucas, Tom, et al.
Publicado: (2026)
LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation
por: Nguyen, Huyen, et al.
Publicado: (2026)
por: Nguyen, Huyen, et al.
Publicado: (2026)
ATANT: An Evaluation Framework for AI Continuity
por: Tanguturi, Samuel Sameer
Publicado: (2026)
por: Tanguturi, Samuel Sameer
Publicado: (2026)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
por: Ghosh, Shubhra, et al.
Publicado: (2025)
por: Ghosh, Shubhra, et al.
Publicado: (2025)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
por: Lai, Junyu, et al.
Publicado: (2025)
por: Lai, Junyu, et al.
Publicado: (2025)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
por: Iqbal, Hasan, et al.
Publicado: (2024)
por: Iqbal, Hasan, et al.
Publicado: (2024)
Evaluating LLM-Based Grant Proposal Review via Structured Perturbations
por: Thorne, William, et al.
Publicado: (2026)
por: Thorne, William, et al.
Publicado: (2026)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
por: Loi, Dario, et al.
Publicado: (2025)
por: Loi, Dario, et al.
Publicado: (2025)
Evaluating the Clinical Safety of LLMs in Response to High-Risk Mental Health Disclosures
por: Shah, Siddharth, et al.
Publicado: (2025)
por: Shah, Siddharth, et al.
Publicado: (2025)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
por: Hashemi, Helia, et al.
Publicado: (2024)
por: Hashemi, Helia, et al.
Publicado: (2024)
Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning
por: Lin, Xinxin, et al.
Publicado: (2026)
por: Lin, Xinxin, et al.
Publicado: (2026)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
por: Hinterleitner, Lukas, et al.
Publicado: (2026)
por: Hinterleitner, Lukas, et al.
Publicado: (2026)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
por: Hu, Junyi, et al.
Publicado: (2024)
por: Hu, Junyi, et al.
Publicado: (2024)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
por: Li, Bowen, et al.
Publicado: (2026)
por: Li, Bowen, et al.
Publicado: (2026)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
por: Sun, Jingyi, et al.
Publicado: (2024)
por: Sun, Jingyi, et al.
Publicado: (2024)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
por: Han, Lifeng, et al.
Publicado: (2016)
por: Han, Lifeng, et al.
Publicado: (2016)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
por: Liu, Fangxin, et al.
Publicado: (2025)
por: Liu, Fangxin, et al.
Publicado: (2025)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
por: Guo, Pengze, et al.
Publicado: (2026)
por: Guo, Pengze, et al.
Publicado: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
por: Sun, Jiachen, et al.
Publicado: (2024)
por: Sun, Jiachen, et al.
Publicado: (2024)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
por: Da, Longchao, et al.
Publicado: (2025)
por: Da, Longchao, et al.
Publicado: (2025)
TiCT: A Synthetically Pre-Trained Foundation Model for Time Series Classification
por: Yeh, Chin-Chia Michael, et al.
Publicado: (2025)
por: Yeh, Chin-Chia Michael, et al.
Publicado: (2025)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
por: Karpurapu, Shanthi, et al.
Publicado: (2024)
por: Karpurapu, Shanthi, et al.
Publicado: (2024)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
por: Mehta, Manisha, et al.
Publicado: (2025)
por: Mehta, Manisha, et al.
Publicado: (2025)
Ejemplares similares
-
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
por: Wang, Xiaohua, et al.
Publicado: (2026) -
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
por: Fostiropoulos, Iordanis, et al.
Publicado: (2026) -
Toward Architecture-Aware Evaluation Metrics for LLM Agents
por: Souza, Débora, et al.
Publicado: (2026) -
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
por: Karinshak, Elise, et al.
Publicado: (2024) -
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
por: Wu, Dekun, et al.
Publicado: (2023)