Gespeichert in:
| Hauptverfasser: | Miller, Justin K, Tang, Wenjia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.08253 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
von: Wang, Liang, et al.
Veröffentlicht: (2026)
von: Wang, Liang, et al.
Veröffentlicht: (2026)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
von: Ivanov, Igor
Veröffentlicht: (2025)
von: Ivanov, Igor
Veröffentlicht: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
von: Redkar, Neel
Veröffentlicht: (2024)
von: Redkar, Neel
Veröffentlicht: (2024)
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
von: Davide, Fabrizio, et al.
Veröffentlicht: (2024)
von: Davide, Fabrizio, et al.
Veröffentlicht: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
KNOW: A Real-World Ontology for Knowledge Capture with Large Language Models
von: Bendiken, Arto
Veröffentlicht: (2024)
von: Bendiken, Arto
Veröffentlicht: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
von: Christop, Iwona, et al.
Veröffentlicht: (2026)
von: Christop, Iwona, et al.
Veröffentlicht: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
Active Context Compression: Autonomous Memory Management in LLM Agents
von: Verma, Nikhil
Veröffentlicht: (2026)
von: Verma, Nikhil
Veröffentlicht: (2026)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
von: Danilevsky, Marina, et al.
Veröffentlicht: (2025)
von: Danilevsky, Marina, et al.
Veröffentlicht: (2025)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
von: Chang, Edward Y.
Veröffentlicht: (2026)
von: Chang, Edward Y.
Veröffentlicht: (2026)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
Evaluating Relational Reasoning in LLMs with REL
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
von: Lai, Junyu, et al.
Veröffentlicht: (2025)
von: Lai, Junyu, et al.
Veröffentlicht: (2025)
EVINCE: Optimizing Multi-LLM Dialogues Using Conditional Statistics and Information Theory
von: Chang, Edward Y.
Veröffentlicht: (2024)
von: Chang, Edward Y.
Veröffentlicht: (2024)
Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
von: Paulsen, Norman
Veröffentlicht: (2025)
von: Paulsen, Norman
Veröffentlicht: (2025)
DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2026)
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2026)
Applying Cognitive Design Patterns to General LLM Agents
von: Wray, Robert E., et al.
Veröffentlicht: (2025)
von: Wray, Robert E., et al.
Veröffentlicht: (2025)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
von: Abramov, Roman, et al.
Veröffentlicht: (2025)
von: Abramov, Roman, et al.
Veröffentlicht: (2025)
LLM Performance Predictors: Learning When to Escalate in Hybrid Human-AI Moderation Systems
von: Bachar, Or, et al.
Veröffentlicht: (2026)
von: Bachar, Or, et al.
Veröffentlicht: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
von: Hasan, Mohammed Rakibul
Veröffentlicht: (2026)
von: Hasan, Mohammed Rakibul
Veröffentlicht: (2026)
From Fake Focus to Real Precision: Confusion-Driven Adversarial Attention Learning in Transformers
von: Liu, Yawei
Veröffentlicht: (2025)
von: Liu, Yawei
Veröffentlicht: (2025)
Evaluating Steering Techniques using Human Similarity Judgments
von: Studdiford, Zach, et al.
Veröffentlicht: (2025)
von: Studdiford, Zach, et al.
Veröffentlicht: (2025)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
von: Hu, Junyi, et al.
Veröffentlicht: (2024)
von: Hu, Junyi, et al.
Veröffentlicht: (2024)
Efficient LLM Safety Evaluation through Multi-Agent Debate
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2026)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2026)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
von: Vera, Patricio
Veröffentlicht: (2026)
von: Vera, Patricio
Veröffentlicht: (2026)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
von: Preuveneers, Jack, et al.
Veröffentlicht: (2025)
von: Preuveneers, Jack, et al.
Veröffentlicht: (2025)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
von: Wang, Liang, et al.
Veröffentlicht: (2026) -
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023) -
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
von: Ivanov, Igor
Veröffentlicht: (2025) -
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025) -
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
von: Redkar, Neel
Veröffentlicht: (2024)