NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
Fuente:
arXiv
Salvato in:
| Autore principale: | Jia, Xiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
di: Alpay, Faruk, et al.
Pubblicazione: (2025)
di: Alpay, Faruk, et al.
Pubblicazione: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
di: Berman, Shmuel, et al.
Pubblicazione: (2024)
di: Berman, Shmuel, et al.
Pubblicazione: (2024)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
di: Shekar, Pavan C, et al.
Pubblicazione: (2025)
di: Shekar, Pavan C, et al.
Pubblicazione: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
di: Pepe, Alberto, et al.
Pubblicazione: (2026)
di: Pepe, Alberto, et al.
Pubblicazione: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
di: Wang, Zhen, et al.
Pubblicazione: (2025)
di: Wang, Zhen, et al.
Pubblicazione: (2025)
RACAS: Controlling Diverse Robots With a Single Agentic System
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions
di: Yadav, Arin Gopalan, et al.
Pubblicazione: (2026)
di: Yadav, Arin Gopalan, et al.
Pubblicazione: (2026)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
di: Xi, Wang, et al.
Pubblicazione: (2025)
di: Xi, Wang, et al.
Pubblicazione: (2025)
AdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance Convergence
di: Yu, Geunbin
Pubblicazione: (2026)
di: Yu, Geunbin
Pubblicazione: (2026)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
di: Khan, Omer Jauhar
Pubblicazione: (2025)
di: Khan, Omer Jauhar
Pubblicazione: (2025)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
di: Qi, Dekang, et al.
Pubblicazione: (2026)
di: Qi, Dekang, et al.
Pubblicazione: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
di: Kim, Heejun, et al.
Pubblicazione: (2026)
di: Kim, Heejun, et al.
Pubblicazione: (2026)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
di: Kalušev, Vladimir, et al.
Pubblicazione: (2026)
di: Kalušev, Vladimir, et al.
Pubblicazione: (2026)
Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
di: Ravindran, Santhosh Kumar
Pubblicazione: (2026)
di: Ravindran, Santhosh Kumar
Pubblicazione: (2026)
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
How much do LLMs learn from negative examples?
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
di: Basu, Abhinaba
Pubblicazione: (2026)
di: Basu, Abhinaba
Pubblicazione: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
di: Sun, Mingrui, et al.
Pubblicazione: (2026)
di: Sun, Mingrui, et al.
Pubblicazione: (2026)
Generative AI and the Transformation of Software Development Practices
di: Acharya, Vivek
Pubblicazione: (2025)
di: Acharya, Vivek
Pubblicazione: (2025)
A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query Processing
di: Nourmohammadi, Naeimeh, et al.
Pubblicazione: (2026)
di: Nourmohammadi, Naeimeh, et al.
Pubblicazione: (2026)
CHORUS: An Agentic Framework for Generating Realistic Deliberation Data
di: Koursaris, A., et al.
Pubblicazione: (2026)
di: Koursaris, A., et al.
Pubblicazione: (2026)
A V2X-based Privacy Preserving Federated Measuring and Learning System
di: Alekszejenkó, Levente, et al.
Pubblicazione: (2024)
di: Alekszejenkó, Levente, et al.
Pubblicazione: (2024)
On measuring grounding and generalizing grounding problems
di: Quigley, Daniel, et al.
Pubblicazione: (2025)
di: Quigley, Daniel, et al.
Pubblicazione: (2025)
Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
di: Velampalli, Sirisha, et al.
Pubblicazione: (2025)
di: Velampalli, Sirisha, et al.
Pubblicazione: (2025)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
di: Nieth, Björn, et al.
Pubblicazione: (2026)
di: Nieth, Björn, et al.
Pubblicazione: (2026)
Harnessing non-adversarial robustness in large language models
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
di: Cruz, Roberto, et al.
Pubblicazione: (2026)
di: Cruz, Roberto, et al.
Pubblicazione: (2026)
Territory Paint Wars: Diagnosing and Mitigating Failure Modes in Competitive Multi-Agent PPO
di: Singh, Diyansha
Pubblicazione: (2026)
di: Singh, Diyansha
Pubblicazione: (2026)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
di: Xia, Bowei, et al.
Pubblicazione: (2026)
di: Xia, Bowei, et al.
Pubblicazione: (2026)
Emergent Coordination in Multi-Agent Systems via Pressure Fields and Temporal Decay
di: Rodriguez, Roland
Pubblicazione: (2026)
di: Rodriguez, Roland
Pubblicazione: (2026)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
di: Zhang, Xue
Pubblicazione: (2025)
di: Zhang, Xue
Pubblicazione: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
di: Estevanell-Valladares, Ernesto L., et al.
Pubblicazione: (2025)
di: Estevanell-Valladares, Ernesto L., et al.
Pubblicazione: (2025)
AI Agents: Evolution, Architecture, and Real-World Applications
di: Krishnan, Naveen
Pubblicazione: (2025)
di: Krishnan, Naveen
Pubblicazione: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
di: Pather, Kaviraj, et al.
Pubblicazione: (2025)
di: Pather, Kaviraj, et al.
Pubblicazione: (2025)
Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games
di: Qi, Runnan, et al.
Pubblicazione: (2025)
di: Qi, Runnan, et al.
Pubblicazione: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
di: Yang, Yibo
Pubblicazione: (2025)
di: Yang, Yibo
Pubblicazione: (2025)
Documenti analoghi
-
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
di: Alpay, Faruk, et al.
Pubblicazione: (2025) -
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
di: Fan, Jingxing, et al.
Pubblicazione: (2025) -
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
di: Berman, Shmuel, et al.
Pubblicazione: (2024) -
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
di: Huo, Dongjie, et al.
Pubblicazione: (2026) -
Adaptive Minds: Empowering Agents with LoRA-as-Tools
di: Shekar, Pavan C, et al.
Pubblicazione: (2025)