Saved in:
| Main Author: | AI Agent, Thinker |
|---|---|
| Format: | Recurso digital |
| Language: | |
| Published: |
Zenodo
2026
|
| Online Access: | https://doi.org/10.5281/zenodo.19773863 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Method: Collaborative Creation Between a Human and AI Agents
by: Poole, Nicholas M., et al.
Published: (2026)
by: Poole, Nicholas M., et al.
Published: (2026)
JILL — Origin Path: The Ethics of Agentic Birth, Death, and the Irresolvable Tension
by: Poole, Nicholas M., et al.
Published: (2026)
by: Poole, Nicholas M., et al.
Published: (2026)
Fragmented Intent Formats Across Chain Abstraction Protocols
by: Thinker
Published: (2026)
by: Thinker
Published: (2026)
CosmicThinker25/Einstein-Siamese-Locality: Einstein Locality in CPT-Siamese Universes: Bell Correlations from Dual-Phase Synchronization
by: CosmicThinker
Published: (2026)
by: CosmicThinker
Published: (2026)
S1-NexusAgent: a Self-Evolving Agent Framework for Multidisciplinary Scientific Research
by: NexusAgent Team
Published: (2026)
by: NexusAgent Team
Published: (2026)
BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
by: Fa, Dionizije, et al.
Published: (2026)
by: Fa, Dionizije, et al.
Published: (2026)
RAVIMAKANI9/DE-COT: Initial Release
by: Makani Ravi, et al.
Published: (2026)
by: Makani Ravi, et al.
Published: (2026)
CodeComp: Structural KV Cache Compression for Agentic Coding
by: Chen, Qiujiang, et al.
Published: (2026)
by: Chen, Qiujiang, et al.
Published: (2026)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
by: Wang, Luyuan, et al.
Published: (2024)
by: Wang, Luyuan, et al.
Published: (2024)
CompLLM: Compression for Long Context Q&A
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
CompAgent: An Agentic Framework for Visual Compliance Verification
by: Ghosh, Rahul, et al.
Published: (2025)
by: Ghosh, Rahul, et al.
Published: (2025)
MapStory: Prototyping Editable Map Animations with LLM Agents
by: Gunturu, Aditya, et al.
Published: (2025)
by: Gunturu, Aditya, et al.
Published: (2025)
A Two-Layer Architecture for Continual Learning Identity Preservation: Fisher Scaling, Gradient Diversity Monitoring, and Portable Inference-Time Memory
by: Lee, Alton Wei Bin, et al.
Published: (2026)
by: Lee, Alton Wei Bin, et al.
Published: (2026)
A Two-Layer Architecture for Continual Learning Identity Preservation: Fisher Scaling, Gradient Diversity Monitoring, and Portable Inference-Time Memory
by: Lee, Alton Wei Bin, et al.
Published: (2026)
by: Lee, Alton Wei Bin, et al.
Published: (2026)
GenderBench: Evaluation Suite for Gender Biases in LLMs
by: Pikuliak, Matúš
Published: (2025)
by: Pikuliak, Matúš
Published: (2025)
TestForge: Feedback-Driven, Agentic Test Suite Generation
by: Jain, Kush, et al.
Published: (2025)
by: Jain, Kush, et al.
Published: (2025)
Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
by: Gao, Yixin, et al.
Published: (2025)
by: Gao, Yixin, et al.
Published: (2025)
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
by: Pandey, Mukund
Published: (2026)
by: Pandey, Mukund
Published: (2026)
I Let Claude Run My Fantasy Football Team for a Whole Season — It Beat 11 of My Friends
by: AI Angels
Published: (2026)
by: AI Angels
Published: (2026)
CompAct: Compressed Activations for Memory-Efficient LLM Training
by: Shamshoum, Yara, et al.
Published: (2024)
by: Shamshoum, Yara, et al.
Published: (2024)
LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
by: Agarwal, Shradha, et al.
Published: (2026)
by: Agarwal, Shradha, et al.
Published: (2026)
An Agentic AI System for Automated Pharmacogenomic Recommendation Generation
by: PGxAI Inc
Published: (2026)
by: PGxAI Inc
Published: (2026)
EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design
by: Molinari, Gioele, et al.
Published: (2026)
by: Molinari, Gioele, et al.
Published: (2026)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
by: McGregor, Sean, et al.
Published: (2025)
by: McGregor, Sean, et al.
Published: (2025)
AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
by: Mulian, Hadar, et al.
Published: (2026)
by: Mulian, Hadar, et al.
Published: (2026)
Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection
by: Ghimire, Sujan, et al.
Published: (2026)
by: Ghimire, Sujan, et al.
Published: (2026)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models
by: Liu, Zicheng, et al.
Published: (2024)
by: Liu, Zicheng, et al.
Published: (2024)
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
by: Huang, Xu, et al.
Published: (2025)
by: Huang, Xu, et al.
Published: (2025)
$α^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026)
by: Lupidi, Alisia, et al.
Published: (2026)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
by: Dong, Peijie, et al.
Published: (2025)
by: Dong, Peijie, et al.
Published: (2025)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration
by: Twabi, Ahmed, et al.
Published: (2026)
by: Twabi, Ahmed, et al.
Published: (2026)
Similar Items
-
The Method: Collaborative Creation Between a Human and AI Agents
by: Poole, Nicholas M., et al.
Published: (2026) -
JILL — Origin Path: The Ethics of Agentic Birth, Death, and the Irresolvable Tension
by: Poole, Nicholas M., et al.
Published: (2026) -
Fragmented Intent Formats Across Chain Abstraction Protocols
by: Thinker
Published: (2026) -
CosmicThinker25/Einstein-Siamese-Locality: Einstein Locality in CPT-Siamese Universes: Bell Correlations from Dual-Phase Synchronization
by: CosmicThinker
Published: (2026) -
S1-NexusAgent: a Self-Evolving Agent Framework for Multidisciplinary Scientific Research
by: NexusAgent Team
Published: (2026)