InnoGym: Benchmarking the Innovation Potential of AI Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jintian, Xu, Kewei, Zheng, Jingsheng, Yu, Zhuoyun, Zhu, Yuqi, Luo, Yujie, Wei, Lanning, Qiao, Shuofei, Du, Lun, Zheng, Da, Deng, Shumin, Chen, Huajun, Zhang, Ningyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
von: Ou, Yixin, et al.
Veröffentlicht: (2025)
von: Ou, Yixin, et al.
Veröffentlicht: (2025)
What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations
von: Luo, Yujie, et al.
Veröffentlicht: (2025)
von: Luo, Yujie, et al.
Veröffentlicht: (2025)
Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
von: Zhu, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2025)
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
von: Qiu, Zhisong, et al.
Veröffentlicht: (2026)
von: Qiu, Zhisong, et al.
Veröffentlicht: (2026)
Can We Predict Before Executing Machine Learning Agents?
von: Zheng, Jingsheng, et al.
Veröffentlicht: (2026)
von: Zheng, Jingsheng, et al.
Veröffentlicht: (2026)
KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
von: Ren, Baochang, et al.
Veröffentlicht: (2025)
von: Ren, Baochang, et al.
Veröffentlicht: (2025)
KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents
von: Zhu, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2024)
Agent Planning with World Knowledge Model
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View
von: Zhang, Jintian, et al.
Veröffentlicht: (2023)
von: Zhang, Jintian, et al.
Veröffentlicht: (2023)
LightMem: Lightweight and Efficient Memory-Augmented Generation
von: Fang, Jizhan, et al.
Veröffentlicht: (2025)
von: Fang, Jizhan, et al.
Veröffentlicht: (2025)
Benchmarking Agentic Workflow Generation
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
Exploring Model Kinship for Merging Large Language Models
von: Hu, Yedi, et al.
Veröffentlicht: (2024)
von: Hu, Yedi, et al.
Veröffentlicht: (2024)
LightThinker: Thinking Step-by-Step Compression
von: Zhang, Jintian, et al.
Veröffentlicht: (2025)
von: Zhang, Jintian, et al.
Veröffentlicht: (2025)
OceanGym: A Benchmark Environment for Underwater Embodied Agents
von: Xue, Yida, et al.
Veröffentlicht: (2025)
von: Xue, Yida, et al.
Veröffentlicht: (2025)
SkillX: Automatically Constructing Skill Knowledge Bases for Agents
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
Memp: Exploring Agent Procedural Memory
von: Fang, Runnan, et al.
Veröffentlicht: (2025)
von: Fang, Runnan, et al.
Veröffentlicht: (2025)
LightThinker++: From Reasoning Compression to Memory Management
von: Zhu, Yuqi, et al.
Veröffentlicht: (2026)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2026)
SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement
von: Fang, Runnan, et al.
Veröffentlicht: (2025)
von: Fang, Runnan, et al.
Veröffentlicht: (2025)
Agentic Knowledgeable Self-awareness
von: Qiao, Shuofei, et al.
Veröffentlicht: (2025)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2025)
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
von: Xu, Haoming, et al.
Veröffentlicht: (2026)
von: Xu, Haoming, et al.
Veröffentlicht: (2026)
SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories
von: Yu, Zhuoyun, et al.
Veröffentlicht: (2026)
von: Yu, Zhuoyun, et al.
Veröffentlicht: (2026)
StructMem: Structured Memory for Long-Horizon Behavior in LLMs
von: Xu, Buqiang, et al.
Veröffentlicht: (2026)
von: Xu, Buqiang, et al.
Veröffentlicht: (2026)
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
von: Qiao, Shuofei, et al.
Veröffentlicht: (2026)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2026)
LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities
von: Zhu, Yuqi, et al.
Veröffentlicht: (2023)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2023)
Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey
von: Zheng, Da, et al.
Veröffentlicht: (2025)
von: Zheng, Da, et al.
Veröffentlicht: (2025)
T2VTree: User-Centered Visual Analytics for Agent-Assisted Thought-to-Video Authoring
von: Zheng, Zhuoyun, et al.
Veröffentlicht: (2026)
von: Zheng, Zhuoyun, et al.
Veröffentlicht: (2026)
LLMs Can Simulate Standardized Patients via Agent Coevolution
von: Du, Zhuoyun, et al.
Veröffentlicht: (2024)
von: Du, Zhuoyun, et al.
Veröffentlicht: (2024)
Aligning Agentic World Models via Knowledgeable Experience Learning
von: Ren, Baochang, et al.
Veröffentlicht: (2026)
von: Ren, Baochang, et al.
Veröffentlicht: (2026)
Making Language Models Better Tool Learners with Execution Feedback
von: Qiao, Shuofei, et al.
Veröffentlicht: (2023)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2023)
DarwinTOD: LLM-driven Lifelong Self-evolution for Task-oriented Dialog Systems
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
PolicySimEval: A Benchmark for Evaluating Policy Outcomes through Agent-Based Simulation
von: Kang, Jiaju, et al.
Veröffentlicht: (2025)
von: Kang, Jiaju, et al.
Veröffentlicht: (2025)
ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
von: Wu, Fangwen, et al.
Veröffentlicht: (2025)
von: Wu, Fangwen, et al.
Veröffentlicht: (2025)
Cognitive Duality for Adaptive Web Agents
von: Liu, Jiarun, et al.
Veröffentlicht: (2025)
von: Liu, Jiarun, et al.
Veröffentlicht: (2025)
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
von: Xu, Shihao, et al.
Veröffentlicht: (2026)
von: Xu, Shihao, et al.
Veröffentlicht: (2026)
Editing Personality for Large Language Models
von: Mao, Shengyu, et al.
Veröffentlicht: (2023)
von: Mao, Shengyu, et al.
Veröffentlicht: (2023)
Benchmarking LLMs' Swarm intelligence
von: Ruan, Kai, et al.
Veröffentlicht: (2025)
von: Ruan, Kai, et al.
Veröffentlicht: (2025)
APS: Bias-Controlled Adaptive Prototype Simulation for Population-Scale LLM Agents
von: Zheng, Quan, et al.
Veröffentlicht: (2026)
von: Zheng, Quan, et al.
Veröffentlicht: (2026)
InstructIE: A Bilingual Instruction-based Information Extraction Dataset
von: Gui, Honghao, et al.
Veröffentlicht: (2023)
von: Gui, Honghao, et al.
Veröffentlicht: (2023)
Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web
von: Nie, Xiaohang, et al.
Veröffentlicht: (2026)
von: Nie, Xiaohang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
von: Ou, Yixin, et al.
Veröffentlicht: (2025) -
What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations
von: Luo, Yujie, et al.
Veröffentlicht: (2025) -
Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
von: Zhu, Yuqi, et al.
Veröffentlicht: (2025) -
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
von: Qiu, Zhisong, et al.
Veröffentlicht: (2026) -
Can We Predict Before Executing Machine Learning Agents?
von: Zheng, Jingsheng, et al.
Veröffentlicht: (2026)