Interactive Evaluation Requires a Design Science
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xuan, Keyang, Song, Peiyang, Lu, Pan, Han, Pengrui, Li, Wenkai, Zhang, Zhenyu, He, Zexue, Hua, Wenyue, Li, Manling, You, Jiaxuan, Weller, Adrian, Wang, Yizhong, Pei, Jiaxin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models
par: Han, Pengrui, et autres
Publié: (2024)
par: Han, Pengrui, et autres
Publié: (2024)
Large Language Model Reasoning Failures
par: Song, Peiyang, et autres
Publié: (2026)
par: Song, Peiyang, et autres
Publié: (2026)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
par: Han, Pengrui, et autres
Publié: (2026)
par: Han, Pengrui, et autres
Publié: (2026)
AcademicEval: Live Long-Context LLM Benchmark
par: Zhang, Haozhen, et autres
Publié: (2025)
par: Zhang, Haozhen, et autres
Publié: (2025)
TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents
par: Yu, Haofei, et autres
Publié: (2025)
par: Yu, Haofei, et autres
Publié: (2025)
SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers
par: Xuan, Keyang, et autres
Publié: (2026)
par: Xuan, Keyang, et autres
Publié: (2026)
Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action
par: Pan, Zhenyu, et autres
Publié: (2024)
par: Pan, Zhenyu, et autres
Publié: (2024)
Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models
par: Pan, Zhenyu, et autres
Publié: (2024)
par: Pan, Zhenyu, et autres
Publié: (2024)
Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems
par: Feng, Tao, et autres
Publié: (2026)
par: Feng, Tao, et autres
Publié: (2026)
Paper Copilot: A Self-Evolving and Efficient LLM System for Personalized Academic Assistance
par: Lin, Guanyu, et autres
Publié: (2024)
par: Lin, Guanyu, et autres
Publié: (2024)
Sotopia-RL: Reward Design for Social Intelligence
par: Yu, Haofei, et autres
Publié: (2025)
par: Yu, Haofei, et autres
Publié: (2025)
ResearchTown: Simulator of Human Research Community
par: Yu, Haofei, et autres
Publié: (2024)
par: Yu, Haofei, et autres
Publié: (2024)
Quantifying Trust: Financial Risk Management for Trustworthy AI Agents
par: Hua, Wenyue, et autres
Publié: (2026)
par: Hua, Wenyue, et autres
Publié: (2026)
OpenP5: An Open-Source Platform for Developing, Training, and Evaluating LLM-based Recommender Systems
par: Xu, Shuyuan, et autres
Publié: (2023)
par: Xu, Shuyuan, et autres
Publié: (2023)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
par: Liang, Chumeng, et autres
Publié: (2025)
par: Liang, Chumeng, et autres
Publié: (2025)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
par: Han, Pengrui, et autres
Publié: (2025)
par: Han, Pengrui, et autres
Publié: (2025)
Why Does New Knowledge Create Messy Ripple Effects in LLMs?
par: Qin, Jiaxin, et autres
Publié: (2024)
par: Qin, Jiaxin, et autres
Publié: (2024)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
par: Feng, Tao, et autres
Publié: (2025)
par: Feng, Tao, et autres
Publié: (2025)
Defensible Design for OpenClaw: Securing Autonomous Tool-Invoking Agents
par: Li, Zongwei, et autres
Publié: (2026)
par: Li, Zongwei, et autres
Publié: (2026)
Improvement of Solid-Fluid Interaction Scheme in Lattice Boltzmann Immiscible Pseudopotential Models
par: Chen, Yizhong, et autres
Publié: (2025)
par: Chen, Yizhong, et autres
Publié: (2025)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
par: Hao, Yijie, et autres
Publié: (2025)
par: Hao, Yijie, et autres
Publié: (2025)
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
par: Hong, Zhaochen, et autres
Publié: (2025)
par: Hong, Zhaochen, et autres
Publié: (2025)
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
par: Juneja, Gurusha, et autres
Publié: (2025)
par: Juneja, Gurusha, et autres
Publié: (2025)
Zangetsu: A Candidate of Isolated, Quiescent, and Backsplash Ultra-Diffuse Galaxy in the COSMOS Field
par: Wei, Leyao, et autres
Publié: (2025)
par: Wei, Leyao, et autres
Publié: (2025)
A Multimodal, Multilingual, and Multidimensional Pipeline for Fine-grained Crowdsourcing Earthquake Damage Evaluation
par: Ma, Zihui, et autres
Publié: (2025)
par: Ma, Zihui, et autres
Publié: (2025)
Optimisation recommendations from a study on parental choice of accompaniment during children's surgery: Focus on family structure and mental health factors
par: Lianwei Zhou, et autres
Publié: (2024)
par: Lianwei Zhou, et autres
Publié: (2024)
GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation
par: Feng, Tao, et autres
Publié: (2025)
par: Feng, Tao, et autres
Publié: (2025)
From Traditional Machine Learning Models to Multimodal Large Models: A Review of Aquaculture
par: Sitao Liu, et autres
Publié: (2025)
par: Sitao Liu, et autres
Publié: (2025)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
par: Li, Xuyuan, et autres
Publié: (2023)
par: Li, Xuyuan, et autres
Publié: (2023)
LiveTradeBench: Seeking Real-World Alpha with Large Language Models
par: Yu, Haofei, et autres
Publié: (2025)
par: Yu, Haofei, et autres
Publié: (2025)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
par: Chen, Chaoran, et autres
Publié: (2025)
par: Chen, Chaoran, et autres
Publié: (2025)
GMTRouter: Personalized LLM Router over Multi-turn User Interactions
par: Xie, Encheng, et autres
Publié: (2025)
par: Xie, Encheng, et autres
Publié: (2025)
Interaction-Aware Vulnerability Detection in Smart Contract Bytecodes
par: Li, Wenkai, et autres
Publié: (2024)
par: Li, Wenkai, et autres
Publié: (2024)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
par: Wang, Qineng, et autres
Publié: (2025)
par: Wang, Qineng, et autres
Publié: (2025)
Hot Electron Dynamics Modulated by Nonequilibrium Phonon Excitations
par: Jiaxuan Xu, et autres
Publié: (2025)
par: Jiaxuan Xu, et autres
Publié: (2025)
Interact3D: Compositional 3D Generation of Interactive Objects
par: Shan, Hui, et autres
Publié: (2026)
par: Shan, Hui, et autres
Publié: (2026)
A Spiral Coverage Path Planning Algorithm for Nonomnidirectional Robots
par: Taogang Hou, et autres
Publié: (2025)
par: Taogang Hou, et autres
Publié: (2025)
Do Code LLMs Understand Design Patterns?
par: Pan, Zhenyu, et autres
Publié: (2025)
par: Pan, Zhenyu, et autres
Publié: (2025)
Modeling Public Perceptions of Science in Media
par: Pei, Jiaxin, et autres
Publié: (2025)
par: Pei, Jiaxin, et autres
Publié: (2025)
Energy spectrum of two-dimensional isotropic rapidly rotating turbulence
par: Li, Peiyang, et autres
Publié: (2024)
par: Li, Peiyang, et autres
Publié: (2024)
Documents similaires
-
In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models
par: Han, Pengrui, et autres
Publié: (2024) -
Large Language Model Reasoning Failures
par: Song, Peiyang, et autres
Publié: (2026) -
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
par: Han, Pengrui, et autres
Publié: (2026) -
AcademicEval: Live Long-Context LLM Benchmark
par: Zhang, Haozhen, et autres
Publié: (2025) -
TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents
par: Yu, Haofei, et autres
Publié: (2025)