AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Hao, Cuvin, Phil, Yu, Xinkai, Yan, Charlotte Ka Yee, Zhang, Jason, Yang, Diyi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DECEPTICON: How Dark Patterns Manipulate Web Agents
di: Cuvin, Phil, et al.
Pubblicazione: (2025)
di: Cuvin, Phil, et al.
Pubblicazione: (2025)
EgoNormia: Benchmarking Physical Social Norm Understanding
di: Rezaei, MohammadHossein, et al.
Pubblicazione: (2025)
di: Rezaei, MohammadHossein, et al.
Pubblicazione: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
Generating Planning Feedback for Open-Ended Programming Exercises with LLMs
di: Demirtaş, Mehmet Arif, et al.
Pubblicazione: (2025)
di: Demirtaş, Mehmet Arif, et al.
Pubblicazione: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
di: Demchak, Nathaniel, et al.
Pubblicazione: (2024)
di: Demchak, Nathaniel, et al.
Pubblicazione: (2024)
O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
di: Mei, Jianbiao, et al.
Pubblicazione: (2025)
di: Mei, Jianbiao, et al.
Pubblicazione: (2025)
AEL: Agent Evolving Learning for Open-Ended Environments
di: Xu, Wujiang, et al.
Pubblicazione: (2026)
di: Xu, Wujiang, et al.
Pubblicazione: (2026)
Basis Vector Metric: A Method for Robust Open-Ended State Change Detection
di: Oprea, David, et al.
Pubblicazione: (2025)
di: Oprea, David, et al.
Pubblicazione: (2025)
AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
di: Li, Yishan, et al.
Pubblicazione: (2026)
di: Li, Yishan, et al.
Pubblicazione: (2026)
ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents
di: Shao, Jie-Jing, et al.
Pubblicazione: (2024)
di: Shao, Jie-Jing, et al.
Pubblicazione: (2024)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
Searching for Privacy Risks in LLM Agents via Simulation
di: Zhang, Yanzhe, et al.
Pubblicazione: (2025)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2025)
O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL
di: Yao, Yi, et al.
Pubblicazione: (2026)
di: Yao, Yi, et al.
Pubblicazione: (2026)
Reverse-Engineered Reasoning for Open-Ended Generation
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
di: Shao, Yijia, et al.
Pubblicazione: (2024)
di: Shao, Yijia, et al.
Pubblicazione: (2024)
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
MCU: An Evaluation Framework for Open-Ended Game Agents
di: Zheng, Xinyue, et al.
Pubblicazione: (2023)
di: Zheng, Xinyue, et al.
Pubblicazione: (2023)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
di: Ma, Rachel, et al.
Pubblicazione: (2025)
di: Ma, Rachel, et al.
Pubblicazione: (2025)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
di: Chen, Hui, et al.
Pubblicazione: (2025)
di: Chen, Hui, et al.
Pubblicazione: (2025)
Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations
di: Yu, Sunny, et al.
Pubblicazione: (2025)
di: Yu, Sunny, et al.
Pubblicazione: (2025)
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
di: Liu, Wanlong, et al.
Pubblicazione: (2026)
di: Liu, Wanlong, et al.
Pubblicazione: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
Towards Open-Ended Discovery for Low-Resource NLP
di: Dossou, Bonaventure F. P., et al.
Pubblicazione: (2025)
di: Dossou, Bonaventure F. P., et al.
Pubblicazione: (2025)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
di: Wang, Jize, et al.
Pubblicazione: (2026)
di: Wang, Jize, et al.
Pubblicazione: (2026)
CooperBench: Why Coding Agents Cannot be Your Teammates Yet
di: Khatua, Arpandeep, et al.
Pubblicazione: (2026)
di: Khatua, Arpandeep, et al.
Pubblicazione: (2026)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
di: Yang, Zixuan, et al.
Pubblicazione: (2026)
di: Yang, Zixuan, et al.
Pubblicazione: (2026)
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2026)
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2026)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
di: Wu, JiaRu, et al.
Pubblicazione: (2025)
di: Wu, JiaRu, et al.
Pubblicazione: (2025)
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
di: An, Subin, et al.
Pubblicazione: (2025)
di: An, Subin, et al.
Pubblicazione: (2025)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
di: Yang, Rui, et al.
Pubblicazione: (2026)
di: Yang, Rui, et al.
Pubblicazione: (2026)
Open-Ended Wargames with Large Language Models
di: Hogan, Daniel P., et al.
Pubblicazione: (2024)
di: Hogan, Daniel P., et al.
Pubblicazione: (2024)
Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation
di: Chen, Nuo, et al.
Pubblicazione: (2026)
di: Chen, Nuo, et al.
Pubblicazione: (2026)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2025)
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2025)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
di: Liu, Zijun, et al.
Pubblicazione: (2023)
di: Liu, Zijun, et al.
Pubblicazione: (2023)
GRLO: Towards Generalizable Reinforcement Learning in Open-Ended Environments from Zero
di: Yin, Shangjian, et al.
Pubblicazione: (2026)
di: Yin, Shangjian, et al.
Pubblicazione: (2026)
The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation
di: Carlsson, Fredrik, et al.
Pubblicazione: (2024)
di: Carlsson, Fredrik, et al.
Pubblicazione: (2024)
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
di: Mitsides, Konstantinos, et al.
Pubblicazione: (2026)
di: Mitsides, Konstantinos, et al.
Pubblicazione: (2026)
Learning Personalized Agents from Human Feedback
di: Liang, Kaiqu, et al.
Pubblicazione: (2026)
di: Liang, Kaiqu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DECEPTICON: How Dark Patterns Manipulate Web Agents
di: Cuvin, Phil, et al.
Pubblicazione: (2025) -
EgoNormia: Benchmarking Physical Social Norm Understanding
di: Rezaei, MohammadHossein, et al.
Pubblicazione: (2025) -
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
di: Ryan, Michael J., et al.
Pubblicazione: (2025) -
Generating Planning Feedback for Open-Ended Programming Exercises with LLMs
di: Demirtaş, Mehmet Arif, et al.
Pubblicazione: (2025) -
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
di: Demchak, Nathaniel, et al.
Pubblicazione: (2024)