Multi-agent AI systems outperform human teams in creativity
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Tiancheng, Jiang, Yixuan, Li, Haotian, Hernández-Orallo, José, Xie, Xing, Collier, Nigel, Stillwell, David, Sun, Luning |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News
por: Hu, Tiancheng, et al.
Publicado: (2025)
por: Hu, Tiancheng, et al.
Publicado: (2025)
Quantifying the Persona Effect in LLM Simulations
por: Hu, Tiancheng, et al.
Publicado: (2024)
por: Hu, Tiancheng, et al.
Publicado: (2024)
Evaluating General-Purpose AI with Psychometrics
por: Wang, Xiting, et al.
Publicado: (2023)
por: Wang, Xiting, et al.
Publicado: (2023)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
por: Hu, Tiancheng, et al.
Publicado: (2025)
por: Hu, Tiancheng, et al.
Publicado: (2025)
Can LLM be a Personalized Judge?
por: Dong, Yijiang River, et al.
Publicado: (2024)
por: Dong, Yijiang River, et al.
Publicado: (2024)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
por: Balkır, Esma, et al.
Publicado: (2026)
por: Balkır, Esma, et al.
Publicado: (2026)
Large Language Models show both individual and collective creativity comparable to humans
por: Sun, Luning, et al.
Publicado: (2024)
por: Sun, Luning, et al.
Publicado: (2024)
Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
por: Dong, Yijiang River, et al.
Publicado: (2026)
por: Dong, Yijiang River, et al.
Publicado: (2026)
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
por: Dong, Yijiang River, et al.
Publicado: (2025)
por: Dong, Yijiang River, et al.
Publicado: (2025)
A Survey on Prompt Tuning
por: Li, Zongqian, et al.
Publicado: (2025)
por: Li, Zongqian, et al.
Publicado: (2025)
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
por: Li, Zongqian, et al.
Publicado: (2025)
por: Li, Zongqian, et al.
Publicado: (2025)
500xCompressor: Generalized Prompt Compression for Large Language Models
por: Li, Zongqian, et al.
Publicado: (2024)
por: Li, Zongqian, et al.
Publicado: (2024)
Prompt Compression for Large Language Models: A Survey
por: Li, Zongqian, et al.
Publicado: (2024)
por: Li, Zongqian, et al.
Publicado: (2024)
Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence
por: Liu, Yinhong, et al.
Publicado: (2024)
por: Liu, Yinhong, et al.
Publicado: (2024)
Confidence Estimation for LLMs in Multi-turn Interactions
por: Zhang, Caiqi, et al.
Publicado: (2026)
por: Zhang, Caiqi, et al.
Publicado: (2026)
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
por: Zhou, Ej, et al.
Publicado: (2025)
por: Zhou, Ej, et al.
Publicado: (2025)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
por: Testini, Irene, et al.
Publicado: (2025)
por: Testini, Irene, et al.
Publicado: (2025)
Augmenting Rating-Scale Measures with Text-Derived Items Using the Information-Determined Scoring (IDS) Framework
por: Watson, Joe, et al.
Publicado: (2025)
por: Watson, Joe, et al.
Publicado: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
por: Hu, Tiancheng, et al.
Publicado: (2025)
por: Hu, Tiancheng, et al.
Publicado: (2025)
COFFEE: A Contrastive Oracle-Free Framework for Event Extraction
por: Zhang, Meiru, et al.
Publicado: (2023)
por: Zhang, Meiru, et al.
Publicado: (2023)
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
por: Xing, Tiancheng, et al.
Publicado: (2025)
por: Xing, Tiancheng, et al.
Publicado: (2025)
Value of Information: A Framework for Human-Agent Communication
por: Dong, Yijiang River, et al.
Publicado: (2026)
por: Dong, Yijiang River, et al.
Publicado: (2026)
Generative Language Models Exhibit Social Identity Biases
por: Hu, Tiancheng, et al.
Publicado: (2023)
por: Hu, Tiancheng, et al.
Publicado: (2023)
Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks
por: Li, Zongqian, et al.
Publicado: (2026)
por: Li, Zongqian, et al.
Publicado: (2026)
Time to Revist Exact Match
por: Abbood, Auss, et al.
Publicado: (2025)
por: Abbood, Auss, et al.
Publicado: (2025)
Attention Instruction: Amplifying Attention in the Middle via Prompting
por: Zhang, Meiru, et al.
Publicado: (2024)
por: Zhang, Meiru, et al.
Publicado: (2024)
TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
por: Liu, Yinhong, et al.
Publicado: (2024)
por: Liu, Yinhong, et al.
Publicado: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
por: Sivapiromrat, Sanhanat, et al.
Publicado: (2025)
por: Sivapiromrat, Sanhanat, et al.
Publicado: (2025)
Conversational Complexity for Assessing Risk in Large Language Models
por: Burden, John, et al.
Publicado: (2024)
por: Burden, John, et al.
Publicado: (2024)
Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
por: Huang, Yupan, et al.
Publicado: (2023)
por: Huang, Yupan, et al.
Publicado: (2023)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
por: Zhou, Lexin, et al.
Publicado: (2025)
por: Zhou, Lexin, et al.
Publicado: (2025)
ReasonGraph: Visualisation of Reasoning Paths
por: Li, Zongqian, et al.
Publicado: (2025)
por: Li, Zongqian, et al.
Publicado: (2025)
LUQ: Long-text Uncertainty Quantification for LLMs
por: Zhang, Caiqi, et al.
Publicado: (2024)
por: Zhang, Caiqi, et al.
Publicado: (2024)
PiVe: Prompting with Iterative Verification Improving Graph-based Generative Capability of LLMs
por: Han, Jiuzhou, et al.
Publicado: (2023)
por: Han, Jiuzhou, et al.
Publicado: (2023)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
por: Hui, Zheng, et al.
Publicado: (2025)
por: Hui, Zheng, et al.
Publicado: (2025)
Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games
por: Ma, Chengdong, et al.
Publicado: (2023)
por: Ma, Chengdong, et al.
Publicado: (2023)
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
por: Zhang, Caiqi, et al.
Publicado: (2025)
por: Zhang, Caiqi, et al.
Publicado: (2025)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
por: Fitz, Stephen, et al.
Publicado: (2025)
por: Fitz, Stephen, et al.
Publicado: (2025)
Ejemplares similares
-
iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News
por: Hu, Tiancheng, et al.
Publicado: (2025) -
Quantifying the Persona Effect in LLM Simulations
por: Hu, Tiancheng, et al.
Publicado: (2024) -
Evaluating General-Purpose AI with Psychometrics
por: Wang, Xiting, et al.
Publicado: (2023) -
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
por: Hu, Tiancheng, et al.
Publicado: (2025) -
Can LLM be a Personalized Judge?
por: Dong, Yijiang River, et al.
Publicado: (2024)