Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
Fuente:
arXiv
Saved in:
| Main Authors: | Schoenegger, Philipp, Tuminauskaite, Indre, Park, Peter S., Tetlock, Philip E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
by: Abels, Axel, et al.
Published: (2025)
by: Abels, Axel, et al.
Published: (2025)
Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition
by: Akben, Mustafa, et al.
Published: (2025)
by: Akben, Mustafa, et al.
Published: (2025)
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
by: Liu, Shang, et al.
Published: (2024)
by: Liu, Shang, et al.
Published: (2024)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
by: Ni, Jinjie, et al.
Published: (2024)
by: Ni, Jinjie, et al.
Published: (2024)
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
by: Li, Yisen, et al.
Published: (2025)
by: Li, Yisen, et al.
Published: (2025)
Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations
by: Wang, Yunzhe, et al.
Published: (2025)
by: Wang, Yunzhe, et al.
Published: (2025)
Prompt Engineering Large Language Models' Forecasting Capabilities
by: Schoenegger, Philipp, et al.
Published: (2025)
by: Schoenegger, Philipp, et al.
Published: (2025)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
by: Karger, Ezra, et al.
Published: (2024)
by: Karger, Ezra, et al.
Published: (2024)
LLMs Can Teach Themselves to Better Predict the Future
by: Turtel, Benjamin, et al.
Published: (2025)
by: Turtel, Benjamin, et al.
Published: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
by: Husom, Erik Johannes, et al.
Published: (2025)
by: Husom, Erik Johannes, et al.
Published: (2025)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
by: Bojic, Ljubisa, et al.
Published: (2026)
by: Bojic, Ljubisa, et al.
Published: (2026)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
by: Kwon, Jea, et al.
Published: (2025)
by: Kwon, Jea, et al.
Published: (2025)
LLMCO2: Advancing Accurate Carbon Footprint Prediction for LLM Inferences
by: Fu, Zhenxiao, et al.
Published: (2024)
by: Fu, Zhenxiao, et al.
Published: (2024)
Improving Academic Skills Assessment with NLP and Ensemble Learning
by: Huang, Xinyi, et al.
Published: (2024)
by: Huang, Xinyi, et al.
Published: (2024)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
by: Ren, Richard, et al.
Published: (2025)
by: Ren, Richard, et al.
Published: (2025)
Exploring Accuracy-Fairness Trade-off in Large Language Models
by: Zhang, Qingquan, et al.
Published: (2024)
by: Zhang, Qingquan, et al.
Published: (2024)
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
by: Zhang, Jifan, et al.
Published: (2024)
by: Zhang, Jifan, et al.
Published: (2024)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction
by: Saha, Anisha, et al.
Published: (2025)
by: Saha, Anisha, et al.
Published: (2025)
Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding
by: Chuang, Yun-Shiuan, et al.
Published: (2025)
by: Chuang, Yun-Shiuan, et al.
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
That's So FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral
by: Steenhuis, Quinten
Published: (2025)
by: Steenhuis, Quinten
Published: (2025)
Human-LLM Coevolution: Evidence from Academic Writing
by: Geng, Mingmeng, et al.
Published: (2025)
by: Geng, Mingmeng, et al.
Published: (2025)
On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
by: Geng, Mingmeng, et al.
Published: (2025)
by: Geng, Mingmeng, et al.
Published: (2025)
Learning to Trust the Crowd: A Multi-Model Consensus Reasoning Engine for Large Language Models
by: Kallem, Pranav
Published: (2026)
by: Kallem, Pranav
Published: (2026)
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
by: Lerner, Emilia Agis, et al.
Published: (2024)
by: Lerner, Emilia Agis, et al.
Published: (2024)
A Multi-LLM Debiasing Framework
by: Owens, Deonna M., et al.
Published: (2024)
by: Owens, Deonna M., et al.
Published: (2024)
The Wisdom of Partisan Crowds: Comparing Collective Intelligence in Humans and LLM-based Agents
by: Chuang, Yun-Shiuan, et al.
Published: (2023)
by: Chuang, Yun-Shiuan, et al.
Published: (2023)
Towards Best Practices for Open Datasets for LLM Training
by: Baack, Stefan, et al.
Published: (2025)
by: Baack, Stefan, et al.
Published: (2025)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
by: Deng, Wenlong, et al.
Published: (2024)
by: Deng, Wenlong, et al.
Published: (2024)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
by: Saha, Rounak, et al.
Published: (2026)
by: Saha, Rounak, et al.
Published: (2026)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
by: Shojaei, Mostafa Faghih, et al.
Published: (2025)
by: Shojaei, Mostafa Faghih, et al.
Published: (2025)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
by: Do, Heejin, et al.
Published: (2026)
by: Do, Heejin, et al.
Published: (2026)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
by: Zhang, Andy K., et al.
Published: (2024)
by: Zhang, Andy K., et al.
Published: (2024)
Similar Items
-
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024) -
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
by: Abels, Axel, et al.
Published: (2025) -
Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition
by: Akben, Mustafa, et al.
Published: (2025) -
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
by: Liu, Shang, et al.
Published: (2024) -
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
by: Ni, Jinjie, et al.
Published: (2024)