AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Davide, Fabrizio, Torre, Pietro, Ercolani, Leonardo, Gaggioli, Andrea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
von: Mullens, Drake, et al.
Veröffentlicht: (2026)
von: Mullens, Drake, et al.
Veröffentlicht: (2026)
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
von: Jin, Ming
Veröffentlicht: (2026)
von: Jin, Ming
Veröffentlicht: (2026)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
von: Ivanov, Igor
Veröffentlicht: (2025)
von: Ivanov, Igor
Veröffentlicht: (2025)
Can LLMs Do Rocket Science? Exploring the Limits of Complex Reasoning with GTOC 12
von: del Campo, Iñaki, et al.
Veröffentlicht: (2026)
von: del Campo, Iñaki, et al.
Veröffentlicht: (2026)
Evaluating Relational Reasoning in LLMs with REL
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
von: Christop, Iwona, et al.
Veröffentlicht: (2026)
von: Christop, Iwona, et al.
Veröffentlicht: (2026)
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
von: Redkar, Neel
Veröffentlicht: (2024)
von: Redkar, Neel
Veröffentlicht: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
von: Xu, Chenjun, et al.
Veröffentlicht: (2025)
von: Xu, Chenjun, et al.
Veröffentlicht: (2025)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
von: Rai, Daking, et al.
Veröffentlicht: (2024)
von: Rai, Daking, et al.
Veröffentlicht: (2024)
Evaluating LLM Metrics Through Real-World Capabilities
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
Argumentatively Coherent Judgmental Forecasting
von: Gorur, Deniz, et al.
Veröffentlicht: (2025)
von: Gorur, Deniz, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
von: Preuveneers, Jack, et al.
Veröffentlicht: (2025)
von: Preuveneers, Jack, et al.
Veröffentlicht: (2025)
KemenkeuGPT: Leveraging a Large Language Model on Indonesia's Government Financial Data and Regulations to Enhance Decision Making
von: Febrian, Gilang Fajar, et al.
Veröffentlicht: (2024)
von: Febrian, Gilang Fajar, et al.
Veröffentlicht: (2024)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
von: Sun, Xiangkun, et al.
Veröffentlicht: (2026)
von: Sun, Xiangkun, et al.
Veröffentlicht: (2026)
KACE: Knowledge-Adaptive Context Engineering for Mathematical Reasoning
von: Parashar, Jayant, et al.
Veröffentlicht: (2026)
von: Parashar, Jayant, et al.
Veröffentlicht: (2026)
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
von: Gu, Yu, et al.
Veröffentlicht: (2024)
von: Gu, Yu, et al.
Veröffentlicht: (2024)
PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction
von: Guo, Xu
Veröffentlicht: (2025)
von: Guo, Xu
Veröffentlicht: (2025)
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
RAudit: A Blind Auditing Protocol for Large Language Model Reasoning
von: Chang, Edward Y., et al.
Veröffentlicht: (2026)
von: Chang, Edward Y., et al.
Veröffentlicht: (2026)
From Extraction to Synthesis: Entangled Heuristics for Agent-Augmented Strategic Reasoning
von: Ghisellini, Renato, et al.
Veröffentlicht: (2025)
von: Ghisellini, Renato, et al.
Veröffentlicht: (2025)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
von: Pathak, Gangesh, et al.
Veröffentlicht: (2025)
von: Pathak, Gangesh, et al.
Veröffentlicht: (2025)
MapAgent: A Hierarchical Agent for Geospatial Reasoning with Dynamic Map Tool Integration
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2025)
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2025)
The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
von: Saha, Soumadeep, et al.
Veröffentlicht: (2025)
von: Saha, Soumadeep, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment
von: Su, Ruoxi, et al.
Veröffentlicht: (2026)
von: Su, Ruoxi, et al.
Veröffentlicht: (2026)
CausalT5K: Diagnosing and Informing Refusal for Trustworthy Causal Reasoning of Skepticism, Sycophancy, Detection-Correction, and Rung Collapse
von: Geng, Longling, et al.
Veröffentlicht: (2026)
von: Geng, Longling, et al.
Veröffentlicht: (2026)
PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI
von: Chaitanya, Keshava, et al.
Veröffentlicht: (2026)
von: Chaitanya, Keshava, et al.
Veröffentlicht: (2026)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
von: Chen, Haolin, et al.
Veröffentlicht: (2024)
von: Chen, Haolin, et al.
Veröffentlicht: (2024)
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
von: Eisenstadt, Roy, et al.
Veröffentlicht: (2025)
von: Eisenstadt, Roy, et al.
Veröffentlicht: (2025)
Learning Software Bug Reports: A Systematic Literature Review
von: Long, Guoming, et al.
Veröffentlicht: (2025)
von: Long, Guoming, et al.
Veröffentlicht: (2025)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
von: Pan, Xinghan
Veröffentlicht: (2025)
von: Pan, Xinghan
Veröffentlicht: (2025)
Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
von: Bucher, Martin Juan José, et al.
Veröffentlicht: (2024)
von: Bucher, Martin Juan José, et al.
Veröffentlicht: (2024)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
von: Mullens, Drake, et al.
Veröffentlicht: (2026) -
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
von: Jin, Ming
Veröffentlicht: (2026) -
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
von: Ivanov, Igor
Veröffentlicht: (2025) -
Can LLMs Do Rocket Science? Exploring the Limits of Complex Reasoning with GTOC 12
von: del Campo, Iñaki, et al.
Veröffentlicht: (2026) -
Evaluating Relational Reasoning in LLMs with REL
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)