Position: Stop Making Unscientific AGI Performance Claims
Fuente:
arXiv
Saved in:
| Main Authors: | Altmeyer, Patrick, Demetriou, Andrew M., Bartlett, Antony, Liem, Cynthia C. S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Last Dependency Crusade: Solving Python Dependency Conflicts with LLMs
by: Bartlett, Antony, et al.
Published: (2025)
by: Bartlett, Antony, et al.
Published: (2025)
Counterfactual Training: Teaching Models Plausible and Actionable Explanations
by: Altmeyer, Patrick, et al.
Published: (2026)
by: Altmeyer, Patrick, et al.
Published: (2026)
Towards Estimating Personal Values in Song Lyrics
by: Demetriou, Andrew M., et al.
Published: (2024)
by: Demetriou, Andrew M., et al.
Published: (2024)
Heuristics and Biases in AI Decision-Making: Implications for Responsible AGI
by: Saeedi, Payam, et al.
Published: (2024)
by: Saeedi, Payam, et al.
Published: (2024)
Position: Stop Acting Like Language Model Agents Are Normal Agents
by: Perrier, Elija, et al.
Published: (2025)
by: Perrier, Elija, et al.
Published: (2025)
Emergence is Overrated: AGI as an Archipelago of Experts
by: Kilov, Daniel
Published: (2026)
by: Kilov, Daniel
Published: (2026)
ARC-AGI-2 Technical Report
by: de Oliveira, Wallyson Lemes, et al.
Published: (2026)
by: de Oliveira, Wallyson Lemes, et al.
Published: (2026)
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
by: Anik, Anirban Saha, et al.
Published: (2025)
by: Anik, Anirban Saha, et al.
Published: (2025)
LightHouse: A Survey of AGI Hallucination
by: Wang, Feng
Published: (2024)
by: Wang, Feng
Published: (2024)
Improving AGI Evaluation: A Data Science Perspective
by: Hawkins, John
Published: (2025)
by: Hawkins, John
Published: (2025)
AI-native Memory: A Pathway from LLMs Towards AGI
by: Shang, Jingbo, et al.
Published: (2024)
by: Shang, Jingbo, et al.
Published: (2024)
Towards AGI A Pragmatic Approach Towards Self Evolving Agent
by: Kar, Indrajit, et al.
Published: (2026)
by: Kar, Indrajit, et al.
Published: (2026)
ClaimGen-CN: A Large-scale Chinese Dataset for Legal Claim Generation
by: Zhou, Siying, et al.
Published: (2025)
by: Zhou, Siying, et al.
Published: (2025)
The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz
by: Noever, David, et al.
Published: (2024)
by: Noever, David, et al.
Published: (2024)
Data Science and Technology Towards AGI Part I: Tiered Data Management
by: Wang, Yudong, et al.
Published: (2026)
by: Wang, Yudong, et al.
Published: (2026)
ClaimBrush: A Novel Framework for Automated Patent Claim Refinement Based on Large Language Models
by: Kawano, Seiya, et al.
Published: (2024)
by: Kawano, Seiya, et al.
Published: (2024)
Optimizing Decomposition for Optimal Claim Verification
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
by: Majurski, Michael, et al.
Published: (2026)
by: Majurski, Michael, et al.
Published: (2026)
Robust Claim Verification Through Fact Detection
by: Jafari, Nazanin, et al.
Published: (2024)
by: Jafari, Nazanin, et al.
Published: (2024)
The Alignment Bottleneck in Decomposition-Based Claim Verification
by: Akhter, Mahmud Elahi, et al.
Published: (2026)
by: Akhter, Mahmud Elahi, et al.
Published: (2026)
Adaptive Stopping for Multi-Turn LLM Reasoning
by: Zhou, Xiaofan, et al.
Published: (2026)
by: Zhou, Xiaofan, et al.
Published: (2026)
Foundation Models to Unlock Real-World Evidence from Nationwide Medical Claims
by: Ma, Fan, et al.
Published: (2026)
by: Ma, Fan, et al.
Published: (2026)
ClaimPKG: Enhancing Claim Verification via Pseudo-Subgraph Generation with Lightweight Specialized LLM
by: Pham, Hoang, et al.
Published: (2025)
by: Pham, Hoang, et al.
Published: (2025)
Towards Efficient Neurally-Guided Program Induction for ARC-AGI
by: Ouellette, Simon
Published: (2024)
by: Ouellette, Simon
Published: (2024)
Fine-grained Claim-level RAG Benchmark for Law
by: Das, Souvick, et al.
Published: (2026)
by: Das, Souvick, et al.
Published: (2026)
Accelerate Creation of Product Claims Using Generative AI
by: Liang, Po-Yu, et al.
Published: (2025)
by: Liang, Po-Yu, et al.
Published: (2025)
Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?
by: Hu, Qisheng, et al.
Published: (2024)
by: Hu, Qisheng, et al.
Published: (2024)
SciClaims: An End-to-End Generative System for Biomedical Claim Analysis
by: Ortega, Raúl, et al.
Published: (2025)
by: Ortega, Raúl, et al.
Published: (2025)
Piecing It All Together: Verifying Multi-Hop Multimodal Claims
by: Wang, Haoran, et al.
Published: (2024)
by: Wang, Haoran, et al.
Published: (2024)
Claim Verification in the Age of Large Language Models: A Survey
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
A Claim Decomposition Benchmark for Long-form Answer Verification
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak
by: Makaiova, Lucia, et al.
Published: (2025)
by: Makaiova, Lucia, et al.
Published: (2025)
Agent-based Automated Claim Matching with Instruction-following LLMs
by: Pisarevskaya, Dina, et al.
Published: (2025)
by: Pisarevskaya, Dina, et al.
Published: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
by: Sundriyal, Megha, et al.
Published: (2023)
by: Sundriyal, Megha, et al.
Published: (2023)
Entropic Claim Resolution: Uncertainty-Driven Evidence Selection for RAG
by: Di Gioia, Davide
Published: (2026)
by: Di Gioia, Davide
Published: (2026)
Evergreen: Efficient Claim Verification for Semantic Aggregates
by: Lee, Alexander W., et al.
Published: (2026)
by: Lee, Alexander W., et al.
Published: (2026)
Ta'keed: The First Generative Fact-Checking System for Arabic Claims
by: Althabiti, Saud, et al.
Published: (2024)
by: Althabiti, Saud, et al.
Published: (2024)
Classifying Human-Generated and AI-Generated Election Claims in Social Media
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
by: Liu, Muxin, et al.
Published: (2026)
by: Liu, Muxin, et al.
Published: (2026)
Step-by-Step Fact Verification System for Medical Claims with Explainable Reasoning
by: Vladika, Juraj, et al.
Published: (2025)
by: Vladika, Juraj, et al.
Published: (2025)
Similar Items
-
The Last Dependency Crusade: Solving Python Dependency Conflicts with LLMs
by: Bartlett, Antony, et al.
Published: (2025) -
Counterfactual Training: Teaching Models Plausible and Actionable Explanations
by: Altmeyer, Patrick, et al.
Published: (2026) -
Towards Estimating Personal Values in Song Lyrics
by: Demetriou, Andrew M., et al.
Published: (2024) -
Heuristics and Biases in AI Decision-Making: Implications for Responsible AGI
by: Saeedi, Payam, et al.
Published: (2024) -
Position: Stop Acting Like Language Model Agents Are Normal Agents
by: Perrier, Elija, et al.
Published: (2025)