Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
Fuente:
arXiv
Guardado en:
| Autores principales: | Pfister, Rolf, Jud, Hansueli |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
por: Xu, Pusheng, et al.
Publicado: (2025)
por: Xu, Pusheng, et al.
Publicado: (2025)
A Representationalist, Functionalist and Naturalistic Conception of Intelligence as a Foundation for AGI
por: Pfister, Rolf
Publicado: (2025)
por: Pfister, Rolf
Publicado: (2025)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
por: Kumar, Deepak, et al.
Publicado: (2025)
por: Kumar, Deepak, et al.
Publicado: (2025)
Counting and Algorithmic Generalization with Transformers
por: Ouellette, Simon, et al.
Publicado: (2023)
por: Ouellette, Simon, et al.
Publicado: (2023)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
por: Arantes, Gabriel M., et al.
Publicado: (2025)
por: Arantes, Gabriel M., et al.
Publicado: (2025)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
por: Cai, Yanan, et al.
Publicado: (2025)
por: Cai, Yanan, et al.
Publicado: (2025)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
por: Barad, Haim, et al.
Publicado: (2023)
por: Barad, Haim, et al.
Publicado: (2023)
PixLift: Accelerating Web Browsing via AI Upscaling
por: Atinafu, Yonas, et al.
Publicado: (2025)
por: Atinafu, Yonas, et al.
Publicado: (2025)
XTC, A Research Platform for Optimizing AI Workload Operators
por: Hugo, Pompougnac, et al.
Publicado: (2025)
por: Hugo, Pompougnac, et al.
Publicado: (2025)
Learning, Potential, and Retention: An Approach for Evaluating Adaptive AI-Enabled Medical Devices
por: Burgon, Alexis, et al.
Publicado: (2026)
por: Burgon, Alexis, et al.
Publicado: (2026)
Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage
por: Ngabonziza, Bernard, et al.
Publicado: (2026)
por: Ngabonziza, Bernard, et al.
Publicado: (2026)
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
por: Atinafu, Yonas, et al.
Publicado: (2026)
por: Atinafu, Yonas, et al.
Publicado: (2026)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
por: Zhao, Qihao, et al.
Publicado: (2024)
por: Zhao, Qihao, et al.
Publicado: (2024)
Reliability by design: quantifying and eliminating fabrication risk in LLMs. From generative to consultative AI: a comparative analysis in the legal domain and lessons for high-stakes knowledge bases
por: Dantart, Alex
Publicado: (2026)
por: Dantart, Alex
Publicado: (2026)
OpenAI o1 System Card
por: OpenAI, et al.
Publicado: (2024)
por: OpenAI, et al.
Publicado: (2024)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
por: Jung, Alexander Louis-Ferdinand, et al.
Publicado: (2024)
por: Jung, Alexander Louis-Ferdinand, et al.
Publicado: (2024)
Artificial Intelligence and its Impact on Academic Performance of students with Disabilities in Nasarawa State University, Keffi
por: Osita, Emmanuel Izuchukwu, et al.
Publicado: (2026)
por: Osita, Emmanuel Izuchukwu, et al.
Publicado: (2026)
Deploying Open-Source Large Language Models: A performance Analysis
por: Bendi-Ouis, Yannis, et al.
Publicado: (2024)
por: Bendi-Ouis, Yannis, et al.
Publicado: (2024)
On the Sustainability of AI Inferences in the Edge
por: Sobhani, Ghazal, et al.
Publicado: (2025)
por: Sobhani, Ghazal, et al.
Publicado: (2025)
Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial
por: Cortes, David, et al.
Publicado: (2025)
por: Cortes, David, et al.
Publicado: (2025)
Photonic Fabric Platform for AI Accelerators
por: Ding, Jing, et al.
Publicado: (2025)
por: Ding, Jing, et al.
Publicado: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
por: Liu, Jiashuo, et al.
Publicado: (2025)
por: Liu, Jiashuo, et al.
Publicado: (2025)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
por: Rani, Pooja, et al.
Publicado: (2025)
por: Rani, Pooja, et al.
Publicado: (2025)
Looking Forward: Challenges and Opportunities in Agentic AI Reliability
por: Xing, Liudong, et al.
Publicado: (2025)
por: Xing, Liudong, et al.
Publicado: (2025)
Information Retrieval in the Age of Generative AI: The RGB Model
por: Garetto, Michele, et al.
Publicado: (2025)
por: Garetto, Michele, et al.
Publicado: (2025)
The Race to Efficiency: A New Perspective on AI Scaling Laws
por: Lu, Chien-Ping
Publicado: (2025)
por: Lu, Chien-Ping
Publicado: (2025)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
por: Georganas, Evangelos, et al.
Publicado: (2025)
por: Georganas, Evangelos, et al.
Publicado: (2025)
Revolutionizing System Reliability: The Role of AI in Predictive Maintenance Strategies
por: Bidollahkhani, Michael, et al.
Publicado: (2024)
por: Bidollahkhani, Michael, et al.
Publicado: (2024)
AutoLALA: Automatic Loop Algebraic Locality Analysis for AI and HPC Kernels
por: Zhu, Yifan, et al.
Publicado: (2026)
por: Zhu, Yifan, et al.
Publicado: (2026)
MoEITS: A Green AI approach for simplifying MoE-LLMs
por: Balderas, Luis, et al.
Publicado: (2026)
por: Balderas, Luis, et al.
Publicado: (2026)
AGI: Artificial General Intelligence for Education
por: Latif, Ehsan, et al.
Publicado: (2023)
por: Latif, Ehsan, et al.
Publicado: (2023)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
por: Almurshed, Osama, et al.
Publicado: (2025)
por: Almurshed, Osama, et al.
Publicado: (2025)
Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
por: Panigrahy, Deepak, et al.
Publicado: (2026)
por: Panigrahy, Deepak, et al.
Publicado: (2026)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
por: Ivanov, Andrei, et al.
Publicado: (2025)
por: Ivanov, Andrei, et al.
Publicado: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
por: Prieto, Pablo, et al.
Publicado: (2025)
por: Prieto, Pablo, et al.
Publicado: (2025)
Tiny-QMoE
por: Cashman, Jack, et al.
Publicado: (2025)
por: Cashman, Jack, et al.
Publicado: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
por: Chai, Yuji, et al.
Publicado: (2025)
por: Chai, Yuji, et al.
Publicado: (2025)
WANDER: An Explainable Decision-Support Framework for HPC
por: Lahiry, Ankur, et al.
Publicado: (2025)
por: Lahiry, Ankur, et al.
Publicado: (2025)
Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
por: Zhao, Youpeng, et al.
Publicado: (2025)
por: Zhao, Youpeng, et al.
Publicado: (2025)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Ejemplares similares
-
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
por: Xu, Pusheng, et al.
Publicado: (2025) -
A Representationalist, Functionalist and Naturalistic Conception of Intelligence as a Foundation for AGI
por: Pfister, Rolf
Publicado: (2025) -
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
por: Kumar, Deepak, et al.
Publicado: (2025) -
Counting and Algorithmic Generalization with Transformers
por: Ouellette, Simon, et al.
Publicado: (2023) -
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
por: Arantes, Gabriel M., et al.
Publicado: (2025)