Measuring the metacognition of AI
Fuente:
arXiv
Salvato in:
| Autori principali: | Servajean, Richard, Servajean, Philippe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GeoPl@ntNet: A Platform for Exploring Essential Biodiversity Variables
di: Picek, Lukas, et al.
Pubblicazione: (2025)
di: Picek, Lukas, et al.
Pubblicazione: (2025)
How to Optimize Multispecies Set Predictions in Presence-Absence Modeling ?
di: Gigot--Léandri, Sébastien, et al.
Pubblicazione: (2026)
di: Gigot--Léandri, Sébastien, et al.
Pubblicazione: (2026)
Impact of population size on early adaptation in rugged fitness landscapes
di: Servajean, Richard, et al.
Pubblicazione: (2022)
di: Servajean, Richard, et al.
Pubblicazione: (2022)
Mapping biodiversity at very-high resolution in Europe
di: Leblanc, César, et al.
Pubblicazione: (2025)
di: Leblanc, César, et al.
Pubblicazione: (2025)
Impact of complex spatial population structure on early and long-term adaptation in rugged fitness landscapes
di: Servajean, Richard, et al.
Pubblicazione: (2024)
di: Servajean, Richard, et al.
Pubblicazione: (2024)
Imagining and building wise machines: The centrality of AI metacognition
di: Johnson, Samuel G. B., et al.
Pubblicazione: (2024)
di: Johnson, Samuel G. B., et al.
Pubblicazione: (2024)
Before you <think>, monitor: Implementing Flavell's metacognitive framework in LLMs
di: Oh, Nick
Pubblicazione: (2025)
di: Oh, Nick
Pubblicazione: (2025)
AI-based Mapping of the Conservation Status of Orchid Assemblages at Global Scale
di: Estopinan, Joaquim, et al.
Pubblicazione: (2024)
di: Estopinan, Joaquim, et al.
Pubblicazione: (2024)
Could you be wrong: Debiasing LLMs using a metacognitive prompt for improving human decision making
di: Hills, Thomas T.
Pubblicazione: (2025)
di: Hills, Thomas T.
Pubblicazione: (2025)
Generative AI as a metacognitive agent: A comparative mixed-method study with human participants on ICF-mimicking exam performance
di: Pavlovic, Jelena, et al.
Pubblicazione: (2024)
di: Pavlovic, Jelena, et al.
Pubblicazione: (2024)
Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees
di: Greenberg, Craig, et al.
Pubblicazione: (2025)
di: Greenberg, Craig, et al.
Pubblicazione: (2025)
Modelling Species Distributions with Deep Learning to Predict Plant Extinction Risk and Assess Climate Change Impacts
di: Estopinan, Joaquim, et al.
Pubblicazione: (2024)
di: Estopinan, Joaquim, et al.
Pubblicazione: (2024)
Measuring AI Alignment with Human Flourishing
di: Hilliard, Elizabeth, et al.
Pubblicazione: (2025)
di: Hilliard, Elizabeth, et al.
Pubblicazione: (2025)
Measuring What Matters: The AI Pluralism Index
di: Mushkani, Rashid
Pubblicazione: (2025)
di: Mushkani, Rashid
Pubblicazione: (2025)
Open-World Evaluations for Measuring Frontier AI Capabilities
di: Kapoor, Sayash, et al.
Pubblicazione: (2026)
di: Kapoor, Sayash, et al.
Pubblicazione: (2026)
Measuring the environmental impact of delivering AI at Google Scale
di: Elsworth, Cooper, et al.
Pubblicazione: (2025)
di: Elsworth, Cooper, et al.
Pubblicazione: (2025)
Don't Measure Once: Measuring Visibility in AI Search (GEO)
di: Schulte, Julius, et al.
Pubblicazione: (2026)
di: Schulte, Julius, et al.
Pubblicazione: (2026)
Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems
di: Johnson, Rebecca L.
Pubblicazione: (2026)
di: Johnson, Rebecca L.
Pubblicazione: (2026)
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
di: Goldfeder, Judah, et al.
Pubblicazione: (2026)
di: Goldfeder, Judah, et al.
Pubblicazione: (2026)
Measuring AI R&D Automation
di: Chan, Alan, et al.
Pubblicazione: (2026)
di: Chan, Alan, et al.
Pubblicazione: (2026)
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
di: Desikan, Prasanna, et al.
Pubblicazione: (2026)
di: Desikan, Prasanna, et al.
Pubblicazione: (2026)
Intentionality is a Design Decision: Measuring Functional Intentionality for Accountable AI Systems
di: Chiappetta, Allessia, et al.
Pubblicazione: (2026)
di: Chiappetta, Allessia, et al.
Pubblicazione: (2026)
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
di: Grey, Markov, et al.
Pubblicazione: (2025)
di: Grey, Markov, et al.
Pubblicazione: (2025)
Measuring AI agent autonomy: Towards a scalable approach with code inspection
di: Cihon, Peter, et al.
Pubblicazione: (2025)
di: Cihon, Peter, et al.
Pubblicazione: (2025)
Measuring AI Reasoning: A Guide for Researchers
di: Nwadike, Munachiso Samuel, et al.
Pubblicazione: (2026)
di: Nwadike, Munachiso Samuel, et al.
Pubblicazione: (2026)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2026)
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2026)
Advancing Explainable AI Toward Human-Like Intelligence: Forging the Path to Artificial Brain
di: Zhou, Yongchen, et al.
Pubblicazione: (2024)
di: Zhou, Yongchen, et al.
Pubblicazione: (2024)
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation
di: von Däniken, Pius, et al.
Pubblicazione: (2024)
di: von Däniken, Pius, et al.
Pubblicazione: (2024)
Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
di: Fischer, Raphael
Pubblicazione: (2025)
di: Fischer, Raphael
Pubblicazione: (2025)
Voice-Enabled AI Agents can Perform Common Scams
di: Fang, Richard, et al.
Pubblicazione: (2024)
di: Fang, Richard, et al.
Pubblicazione: (2024)
Measuring AI Ability to Complete Long Software Tasks
di: Kwa, Thomas, et al.
Pubblicazione: (2025)
di: Kwa, Thomas, et al.
Pubblicazione: (2025)
Feedback Forensics: A Toolkit to Measure AI Personality
di: Findeis, Arduin, et al.
Pubblicazione: (2025)
di: Findeis, Arduin, et al.
Pubblicazione: (2025)
Using AI to Measure Parkinson's Disease Severity at Home
di: Islam, Md Saiful, et al.
Pubblicazione: (2023)
di: Islam, Md Saiful, et al.
Pubblicazione: (2023)
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
di: Shah, Cheril, et al.
Pubblicazione: (2025)
di: Shah, Cheril, et al.
Pubblicazione: (2025)
Measuring AI Diffusion: A Population-Normalized Metric for Tracking Global AI Usage
di: Misra, Amit, et al.
Pubblicazione: (2025)
di: Misra, Amit, et al.
Pubblicazione: (2025)
Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure
di: Yin, Michelle, et al.
Pubblicazione: (2026)
di: Yin, Michelle, et al.
Pubblicazione: (2026)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
di: Young, Richard J.
Pubblicazione: (2026)
di: Young, Richard J.
Pubblicazione: (2026)
AI in Money Matters
di: Tchatchoua, Nadine Sandjo, et al.
Pubblicazione: (2025)
di: Tchatchoua, Nadine Sandjo, et al.
Pubblicazione: (2025)
Applying the maximum entropy principle to neural networks enhances multi‐species distribution models
di: Maxime Ryckewaert, et al.
Pubblicazione: (2026)
di: Maxime Ryckewaert, et al.
Pubblicazione: (2026)
Documenti analoghi
-
GeoPl@ntNet: A Platform for Exploring Essential Biodiversity Variables
di: Picek, Lukas, et al.
Pubblicazione: (2025) -
How to Optimize Multispecies Set Predictions in Presence-Absence Modeling ?
di: Gigot--Léandri, Sébastien, et al.
Pubblicazione: (2026) -
Impact of population size on early adaptation in rugged fitness landscapes
di: Servajean, Richard, et al.
Pubblicazione: (2022) -
Mapping biodiversity at very-high resolution in Europe
di: Leblanc, César, et al.
Pubblicazione: (2025) -
Impact of complex spatial population structure on early and long-term adaptation in rugged fitness landscapes
di: Servajean, Richard, et al.
Pubblicazione: (2024)