Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Gringras, David, Salahshoor, Misha |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024)
by: Kanepajs, Artūrs, et al.
Published: (2024)
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
AuditWen:An Open-Source Large Language Model for Audit
by: Huang, Jiajia, et al.
Published: (2024)
by: Huang, Jiajia, et al.
Published: (2024)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
by: Maurya, Sneha, et al.
Published: (2026)
by: Maurya, Sneha, et al.
Published: (2026)
AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course
by: Woo, David James, et al.
Published: (2026)
by: Woo, David James, et al.
Published: (2026)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models
by: Venkata, Pruthvinath Jeripity
Published: (2026)
by: Venkata, Pruthvinath Jeripity
Published: (2026)
MIRA: A Bilingual Benchmark for Medical Information Response Audit
by: Xu, Mengyu, et al.
Published: (2026)
by: Xu, Mengyu, et al.
Published: (2026)
PRISM: A Methodology for Auditing Biases in Large Language Models
by: Azzopardi, Leif, et al.
Published: (2024)
by: Azzopardi, Leif, et al.
Published: (2024)
Human-AI Collaboration or Academic Misconduct? Measuring AI Use in Student Writing Through Stylometric Evidence
by: Oliveira, Eduardo Araujo, et al.
Published: (2025)
by: Oliveira, Eduardo Araujo, et al.
Published: (2025)
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
by: Najjar, Ayat A., et al.
Published: (2025)
by: Najjar, Ayat A., et al.
Published: (2025)
Evaluation Framework for AI Systems in "the Wild"
by: Jabbour, Sarah, et al.
Published: (2025)
by: Jabbour, Sarah, et al.
Published: (2025)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
Frontier AI systems have surpassed the self-replicating red line
by: Pan, Xudong, et al.
Published: (2024)
by: Pan, Xudong, et al.
Published: (2024)
AuditGPT: Auditing Smart Contracts with ChatGPT
by: Xia, Shihao, et al.
Published: (2024)
by: Xia, Shihao, et al.
Published: (2024)
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI
by: Gautam, Sanjana, et al.
Published: (2024)
by: Gautam, Sanjana, et al.
Published: (2024)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review
by: Weng, Yixuan, et al.
Published: (2026)
by: Weng, Yixuan, et al.
Published: (2026)
The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events
by: Gunjan, et al.
Published: (2026)
by: Gunjan, et al.
Published: (2026)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
Human-Centred LLM Privacy Audits: Findings and Frictions
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
by: Qiu, Peiran, et al.
Published: (2025)
by: Qiu, Peiran, et al.
Published: (2025)
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
by: Lee, Kyeongryul, et al.
Published: (2025)
by: Lee, Kyeongryul, et al.
Published: (2025)
PAIR-SAFE: A Paired-Agent Approach for Runtime Auditing and Refining AI-Mediated Mental Health Support
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study
by: Silva, Jhessica, et al.
Published: (2025)
by: Silva, Jhessica, et al.
Published: (2025)
Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education
by: Chamberland, Jean-Francois, et al.
Published: (2025)
by: Chamberland, Jean-Francois, et al.
Published: (2025)
Measuring Political Preferences in AI Systems: An Integrative Approach
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
by: Shao, Yijia, et al.
Published: (2025)
by: Shao, Yijia, et al.
Published: (2025)
The Biggest Risk of Embodied AI is Governance Lag
by: Liu, Shaoshan
Published: (2026)
by: Liu, Shaoshan
Published: (2026)
Conformity and Social Impact on AI Agents
by: Bellina, Alessandro, et al.
Published: (2026)
by: Bellina, Alessandro, et al.
Published: (2026)
Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants
by: Borges, Beatriz, et al.
Published: (2024)
by: Borges, Beatriz, et al.
Published: (2024)
Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges
by: El-boghdadi, Hatem M., et al.
Published: (2026)
by: El-boghdadi, Hatem M., et al.
Published: (2026)
A vibe coding learning design to enhance EFL students' talking to, through, and about AI
by: Woo, David James, et al.
Published: (2025)
by: Woo, David James, et al.
Published: (2025)
Similar Items
-
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026) -
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024) -
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
by: Gringras, David
Published: (2026) -
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
by: Mushtaq, Abdullah, et al.
Published: (2025) -
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
by: Venkit, Pranav Narayanan, et al.
Published: (2025)