Guardado en:
| Autor principal: | Frohn, Scott |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2604.26954 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning
por: Yang, Xia, et al.
Publicado: (2026)
por: Yang, Xia, et al.
Publicado: (2026)
Artificial Effort
por: Belotti, Federico, et al.
Publicado: (2026)
por: Belotti, Federico, et al.
Publicado: (2026)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
por: Dimino, Fabrizio, et al.
Publicado: (2026)
por: Dimino, Fabrizio, et al.
Publicado: (2026)
MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction
por: Su, Xinchun, et al.
Publicado: (2026)
por: Su, Xinchun, et al.
Publicado: (2026)
How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI
por: Wilson, Sophia N., et al.
Publicado: (2026)
por: Wilson, Sophia N., et al.
Publicado: (2026)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
por: Tang, Zeyu, et al.
Publicado: (2026)
por: Tang, Zeyu, et al.
Publicado: (2026)
The Global Landscape of Environmental AI Regulation: From the Cost of Reasoning to a Right to Green AI
por: Ebert, Kai, et al.
Publicado: (2026)
por: Ebert, Kai, et al.
Publicado: (2026)
iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries
por: Coscia, Adam, et al.
Publicado: (2024)
por: Coscia, Adam, et al.
Publicado: (2024)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
por: Mohamadi, Alireza, et al.
Publicado: (2025)
por: Mohamadi, Alireza, et al.
Publicado: (2025)
Automated Formative Feedback for Short-form Writing: An LLM-Driven Approach and Adoption Analysis
por: Tavares, Tiago Fernandes, et al.
Publicado: (2025)
por: Tavares, Tiago Fernandes, et al.
Publicado: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
por: Schoenegger, Philipp, et al.
Publicado: (2024)
por: Schoenegger, Philipp, et al.
Publicado: (2024)
How Do Companies Manage the Environmental Sustainability of AI? An Interview Study About Green AI Efforts and Regulations
por: Sampatsing, Ashmita, et al.
Publicado: (2025)
por: Sampatsing, Ashmita, et al.
Publicado: (2025)
Effort-aware Fairness: Incorporating a Philosophy-informed, Human-centered Notion of Effort into Algorithmic Fairness Metrics
por: Nguyen, Tin Trung, et al.
Publicado: (2025)
por: Nguyen, Tin Trung, et al.
Publicado: (2025)
CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback
por: Tseng, En-Qi, et al.
Publicado: (2025)
por: Tseng, En-Qi, et al.
Publicado: (2025)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
por: Walsh, Cole, et al.
Publicado: (2026)
por: Walsh, Cole, et al.
Publicado: (2026)
Artificial Intelligence Ecosystem for Automating Self-Directed Teaching
por: Gotavade, Tejas Satish
Publicado: (2024)
por: Gotavade, Tejas Satish
Publicado: (2024)
Beyond Static Question Banks: Dynamic Knowledge Expansion via LLM-Automated Graph Construction and Adaptive Generation
por: Wang, Yingquan, et al.
Publicado: (2026)
por: Wang, Yingquan, et al.
Publicado: (2026)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
por: Chang, Chia-Hsuan, et al.
Publicado: (2024)
por: Chang, Chia-Hsuan, et al.
Publicado: (2024)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
por: Schoenegger, Philipp, et al.
Publicado: (2024)
por: Schoenegger, Philipp, et al.
Publicado: (2024)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
por: Jahara, Fatima, et al.
Publicado: (2025)
por: Jahara, Fatima, et al.
Publicado: (2025)
Struggle Premium : How Human Effort and Imperfection Drive Perceived Value in the Age of AI
por: Sultana, Nazneen, et al.
Publicado: (2026)
por: Sultana, Nazneen, et al.
Publicado: (2026)
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
por: DeLucia, Alexandra, et al.
Publicado: (2026)
por: DeLucia, Alexandra, et al.
Publicado: (2026)
Content Moderation by LLM: From Accuracy to Legitimacy
por: Huang, Tao
Publicado: (2024)
por: Huang, Tao
Publicado: (2024)
Antisocial Analagous Behavior, Alignment and Human Impact of Google AI Systems: Evaluating through the lens of modified Antisocial Behavior Criteria by Human Interaction, Independent LLM Analysis, and AI Self-Reflection
por: Ogilvie, Alan D.
Publicado: (2024)
por: Ogilvie, Alan D.
Publicado: (2024)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
por: Bojic, Ljubisa, et al.
Publicado: (2026)
por: Bojic, Ljubisa, et al.
Publicado: (2026)
CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations
por: Yu, Xiangning, et al.
Publicado: (2026)
por: Yu, Xiangning, et al.
Publicado: (2026)
Measuring Political Stance and Consistency in Large Language Models
por: Alali, Salah Feras, et al.
Publicado: (2026)
por: Alali, Salah Feras, et al.
Publicado: (2026)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
por: Li, Miles Q., et al.
Publicado: (2026)
por: Li, Miles Q., et al.
Publicado: (2026)
Practical Principles for AI Cost and Compute Accounting
por: Casper, Stephen, et al.
Publicado: (2025)
por: Casper, Stephen, et al.
Publicado: (2025)
The Cost-Benefit of Interdisciplinarity in AI for Mental Health
por: Drakos, Katerina, et al.
Publicado: (2025)
por: Drakos, Katerina, et al.
Publicado: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
por: Husom, Erik Johannes, et al.
Publicado: (2025)
por: Husom, Erik Johannes, et al.
Publicado: (2025)
The Shady Light of Art Automation
por: Grba, Dejan
Publicado: (2025)
por: Grba, Dejan
Publicado: (2025)
Combining Cost-Constrained Runtime Monitors for AI Safety
por: Hua, Tim Tian, et al.
Publicado: (2025)
por: Hua, Tim Tian, et al.
Publicado: (2025)
A Review of Generative AI in Computer Science Education: Challenges and Opportunities in Accuracy, Authenticity, and Assessment
por: Reihanian, Iman, et al.
Publicado: (2025)
por: Reihanian, Iman, et al.
Publicado: (2025)
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
por: Cheng, Myra, et al.
Publicado: (2024)
por: Cheng, Myra, et al.
Publicado: (2024)
Measuring AI R&D Automation
por: Chan, Alan, et al.
Publicado: (2026)
por: Chan, Alan, et al.
Publicado: (2026)
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
por: Wang, Lichao, et al.
Publicado: (2026)
por: Wang, Lichao, et al.
Publicado: (2026)
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models
por: Frisch, Ivar, et al.
Publicado: (2024)
por: Frisch, Ivar, et al.
Publicado: (2024)
The Social Impact of Generative LLM-Based AI
por: Xie, Yu, et al.
Publicado: (2024)
por: Xie, Yu, et al.
Publicado: (2024)
Alternative Fairness and Accuracy Optimization in Criminal Justice
por: Wu, Shaolong, et al.
Publicado: (2025)
por: Wu, Shaolong, et al.
Publicado: (2025)
Ejemplares similares
-
Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning
por: Yang, Xia, et al.
Publicado: (2026) -
Artificial Effort
por: Belotti, Federico, et al.
Publicado: (2026) -
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
por: Dimino, Fabrizio, et al.
Publicado: (2026) -
MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction
por: Su, Xinchun, et al.
Publicado: (2026) -
How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI
por: Wilson, Sophia N., et al.
Publicado: (2026)