Identifying, Evaluating, and Mitigating Risks of AI Thought Partnerships
Fuente:
arXiv
Guardado en:
| Autores principales: | Oktar, Kerem, Collins, Katherine M., Hernandez-Orallo, Jose, Coyle, Diane, Cave, Stephen, Weller, Adrian, Sucholutsky, Ilia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Revisiting Rogers' Paradox in the Context of Human-AI Interaction
por: Collins, Katherine M., et al.
Publicado: (2025)
por: Collins, Katherine M., et al.
Publicado: (2025)
On Benchmarking Human-Like Intelligence in Machines
por: Ying, Lance, et al.
Publicado: (2025)
por: Ying, Lance, et al.
Publicado: (2025)
Medical Model Synthesis Architectures: A Case Study
por: Collins, Katherine M., et al.
Publicado: (2026)
por: Collins, Katherine M., et al.
Publicado: (2026)
Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey
por: Smith, Victoria, et al.
Publicado: (2023)
por: Smith, Victoria, et al.
Publicado: (2023)
Modulating Language Model Experiences through Frictions
por: Collins, Katherine M., et al.
Publicado: (2024)
por: Collins, Katherine M., et al.
Publicado: (2024)
Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison
por: Yang, Tiancheng, et al.
Publicado: (2026)
por: Yang, Tiancheng, et al.
Publicado: (2026)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
por: Testini, Irene, et al.
Publicado: (2025)
por: Testini, Irene, et al.
Publicado: (2025)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
por: Burden, John, et al.
Publicado: (2025)
por: Burden, John, et al.
Publicado: (2025)
Identifying and Mitigating the Security Risks of Generative AI
por: Barrett, Clark, et al.
Publicado: (2023)
por: Barrett, Clark, et al.
Publicado: (2023)
Learning Human-like Representations to Enable Learning Human Values
por: Wynn, Andrea, et al.
Publicado: (2023)
por: Wynn, Andrea, et al.
Publicado: (2023)
Using LLMs to Advance the Cognitive Science of Collectives
por: Sucholutsky, Ilia, et al.
Publicado: (2025)
por: Sucholutsky, Ilia, et al.
Publicado: (2025)
Conversational Complexity for Assessing Risk in Large Language Models
por: Burden, John, et al.
Publicado: (2024)
por: Burden, John, et al.
Publicado: (2024)
Why Human Guidance Matters in Collaborative Vibe Coding
por: Hu, Haoyu, et al.
Publicado: (2026)
por: Hu, Haoyu, et al.
Publicado: (2026)
What should an AI assessor optimise for?
por: Romero-Alvarado, Daniel, et al.
Publicado: (2025)
por: Romero-Alvarado, Daniel, et al.
Publicado: (2025)
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
por: Ying, Lance, et al.
Publicado: (2026)
por: Ying, Lance, et al.
Publicado: (2026)
Evaluating General-Purpose AI with Psychometrics
por: Wang, Xiting, et al.
Publicado: (2023)
por: Wang, Xiting, et al.
Publicado: (2023)
Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs
por: Mieczkowski, Elizabeth, et al.
Publicado: (2026)
por: Mieczkowski, Elizabeth, et al.
Publicado: (2026)
What is a Number, That a Large Language Model May Know It?
por: Marjieh, Raja, et al.
Publicado: (2025)
por: Marjieh, Raja, et al.
Publicado: (2025)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
por: Zaman, Kerem, et al.
Publicado: (2025)
por: Zaman, Kerem, et al.
Publicado: (2025)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
por: Rutar, Danaja, et al.
Publicado: (2025)
por: Rutar, Danaja, et al.
Publicado: (2025)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
por: Liu, Ryan, et al.
Publicado: (2024)
por: Liu, Ryan, et al.
Publicado: (2024)
Are Large Language Models Sensitive to the Motives Behind Communication?
por: Wu, Addison J., et al.
Publicado: (2025)
por: Wu, Addison J., et al.
Publicado: (2025)
Estimation of Concept Explanations Should be Uncertainty Aware
por: Piratla, Vihari, et al.
Publicado: (2023)
por: Piratla, Vihari, et al.
Publicado: (2023)
Mitigating Shortcut Learning with InterpoLated Learning
por: Korakakis, Michalis, et al.
Publicado: (2025)
por: Korakakis, Michalis, et al.
Publicado: (2025)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
por: Fitz, Stephen, et al.
Publicado: (2025)
por: Fitz, Stephen, et al.
Publicado: (2025)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
por: Voudouris, Konstantinos, et al.
Publicado: (2026)
por: Voudouris, Konstantinos, et al.
Publicado: (2026)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
por: Balkır, Esma, et al.
Publicado: (2026)
por: Balkır, Esma, et al.
Publicado: (2026)
Building Machines that Learn and Think with People
por: Collins, Katherine M., et al.
Publicado: (2024)
por: Collins, Katherine M., et al.
Publicado: (2024)
Do Large Language Models Mentalize When They Teach?
por: Harootonian, Sevan K., et al.
Publicado: (2026)
por: Harootonian, Sevan K., et al.
Publicado: (2026)
Empathy in Explanation
por: Collins, Katherine M., et al.
Publicado: (2025)
por: Collins, Katherine M., et al.
Publicado: (2025)
Learning to Receive Help: Intervention-Aware Concept Embedding Models
por: Zarlenga, Mateo Espinosa, et al.
Publicado: (2023)
por: Zarlenga, Mateo Espinosa, et al.
Publicado: (2023)
Analyzing the Roles of Language and Vision in Learning from Limited Data
por: Chen, Allison, et al.
Publicado: (2024)
por: Chen, Allison, et al.
Publicado: (2024)
Tracing Thought: Using Chain-of-Thought Reasoning to Identify the LLM Behind AI-Generated Text
por: Agrahari, Shifali, et al.
Publicado: (2025)
por: Agrahari, Shifali, et al.
Publicado: (2025)
Mapping AI Risk Mitigations: Evidence Scan and Preliminary AI Risk Mitigation Taxonomy
por: Saeri, Alexander K., et al.
Publicado: (2025)
por: Saeri, Alexander K., et al.
Publicado: (2025)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
Concept Alignment
por: Rane, Sunayana, et al.
Publicado: (2024)
por: Rane, Sunayana, et al.
Publicado: (2024)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
Beyond the high score: Prosocial ability profiles of multi-agent populations
por: Tesic, Marko, et al.
Publicado: (2025)
por: Tesic, Marko, et al.
Publicado: (2025)
Multi-agent AI systems outperform human teams in creativity
por: Hu, Tiancheng, et al.
Publicado: (2026)
por: Hu, Tiancheng, et al.
Publicado: (2026)
Large Language Models Assume People are More Rational than We Really are
por: Liu, Ryan, et al.
Publicado: (2024)
por: Liu, Ryan, et al.
Publicado: (2024)
Ejemplares similares
-
Revisiting Rogers' Paradox in the Context of Human-AI Interaction
por: Collins, Katherine M., et al.
Publicado: (2025) -
On Benchmarking Human-Like Intelligence in Machines
por: Ying, Lance, et al.
Publicado: (2025) -
Medical Model Synthesis Architectures: A Case Study
por: Collins, Katherine M., et al.
Publicado: (2026) -
Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey
por: Smith, Victoria, et al.
Publicado: (2023) -
Modulating Language Model Experiences through Frictions
por: Collins, Katherine M., et al.
Publicado: (2024)