CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shaikh, Ammar, Dandekar, Raj Abhijit, Panat, Sreedath, Dandekar, Rajat |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Passive Viewing: A Pilot Study of a Hybrid Learning Platform Augmenting Video Lectures with Conversational AI
by: Abraar, Mohammed, et al.
Published: (2026)
by: Abraar, Mohammed, et al.
Published: (2026)
Decoders Laugh as Loud as Encoders
by: Borodach, Eli, et al.
Published: (2025)
by: Borodach, Eli, et al.
Published: (2025)
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Vision-Language Models display a strong gender bias
by: Konavoor, Aiswarya, et al.
Published: (2025)
by: Konavoor, Aiswarya, et al.
Published: (2025)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
by: Dawson, Fiifi, et al.
Published: (2024)
by: Dawson, Fiifi, et al.
Published: (2024)
Simulating Misinformation Propagation in Social Networks using Large Language Models
by: Maurya, Raj Gaurav, et al.
Published: (2025)
by: Maurya, Raj Gaurav, et al.
Published: (2025)
Scientific machine learning in ecological systems: A study on the predator-prey dynamics
by: Devgupta, Ranabir, et al.
Published: (2024)
by: Devgupta, Ranabir, et al.
Published: (2024)
Modeling chaotic Lorenz ODE System using Scientific Machine Learning
by: Kashyap, Sameera S, et al.
Published: (2024)
by: Kashyap, Sameera S, et al.
Published: (2024)
NanoVLMs: How small can we go and still make coherent Vision Language Models?
by: Agarwalla, Mukund, et al.
Published: (2025)
by: Agarwalla, Mukund, et al.
Published: (2025)
A comparative study of NeuralODE and Universal ODE approaches to solving Chandrasekhar White Dwarf equation
by: Martinez, Raymundo Vazquez, et al.
Published: (2024)
by: Martinez, Raymundo Vazquez, et al.
Published: (2024)
Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Forecasting N-Body Dynamics: A Comparative Study of Neural Ordinary Differential Equations and Universal Differential Equations
by: S, Suriya R, et al.
Published: (2025)
by: S, Suriya R, et al.
Published: (2025)
A Scientific Machine Learning Approach for Predicting and Forecasting Battery Degradation in Electric Vehicles
by: Murgai, Sharv, et al.
Published: (2024)
by: Murgai, Sharv, et al.
Published: (2024)
HULLMI: Human vs LLM identification with explainability
by: Joshi, Prathamesh Dinesh, et al.
Published: (2024)
by: Joshi, Prathamesh Dinesh, et al.
Published: (2024)
EARS-UDE: Evaluating Auditory Response in Sensory Overload with Universal Differential Equations
by: Salunke, Miheer, et al.
Published: (2025)
by: Salunke, Miheer, et al.
Published: (2025)
Adaptive tumor growth forecasting via neural & universal ODEs
by: Subramanian, Kavya, et al.
Published: (2025)
by: Subramanian, Kavya, et al.
Published: (2025)
Regional Tiny Stories: Using Small Models to Compare Language Learning and Tokenizer Performance
by: Patil, Nirvan, et al.
Published: (2025)
by: Patil, Nirvan, et al.
Published: (2025)
BULL-ODE: Bullwhip Learning with Neural ODEs and Universal Differential Equations under Stochastic Demand
by: Naik, Nachiket N., et al.
Published: (2025)
by: Naik, Nachiket N., et al.
Published: (2025)
A study of Universal ODE approaches to predicting soil organic carbon
by: V. V, Satyanarayana Raju G., et al.
Published: (2025)
by: V. V, Satyanarayana Raju G., et al.
Published: (2025)
Three methods, one problem: Classical and AI approaches to no-three-in-line
by: Ramanathan, Pranav, et al.
Published: (2025)
by: Ramanathan, Pranav, et al.
Published: (2025)
To Bias or Not to Bias: Detecting bias in News with bias-detector
by: Ghosh, Himel, et al.
Published: (2025)
by: Ghosh, Himel, et al.
Published: (2025)
The role of large language models in UI/UX design: A systematic literature review
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Muon: Training and Trade-offs with Latent Attention and MoE
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Rehearsal: Simulating Conflict to Teach Conflict Resolution
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models
by: Peters, Uwe, et al.
Published: (2026)
by: Peters, Uwe, et al.
Published: (2026)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
"What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer Use
by: Sapkota, Shardul, et al.
Published: (2026)
by: Sapkota, Shardul, et al.
Published: (2026)
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
by: Lam, Michelle S., et al.
Published: (2025)
by: Lam, Michelle S., et al.
Published: (2025)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
Creating General User Models from Computer Use
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
HICode: Hierarchical Inductive Coding with LLMs
by: Zhong, Mian, et al.
Published: (2025)
by: Zhong, Mian, et al.
Published: (2025)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
by: Liu, Shuyu, et al.
Published: (2025)
by: Liu, Shuyu, et al.
Published: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025)
by: Liu, Hongtao, et al.
Published: (2025)
Can LLMs Generate Visualizations with Dataless Prompts?
by: Coelho, Darius, et al.
Published: (2024)
by: Coelho, Darius, et al.
Published: (2024)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026)
by: Hoq, Adnan, et al.
Published: (2026)
Offscript: Automated Auditing of Instruction Adherence in LLMs
by: Clark, Nicholas, et al.
Published: (2025)
by: Clark, Nicholas, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Similar Items
-
Beyond Passive Viewing: A Pilot Study of a Hybrid Learning Platform Augmenting Video Lectures with Conversational AI
by: Abraar, Mohammed, et al.
Published: (2026) -
Decoders Laugh as Loud as Encoders
by: Borodach, Eli, et al.
Published: (2025) -
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025) -
Vision-Language Models display a strong gender bias
by: Konavoor, Aiswarya, et al.
Published: (2025) -
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
by: Dawson, Fiifi, et al.
Published: (2024)