Uncovering Biases with Reflective Large Language Models
Fuente:
arXiv
Salvato in:
| Autore principale: | Chang, Edward Y. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Benchmarking Cognitive Biases in Large Language Models as Evaluators
di: Koo, Ryan, et al.
Pubblicazione: (2023)
di: Koo, Ryan, et al.
Pubblicazione: (2023)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2024)
di: Chang, Edward Y.
Pubblicazione: (2024)
SocraSynth: Multi-LLM Reasoning with Conditional Statistics
di: Chang, Edward Y.
Pubblicazione: (2024)
di: Chang, Edward Y.
Pubblicazione: (2024)
Uncovering Latent Human Wellbeing in Language Model Embeddings
di: Freire, Pedro, et al.
Pubblicazione: (2024)
di: Freire, Pedro, et al.
Pubblicazione: (2024)
Internal Reasoning vs. External Control: A Thermodynamic Analysis of Sycophancy in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2025)
di: Chang, Edward Y.
Pubblicazione: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Danoliteracy of Generative Large Language Models
di: Holm, Søren Vejlgaard, et al.
Pubblicazione: (2024)
di: Holm, Søren Vejlgaard, et al.
Pubblicazione: (2024)
Random-Set Large Language Models
di: Mubashar, Muhammad, et al.
Pubblicazione: (2025)
di: Mubashar, Muhammad, et al.
Pubblicazione: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
di: Fadli, Samih
Pubblicazione: (2025)
di: Fadli, Samih
Pubblicazione: (2025)
SWI: Speaking with Intent in Large Language Models
di: Yin, Yuwei, et al.
Pubblicazione: (2025)
di: Yin, Yuwei, et al.
Pubblicazione: (2025)
Self-Supervised Position Debiasing for Large Language Models
di: Liu, Zhongkun, et al.
Pubblicazione: (2024)
di: Liu, Zhongkun, et al.
Pubblicazione: (2024)
Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
di: Zhang, Liyi, et al.
Pubblicazione: (2025)
di: Zhang, Liyi, et al.
Pubblicazione: (2025)
HAMMER: Hamiltonian Curiosity Augmented Large Language Model Reinforcement
di: Yang, Ming, et al.
Pubblicazione: (2025)
di: Yang, Ming, et al.
Pubblicazione: (2025)
RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models
di: Lin, Tianqianjin, et al.
Pubblicazione: (2025)
di: Lin, Tianqianjin, et al.
Pubblicazione: (2025)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
di: Quevedo, Ernesto, et al.
Pubblicazione: (2024)
di: Quevedo, Ernesto, et al.
Pubblicazione: (2024)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
di: Kurtic, Eldar, et al.
Pubblicazione: (2024)
di: Kurtic, Eldar, et al.
Pubblicazione: (2024)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
di: Gilhuly, Colleen, et al.
Pubblicazione: (2025)
di: Gilhuly, Colleen, et al.
Pubblicazione: (2025)
External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
di: Yan, Jian
Pubblicazione: (2025)
di: Yan, Jian
Pubblicazione: (2025)
Concept Navigation and Classification via Open-Source Large Language Model Processing
di: Kubli, Maël
Pubblicazione: (2025)
di: Kubli, Maël
Pubblicazione: (2025)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
di: Song, Chenyang, et al.
Pubblicazione: (2024)
di: Song, Chenyang, et al.
Pubblicazione: (2024)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
di: Gao, Yutong, et al.
Pubblicazione: (2026)
di: Gao, Yutong, et al.
Pubblicazione: (2026)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
di: Labadie-Tamayo, Roberto, et al.
Pubblicazione: (2025)
di: Labadie-Tamayo, Roberto, et al.
Pubblicazione: (2025)
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
di: Simplício, Afonso, et al.
Pubblicazione: (2026)
di: Simplício, Afonso, et al.
Pubblicazione: (2026)
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
di: Lee, Myeonghwa, et al.
Pubblicazione: (2024)
di: Lee, Myeonghwa, et al.
Pubblicazione: (2024)
Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study
di: Xu, Xiaonan, et al.
Pubblicazione: (2026)
di: Xu, Xiaonan, et al.
Pubblicazione: (2026)
AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
di: Wang, Fali, et al.
Pubblicazione: (2025)
di: Wang, Fali, et al.
Pubblicazione: (2025)
Few-Shot Optimization for Sensor Data Using Large Language Models: A Case Study on Fatigue Detection
di: Ronando, Elsen, et al.
Pubblicazione: (2025)
di: Ronando, Elsen, et al.
Pubblicazione: (2025)
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
di: Zhang, Gongbo, et al.
Pubblicazione: (2026)
di: Zhang, Gongbo, et al.
Pubblicazione: (2026)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
di: Adapala, Sai Teja Reddy
Pubblicazione: (2025)
di: Adapala, Sai Teja Reddy
Pubblicazione: (2025)
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
di: Cui, Sasha, et al.
Pubblicazione: (2025)
di: Cui, Sasha, et al.
Pubblicazione: (2025)
Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
di: Hassell, Jackson, et al.
Pubblicazione: (2025)
di: Hassell, Jackson, et al.
Pubblicazione: (2025)
Improving Language Models with Intentional Analysis
di: Yin, Yuwei, et al.
Pubblicazione: (2025)
di: Yin, Yuwei, et al.
Pubblicazione: (2025)
Do Biased Models Have Biased Thoughts?
di: Rajwal, Swati, et al.
Pubblicazione: (2025)
di: Rajwal, Swati, et al.
Pubblicazione: (2025)
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
di: Salla, Rohit Kumar, et al.
Pubblicazione: (2025)
di: Salla, Rohit Kumar, et al.
Pubblicazione: (2025)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
di: Yordanov, Yordan, et al.
Pubblicazione: (2026)
di: Yordanov, Yordan, et al.
Pubblicazione: (2026)
Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models
di: Hu, Zhibo, et al.
Pubblicazione: (2026)
di: Hu, Zhibo, et al.
Pubblicazione: (2026)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
di: Herbst, Jeremy, et al.
Pubblicazione: (2026)
di: Herbst, Jeremy, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Benchmarking Cognitive Biases in Large Language Models as Evaluators
di: Koo, Ryan, et al.
Pubblicazione: (2023) -
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2024) -
SocraSynth: Multi-LLM Reasoning with Conditional Statistics
di: Chang, Edward Y.
Pubblicazione: (2024) -
Uncovering Latent Human Wellbeing in Language Model Embeddings
di: Freire, Pedro, et al.
Pubblicazione: (2024) -
Internal Reasoning vs. External Control: A Thermodynamic Analysis of Sycophancy in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2025)