Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Anantaprayoon, Panatchakorn, Kaneko, Masahiro, Okazaki, Naoaki |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2023)
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2023)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
di: Oi, Masanari, et al.
Pubblicazione: (2024)
di: Oi, Masanari, et al.
Pubblicazione: (2024)
Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2025)
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2025)
Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2026)
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2026)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
di: Ohi, Masanari, et al.
Pubblicazione: (2024)
di: Ohi, Masanari, et al.
Pubblicazione: (2024)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
di: Hida, Rem, et al.
Pubblicazione: (2024)
di: Hida, Rem, et al.
Pubblicazione: (2024)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
Mitigating Social Biases in Language Models through Unlearning
di: Dige, Omkar, et al.
Pubblicazione: (2024)
di: Dige, Omkar, et al.
Pubblicazione: (2024)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
di: Christian, Brian, et al.
Pubblicazione: (2026)
di: Christian, Brian, et al.
Pubblicazione: (2026)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
di: Kaneko, Masahiro, et al.
Pubblicazione: (2023)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2023)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
di: Loem, Mengsay, et al.
Pubblicazione: (2023)
di: Loem, Mengsay, et al.
Pubblicazione: (2023)
Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments
di: Sumita, Yasuaki, et al.
Pubblicazione: (2024)
di: Sumita, Yasuaki, et al.
Pubblicazione: (2024)
Building a Large Japanese Web Corpus for Large Language Models
di: Okazaki, Naoaki, et al.
Pubblicazione: (2024)
di: Okazaki, Naoaki, et al.
Pubblicazione: (2024)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
di: Miyamoto, Sora, et al.
Pubblicazione: (2026)
di: Miyamoto, Sora, et al.
Pubblicazione: (2026)
Balancing Rigor and Utility: Mitigating Cognitive Biases in Large Language Models for Multiple-Choice Questions
di: Zhong, Hanyang, et al.
Pubblicazione: (2024)
di: Zhong, Hanyang, et al.
Pubblicazione: (2024)
Large Language Models are Biased Because They Are Large Language Models
di: Resnik, Philip
Pubblicazione: (2024)
di: Resnik, Philip
Pubblicazione: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
di: Kaneko, Masahiro
Pubblicazione: (2026)
di: Kaneko, Masahiro
Pubblicazione: (2026)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
di: Kim, Jaekyeom, et al.
Pubblicazione: (2024)
di: Kim, Jaekyeom, et al.
Pubblicazione: (2024)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
di: Shirafuji, Daiki, et al.
Pubblicazione: (2024)
di: Shirafuji, Daiki, et al.
Pubblicazione: (2024)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
LLM Output Detectability and Task Performance Can be Jointly Optimized
di: Saito, Koshiro, et al.
Pubblicazione: (2026)
di: Saito, Koshiro, et al.
Pubblicazione: (2026)
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
di: Wu, Addison J., et al.
Pubblicazione: (2025)
di: Wu, Addison J., et al.
Pubblicazione: (2025)
FairPy: A Toolkit for Evaluation of Prediction Biases and their Mitigation in Large Language Models
di: Viswanath, Hrishikesh, et al.
Pubblicazione: (2023)
di: Viswanath, Hrishikesh, et al.
Pubblicazione: (2023)
Multilingual Performance Biases of Large Language Models in Education
di: Gupta, Vansh, et al.
Pubblicazione: (2025)
di: Gupta, Vansh, et al.
Pubblicazione: (2025)
Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
di: Huang, Jianheng, et al.
Pubblicazione: (2024)
di: Huang, Jianheng, et al.
Pubblicazione: (2024)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
di: Maeda, Koki, et al.
Pubblicazione: (2026)
di: Maeda, Koki, et al.
Pubblicazione: (2026)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
di: Hao, Yijie, et al.
Pubblicazione: (2025)
di: Hao, Yijie, et al.
Pubblicazione: (2025)
Large Language Models Cannot Self-Correct Reasoning Yet
di: Huang, Jie, et al.
Pubblicazione: (2023)
di: Huang, Jie, et al.
Pubblicazione: (2023)
Large Language Models have Intrinsic Self-Correction Ability
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
di: Oba, Daisuke, et al.
Pubblicazione: (2026)
di: Oba, Daisuke, et al.
Pubblicazione: (2026)
Knowledge-Aware Self-Correction in Language Models via Structured Memory Graphs
di: Saha, Swayamjit
Pubblicazione: (2025)
di: Saha, Swayamjit
Pubblicazione: (2025)
IRepair: An Intent-Aware Approach to Repair Data-Driven Errors in Large Language Models
di: Imtiaz, Sayem Mohammad, et al.
Pubblicazione: (2025)
di: Imtiaz, Sayem Mohammad, et al.
Pubblicazione: (2025)
Self-Correcting Large Language Models: Generation vs. Multiple Choice
di: Rahmani, Hossein A., et al.
Pubblicazione: (2025)
di: Rahmani, Hossein A., et al.
Pubblicazione: (2025)
Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models
di: Zhang, Che, et al.
Pubblicazione: (2024)
di: Zhang, Che, et al.
Pubblicazione: (2024)
CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models
di: Wu, Yuning, et al.
Pubblicazione: (2026)
di: Wu, Yuning, et al.
Pubblicazione: (2026)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
di: Wan, Yixin, et al.
Pubblicazione: (2024)
di: Wan, Yixin, et al.
Pubblicazione: (2024)
Empathetic Cascading Networks: A Multi-Stage Prompting Technique for Reducing Social Biases in Large Language Models
di: Xin, Wangjiaxuan
Pubblicazione: (2025)
di: Xin, Wangjiaxuan
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2023) -
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
di: Oi, Masanari, et al.
Pubblicazione: (2024) -
Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2025) -
Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2026) -
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
di: Ohi, Masanari, et al.
Pubblicazione: (2024)