KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Seorin, Lee, Dongyoung, Lee, Jaejin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
par: Son, Youngjun, et autres
Publié: (2025)
par: Son, Youngjun, et autres
Publié: (2025)
Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs
par: Ji, SeungWon, et autres
Publié: (2025)
par: Ji, SeungWon, et autres
Publié: (2025)
Models Know Models Best: Evaluation via Model-Preferred Formats
par: Lee, Joonhak, et autres
Publié: (2026)
par: Lee, Joonhak, et autres
Publié: (2026)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
par: Han, Hojae, et autres
Publié: (2024)
par: Han, Hojae, et autres
Publié: (2024)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
par: Cho, Gyeongje, et autres
Publié: (2025)
par: Cho, Gyeongje, et autres
Publié: (2025)
Red-Teaming for Inducing Societal Bias in Large Language Models
par: Luo, Chu Fei, et autres
Publié: (2024)
par: Luo, Chu Fei, et autres
Publié: (2024)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
par: Lee, Young-Jun, et autres
Publié: (2025)
par: Lee, Young-Jun, et autres
Publié: (2025)
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
par: Chae, Kyubyung, et autres
Publié: (2025)
par: Chae, Kyubyung, et autres
Publié: (2025)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
par: Hahm, Sungeun, et autres
Publié: (2025)
par: Hahm, Sungeun, et autres
Publié: (2025)
Reasoning by Commented Code for Table Question Answering
par: Pyo, Seho, et autres
Publié: (2026)
par: Pyo, Seho, et autres
Publié: (2026)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
par: Kim, Dongyoung, et autres
Publié: (2024)
par: Kim, Dongyoung, et autres
Publié: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
par: Sathe, Ashutosh, et autres
Publié: (2024)
par: Sathe, Ashutosh, et autres
Publié: (2024)
Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
par: Lee, Nakyung, et autres
Publié: (2025)
par: Lee, Nakyung, et autres
Publié: (2025)
CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model
par: Go, Dongyoung, et autres
Publié: (2024)
par: Go, Dongyoung, et autres
Publié: (2024)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
par: Park, Seongmin, et autres
Publié: (2024)
par: Park, Seongmin, et autres
Publié: (2024)
On The Conceptualization and Societal Impact of Cross-Cultural Bias
par: Bhandari, Vitthal
Publié: (2025)
par: Bhandari, Vitthal
Publié: (2025)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
par: Han, Hojae, et autres
Publié: (2026)
par: Han, Hojae, et autres
Publié: (2026)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
par: Jung, Sungmok, et autres
Publié: (2026)
par: Jung, Sungmok, et autres
Publié: (2026)
Instructive Decoding: Instruction-Tuned Large Language Models are Self-Refiner from Noisy Instructions
par: Kim, Taehyeon, et autres
Publié: (2023)
par: Kim, Taehyeon, et autres
Publié: (2023)
Choices Speak Louder than Questions
par: Cho, Gyeongje, et autres
Publié: (2025)
par: Cho, Gyeongje, et autres
Publié: (2025)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
par: So, Yeonkyoung, et autres
Publié: (2025)
par: So, Yeonkyoung, et autres
Publié: (2025)
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
par: Gamboa, Lance Calvin Lim, et autres
Publié: (2025)
par: Gamboa, Lance Calvin Lim, et autres
Publié: (2025)
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
par: Yoo, Jaeseok, et autres
Publié: (2024)
par: Yoo, Jaeseok, et autres
Publié: (2024)
Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT
par: Lee, Wonhyuk, et autres
Publié: (2025)
par: Lee, Wonhyuk, et autres
Publié: (2025)
Detecting Bias in Large Language Models: Fine-tuned KcBERT
par: Lee, J. K., et autres
Publié: (2024)
par: Lee, J. K., et autres
Publié: (2024)
Social Bias in Multilingual Language Models: A Survey
par: Gamboa, Lance Calvin Lim, et autres
Publié: (2025)
par: Gamboa, Lance Calvin Lim, et autres
Publié: (2025)
Learning to Reduce: Optimal Representations of Structured Data in Prompting Large Language Models
par: Lee, Younghun, et autres
Publié: (2024)
par: Lee, Younghun, et autres
Publié: (2024)
Learning to Reduce: Towards Improving Performance of Large Language Models on Structured Data
par: Lee, Younghun, et autres
Publié: (2024)
par: Lee, Younghun, et autres
Publié: (2024)
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads
par: Yang, Yi, et autres
Publié: (2023)
par: Yang, Yi, et autres
Publié: (2023)
Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
par: Gamboa, Lance Calvin Lim, et autres
Publié: (2026)
par: Gamboa, Lance Calvin Lim, et autres
Publié: (2026)
Interpreting Bias in Large Language Models: A Feature-Based Approach
par: Prakash, Nirmalendu, et autres
Publié: (2024)
par: Prakash, Nirmalendu, et autres
Publié: (2024)
Structured Language Generation Model: Loss Calibration and Formatted Decoding for Robust Structure Prediction and Knowledge Retrieval
par: Lee, Minho, et autres
Publié: (2024)
par: Lee, Minho, et autres
Publié: (2024)
Small Language Models are Equation Reasoners
par: Kim, Bumjun, et autres
Publié: (2024)
par: Kim, Bumjun, et autres
Publié: (2024)
Large Language Models as Search Engines: Societal Challenges
par: Sadeddine, Zacchary, et autres
Publié: (2025)
par: Sadeddine, Zacchary, et autres
Publié: (2025)
Refining Time Series Anomaly Detectors using Large Language Models
par: Yang, Alan, et autres
Publié: (2025)
par: Yang, Alan, et autres
Publié: (2025)
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
par: Kim, Jungwoo, et autres
Publié: (2025)
par: Kim, Jungwoo, et autres
Publié: (2025)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
par: Kim, Jinpyo, et autres
Publié: (2025)
par: Kim, Jinpyo, et autres
Publié: (2025)
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
par: Yu, Sangwon, et autres
Publié: (2024)
par: Yu, Sangwon, et autres
Publié: (2024)
Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
par: Park, Bumjin, et autres
Publié: (2025)
par: Park, Bumjin, et autres
Publié: (2025)
Selective Generation for Controllable Language Models
par: Lee, Minjae, et autres
Publié: (2023)
par: Lee, Minjae, et autres
Publié: (2023)
Documents similaires
-
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
par: Son, Youngjun, et autres
Publié: (2025) -
Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs
par: Ji, SeungWon, et autres
Publié: (2025) -
Models Know Models Best: Evaluation via Model-Preferred Formats
par: Lee, Joonhak, et autres
Publié: (2026) -
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
par: Han, Hojae, et autres
Publié: (2024) -
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
par: Cho, Gyeongje, et autres
Publié: (2025)