On the Alignment of Large Language Models with Global Human Opinion
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yang, Kaneko, Masahiro, Chu, Chenhui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
In-Contextual Gender Bias Suppression for Large Language Models
di: Oba, Daisuke, et al.
Pubblicazione: (2023)
di: Oba, Daisuke, et al.
Pubblicazione: (2023)
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
di: Hida, Rem, et al.
Pubblicazione: (2024)
di: Hida, Rem, et al.
Pubblicazione: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2025)
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2025)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2023)
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2023)
Understanding the Prompt Sensitivity
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
When Large Language Models Meet Speech: A Survey on Integration Approaches
di: Yang, Zhengdong, et al.
Pubblicazione: (2025)
di: Yang, Zhengdong, et al.
Pubblicazione: (2025)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
di: Oi, Masanari, et al.
Pubblicazione: (2024)
di: Oi, Masanari, et al.
Pubblicazione: (2024)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
di: Kaneko, Masahiro, et al.
Pubblicazione: (2023)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2023)
Understanding The Effect Of Temperature On Alignment With Human Opinions
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
di: Zhong, Chengzhi, et al.
Pubblicazione: (2025)
di: Zhong, Chengzhi, et al.
Pubblicazione: (2025)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
di: Kaneko, Masahiro
Pubblicazione: (2026)
di: Kaneko, Masahiro
Pubblicazione: (2026)
How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
di: Takenami, Yoshiki, et al.
Pubblicazione: (2025)
di: Takenami, Yoshiki, et al.
Pubblicazione: (2025)
Large Language Models Lack Understanding of Character Composition of Words
di: Shin, Andrew, et al.
Pubblicazione: (2024)
di: Shin, Andrew, et al.
Pubblicazione: (2024)
MM-LLMs: Recent Advances in MultiModal Large Language Models
di: Zhang, Duzhen, et al.
Pubblicazione: (2024)
di: Zhang, Duzhen, et al.
Pubblicazione: (2024)
Exploring Vulnerabilities and Protections in Large Language Models: A Survey
di: Liu, Frank Weizhen, et al.
Pubblicazione: (2024)
di: Liu, Frank Weizhen, et al.
Pubblicazione: (2024)
Personality Alignment of Large Language Models
di: Zhu, Minjun, et al.
Pubblicazione: (2024)
di: Zhu, Minjun, et al.
Pubblicazione: (2024)
Rapidly Developing High-quality Instruction Data and Evaluation Benchmark for Large Language Models with Minimal Human Effort: A Case Study on Japanese
di: Sun, Yikun, et al.
Pubblicazione: (2024)
di: Sun, Yikun, et al.
Pubblicazione: (2024)
Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models
di: Tuck, Bryan E., et al.
Pubblicazione: (2025)
di: Tuck, Bryan E., et al.
Pubblicazione: (2025)
Fine-Grained Interpretation of Political Opinions in Large Language Models
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
Towards Atoms of Large Language Models
di: Hu, Chenhui, et al.
Pubblicazione: (2025)
di: Hu, Chenhui, et al.
Pubblicazione: (2025)
Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation
di: Li, Yihang, et al.
Pubblicazione: (2026)
di: Li, Yihang, et al.
Pubblicazione: (2026)
Aligning Large Language Models with Human Opinions through Persona Selection and Value--Belief--Norm Reasoning
di: Long, Do Xuan, et al.
Pubblicazione: (2023)
di: Long, Do Xuan, et al.
Pubblicazione: (2023)
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
di: Koike, Ryuto, et al.
Pubblicazione: (2025)
di: Koike, Ryuto, et al.
Pubblicazione: (2025)
The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language Models
di: Zhang, Shuo, et al.
Pubblicazione: (2023)
di: Zhang, Shuo, et al.
Pubblicazione: (2023)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
di: Ohi, Masanari, et al.
Pubblicazione: (2024)
di: Ohi, Masanari, et al.
Pubblicazione: (2024)
Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models
di: Hu, Chenhui, et al.
Pubblicazione: (2024)
di: Hu, Chenhui, et al.
Pubblicazione: (2024)
Can Large Language Models be Effective Online Opinion Miners?
di: Heo, Ryang, et al.
Pubblicazione: (2025)
di: Heo, Ryang, et al.
Pubblicazione: (2025)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
di: Niwa, Ayana, et al.
Pubblicazione: (2025)
di: Niwa, Ayana, et al.
Pubblicazione: (2025)
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
di: Shiotani, Taihei, et al.
Pubblicazione: (2026)
The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
Beyond the Resumé: A Rubric-Aware Automatic Interview System for Information Elicitation
di: Stuart, Harry, et al.
Pubblicazione: (2026)
di: Stuart, Harry, et al.
Pubblicazione: (2026)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
di: Koike, Ryuto, et al.
Pubblicazione: (2023)
Eagle: Ethical Dataset Given from Real Interactions
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
di: Loem, Mengsay, et al.
Pubblicazione: (2023)
di: Loem, Mengsay, et al.
Pubblicazione: (2023)
Documenti analoghi
-
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
di: Kaneko, Masahiro, et al.
Pubblicazione: (2024) -
In-Contextual Gender Bias Suppression for Large Language Models
di: Oba, Daisuke, et al.
Pubblicazione: (2023) -
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
di: Liu, Yang, et al.
Pubblicazione: (2025) -
Social Bias Evaluation for Large Language Models Requires Prompt Variations
di: Hida, Rem, et al.
Pubblicazione: (2024) -
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)