DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
Fuente:
arXiv
Salvato in:
| Autori principali: | Chiu, Yu Ying, Jiang, Liwei, Choi, Yejin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
di: Chiu, Yu Ying, et al.
Pubblicazione: (2025)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2025)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
di: Jung, Jaehun, et al.
Pubblicazione: (2023)
di: Jung, Jaehun, et al.
Pubblicazione: (2023)
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
di: Liu, Jiacheng, et al.
Pubblicazione: (2023)
di: Liu, Jiacheng, et al.
Pubblicazione: (2023)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
di: Ethayarajh, Kawin, et al.
Pubblicazione: (2021)
di: Ethayarajh, Kawin, et al.
Pubblicazione: (2021)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
di: Shan, Zikang, et al.
Pubblicazione: (2026)
di: Shan, Zikang, et al.
Pubblicazione: (2026)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2025)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2025)
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
di: Dai, Hui, et al.
Pubblicazione: (2024)
di: Dai, Hui, et al.
Pubblicazione: (2024)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
di: Sclar, Melanie, et al.
Pubblicazione: (2023)
di: Sclar, Melanie, et al.
Pubblicazione: (2023)
Generative Value Conflicts Reveal LLM Priorities
di: Liu, Andy, et al.
Pubblicazione: (2025)
di: Liu, Andy, et al.
Pubblicazione: (2025)
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2025)
di: Liu, Zhenhua, et al.
Pubblicazione: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
di: Zhang, Wenbo, et al.
Pubblicazione: (2025)
di: Zhang, Wenbo, et al.
Pubblicazione: (2025)
Using LLMs to Model the Beliefs and Preferences of Targeted Populations
di: Namikoshi, Keiichi, et al.
Pubblicazione: (2024)
di: Namikoshi, Keiichi, et al.
Pubblicazione: (2024)
RouteLLM: Learning to Route LLMs with Preference Data
di: Ong, Isaac, et al.
Pubblicazione: (2024)
di: Ong, Isaac, et al.
Pubblicazione: (2024)
Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
di: Long, Quanyu, et al.
Pubblicazione: (2026)
di: Long, Quanyu, et al.
Pubblicazione: (2026)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
di: Sclar, Melanie, et al.
Pubblicazione: (2024)
di: Sclar, Melanie, et al.
Pubblicazione: (2024)
Towards Execution-Grounded Automated AI Research
di: Si, Chenglei, et al.
Pubblicazione: (2026)
di: Si, Chenglei, et al.
Pubblicazione: (2026)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
di: Kawakami, Wataru, et al.
Pubblicazione: (2025)
di: Kawakami, Wataru, et al.
Pubblicazione: (2025)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
di: Turk, Matt
Pubblicazione: (2026)
di: Turk, Matt
Pubblicazione: (2026)
CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
di: Li, Junsong, et al.
Pubblicazione: (2025)
di: Li, Junsong, et al.
Pubblicazione: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
di: Lai, Xin, et al.
Pubblicazione: (2024)
di: Lai, Xin, et al.
Pubblicazione: (2024)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
di: Xie, Shuo, et al.
Pubblicazione: (2024)
di: Xie, Shuo, et al.
Pubblicazione: (2024)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
di: Dang, John, et al.
Pubblicazione: (2024)
di: Dang, John, et al.
Pubblicazione: (2024)
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
di: Sun, Ao, et al.
Pubblicazione: (2026)
di: Sun, Ao, et al.
Pubblicazione: (2026)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
di: Wu, Fang, et al.
Pubblicazione: (2025)
di: Wu, Fang, et al.
Pubblicazione: (2025)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
di: Zhang, Yuheng, et al.
Pubblicazione: (2025)
di: Zhang, Yuheng, et al.
Pubblicazione: (2025)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
di: Feng, Guhao, et al.
Pubblicazione: (2024)
di: Feng, Guhao, et al.
Pubblicazione: (2024)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
di: Seo, Wooseok, et al.
Pubblicazione: (2025)
di: Seo, Wooseok, et al.
Pubblicazione: (2025)
Tool Preferences in Agentic LLMs are Unreliable
di: Faghih, Kazem, et al.
Pubblicazione: (2025)
di: Faghih, Kazem, et al.
Pubblicazione: (2025)
Exploring Design Choices for Building Language-Specific LLMs
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
di: Misra, Dipendra, et al.
Pubblicazione: (2026)
di: Misra, Dipendra, et al.
Pubblicazione: (2026)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
di: Liu, Mingjie, et al.
Pubblicazione: (2025)
di: Liu, Mingjie, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
di: Chiu, Yu Ying, et al.
Pubblicazione: (2025) -
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024) -
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
di: Jung, Jaehun, et al.
Pubblicazione: (2023) -
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
di: Liu, Jiacheng, et al.
Pubblicazione: (2023) -
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
di: Deng, Yuntian, et al.
Pubblicazione: (2024)