To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Fangyi, Seedat, Nabeel, Schwarz, Jonathan Richard, Bean, Andrew M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
von: Bean, Andrew M., et al.
Veröffentlicht: (2025)
von: Bean, Andrew M., et al.
Veröffentlicht: (2025)
Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
von: Yu, Fangyi, et al.
Veröffentlicht: (2025)
von: Yu, Fangyi, et al.
Veröffentlicht: (2025)
Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
Large Language Models to Enhance Bayesian Optimization
von: Liu, Tennison, et al.
Veröffentlicht: (2024)
von: Liu, Tennison, et al.
Veröffentlicht: (2024)
From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards
von: Qian, Chen, et al.
Veröffentlicht: (2026)
von: Qian, Chen, et al.
Veröffentlicht: (2026)
Do Large Language Models Align with Core Mental Health Counseling Competencies?
von: Nguyen, Viet Cuong, et al.
Veröffentlicht: (2024)
von: Nguyen, Viet Cuong, et al.
Veröffentlicht: (2024)
DC-Check: A Data-Centric AI checklist to guide the development of reliable machine learning systems
von: Seedat, Nabeel, et al.
Veröffentlicht: (2022)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2022)
You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling
von: Seedat, Nabeel, et al.
Veröffentlicht: (2024)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2024)
Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World Environments
von: Rauba, Paulius, et al.
Veröffentlicht: (2024)
von: Rauba, Paulius, et al.
Veröffentlicht: (2024)
When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
von: Yu, Fangyi
Veröffentlicht: (2025)
von: Yu, Fangyi
Veröffentlicht: (2025)
DAGnosis: Localized Identification of Data Inconsistencies using Structures
von: Huynh, Nicolas, et al.
Veröffentlicht: (2024)
von: Huynh, Nicolas, et al.
Veröffentlicht: (2024)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
von: Seedat, Nabeel, et al.
Veröffentlicht: (2023)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2023)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
von: Sun, Hao, et al.
Veröffentlicht: (2023)
von: Sun, Hao, et al.
Veröffentlicht: (2023)
Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings
von: Schwalbe, Gesina, et al.
Veröffentlicht: (2026)
von: Schwalbe, Gesina, et al.
Veröffentlicht: (2026)
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
von: Tak, Ala N., et al.
Veröffentlicht: (2026)
von: Tak, Ala N., et al.
Veröffentlicht: (2026)
CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives
von: Lee, Ayoung, et al.
Veröffentlicht: (2025)
von: Lee, Ayoung, et al.
Veröffentlicht: (2025)
Measuring and Aligning Abstraction in Vision-Language Models with Medical Taxonomies
von: Schaper, Ben, et al.
Veröffentlicht: (2026)
von: Schaper, Ben, et al.
Veröffentlicht: (2026)
Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback
von: Saveliev, Evgeny S., et al.
Veröffentlicht: (2026)
von: Saveliev, Evgeny S., et al.
Veröffentlicht: (2026)
Do Language Models Align with Brains? Prediction Scores Are Not Enough
von: Jia, Xiao
Veröffentlicht: (2026)
von: Jia, Xiao
Veröffentlicht: (2026)
Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
RT-H: Action Hierarchies Using Language
von: Belkhale, Suneel, et al.
Veröffentlicht: (2024)
von: Belkhale, Suneel, et al.
Veröffentlicht: (2024)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
von: Kabir, Muhammad Rafsan, et al.
Veröffentlicht: (2024)
von: Kabir, Muhammad Rafsan, et al.
Veröffentlicht: (2024)
Generative Models, Humans, Predictive Models: Who Is Worse at High-Stakes Decision Making?
von: Mallari, Keri, et al.
Veröffentlicht: (2024)
von: Mallari, Keri, et al.
Veröffentlicht: (2024)
Large Language Models are Highly Aligned with Human Ratings of Emotional Stimuli
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
AI Knows What's Wrong But Cannot Fix It: Helicoid Dynamics in Frontier LLMs Under High-Stakes Decisions
von: Jadad, Alejandro R
Veröffentlicht: (2026)
von: Jadad, Alejandro R
Veröffentlicht: (2026)
The Missing Memory Hierarchy: Demand Paging for LLM Context Windows
von: Mason, Tony
Veröffentlicht: (2026)
von: Mason, Tony
Veröffentlicht: (2026)
Fairness in Federated Learning: Fairness for Whom?
von: Taik, Afaf, et al.
Veröffentlicht: (2025)
von: Taik, Afaf, et al.
Veröffentlicht: (2025)
Language Models as Hierarchy Encoders
von: He, Yuan, et al.
Veröffentlicht: (2024)
von: He, Yuan, et al.
Veröffentlicht: (2024)
Model Cards for AI Teammates: Comparing Human-AI Team Familiarization Methods for High-Stakes Environments
von: Bowers, Ryan, et al.
Veröffentlicht: (2025)
von: Bowers, Ryan, et al.
Veröffentlicht: (2025)
Artificial Intelligence / Human Intelligence: Who Controls Whom?
von: Jacquemot, Charlotte
Veröffentlicht: (2025)
von: Jacquemot, Charlotte
Veröffentlicht: (2025)
Evaluating the role of `Constitutions' for learning from AI feedback
von: Redgate, Saskia, et al.
Veröffentlicht: (2024)
von: Redgate, Saskia, et al.
Veröffentlicht: (2024)
ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
von: Chen, Shengzhuang, et al.
Veröffentlicht: (2025)
von: Chen, Shengzhuang, et al.
Veröffentlicht: (2025)
Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures
von: Schwarz, Dominik
Veröffentlicht: (2025)
von: Schwarz, Dominik
Veröffentlicht: (2025)
Status Hierarchies in Language Models
von: Barkett, Emilio
Veröffentlicht: (2026)
von: Barkett, Emilio
Veröffentlicht: (2026)
Balancing Fidelity and Plasticity: Aligning Mixed-Precision Fine-Tuning with Linguistic Hierarchies
von: Zhou, Changhai, et al.
Veröffentlicht: (2025)
von: Zhou, Changhai, et al.
Veröffentlicht: (2025)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Whom to Trust? Elective Learning for Distributed Gaussian Process Regression
von: Yang, Zewen, et al.
Veröffentlicht: (2024)
von: Yang, Zewen, et al.
Veröffentlicht: (2024)
The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future
von: Saleem, Summra, et al.
Veröffentlicht: (2025)
von: Saleem, Summra, et al.
Veröffentlicht: (2025)
Large Language Models for Medical Forecasting -- Foresight 2
von: Kraljevic, Zeljko, et al.
Veröffentlicht: (2024)
von: Kraljevic, Zeljko, et al.
Veröffentlicht: (2024)
Brevity Constraints Reverse Performance Hierarchies in Language Models
von: Hakim, MD Azizul
Veröffentlicht: (2026)
von: Hakim, MD Azizul
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
von: Bean, Andrew M., et al.
Veröffentlicht: (2025) -
Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
von: Yu, Fangyi, et al.
Veröffentlicht: (2025) -
Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification
von: Zhang, Yichi, et al.
Veröffentlicht: (2026) -
Large Language Models to Enhance Bayesian Optimization
von: Liu, Tennison, et al.
Veröffentlicht: (2024) -
From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards
von: Qian, Chen, et al.
Veröffentlicht: (2026)