Strong and weak alignment of large language models with human values
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khamassi, Mehdi, Nahon, Marceau, Chatila, Raja |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025)
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025)
Semantic Deception: When Reasoning Models Can't Compute an Addition
von: de Leeuw, Nathaniël, et al.
Veröffentlicht: (2025)
von: de Leeuw, Nathaniël, et al.
Veröffentlicht: (2025)
Are they human? Detecting large language models by probing human memory constraints
von: Schug, Simon, et al.
Veröffentlicht: (2026)
von: Schug, Simon, et al.
Veröffentlicht: (2026)
Dissociating language and thought in large language models
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
von: Zhong, Shu, et al.
Veröffentlicht: (2024)
von: Zhong, Shu, et al.
Veröffentlicht: (2024)
Comparing large language models and human programmers for generating programming code
von: Hou, Wenpin, et al.
Veröffentlicht: (2024)
von: Hou, Wenpin, et al.
Veröffentlicht: (2024)
Post-training makes large language models less human-like
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
A closer look at how large language models trust humans: patterns and biases
von: Lerman, Valeria, et al.
Veröffentlicht: (2025)
von: Lerman, Valeria, et al.
Veröffentlicht: (2025)
Language models align with human judgments on key grammatical constructions
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
On the attribution of confidence to large language models
von: Keeling, Geoff, et al.
Veröffentlicht: (2024)
von: Keeling, Geoff, et al.
Veröffentlicht: (2024)
Superhuman performance of a large language model on the reasoning tasks of a physician
von: Brodeur, Peter G., et al.
Veröffentlicht: (2024)
von: Brodeur, Peter G., et al.
Veröffentlicht: (2024)
Multi-round jailbreak attack on large language models
von: Zhou, Yihua, et al.
Veröffentlicht: (2024)
von: Zhou, Yihua, et al.
Veröffentlicht: (2024)
The 20 questions game to distinguish large language models
von: Richardeau, Gurvan, et al.
Veröffentlicht: (2024)
von: Richardeau, Gurvan, et al.
Veröffentlicht: (2024)
Quantifying non deterministic drift in large language models
von: Nicholson, Claire
Veröffentlicht: (2026)
von: Nicholson, Claire
Veröffentlicht: (2026)
Can large language models build causal graphs?
von: Long, Stephanie, et al.
Veröffentlicht: (2023)
von: Long, Stephanie, et al.
Veröffentlicht: (2023)
Response: Emergent analogical reasoning in large language models
von: Hodel, Damian, et al.
Veröffentlicht: (2023)
von: Hodel, Damian, et al.
Veröffentlicht: (2023)
Representation in large language models
von: Yetman, Cameron
Veröffentlicht: (2025)
von: Yetman, Cameron
Veröffentlicht: (2025)
Failure of contextual invariance in large language models
von: Kumar, Sagar, et al.
Veröffentlicht: (2026)
von: Kumar, Sagar, et al.
Veröffentlicht: (2026)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
von: Nusrat, Humza, et al.
Veröffentlicht: (2025)
von: Nusrat, Humza, et al.
Veröffentlicht: (2025)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
von: Diaz-Garcia, Jose A., et al.
Veröffentlicht: (2025)
von: Diaz-Garcia, Jose A., et al.
Veröffentlicht: (2025)
Correcting misinformation on social media with a large language model
von: Zhou, Xinyi, et al.
Veröffentlicht: (2024)
von: Zhou, Xinyi, et al.
Veröffentlicht: (2024)
A review on the use of large language models as virtual tutors
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
Evaluating large language models in medical applications: a survey
von: Chen, Xiaolan, et al.
Veröffentlicht: (2024)
von: Chen, Xiaolan, et al.
Veröffentlicht: (2024)
MathDivide: Improved mathematical reasoning by large language models
von: Srivastava, Saksham Sahai, et al.
Veröffentlicht: (2024)
von: Srivastava, Saksham Sahai, et al.
Veröffentlicht: (2024)
Streamlining evidence based clinical recommendations with large language models
von: Li, Dubai, et al.
Veröffentlicht: (2025)
von: Li, Dubai, et al.
Veröffentlicht: (2025)
Re-evaluating Theory of Mind evaluation in large language models
von: Hu, Jennifer, et al.
Veröffentlicht: (2025)
von: Hu, Jennifer, et al.
Veröffentlicht: (2025)
Disentangling generalization and memorization in large language models using chess
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
von: Amouyal, Samuel Joseph, et al.
Veröffentlicht: (2025)
von: Amouyal, Samuel Joseph, et al.
Veröffentlicht: (2025)
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
AI-AI Bias: large language models favor communications generated by large language models
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
Alignment faking in large language models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
Optimizing watermarks for large language models
von: Wouters, Bram
Veröffentlicht: (2023)
von: Wouters, Bram
Veröffentlicht: (2023)
Cognitive models can reveal interpretable value trade-offs in language models
von: Murthy, Sonia K., et al.
Veröffentlicht: (2025)
von: Murthy, Sonia K., et al.
Veröffentlicht: (2025)
Uncovering inequalities in new knowledge learning by large language models across different languages
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Entry-level guide to the use of large language models for medical research
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
Facilitating large language model Russian adaptation with Learned Embedding Propagation
von: Tikhomirov, Mikhail, et al.
Veröffentlicht: (2024)
von: Tikhomirov, Mikhail, et al.
Veröffentlicht: (2024)
MacBehaviour: An R package for behavioural experimentation on large language models
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
Hyacinth6B: A large language model for Traditional Chinese
von: Song, Chih-Wei, et al.
Veröffentlicht: (2024)
von: Song, Chih-Wei, et al.
Veröffentlicht: (2024)
Can large language models understand uncommon meanings of common words?
von: Wu, Jinyang, et al.
Veröffentlicht: (2024)
von: Wu, Jinyang, et al.
Veröffentlicht: (2024)
Generating bilingual example sentences with large language models as lexicography assistants
von: Merx, Raphael, et al.
Veröffentlicht: (2024)
von: Merx, Raphael, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025) -
Semantic Deception: When Reasoning Models Can't Compute an Addition
von: de Leeuw, Nathaniël, et al.
Veröffentlicht: (2025) -
Are they human? Detecting large language models by probing human memory constraints
von: Schug, Simon, et al.
Veröffentlicht: (2026) -
Dissociating language and thought in large language models
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023) -
TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
von: Zhong, Shu, et al.
Veröffentlicht: (2024)