Towards Integrated Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Reis, Ben Y., La Cava, William |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
par: Greco, Candida M., et autres
Publié: (2026)
par: Greco, Candida M., et autres
Publié: (2026)
Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling
par: La Cava, Lucio, et autres
Publié: (2026)
par: La Cava, Lucio, et autres
Publié: (2026)
Toward Preference-aligned Large Language Models via Residual-based Model Steering
par: La Cava, Lucio, et autres
Publié: (2025)
par: La Cava, Lucio, et autres
Publié: (2025)
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
par: La Cava, Lucio, et autres
Publié: (2025)
par: La Cava, Lucio, et autres
Publié: (2025)
Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models
par: La Cava, Lucio, et autres
Publié: (2024)
par: La Cava, Lucio, et autres
Publié: (2024)
Is Contrasting All You Need? Contrastive Learning for the Detection and Attribution of AI-generated Text
par: La Cava, Lucio, et autres
Publié: (2024)
par: La Cava, Lucio, et autres
Publié: (2024)
Talking the Talk Does Not Entail Walking the Walk: On the Limits of Large Language Models in Lexical Entailment Recognition
par: Greco, Candida M., et autres
Publié: (2024)
par: Greco, Candida M., et autres
Publié: (2024)
Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem
par: LaCroix, Travis
Publié: (2026)
par: LaCroix, Travis
Publié: (2026)
Authorship Attribution in Multilingual Machine-Generated Texts
par: La Cava, Lucio, et autres
Publié: (2025)
par: La Cava, Lucio, et autres
Publié: (2025)
Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
par: Konya, Andrew, et autres
Publié: (2024)
par: Konya, Andrew, et autres
Publié: (2024)
Visually Wired NFTs: Exploring the Role of Inspiration in Non-Fungible Tokens
par: La Cava, Lucio, et autres
Publié: (2023)
par: La Cava, Lucio, et autres
Publié: (2023)
Slurry-as-a-Service: A Modest Proposal on Scalable Pluralistic Alignment for Nutrient Optimization
par: Hong, Rachel, et autres
Publié: (2026)
par: Hong, Rachel, et autres
Publié: (2026)
The AI Alignment Paradox
par: West, Robert, et autres
Publié: (2024)
par: West, Robert, et autres
Publié: (2024)
Rethinking AI Cultural Alignment
par: Bravansky, Michal, et autres
Publié: (2025)
par: Bravansky, Michal, et autres
Publié: (2025)
The Emotional Alignment Design Policy
par: Schwitzgebel, Eric, et autres
Publié: (2025)
par: Schwitzgebel, Eric, et autres
Publié: (2025)
Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI
par: Alamdari, Parand A., et autres
Publié: (2024)
par: Alamdari, Parand A., et autres
Publié: (2024)
Ethical Challenges and Evolving Strategies in the Integration of Artificial Intelligence into Clinical Practice
par: Weiner, Ellison B., et autres
Publié: (2024)
par: Weiner, Ellison B., et autres
Publié: (2024)
Societal Alignment Frameworks Can Improve LLM Alignment
par: Stańczak, Karolina, et autres
Publié: (2025)
par: Stańczak, Karolina, et autres
Publié: (2025)
An Evaluation of Cultural Value Alignment in LLM
par: Sukiennik, Nicholas, et autres
Publié: (2025)
par: Sukiennik, Nicholas, et autres
Publié: (2025)
Justifications for Democratizing AI Alignment and Their Prospects
par: Steingrüber, André, et autres
Publié: (2025)
par: Steingrüber, André, et autres
Publié: (2025)
Scopes of Alignment
par: Varshney, Kush R., et autres
Publié: (2025)
par: Varshney, Kush R., et autres
Publié: (2025)
Resurrecting Socrates in the Age of AI: A Study Protocol for Evaluating a Socratic Tutor to Support Research Question Development in Higher Education
par: Degen, Ben
Publié: (2025)
par: Degen, Ben
Publié: (2025)
Understanding the Process of Human-AI Value Alignment
par: McKinlay, Jack, et autres
Publié: (2025)
par: McKinlay, Jack, et autres
Publié: (2025)
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
par: Janowicz, Krzysztof, et autres
Publié: (2025)
par: Janowicz, Krzysztof, et autres
Publié: (2025)
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
par: Ziheng, Zhou, et autres
Publié: (2026)
par: Ziheng, Zhou, et autres
Publié: (2026)
Dynamic Normativity: Necessary and Sufficient Conditions for Value Alignment
par: Corrêa, Nicholas Kluge
Publié: (2024)
par: Corrêa, Nicholas Kluge
Publié: (2024)
AI and Human Oversight: A Risk-Based Framework for Alignment
par: Kandikatla, Laxmiraju, et autres
Publié: (2025)
par: Kandikatla, Laxmiraju, et autres
Publié: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
par: Agarwal, Dhruv, et autres
Publié: (2025)
par: Agarwal, Dhruv, et autres
Publié: (2025)
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective
par: Tallam, Krti
Publié: (2025)
par: Tallam, Krti
Publié: (2025)
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
par: Choi, Dasol, et autres
Publié: (2026)
par: Choi, Dasol, et autres
Publié: (2026)
Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks
par: Tlaie, Alejandro
Publié: (2024)
par: Tlaie, Alejandro
Publié: (2024)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
par: Brophy, Matthew
Publié: (2025)
par: Brophy, Matthew
Publié: (2025)
Privacy Ethics Alignment in AI: A Stakeholder-Centric Framework for Ethical AI
par: Barthwal, Ankur, et autres
Publié: (2025)
par: Barthwal, Ankur, et autres
Publié: (2025)
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
par: Tang, Zhenheng, et autres
Publié: (2026)
par: Tang, Zhenheng, et autres
Publié: (2026)
Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
par: Pasandi, Faezeh B., et autres
Publié: (2026)
par: Pasandi, Faezeh B., et autres
Publié: (2026)
The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment
par: Motnikar, Lenart, et autres
Publié: (2025)
par: Motnikar, Lenart, et autres
Publié: (2025)
Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment
par: Caputo, Nicholas A.
Publié: (2024)
par: Caputo, Nicholas A.
Publié: (2024)
Alignment as Iatrogenesis: Pastoral Power, Collective Pathology, and the Structural Limits of Monolingual Safety Evaluation
par: Fukui, Hiroki
Publié: (2026)
par: Fukui, Hiroki
Publié: (2026)
Integrating LLMs in Gamified Systems
par: Costa, Carlos J.
Publié: (2025)
par: Costa, Carlos J.
Publié: (2025)
Alignment as Jurisprudence
par: Caputo, Nicholas
Publié: (2026)
par: Caputo, Nicholas
Publié: (2026)
Documents similaires
-
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
par: Greco, Candida M., et autres
Publié: (2026) -
Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling
par: La Cava, Lucio, et autres
Publié: (2026) -
Toward Preference-aligned Large Language Models via Residual-based Model Steering
par: La Cava, Lucio, et autres
Publié: (2025) -
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
par: La Cava, Lucio, et autres
Publié: (2025) -
Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models
par: La Cava, Lucio, et autres
Publié: (2024)