A Roadmap to Pluralistic Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Sorensen, Taylor, Moore, Jared, Fisher, Jillian, Gordon, Mitchell, Mireshghallah, Niloofar, Rytting, Christopher Michael, Ye, Andre, Jiang, Liwei, Lu, Ximing, Dziri, Nouha, Althoff, Tim, Choi, Yejin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
por: Sorensen, Taylor, et al.
Publicado: (2025)
por: Sorensen, Taylor, et al.
Publicado: (2025)
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
por: Feng, Shangbin, et al.
Publicado: (2024)
por: Feng, Shangbin, et al.
Publicado: (2024)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
por: Sorensen, Taylor, et al.
Publicado: (2023)
por: Sorensen, Taylor, et al.
Publicado: (2023)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
por: Ravichander, Abhilasha, et al.
Publicado: (2025)
por: Ravichander, Abhilasha, et al.
Publicado: (2025)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
por: Lu, Ximing, et al.
Publicado: (2024)
por: Lu, Ximing, et al.
Publicado: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
por: Jiang, Liwei, et al.
Publicado: (2024)
por: Jiang, Liwei, et al.
Publicado: (2024)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
por: Jung, Jaehun, et al.
Publicado: (2023)
por: Jung, Jaehun, et al.
Publicado: (2023)
StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements
por: Fisher, Jillian, et al.
Publicado: (2024)
por: Fisher, Jillian, et al.
Publicado: (2024)
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
por: Fisher, Jillian, et al.
Publicado: (2024)
por: Fisher, Jillian, et al.
Publicado: (2024)
Similarity-Guided Diffusion for Contrastive Sequential Recommendation
por: Choi, Jinkyeong, et al.
Publicado: (2025)
por: Choi, Jinkyeong, et al.
Publicado: (2025)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
por: Li, Huihan, et al.
Publicado: (2024)
por: Li, Huihan, et al.
Publicado: (2024)
Can Language Models Reason about Individualistic Human Values and Preferences?
por: Jiang, Liwei, et al.
Publicado: (2024)
por: Jiang, Liwei, et al.
Publicado: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
por: Rao, Kavel, et al.
Publicado: (2023)
por: Rao, Kavel, et al.
Publicado: (2023)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
por: Qiu, Linlu, et al.
Publicado: (2023)
por: Qiu, Linlu, et al.
Publicado: (2023)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
por: Mireshghallah, Niloofar, et al.
Publicado: (2024)
por: Mireshghallah, Niloofar, et al.
Publicado: (2024)
Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
por: Sorensen, Taylor, et al.
Publicado: (2025)
por: Sorensen, Taylor, et al.
Publicado: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
por: Han, Seungju, et al.
Publicado: (2024)
por: Han, Seungju, et al.
Publicado: (2024)
Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
por: Choi, Yejin, et al.
Publicado: (2025)
por: Choi, Yejin, et al.
Publicado: (2025)
A Recommender System for NFT Collectibles with Item Feature
por: Choi, Minjoo, et al.
Publicado: (2024)
por: Choi, Minjoo, et al.
Publicado: (2024)
An Ecosystem for Personal Knowledge Graphs: A Survey and Research Roadmap
por: Skjæveland, Martin G., et al.
Publicado: (2023)
por: Skjæveland, Martin G., et al.
Publicado: (2023)
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
por: Liu, Jiacheng, et al.
Publicado: (2024)
por: Liu, Jiacheng, et al.
Publicado: (2024)
Towards A Tri-View Diffusion Framework for Recommendation
por: Chen, Ximing, et al.
Publicado: (2025)
por: Chen, Ximing, et al.
Publicado: (2025)
Gaussian Mixture Flow Matching with Domain Alignment for Multi-Domain Sequential Recommendation
por: Ye, Xiaoxin, et al.
Publicado: (2025)
por: Ye, Xiaoxin, et al.
Publicado: (2025)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
por: Baheti, Ashutosh, et al.
Publicado: (2024)
por: Baheti, Ashutosh, et al.
Publicado: (2024)
Align$^3$GR: Unified Multi-Level Alignment for LLM-based Generative Recommendation
por: Ye, Wencai, et al.
Publicado: (2025)
por: Ye, Wencai, et al.
Publicado: (2025)
Adversarial Alignment and Disentanglement for Cross-Domain CTR Prediction with Domain-Encompassing Features
por: He, Junyou, et al.
Publicado: (2026)
por: He, Junyou, et al.
Publicado: (2026)
Multi-objective Learning to Rank by Model Distillation
por: Tang, Jie, et al.
Publicado: (2024)
por: Tang, Jie, et al.
Publicado: (2024)
AI Co-Scientist for Ranking: Discovering Novel Search Ranking Models alongside LLM-based AI Agents with Cloud Computing Access
por: Wu, Liwei, et al.
Publicado: (2026)
por: Wu, Liwei, et al.
Publicado: (2026)
FinAI Data Assistant: LLM-based Financial Database Query Processing with the OpenAI Function Calling API
por: Kim, Juhyeong, et al.
Publicado: (2025)
por: Kim, Juhyeong, et al.
Publicado: (2025)
A Survey on Sequential Recommendation
por: Pan, Liwei, et al.
Publicado: (2024)
por: Pan, Liwei, et al.
Publicado: (2024)
Beyond Pairwise Learning-To-Rank At Airbnb
por: Haldar, Malay, et al.
Publicado: (2025)
por: Haldar, Malay, et al.
Publicado: (2025)
Cold-Start Recommendation towards the Era of Large Language Models (LLMs): A Comprehensive Survey and Roadmap
por: Zhang, Weizhi, et al.
Publicado: (2025)
por: Zhang, Weizhi, et al.
Publicado: (2025)
Position: Privacy Is Not Just Memorization!
por: Mireshghallah, Niloofar, et al.
Publicado: (2025)
por: Mireshghallah, Niloofar, et al.
Publicado: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
por: Naseh, Ali, et al.
Publicado: (2025)
por: Naseh, Ali, et al.
Publicado: (2025)
IRA: Adaptive Interest-aware Representation and Alignment for Personalized Multi-interest Retrieval
por: Lee, Youngjune, et al.
Publicado: (2025)
por: Lee, Youngjune, et al.
Publicado: (2025)
LLM-Alignment Live-Streaming Recommendation
por: Liu, Yueyang, et al.
Publicado: (2025)
por: Liu, Yueyang, et al.
Publicado: (2025)
Collaborative Semantic Alignment in Recommendation Systems
por: Wang, Chen, et al.
Publicado: (2023)
por: Wang, Chen, et al.
Publicado: (2023)
Quantizing Intent: Cross-Domain Semantic IDs from Organic Activity for Industrial Ranking
por: Choi, Julie, et al.
Publicado: (2026)
por: Choi, Julie, et al.
Publicado: (2026)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
por: Jiang, Liwei, et al.
Publicado: (2025)
por: Jiang, Liwei, et al.
Publicado: (2025)
Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models
por: Choi, Jaekeol
Publicado: (2024)
por: Choi, Jaekeol
Publicado: (2024)
Ejemplares similares
-
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
por: Sorensen, Taylor, et al.
Publicado: (2025) -
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
por: Feng, Shangbin, et al.
Publicado: (2024) -
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
por: Sorensen, Taylor, et al.
Publicado: (2023) -
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
por: Ravichander, Abhilasha, et al.
Publicado: (2025) -
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
por: Lu, Ximing, et al.
Publicado: (2024)