Clone-Robust AI Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Procaccia, Ariel D., Schiffer, Benjamin, Zhang, Shirley |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Finding Common Ground in a Sea of Alternatives
por: Chooi, Jay, et al.
Publicado: (2026)
por: Chooi, Jay, et al.
Publicado: (2026)
Honor Among Bandits: No-Regret Learning for Online Fair Division
por: Procaccia, Ariel D., et al.
Publicado: (2024)
por: Procaccia, Ariel D., et al.
Publicado: (2024)
Metritocracy: Representative Metrics for Lite Benchmarks
por: Procaccia, Ariel, et al.
Publicado: (2025)
por: Procaccia, Ariel, et al.
Publicado: (2025)
Axioms for AI Alignment from Human Feedback
por: Ge, Luise, et al.
Publicado: (2024)
por: Ge, Luise, et al.
Publicado: (2024)
Multi-Apartment Rent Division
por: Procaccia, Ariel D., et al.
Publicado: (2024)
por: Procaccia, Ariel D., et al.
Publicado: (2024)
Generative Social Choice: The Next Generation
por: Boehmer, Niclas, et al.
Publicado: (2025)
por: Boehmer, Niclas, et al.
Publicado: (2025)
Policy Aggregation
por: Alamdari, Parand A., et al.
Publicado: (2024)
por: Alamdari, Parand A., et al.
Publicado: (2024)
Adaptive Contracts for Cost-Effective AI Delegation
por: Saig, Eden, et al.
Publicado: (2026)
por: Saig, Eden, et al.
Publicado: (2026)
Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies
por: Assos, Angelos, et al.
Publicado: (2025)
por: Assos, Angelos, et al.
Publicado: (2025)
Improved Regret Bounds for Online Fair Division with Bandit Learning
por: Schiffer, Benjamin, et al.
Publicado: (2025)
por: Schiffer, Benjamin, et al.
Publicado: (2025)
Generative Social Choice
por: Fish, Sara, et al.
Publicado: (2023)
por: Fish, Sara, et al.
Publicado: (2023)
Strategic Classification With Externalities
por: Hossain, Safwan, et al.
Publicado: (2024)
por: Hossain, Safwan, et al.
Publicado: (2024)
Learning Social Welfare Functions
por: Pardeshi, Kanad Shrikar, et al.
Publicado: (2024)
por: Pardeshi, Kanad Shrikar, et al.
Publicado: (2024)
Geometry Meets Incentives: Sample-Efficient Incentivized Exploration with Linear Contexts
por: Schiffer, Benjamin, et al.
Publicado: (2025)
por: Schiffer, Benjamin, et al.
Publicado: (2025)
Strategic Candidacy in Generative AI Arenas
por: Hays, Chris, et al.
Publicado: (2026)
por: Hays, Chris, et al.
Publicado: (2026)
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
por: Cai, Yang, et al.
Publicado: (2026)
por: Cai, Yang, et al.
Publicado: (2026)
Governing AI Forgetting: Auditing for Machine Unlearning Compliance
por: Lin, Qinqi, et al.
Publicado: (2026)
por: Lin, Qinqi, et al.
Publicado: (2026)
A Simple, Solid, and Reproducible Baseline for Bridge Bidding AI
por: Kita, Haruka, et al.
Publicado: (2024)
por: Kita, Haruka, et al.
Publicado: (2024)
On the Impact of the Utility in Semivalue-based Data Valuation
por: Tamine, Mélissa, et al.
Publicado: (2025)
por: Tamine, Mélissa, et al.
Publicado: (2025)
Robust Deep Monte Carlo Counterfactual Regret Minimization: Addressing Theoretical Risks in Neural Fictitious Self-Play
por: Jaafari, Zakaria El
Publicado: (2025)
por: Jaafari, Zakaria El
Publicado: (2025)
Independence of Approximate Clones
por: Delemazure, Théo
Publicado: (2026)
por: Delemazure, Théo
Publicado: (2026)
Representative Social Choice: From Learning Theory to AI Alignment
por: Qiu, Tianyi
Publicado: (2024)
por: Qiu, Tianyi
Publicado: (2024)
Language Alignment via Nash-learning and Adaptive feedback
por: Azarafrooz, Ari, et al.
Publicado: (2024)
por: Azarafrooz, Ari, et al.
Publicado: (2024)
Clone-Robust Weights in Metric Spaces: Handling Redundancy Bias for Benchmark Aggregation
por: Berriaud, Damien, et al.
Publicado: (2025)
por: Berriaud, Damien, et al.
Publicado: (2025)
Alignment as Institutional Design: From Behavioral Correction to Transaction Structure in Intelligent Systems
por: Chai, Rui
Publicado: (2026)
por: Chai, Rui
Publicado: (2026)
DU-Shapley: A Shapley Value Proxy for Efficient Dataset Valuation
por: Garrido-Lucero, Felipe, et al.
Publicado: (2023)
por: Garrido-Lucero, Felipe, et al.
Publicado: (2023)
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
por: Li, Sijia, et al.
Publicado: (2026)
por: Li, Sijia, et al.
Publicado: (2026)
Optimal Correlated Equilibria in General-Sum Extensive-Form Games: Fixed-Parameter Algorithms, Hardness, and Two-Sided Column-Generation
por: Zhang, Brian, et al.
Publicado: (2022)
por: Zhang, Brian, et al.
Publicado: (2022)
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
por: Park, Chanwoo, et al.
Publicado: (2024)
por: Park, Chanwoo, et al.
Publicado: (2024)
Spot Check Equivalence: an Interpretable Metric for Information Elicitation Mechanisms
por: Xu, Shengwei, et al.
Publicado: (2024)
por: Xu, Shengwei, et al.
Publicado: (2024)
Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis
por: Nayebi, Aran
Publicado: (2025)
por: Nayebi, Aran
Publicado: (2025)
Recommending Best Paper Awards for ML/AI Conferences via the Isotonic Mechanism
por: Wen, Garrett G., et al.
Publicado: (2026)
por: Wen, Garrett G., et al.
Publicado: (2026)
PerfectDou: Dominating DouDizhu with Perfect Information Distillation
por: Yang, Guan, et al.
Publicado: (2022)
por: Yang, Guan, et al.
Publicado: (2022)
Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables
por: Chen, Yang, et al.
Publicado: (2025)
por: Chen, Yang, et al.
Publicado: (2025)
Efficient Last-iterate Convergence Algorithms in Solving Games
por: Meng, Linjian, et al.
Publicado: (2023)
por: Meng, Linjian, et al.
Publicado: (2023)
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
por: La Malfa, Gabriele, et al.
Publicado: (2026)
por: La Malfa, Gabriele, et al.
Publicado: (2026)
Socially-Weighted Alignment: A Game-Theoretic Framework for Multi-Agent LLM Systems
por: Mumcu, Furkan, et al.
Publicado: (2026)
por: Mumcu, Furkan, et al.
Publicado: (2026)
Mechanism-Based Intelligence (MBI): Differentiable Incentives for Rational Coordination and Guaranteed Alignment in Multi-Agent Systems
por: Grassi, Stefano
Publicado: (2025)
por: Grassi, Stefano
Publicado: (2025)
Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
por: Wolf, Will
Publicado: (2025)
por: Wolf, Will
Publicado: (2025)
LLM-Auction: Generative Auction towards LLM-Native Advertising
por: Zhao, Chujie, et al.
Publicado: (2025)
por: Zhao, Chujie, et al.
Publicado: (2025)
Ejemplares similares
-
Finding Common Ground in a Sea of Alternatives
por: Chooi, Jay, et al.
Publicado: (2026) -
Honor Among Bandits: No-Regret Learning for Online Fair Division
por: Procaccia, Ariel D., et al.
Publicado: (2024) -
Metritocracy: Representative Metrics for Lite Benchmarks
por: Procaccia, Ariel, et al.
Publicado: (2025) -
Axioms for AI Alignment from Human Feedback
por: Ge, Luise, et al.
Publicado: (2024) -
Multi-Apartment Rent Division
por: Procaccia, Ariel D., et al.
Publicado: (2024)