Consensus Sampling for Safer Generative AI
Fuente:
arXiv
Saved in:
| Main Authors: | Kalai, Adam Tauman, Kalai, Yael Tauman, Zamir, Or |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
by: Zelikman, Eric, et al.
Published: (2023)
by: Zelikman, Eric, et al.
Published: (2023)
Calibrated Language Models Must Hallucinate
by: Kalai, Adam Tauman, et al.
Published: (2023)
by: Kalai, Adam Tauman, et al.
Published: (2023)
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
Compiling Any $\mathsf{MIP}^{*}$ into a (Succinct) Classical Interactive Argument
by: Huang, Andrew, et al.
Published: (2025)
by: Huang, Andrew, et al.
Published: (2025)
Parallel Repetition for Post-Quantum Arguments
by: Huang, Andrew, et al.
Published: (2025)
by: Huang, Andrew, et al.
Published: (2025)
On Non-interactive Evaluation of Animal Communication Translators
by: Paradise, Orr, et al.
Published: (2025)
by: Paradise, Orr, et al.
Published: (2025)
Do Language Models Know When They're Hallucinating References?
by: Agrawal, Ayush, et al.
Published: (2023)
by: Agrawal, Ayush, et al.
Published: (2023)
How to Classically Verify a Quantum Cat without Killing It
by: Kalai, Yael Tauman, et al.
Published: (2026)
by: Kalai, Yael Tauman, et al.
Published: (2026)
Classical Commitments to Quantum States
by: Gunn, Sam, et al.
Published: (2024)
by: Gunn, Sam, et al.
Published: (2024)
Efficiently Batching Unambiguous Interactive Proofs
by: Berger, Bonnie, et al.
Published: (2025)
by: Berger, Bonnie, et al.
Published: (2025)
Why Language Models Hallucinate
by: Kalai, Adam Tauman, et al.
Published: (2025)
by: Kalai, Adam Tauman, et al.
Published: (2025)
Source Attribution in Retrieval-Augmented Generation
by: Nematov, Ikhtiyor, et al.
Published: (2025)
by: Nematov, Ikhtiyor, et al.
Published: (2025)
SleepPPG-Net2: Deep learning generalization for sleep staging from photoplethysmography
by: Attia, Shirel, et al.
Published: (2024)
by: Attia, Shirel, et al.
Published: (2024)
Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange
by: Vaikuntanathan, Vinod, et al.
Published: (2026)
by: Vaikuntanathan, Vinod, et al.
Published: (2026)
First-Person Fairness in Chatbots
by: Eloundou, Tyna, et al.
Published: (2024)
by: Eloundou, Tyna, et al.
Published: (2024)
Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
On Safer Reinforcement Learning for Sedation and Analgesia in Intensive Care
by: Romero-Hernandez, Joel, et al.
Published: (2026)
by: Romero-Hernandez, Joel, et al.
Published: (2026)
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning
by: Hazra, Somnath, et al.
Published: (2025)
by: Hazra, Somnath, et al.
Published: (2025)
Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
by: Yamabe, Shojiro, et al.
Published: (2025)
by: Yamabe, Shojiro, et al.
Published: (2025)
FactsR: A Safer Method for Producing High Quality Healthcare Documentation
by: Hansen, Victor Petrén Bach, et al.
Published: (2025)
by: Hansen, Victor Petrén Bach, et al.
Published: (2025)
Generalizing Differentially Private Decentralized Deep Learning with Multi-Agent Consensus
by: Bayrooti, Jasmine, et al.
Published: (2023)
by: Bayrooti, Jasmine, et al.
Published: (2023)
Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble
by: Kumthekar, Amit, et al.
Published: (2025)
by: Kumthekar, Amit, et al.
Published: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
by: K, Swaminathan S, et al.
Published: (2026)
by: K, Swaminathan S, et al.
Published: (2026)
Network Level Evaluation of Hangup Susceptibility of HRGCs using Deep Learning and Sensing Techniques: A Goal Towards Safer Future
by: Chatterjee, Kaustav, et al.
Published: (2025)
by: Chatterjee, Kaustav, et al.
Published: (2025)
Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
by: Pandit, Kartik, et al.
Published: (2025)
by: Pandit, Kartik, et al.
Published: (2025)
Aligning AI Agents via Information-Directed Sampling
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
Vector Field Oriented Diffusion Model for Crystal Material Generation
by: Klipfel, Astrid, et al.
Published: (2023)
by: Klipfel, Astrid, et al.
Published: (2023)
Theoretical Convergence of SMOTE-Generated Samples
by: Kamalov, Firuz, et al.
Published: (2026)
by: Kamalov, Firuz, et al.
Published: (2026)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
by: Barad, Haim, et al.
Published: (2023)
by: Barad, Haim, et al.
Published: (2023)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
by: Jiang, Shan, et al.
Published: (2026)
by: Jiang, Shan, et al.
Published: (2026)
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
by: Su, Zelal, et al.
Published: (2026)
by: Su, Zelal, et al.
Published: (2026)
Physics-Enhanced Deep Learning for Proactive Thermal Runaway Forecasting in Li-Ion Batteries
by: Khan, Salman, et al.
Published: (2026)
by: Khan, Salman, et al.
Published: (2026)
Synthetic Data Generation for Augmenting Small Samples
by: Liu, Dan, et al.
Published: (2025)
by: Liu, Dan, et al.
Published: (2025)
Scalable Equilibrium Sampling with Sequential Boltzmann Generators
by: Tan, Charlie B., et al.
Published: (2025)
by: Tan, Charlie B., et al.
Published: (2025)
CALM: Consensus-Aware Localized Merging for Multi-Task Learning
by: Yan, Kunda, et al.
Published: (2025)
by: Yan, Kunda, et al.
Published: (2025)
NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning
by: Jerge, Michael, et al.
Published: (2026)
by: Jerge, Michael, et al.
Published: (2026)
A Probabilistic Consensus-Driven Approach for Robust Counterfactual Explanations
by: Kostrzewa, Marcin, et al.
Published: (2026)
by: Kostrzewa, Marcin, et al.
Published: (2026)
From Generative AI to Innovative AI: An Evolutionary Roadmap
by: Mohammadabadi, Seyed Mahmoud Sajjadi
Published: (2025)
by: Mohammadabadi, Seyed Mahmoud Sajjadi
Published: (2025)
Effective Sample Size and Generalization Bounds for Temporal Networks
by: Gahtan, Barak, et al.
Published: (2025)
by: Gahtan, Barak, et al.
Published: (2025)
Similar Items
-
Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
by: Zelikman, Eric, et al.
Published: (2023) -
Calibrated Language Models Must Hallucinate
by: Kalai, Adam Tauman, et al.
Published: (2023) -
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
by: Suzgun, Mirac, et al.
Published: (2024) -
Compiling Any $\mathsf{MIP}^{*}$ into a (Succinct) Classical Interactive Argument
by: Huang, Andrew, et al.
Published: (2025) -
Parallel Repetition for Post-Quantum Arguments
by: Huang, Andrew, et al.
Published: (2025)