A Voter-Based Stochastic Rejection-Method Framework for Asymptotically Safe Language Model Outputs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Watts, Jake R., Sokol, Joel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast Controlled Generation from Language Models with Adaptive Weighted Rejection Sampling
von: Lipkin, Benjamin, et al.
Veröffentlicht: (2025)
von: Lipkin, Benjamin, et al.
Veröffentlicht: (2025)
SLOT: Structuring the Output of Large Language Models
von: Wang, Darren Yow-Bang, et al.
Veröffentlicht: (2025)
von: Wang, Darren Yow-Bang, et al.
Veröffentlicht: (2025)
Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
von: Zollo, Thomas P., et al.
Veröffentlicht: (2023)
von: Zollo, Thomas P., et al.
Veröffentlicht: (2023)
An Evaluation on Large Language Model Outputs: Discourse and Memorization
von: de Wynter, Adrian, et al.
Veröffentlicht: (2023)
von: de Wynter, Adrian, et al.
Veröffentlicht: (2023)
Leviathan: Decoupling Input and Output Representations in Language Models
von: Batley, Reza T., et al.
Veröffentlicht: (2026)
von: Batley, Reza T., et al.
Veröffentlicht: (2026)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
von: Jiang, Mingjian, et al.
Veröffentlicht: (2024)
von: Jiang, Mingjian, et al.
Veröffentlicht: (2024)
Constrained Adaptive Rejection Sampling
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
von: Zhao, Stephen, et al.
Veröffentlicht: (2025)
von: Zhao, Stephen, et al.
Veröffentlicht: (2025)
Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
von: Golowich, Noah, et al.
Veröffentlicht: (2026)
von: Golowich, Noah, et al.
Veröffentlicht: (2026)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
von: Fadeeva, Ekaterina, et al.
Veröffentlicht: (2024)
von: Fadeeva, Ekaterina, et al.
Veröffentlicht: (2024)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
von: Su, Hong
Veröffentlicht: (2025)
von: Su, Hong
Veröffentlicht: (2025)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
Phase Transitions in the Output Distribution of Large Language Models
von: Arnold, Julian, et al.
Veröffentlicht: (2024)
von: Arnold, Julian, et al.
Veröffentlicht: (2024)
AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2024)
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2024)
Pretrained Generative Language Models as General Learning Frameworks for Sequence-Based Tasks
von: Fauber, Ben
Veröffentlicht: (2024)
von: Fauber, Ben
Veröffentlicht: (2024)
A Markov Categorical Framework for Language Modeling
von: Zhang, Yifan
Veröffentlicht: (2025)
von: Zhang, Yifan
Veröffentlicht: (2025)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
Do Language Models Have Bayesian Brains? Distinguishing Stochastic and Deterministic Decision Patterns within Large Language Models
von: Cui, Andrea Yaoyun, et al.
Veröffentlicht: (2025)
von: Cui, Andrea Yaoyun, et al.
Veröffentlicht: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
Stochastic Parrots or ICU Experts? Large Language Models in Critical Care Medicine: A Scoping Review
von: Shi, Tongyue, et al.
Veröffentlicht: (2024)
von: Shi, Tongyue, et al.
Veröffentlicht: (2024)
RevOrder: A Novel Method for Enhanced Arithmetic in Language Models
von: Shen, Si, et al.
Veröffentlicht: (2024)
von: Shen, Si, et al.
Veröffentlicht: (2024)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
von: Dhaini, Mahdi, et al.
Veröffentlicht: (2025)
von: Dhaini, Mahdi, et al.
Veröffentlicht: (2025)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
Building Safe and Deployable Clinical Natural Language Processing under Temporal Leakage Constraints
von: Cho, Ha Na, et al.
Veröffentlicht: (2026)
von: Cho, Ha Na, et al.
Veröffentlicht: (2026)
Understanding Token Probability Encoding in Output Embeddings
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
Output Embedding Centering for Stable LLM Pretraining
von: Stollenwerk, Felix, et al.
Veröffentlicht: (2026)
von: Stollenwerk, Felix, et al.
Veröffentlicht: (2026)
SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
von: Chen, Yurun, et al.
Veröffentlicht: (2026)
von: Chen, Yurun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Fast Controlled Generation from Language Models with Adaptive Weighted Rejection Sampling
von: Lipkin, Benjamin, et al.
Veröffentlicht: (2025) -
SLOT: Structuring the Output of Large Language Models
von: Wang, Darren Yow-Bang, et al.
Veröffentlicht: (2025) -
Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
von: Zollo, Thomas P., et al.
Veröffentlicht: (2023) -
An Evaluation on Large Language Model Outputs: Discourse and Memorization
von: de Wynter, Adrian, et al.
Veröffentlicht: (2023) -
Leviathan: Decoupling Input and Output Representations in Language Models
von: Batley, Reza T., et al.
Veröffentlicht: (2026)