This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wolf, Lorenz, Yoon, Sangwoong, Bogunovic, Ilija |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
Robust Multi-Objective Controlled Decoding of Large Language Models
von: Son, Seongho, et al.
Veröffentlicht: (2025)
von: Son, Seongho, et al.
Veröffentlicht: (2025)
RSPO: Regularized Self-Play Alignment of Large Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2026)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2026)
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
von: Ziomek, Juliusz, et al.
Veröffentlicht: (2026)
von: Ziomek, Juliusz, et al.
Veröffentlicht: (2026)
Should You Use Your Large Language Model to Explore or Exploit?
von: Harris, Keegan, et al.
Veröffentlicht: (2025)
von: Harris, Keegan, et al.
Veröffentlicht: (2025)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
von: Kumar, Sachin
Veröffentlicht: (2026)
von: Kumar, Sachin
Veröffentlicht: (2026)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
von: Yuan, Danlong, et al.
Veröffentlicht: (2025)
von: Yuan, Danlong, et al.
Veröffentlicht: (2025)
An Assessment of Model-On-Model Deception
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
Adversarially Robust Decision Transformer
von: Tang, Xiaohang, et al.
Veröffentlicht: (2024)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2024)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
von: Mavi, John, et al.
Veröffentlicht: (2024)
von: Mavi, John, et al.
Veröffentlicht: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Pastiche Novel Generation Creating: Fan Fiction You Love in Your Favorite Author's Style
von: Han, Xueran, et al.
Veröffentlicht: (2025)
von: Han, Xueran, et al.
Veröffentlicht: (2025)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
von: Liu, Yanjiang, et al.
Veröffentlicht: (2026)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2026)
Post-training Large Language Models for Diverse High-Quality Responses
von: Chen, Yilei, et al.
Veröffentlicht: (2025)
von: Chen, Yilei, et al.
Veröffentlicht: (2025)
Your AI, Not Your View: The Bias of LLMs in Investment Analysis
von: Lee, Hoyoung, et al.
Veröffentlicht: (2025)
von: Lee, Hoyoung, et al.
Veröffentlicht: (2025)
To Tell The Truth: Language of Deception and Language Models
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
Syntriever: How to Train Your Retriever with Synthetic Data from LLMs
von: Kim, Minsang, et al.
Veröffentlicht: (2025)
von: Kim, Minsang, et al.
Veröffentlicht: (2025)
Protecting Your LLMs with Information Bottleneck
von: Liu, Zichuan, et al.
Veröffentlicht: (2024)
von: Liu, Zichuan, et al.
Veröffentlicht: (2024)
Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification
von: Demir, M. Mikail, et al.
Veröffentlicht: (2026)
von: Demir, M. Mikail, et al.
Veröffentlicht: (2026)
Overton Pluralistic Reinforcement Learning for Large Language Models
von: Fu, Yu, et al.
Veröffentlicht: (2026)
von: Fu, Yu, et al.
Veröffentlicht: (2026)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
Exploring the Latest LLMs for Leaderboard Extraction
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
von: Li, Houyi, et al.
Veröffentlicht: (2025)
von: Li, Houyi, et al.
Veröffentlicht: (2025)
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication
von: Ahmad, Arif, et al.
Veröffentlicht: (2024)
von: Ahmad, Arif, et al.
Veröffentlicht: (2024)
YouTube Comments Decoded: Leveraging LLMs for Low Resource Language Classification
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs
von: Kansal, Yuval, et al.
Veröffentlicht: (2025)
von: Kansal, Yuval, et al.
Veröffentlicht: (2025)
Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs
von: Bouchard, Dylan
Veröffentlicht: (2024)
von: Bouchard, Dylan
Veröffentlicht: (2024)
Weaker LLMs' Opinions Also Matter: Mixture of Opinions Enhances LLM's Mathematical Reasoning
von: Chen, Yanan, et al.
Veröffentlicht: (2025)
von: Chen, Yanan, et al.
Veröffentlicht: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025) -
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026) -
Robust Multi-Objective Controlled Decoding of Large Language Models
von: Son, Seongho, et al.
Veröffentlicht: (2025) -
RSPO: Regularized Self-Play Alignment of Large Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025) -
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2026)