Control Tax: The Price of Keeping AI in Check
Fuente:
arXiv
Saved in:
| Main Authors: | Terekhov, Mikhail, Liu, Zhen Ning David, Gulcehre, Caglar, Albanie, Samuel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
by: Terekhov, Mikhail, et al.
Published: (2024)
by: Terekhov, Mikhail, et al.
Published: (2024)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025)
by: Terekhov, Mikhail, et al.
Published: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
Simple Hierarchical Planning with Diffusion
by: Chen, Chang, et al.
Published: (2024)
by: Chen, Chang, et al.
Published: (2024)
PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer
by: Chen, Chang, et al.
Published: (2024)
by: Chen, Chang, et al.
Published: (2024)
Learning pure quantum states (almost) without regret
by: Lumbreras, Josep, et al.
Published: (2024)
by: Lumbreras, Josep, et al.
Published: (2024)
Self-Recognition in Language Models
by: Davidson, Tim R., et al.
Published: (2024)
by: Davidson, Tim R., et al.
Published: (2024)
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
by: Karzanov, Daniil, et al.
Published: (2025)
by: Karzanov, Daniil, et al.
Published: (2025)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
by: Klein, Lars, et al.
Published: (2024)
by: Klein, Lars, et al.
Published: (2024)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
by: Deschenaux, Justin, et al.
Published: (2026)
by: Deschenaux, Justin, et al.
Published: (2026)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
Promises, Outlooks and Challenges of Diffusion Language Modeling
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
Keeping Medical AI Healthy and Trustworthy: A Review of Detection and Correction Methods for System Degradation
by: Guan, Hao, et al.
Published: (2025)
by: Guan, Hao, et al.
Published: (2025)
The Impact of Post-training on Data Contamination
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2026)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2026)
Contextual Bandit Optimization with Pre-Trained Neural Networks
by: Terekhov, Mikhail
Published: (2025)
by: Terekhov, Mikhail
Published: (2025)
Integrating Attention-Enhanced LSTM and Particle Swarm Optimization for Dynamic Pricing and Replenishment Strategies in Fresh Food Supermarkets
by: Liu, Xianchen, et al.
Published: (2025)
by: Liu, Xianchen, et al.
Published: (2025)
Sanity Checks for Explanation Uncertainty
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
Keep Rehearsing and Refining: Lifelong Learning Vehicle Routing under Continually Drifting Tasks
by: Pei, Jiyuan, et al.
Published: (2026)
by: Pei, Jiyuan, et al.
Published: (2026)
Partition Generative Modeling: Masked Modeling Without Masks
by: Deschenaux, Justin, et al.
Published: (2025)
by: Deschenaux, Justin, et al.
Published: (2025)
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
by: Matrenok, Simon, et al.
Published: (2025)
by: Matrenok, Simon, et al.
Published: (2025)
Taxon: Hierarchical Tax Code Prediction with Semantically Aligned LLM Expert Guidance
by: Li, Jihang, et al.
Published: (2026)
by: Li, Jihang, et al.
Published: (2026)
Active Learning for Continual Learning: Keeping the Past Alive in the Present
by: Park, Jaehyun, et al.
Published: (2025)
by: Park, Jaehyun, et al.
Published: (2025)
Checking extracted rules in Neural Networks
by: Wurm, Adrian
Published: (2025)
by: Wurm, Adrian
Published: (2025)
Sanity Checks for Agentic Data Science
by: Rewolinski, Zachary T., et al.
Published: (2026)
by: Rewolinski, Zachary T., et al.
Published: (2026)
seeBias: A Comprehensive Tool for Assessing and Visualizing AI Fairness
by: Ning, Yilin, et al.
Published: (2025)
by: Ning, Yilin, et al.
Published: (2025)
Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
by: Liu, Zhihan, et al.
Published: (2026)
by: Liu, Zhihan, et al.
Published: (2026)
Training with Confidence: Catching Silent Errors in Deep Learning Training with Automated Proactive Checks
by: Jiang, Yuxuan, et al.
Published: (2025)
by: Jiang, Yuxuan, et al.
Published: (2025)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
by: Liu, Mingyi
Published: (2026)
by: Liu, Mingyi
Published: (2026)
What Is the Alignment Tax?
by: Young, Robin
Published: (2026)
by: Young, Robin
Published: (2026)
Regression Models Meet Foundation Models: A Hybrid-AI Approach to Practical Electricity Price Forecasting
by: Qiu, Yunzhong, et al.
Published: (2026)
by: Qiu, Yunzhong, et al.
Published: (2026)
Embedding Knowledge Graphs in Degenerate Clifford Algebras
by: Teyou, Louis Mozart Kamdem, et al.
Published: (2024)
by: Teyou, Louis Mozart Kamdem, et al.
Published: (2024)
Embedding Knowledge Graph in Function Spaces
by: Teyou, Louis Mozart Kamdem, et al.
Published: (2024)
by: Teyou, Louis Mozart Kamdem, et al.
Published: (2024)
The Diffusion Duality, Chapter II: $Ψ$-Samplers
by: Deschenaux, Justin, et al.
Published: (2026)
by: Deschenaux, Justin, et al.
Published: (2026)
TaxDistill: Improving Metagenomic Taxonomic Annotation via Distilled Genomic Foundation Models
by: Ye, Rongye, et al.
Published: (2026)
by: Ye, Rongye, et al.
Published: (2026)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
by: Kiyani, Shayan, et al.
Published: (2026)
by: Kiyani, Shayan, et al.
Published: (2026)
Similar Items
-
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
by: Terekhov, Mikhail, et al.
Published: (2024) -
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025) -
The Role of Deep Learning Regularizations on Actors in Offline RL
by: Tarasov, Denis, et al.
Published: (2024) -
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024) -
Simple Hierarchical Planning with Diffusion
by: Chen, Chang, et al.
Published: (2024)