Gespeichert in:
| Hauptverfasser: | Joglekar, Manas, Chen, Jeremy, Wu, Gabriel, Yosinski, Jason, Wang, Jasmine, Barak, Boaz, Glaese, Amelia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2512.08093 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deliberative Alignment: Reasoning Enables Safer Language Models
von: Guan, Melody Y., et al.
Veröffentlicht: (2024)
von: Guan, Melody Y., et al.
Veröffentlicht: (2024)
Automatic Stability and Recovery for Neural Network Training
von: Or, Barak
Veröffentlicht: (2026)
von: Or, Barak
Veröffentlicht: (2026)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
Gradient-Free Training of Quantized Neural Networks
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
MESSI: A Multi-Elevation Semantic Segmentation Image Dataset of an Urban Environment
von: Pinkovich, Barak, et al.
Veröffentlicht: (2025)
von: Pinkovich, Barak, et al.
Veröffentlicht: (2025)
Distinguishing the Knowable from the Unknowable with Language Models
von: Ahdritz, Gustaf, et al.
Veröffentlicht: (2024)
von: Ahdritz, Gustaf, et al.
Veröffentlicht: (2024)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
Generative Neural Reparameterization for Differentiable PDE-constrained Optimization
von: Joglekar, Archis S.
Veröffentlicht: (2024)
von: Joglekar, Archis S.
Veröffentlicht: (2024)
CarSpeedNet: Learning-Based Speed Estimation from Accelerometer-Only Inertial Sensing
von: Or, Barak
Veröffentlicht: (2024)
von: Or, Barak
Veröffentlicht: (2024)
Kalman-Inspired Runtime Stability and Recovery in Hybrid Reasoning Systems
von: Or, Barak
Veröffentlicht: (2026)
von: Or, Barak
Veröffentlicht: (2026)
Think Before You Lie: How Reasoning Leads to Honesty
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
von: Miceli-Barone, Antonio Valerio, et al.
Veröffentlicht: (2026)
von: Miceli-Barone, Antonio Valerio, et al.
Veröffentlicht: (2026)
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025)
von: Ren, Richard, et al.
Veröffentlicht: (2025)
GAC: Stabilizing Asynchronous RL Training for LLMs via Gradient Alignment Control
von: Xu, Haofeng, et al.
Veröffentlicht: (2026)
von: Xu, Haofeng, et al.
Veröffentlicht: (2026)
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
von: Zhang, Junru, et al.
Veröffentlicht: (2025)
von: Zhang, Junru, et al.
Veröffentlicht: (2025)
RxnNano:Training Compact LLMs for Chemical Reaction and Retrosynthesis Prediction via Hierarchical Curriculum Learning
von: Li, Ran, et al.
Veröffentlicht: (2026)
von: Li, Ran, et al.
Veröffentlicht: (2026)
Effective Sample Size and Generalization Bounds for Temporal Networks
von: Gahtan, Barak, et al.
Veröffentlicht: (2025)
von: Gahtan, Barak, et al.
Veröffentlicht: (2025)
Knowledge Integration Strategies in Autonomous Vehicle Prediction and Planning: A Comprehensive Survey
von: Manas, Kumar, et al.
Veröffentlicht: (2025)
von: Manas, Kumar, et al.
Veröffentlicht: (2025)
Scaling Data-Constrained Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
von: Zeng, Yirong, et al.
Veröffentlicht: (2025)
von: Zeng, Yirong, et al.
Veröffentlicht: (2025)
Efficient Representations are Controllable Representations
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
von: Xu, Hefei, et al.
Veröffentlicht: (2025)
von: Xu, Hefei, et al.
Veröffentlicht: (2025)
Stress Testing Deliberative Alignment for Anti-Scheming Training
von: Schoen, Bronson, et al.
Veröffentlicht: (2025)
von: Schoen, Bronson, et al.
Veröffentlicht: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
Learning Dynamics of RNNs in Closed-Loop Environments
von: Ger, Yoav, et al.
Veröffentlicht: (2025)
von: Ger, Yoav, et al.
Veröffentlicht: (2025)
Learning reveals invisible structure in low-rank RNNs
von: Ger, Yoav, et al.
Veröffentlicht: (2026)
von: Ger, Yoav, et al.
Veröffentlicht: (2026)
Thoth: Mid-Training Bridges LLMs to Time Series Understanding
von: Lin, Jiafeng, et al.
Veröffentlicht: (2026)
von: Lin, Jiafeng, et al.
Veröffentlicht: (2026)
A Hybrid Adaptive Velocity Aided Navigation Filter with Application to INS/DVL Fusion
von: Or, Barak, et al.
Veröffentlicht: (2022)
von: Or, Barak, et al.
Veröffentlicht: (2022)
Recent Trends in Modelling the Continuous Time Series using Deep Learning: A Survey
von: Habiba, Mansura, et al.
Veröffentlicht: (2024)
von: Habiba, Mansura, et al.
Veröffentlicht: (2024)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
Test-Time Training on Graphs with Large Language Models (LLMs)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
von: Yang, Kai, et al.
Veröffentlicht: (2025)
von: Yang, Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Deliberative Alignment: Reasoning Enables Safer Language Models
von: Guan, Melody Y., et al.
Veröffentlicht: (2024) -
Automatic Stability and Recovery for Neural Network Training
von: Or, Barak
Veröffentlicht: (2026) -
Preference Learning with Lie Detectors can Induce Honesty or Evasion
von: Cundy, Chris, et al.
Veröffentlicht: (2025) -
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
von: Gu, Renjie, et al.
Veröffentlicht: (2026) -
Gradient-Free Training of Quantized Neural Networks
von: Cohen, Noa, et al.
Veröffentlicht: (2024)