Saved in:
| Main Authors: | Han, Andy, Fujimoto, Kristina, Shah, Avidan, Nguyen, Kiet, Xu, Kai, Yueh-Han, Chen, Sucholutsky, Ilia, Angell, Rico |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.21834 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
by: Shah, Avidan, et al.
Published: (2026)
by: Shah, Avidan, et al.
Published: (2026)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
by: Yueh-Han, Chen, et al.
Published: (2025)
by: Yueh-Han, Chen, et al.
Published: (2025)
Polynomial Precision Dependence Solutions to Alignment Research Center Matrix Completion Problems
by: Angell, Rico
Published: (2024)
by: Angell, Rico
Published: (2024)
Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison
by: Yang, Tiancheng, et al.
Published: (2026)
by: Yang, Tiancheng, et al.
Published: (2026)
Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching
by: Angell, Rico, et al.
Published: (2023)
by: Angell, Rico, et al.
Published: (2023)
Learning Human-like Representations to Enable Learning Human Values
by: Wynn, Andrea, et al.
Published: (2023)
by: Wynn, Andrea, et al.
Published: (2023)
Revisiting Rogers' Paradox in the Context of Human-AI Interaction
by: Collins, Katherine M., et al.
Published: (2025)
by: Collins, Katherine M., et al.
Published: (2025)
Jailbreak Transferability Emerges from Shared Representations
by: Angell, Rico, et al.
Published: (2025)
by: Angell, Rico, et al.
Published: (2025)
Analyzing the Roles of Language and Vision in Learning from Limited Data
by: Chen, Allison, et al.
Published: (2024)
by: Chen, Allison, et al.
Published: (2024)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
by: Yueh-Han, Chen, et al.
Published: (2025)
by: Yueh-Han, Chen, et al.
Published: (2025)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
by: Bai, Xuechunzi, et al.
Published: (2024)
by: Bai, Xuechunzi, et al.
Published: (2024)
What is a Number, That a Large Language Model May Know It?
by: Marjieh, Raja, et al.
Published: (2025)
by: Marjieh, Raja, et al.
Published: (2025)
Training Agents to Self-Report Misbehavior
by: Lee, Bruce W., et al.
Published: (2026)
by: Lee, Bruce W., et al.
Published: (2026)
Efficient Mitigation of Bus Bunching through Setter-Based Curriculum Learning
by: Shah, Avidan, et al.
Published: (2024)
by: Shah, Avidan, et al.
Published: (2024)
Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models
by: Robinson, Sasha, et al.
Published: (2026)
by: Robinson, Sasha, et al.
Published: (2026)
Party Change in Chile in Comparative Perspective
by: Alan Angell
Published: (2003)
by: Alan Angell
Published: (2003)
φ-Geometry in Droplet Deformation: Full-Range Experimental Evidence Supporting HWFT
by: Angell, Harald
Published: (2025)
by: Angell, Harald
Published: (2025)
The φ-Field Structure of the Prime Numbers,
by: Angell, Harald
Published: (2025)
by: Angell, Harald
Published: (2025)
The Zero Thought Experiment: A Philosophical Prelude to the Zero Coherence Law
by: Angell, Harald
Published: (2025)
by: Angell, Harald
Published: (2025)
Las Dimensiones Internacionales del Golpe de Estado Chileno
by: Alan Angell
Published: (2013)
by: Alan Angell
Published: (2013)
Reforma educativa y política en Chile
by: Alan Angell
Published: (1997)
by: Alan Angell
Published: (1997)
States, cities, and border control: Do sub‐state collectives have a right to protect vulnerable people on the move?
by: Kim Angell
Published: (2025)
by: Kim Angell
Published: (2025)
Adaptive Language-Guided Abstraction from Contrastive Explanations
by: Peng, Andi, et al.
Published: (2024)
by: Peng, Andi, et al.
Published: (2024)
TorchOpera: A Compound AI System for LLM Safety
by: Han, Shanshan, et al.
Published: (2024)
by: Han, Shanshan, et al.
Published: (2024)
Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge Retention
by: Lin, Wenye, et al.
Published: (2026)
by: Lin, Wenye, et al.
Published: (2026)
Does the advocacy for universal videolaryngoscopy have robust scientific support?
by: Alexander Avidan
Published: (2025)
by: Alexander Avidan
Published: (2025)
Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs
by: Mieczkowski, Elizabeth, et al.
Published: (2026)
by: Mieczkowski, Elizabeth, et al.
Published: (2026)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
Measuring LLM Novelty As The Frontier Of Original And High-Quality Output
by: Padmakumar, Vishakh, et al.
Published: (2025)
by: Padmakumar, Vishakh, et al.
Published: (2025)
Quantifying Knowledge Distillation Using Partial Information Decomposition
by: Dissanayake, Pasan, et al.
Published: (2024)
by: Dissanayake, Pasan, et al.
Published: (2024)
Concept Alignment
by: Rane, Sunayana, et al.
Published: (2024)
by: Rane, Sunayana, et al.
Published: (2024)
Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with People
by: Huang, Dun-Ming, et al.
Published: (2024)
by: Huang, Dun-Ming, et al.
Published: (2024)
Language Model Teams as Distributed Systems
by: Mieczkowski, Elizabeth, et al.
Published: (2026)
by: Mieczkowski, Elizabeth, et al.
Published: (2026)
Towards Formalizing Spuriousness of Biased Datasets Using Partial Information Decomposition
by: Halder, Barproda, et al.
Published: (2024)
by: Halder, Barproda, et al.
Published: (2024)
Large Language Models Assume People are More Rational than We Really are
by: Liu, Ryan, et al.
Published: (2024)
by: Liu, Ryan, et al.
Published: (2024)
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases
by: Nguyen, Sang Quang, et al.
Published: (2025)
by: Nguyen, Sang Quang, et al.
Published: (2025)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
by: Zhang, Yi-Kai, et al.
Published: (2025)
by: Zhang, Yi-Kai, et al.
Published: (2025)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
by: Karger, Ezra, et al.
Published: (2024)
by: Karger, Ezra, et al.
Published: (2024)
Learning with Language-Guided State Abstractions
by: Peng, Andi, et al.
Published: (2024)
by: Peng, Andi, et al.
Published: (2024)
Similar Items
-
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
by: Shah, Avidan, et al.
Published: (2026) -
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
by: Yueh-Han, Chen, et al.
Published: (2025) -
Polynomial Precision Dependence Solutions to Alignment Research Center Matrix Completion Problems
by: Angell, Rico
Published: (2024) -
Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison
by: Yang, Tiancheng, et al.
Published: (2026) -
Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching
by: Angell, Rico, et al.
Published: (2023)