Evaluating Language Model Reasoning about Confidential Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sam, Dylan, Robey, Alexander, Zou, Andy, Fredrikson, Matt, Kolter, J. Zico |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Safety Pretraining: Toward the Next Generation of Safe AI
von: Maini, Pratyush, et al.
Veröffentlicht: (2025)
von: Maini, Pratyush, et al.
Veröffentlicht: (2025)
When Should We Introduce Safety Interventions During Pretraining?
von: Sam, Dylan, et al.
Veröffentlicht: (2026)
von: Sam, Dylan, et al.
Veröffentlicht: (2026)
Finetuning CLIP to Reason about Pairwise Differences
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
Predicting the Performance of Black-box LLMs through Follow-up Queries
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
Adversarial Attacks on Robotic Vision Language Action Models
von: Jones, Eliot Krzysztof, et al.
Veröffentlicht: (2025)
von: Jones, Eliot Krzysztof, et al.
Veröffentlicht: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
von: Huang, Benhao, et al.
Veröffentlicht: (2026)
von: Huang, Benhao, et al.
Veröffentlicht: (2026)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
Bayesian Neural Networks with Domain Knowledge Priors
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
Weight Ensembling Improves Reasoning in Language Models
von: Dang, Xingyu, et al.
Veröffentlicht: (2025)
von: Dang, Xingyu, et al.
Veröffentlicht: (2025)
Improving Alignment and Robustness with Circuit Breakers
von: Zou, Andy, et al.
Veröffentlicht: (2024)
von: Zou, Andy, et al.
Veröffentlicht: (2024)
Massive Activations in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks
von: Kim, Eungyeup, et al.
Veröffentlicht: (2026)
von: Kim, Eungyeup, et al.
Veröffentlicht: (2026)
Why is SAM Robust to Label Noise?
von: Baek, Christina, et al.
Veröffentlicht: (2024)
von: Baek, Christina, et al.
Veröffentlicht: (2024)
Mimetic Initialization of MLPs
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
A Simple and Effective Pruning Approach for Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
One-Step Diffusion Distillation via Deep Equilibrium Models
von: Geng, Zhengyang, et al.
Veröffentlicht: (2023)
von: Geng, Zhengyang, et al.
Veröffentlicht: (2023)
Forcing Diffuse Distributions out of Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
LipNeXt: Scaling up Lipschitz-based Certified Robustness to Billion-parameter Models
von: Hu, Kai, et al.
Veröffentlicht: (2026)
von: Hu, Kai, et al.
Veröffentlicht: (2026)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
An Axiomatic Approach to Model-Agnostic Concept Explanations
von: Feng, Zhili, et al.
Veröffentlicht: (2024)
von: Feng, Zhili, et al.
Veröffentlicht: (2024)
Antidistillation Fingerprinting
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Diffusing Differentiable Representations
von: Savani, Yash, et al.
Veröffentlicht: (2024)
von: Savani, Yash, et al.
Veröffentlicht: (2024)
Generative Posterior Networks for Approximately Bayesian Epistemic Uncertainty Estimation
von: Roderick, Melrose, et al.
Veröffentlicht: (2023)
von: Roderick, Melrose, et al.
Veröffentlicht: (2023)
Understanding Hallucinations in Diffusion Models through Mode Interpolation
von: Aithal, Sumukh K, et al.
Veröffentlicht: (2024)
von: Aithal, Sumukh K, et al.
Veröffentlicht: (2024)
Algorithms for Adversarially Robust Deep Learning
von: Robey, Alexander
Veröffentlicht: (2025)
von: Robey, Alexander
Veröffentlicht: (2025)
Rethinking Distance Metrics for Counterfactual Explainability
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
The Mixing method: low-rank coordinate descent for semidefinite programming with diagonal constraints
von: Wang, Po-Wei, et al.
Veröffentlicht: (2017)
von: Wang, Po-Wei, et al.
Veröffentlicht: (2017)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
Mimetic Initialization Helps State Space Models Learn to Recall
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
von: Finzi, Marc, et al.
Veröffentlicht: (2026)
von: Finzi, Marc, et al.
Veröffentlicht: (2026)
Predicting the Performance of Foundation Models via Agreement-on-the-Line
von: Saxena, Rahul, et al.
Veröffentlicht: (2024)
von: Saxena, Rahul, et al.
Veröffentlicht: (2024)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
von: Jiang, Yiding, et al.
Veröffentlicht: (2024)
von: Jiang, Yiding, et al.
Veröffentlicht: (2024)
Consistency Models Made Easy
von: Geng, Zhengyang, et al.
Veröffentlicht: (2024)
von: Geng, Zhengyang, et al.
Veröffentlicht: (2024)
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
von: Zhai, Runtian, et al.
Veröffentlicht: (2023)
von: Zhai, Runtian, et al.
Veröffentlicht: (2023)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Safety Pretraining: Toward the Next Generation of Safe AI
von: Maini, Pratyush, et al.
Veröffentlicht: (2025) -
When Should We Introduce Safety Interventions During Pretraining?
von: Sam, Dylan, et al.
Veröffentlicht: (2026) -
Finetuning CLIP to Reason about Pairwise Differences
von: Sam, Dylan, et al.
Veröffentlicht: (2024) -
Predicting the Performance of Black-box LLMs through Follow-up Queries
von: Sam, Dylan, et al.
Veröffentlicht: (2025) -
Transferable Adversarial Attacks on Black-Box Vision-Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)