Representation Engineering: A Top-Down Approach to AI Transparency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Andy, Phan, Long, Chen, Sarah, Campbell, James, Guo, Phillip, Ren, Richard, Pan, Alexander, Yin, Xuwang, Mazeika, Mantas, Dombrowski, Ann-Kathrin, Goel, Shashwat, Li, Nathaniel, Byun, Michael J., Wang, Zifan, Mallen, Alex, Basart, Steven, Koyejo, Sanmi, Song, Dawn, Fredrikson, Matt, Kolter, J. Zico, Hendrycks, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
TextQuests: How Good are LLMs at Text-Based Video Games?
von: Phan, Long, et al.
Veröffentlicht: (2025)
von: Phan, Long, et al.
Veröffentlicht: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024)
von: Ren, Richard, et al.
Veröffentlicht: (2024)
Improving Alignment and Robustness with Circuit Breakers
von: Zou, Andy, et al.
Veröffentlicht: (2024)
von: Zou, Andy, et al.
Veröffentlicht: (2024)
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025)
von: Ren, Richard, et al.
Veröffentlicht: (2025)
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
Mimetic Initialization of MLPs
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
Predicting the Performance of Black-box LLMs through Follow-up Queries
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
Discovering Implicit Large Language Model Alignment Objectives
von: Chen, Edward, et al.
Veröffentlicht: (2026)
von: Chen, Edward, et al.
Veröffentlicht: (2026)
Tamper-Resistant Safeguards for Open-Weight LLMs
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
von: Zou, Andy, et al.
Veröffentlicht: (2025)
von: Zou, Andy, et al.
Veröffentlicht: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
One-Step Diffusion Distillation via Deep Equilibrium Models
von: Geng, Zhengyang, et al.
Veröffentlicht: (2023)
von: Geng, Zhengyang, et al.
Veröffentlicht: (2023)
Diffusing Differentiable Representations
von: Savani, Yash, et al.
Veröffentlicht: (2024)
von: Savani, Yash, et al.
Veröffentlicht: (2024)
Reasoning Models Don't Just Think Longer, They Move Differently
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
Compressed Sensing for Capability Localization in Large Language Models
von: Bair, Anna, et al.
Veröffentlicht: (2026)
von: Bair, Anna, et al.
Veröffentlicht: (2026)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
Massive Activations in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
von: Zhu, Junzhe, et al.
Veröffentlicht: (2023)
von: Zhu, Junzhe, et al.
Veröffentlicht: (2023)
Idiosyncrasies in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2025)
von: Sun, Mingjie, et al.
Veröffentlicht: (2025)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
von: Maini, Pratyush, et al.
Veröffentlicht: (2023)
von: Maini, Pratyush, et al.
Veröffentlicht: (2023)
Forcing Diffuse Distributions out of Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
SpecEval: Evaluating Model Adherence to Behavior Specifications
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2025)
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2025)
A Simple and Effective Pruning Approach for Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Why Do Safety Guardrails Degrade Across Languages?
von: Zhang, Max, et al.
Veröffentlicht: (2026)
von: Zhang, Max, et al.
Veröffentlicht: (2026)
Logits are All We Need to Adapt Closed Models
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
Aggressive Compression Enables LLM Weight Theft
von: Brown, Davis, et al.
Veröffentlicht: (2026)
von: Brown, Davis, et al.
Veröffentlicht: (2026)
Mimetic Initialization Helps State Space Models Learn to Recall
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
Finetuning CLIP to Reason about Pairwise Differences
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
von: Sam, Dylan, et al.
Veröffentlicht: (2024)
From Variance to Veracity: Unbundling and Mitigating Gradient Variance in Differentiable Bundle Adjustment Layers
von: Gurumurthy, Swaminathan, et al.
Veröffentlicht: (2024)
von: Gurumurthy, Swaminathan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024) -
TextQuests: How Good are LLMs at Text-Based Video Games?
von: Phan, Long, et al.
Veröffentlicht: (2025) -
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024) -
Improving Alignment and Robustness with Circuit Breakers
von: Zou, Andy, et al.
Veröffentlicht: (2024) -
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)