Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Saglam, Baturay, Kassianik, Paul, Nelson, Blaine, Weerawardhena, Sajana, Singer, Yaron, Karbasi, Amin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Learning Task Representations from In-Context Learning
by: Saglam, Baturay, et al.
Published: (2025)
by: Saglam, Baturay, et al.
Published: (2025)
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Test-Time Detoxification without Training or Learning Anything
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
by: Yang, Zhuoran, et al.
Published: (2026)
by: Yang, Zhuoran, et al.
Published: (2026)
Adversarial Reasoning at Jailbreaking Time
by: Sabbaghi, Mahdi, et al.
Published: (2025)
by: Sabbaghi, Mahdi, et al.
Published: (2025)
Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
by: Vijay, Supriti, et al.
Published: (2025)
by: Vijay, Supriti, et al.
Published: (2025)
(Im)possibility of Automated Hallucination Detection in Large Language Models
by: Karbasi, Amin, et al.
Published: (2025)
by: Karbasi, Amin, et al.
Published: (2025)
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
by: Lee, Jane H., et al.
Published: (2025)
by: Lee, Jane H., et al.
Published: (2025)
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
by: Weerawardhena, Sajana, et al.
Published: (2025)
by: Weerawardhena, Sajana, et al.
Published: (2025)
Corrector Sampling in Language Models
by: Gat, Itai, et al.
Published: (2025)
by: Gat, Itai, et al.
Published: (2025)
Extracting Memorized Training Data via Decomposition
by: Su, Ellen, et al.
Published: (2024)
by: Su, Ellen, et al.
Published: (2024)
Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
by: Kassianik, Paul, et al.
Published: (2025)
by: Kassianik, Paul, et al.
Published: (2025)
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
by: Mura, Raffaele, et al.
Published: (2025)
by: Mura, Raffaele, et al.
Published: (2025)
Compatible Gradient Approximations for Actor-Critic Algorithms
by: Saglam, Baturay, et al.
Published: (2024)
by: Saglam, Baturay, et al.
Published: (2024)
On the Origins of Linear Representations in Large Language Models
by: Jiang, Yibo, et al.
Published: (2024)
by: Jiang, Yibo, et al.
Published: (2024)
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Multi-Lingual Malaysian Embedding: Leveraging Large Language Models for Semantic Representations
by: Zolkepli, Husein, et al.
Published: (2024)
by: Zolkepli, Husein, et al.
Published: (2024)
The Linear Representation Hypothesis and the Geometry of Large Language Models
by: Park, Kiho, et al.
Published: (2023)
by: Park, Kiho, et al.
Published: (2023)
Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling
by: Amin, Adil
Published: (2026)
by: Amin, Adil
Published: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Extracting and Encoding: Leveraging Large Language Models and Medical Knowledge to Enhance Radiological Text Representation
by: Messina, Pablo, et al.
Published: (2024)
by: Messina, Pablo, et al.
Published: (2024)
Differentially Private Steering for Large Language Model Alignment
by: Goel, Anmol, et al.
Published: (2025)
by: Goel, Anmol, et al.
Published: (2025)
Discovering Implicit Large Language Model Alignment Objectives
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
SeMe: Training-Free Language Model Merging via Semantic Alignment
by: Gu, Jian, et al.
Published: (2025)
by: Gu, Jian, et al.
Published: (2025)
Linear Dynamics in the RLVR Training of Large Language Models
by: Wang, Tianle, et al.
Published: (2026)
by: Wang, Tianle, et al.
Published: (2026)
Lizard: An Efficient Linearization Framework for Large Language Models
by: Van Nguyen, Chien, et al.
Published: (2025)
by: Van Nguyen, Chien, et al.
Published: (2025)
Efficient Alignment of Large Language Models via Data Sampling
by: Khera, Amrit, et al.
Published: (2024)
by: Khera, Amrit, et al.
Published: (2024)
A Survey on Training-free Alignment of Large Language Models
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
Semantic Structure of Feature Space in Large Language Models
by: Kozlowski, Austin C., et al.
Published: (2026)
by: Kozlowski, Austin C., et al.
Published: (2026)
Gödel Test: Can Large Language Models Solve Easy Conjectures?
by: Feldman, Moran, et al.
Published: (2025)
by: Feldman, Moran, et al.
Published: (2025)
The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next
by: Amin, Adil
Published: (2026)
by: Amin, Adil
Published: (2026)
Linear Representations of Political Perspective Emerge in Large Language Models
by: Kim, Junsol, et al.
Published: (2025)
by: Kim, Junsol, et al.
Published: (2025)
The Geometry of Tokens in Internal Representations of Large Language Models
by: Viswanathan, Karthik, et al.
Published: (2025)
by: Viswanathan, Karthik, et al.
Published: (2025)
Graph Linearization Methods for Reasoning on Graphs with Large Language Models
by: Xypolopoulos, Christos, et al.
Published: (2024)
by: Xypolopoulos, Christos, et al.
Published: (2024)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
by: Ji, Xiaotong, et al.
Published: (2025)
by: Ji, Xiaotong, et al.
Published: (2025)
MULTIVERSE: Exposing Large Language Model Alignment Problems in Diverse Worlds
by: Jin, Xiaolong, et al.
Published: (2024)
by: Jin, Xiaolong, et al.
Published: (2024)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
by: Lake, Thom, et al.
Published: (2024)
by: Lake, Thom, et al.
Published: (2024)
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
by: Li, Chen, et al.
Published: (2026)
by: Li, Chen, et al.
Published: (2026)
Similar Items
-
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023) -
Learning Task Representations from In-Context Learning
by: Saglam, Baturay, et al.
Published: (2025) -
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026) -
Test-Time Detoxification without Training or Learning Anything
by: Saglam, Baturay, et al.
Published: (2026) -
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
by: Yang, Zhuoran, et al.
Published: (2026)