Characterizing Linear Alignment Across Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gorbett, Matt, Jana, Suman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Model Disagreement as a Label-Free Correctness Signal
by: Gorbett, Matt, et al.
Published: (2026)
by: Gorbett, Matt, et al.
Published: (2026)
Label-Free Reinforcement Learning via Cross-Model Entropy
by: Gorbett, Matt, et al.
Published: (2026)
by: Gorbett, Matt, et al.
Published: (2026)
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
by: Gorbett, Matt, et al.
Published: (2024)
by: Gorbett, Matt, et al.
Published: (2024)
Cross-Silo Federated Learning Across Divergent Domains with Iterative Parameter Alignment
by: Gorbett, Matt, et al.
Published: (2023)
by: Gorbett, Matt, et al.
Published: (2023)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes
by: Jeoung, Sullam, et al.
Published: (2025)
by: Jeoung, Sullam, et al.
Published: (2025)
Systematic Characterization of the Effectiveness of Alignment in Large Language Models for Categorical Decisions
by: Kohane, Isaac
Published: (2024)
by: Kohane, Isaac
Published: (2024)
Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
by: Weller, Niklas, et al.
Published: (2026)
by: Weller, Niklas, et al.
Published: (2026)
EffiPair: Improving the Efficiency of LLM-generated Code with Relative Contrastive Feedback
by: Hajizadeh, Samira, et al.
Published: (2026)
by: Hajizadeh, Samira, et al.
Published: (2026)
Do Language Models Reason Across Languages?
by: Meng, Yan, et al.
Published: (2026)
by: Meng, Yan, et al.
Published: (2026)
On Diversified Preferences of Large Language Model Alignment
by: Zeng, Dun, et al.
Published: (2023)
by: Zeng, Dun, et al.
Published: (2023)
Dialogical Reasoning Across AI Architectures: A Multi-Model Framework for Testing AI Alignment Strategies
by: Cox, Gray
Published: (2026)
by: Cox, Gray
Published: (2026)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
by: Rajabi, Navid, et al.
Published: (2025)
by: Rajabi, Navid, et al.
Published: (2025)
Representation Alignment Rests on Linear Structure
by: Bangachev, Kiril, et al.
Published: (2026)
by: Bangachev, Kiril, et al.
Published: (2026)
Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal Transport
by: Shah, Shaan, et al.
Published: (2025)
by: Shah, Shaan, et al.
Published: (2025)
Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
by: Kaltenpoth, Sascha, et al.
Published: (2025)
by: Kaltenpoth, Sascha, et al.
Published: (2025)
Linear Spatial World Models Emerge in Large Language Models
by: Tehenan, Matthieu, et al.
Published: (2025)
by: Tehenan, Matthieu, et al.
Published: (2025)
On the Limitations of Steering in Language Model Alignment
by: Niranjan, Chebrolu, et al.
Published: (2025)
by: Niranjan, Chebrolu, et al.
Published: (2025)
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
by: Martorell, Nicolas, et al.
Published: (2026)
by: Martorell, Nicolas, et al.
Published: (2026)
Large Language Model's Multi-Capability Alignment in Biomedical Domain
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Towards Complex Ontology Alignment using Large Language Models
by: Amini, Reihaneh, et al.
Published: (2024)
by: Amini, Reihaneh, et al.
Published: (2024)
Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models
by: Lu, Shule, et al.
Published: (2026)
by: Lu, Shule, et al.
Published: (2026)
Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models
by: Lu, Shule, et al.
Published: (2026)
by: Lu, Shule, et al.
Published: (2026)
An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
by: Ming, Chenlin, et al.
Published: (2025)
by: Ming, Chenlin, et al.
Published: (2025)
Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier
by: Badrinath, Anirudhan, et al.
Published: (2024)
by: Badrinath, Anirudhan, et al.
Published: (2024)
Fundamental Limitations of Alignment in Large Language Models
by: Wolf, Yotam, et al.
Published: (2023)
by: Wolf, Yotam, et al.
Published: (2023)
Benchmarking Distributional Alignment of Large Language Models
by: Meister, Nicole, et al.
Published: (2024)
by: Meister, Nicole, et al.
Published: (2024)
Doubly Robust Alignment for Large Language Models
by: Xu, Erhan, et al.
Published: (2025)
by: Xu, Erhan, et al.
Published: (2025)
Iterative Graph Alignment
by: Yu, Fangyuan, et al.
Published: (2024)
by: Yu, Fangyuan, et al.
Published: (2024)
Leveraging Open-Source Large Language Models for encoding Social Determinants of Health using an Intelligent Router
by: Goel, Akul, et al.
Published: (2024)
by: Goel, Akul, et al.
Published: (2024)
Recursive Symbolic Consciousness: A Formal Model of Emergent Intelligence Across Minds and Machines
by: Goudy, Anastasia
Published: (2025)
by: Goudy, Anastasia
Published: (2025)
ItinBench: Benchmarking Planning Across Multiple Cognitive Dimensions with Large Language Models
by: Wang, Tianlong, et al.
Published: (2026)
by: Wang, Tianlong, et al.
Published: (2026)
Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages
by: Tamang, S., et al.
Published: (2024)
by: Tamang, S., et al.
Published: (2024)
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
by: Fu, Shi, et al.
Published: (2026)
by: Fu, Shi, et al.
Published: (2026)
Scaling Item-to-Standard Alignment with Large Language Models: Accuracy, Limits, and Solutions
by: Karimi-Malekabadi, Farzan, et al.
Published: (2025)
by: Karimi-Malekabadi, Farzan, et al.
Published: (2025)
Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models
by: Jiang, Xinyan, et al.
Published: (2025)
by: Jiang, Xinyan, et al.
Published: (2025)
Constrained Language Model Policy Optimization via Risk-aware Stepwise Alignment
by: Zhang, Lijun, et al.
Published: (2025)
by: Zhang, Lijun, et al.
Published: (2025)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
by: Kim, Dongyoung, et al.
Published: (2026)
by: Kim, Dongyoung, et al.
Published: (2026)
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
by: Fukui, Hiroki
Published: (2026)
by: Fukui, Hiroki
Published: (2026)
Similar Items
-
Cross-Model Disagreement as a Label-Free Correctness Signal
by: Gorbett, Matt, et al.
Published: (2026) -
Label-Free Reinforcement Learning via Cross-Model Entropy
by: Gorbett, Matt, et al.
Published: (2026) -
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
by: Gorbett, Matt, et al.
Published: (2024) -
Cross-Silo Federated Learning Across Divergent Domains with Iterative Parameter Alignment
by: Gorbett, Matt, et al.
Published: (2023) -
Atlas-Alignment: Making Interpretability Transferable Across Language Models
by: Puri, Bruno, et al.
Published: (2025)