Does Representation Intervention Really Identify Desired Concepts and Elicit Alignment?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Hongzheng, Chen, Yongqiang, Qin, Zeyu, Liu, Tongliang, Xiao, Chaowei, Zhang, Kun, Han, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CausalEvolve: Towards Open-Ended Discovery with Causal Scratchpad
von: Chen, Yongqiang, et al.
Veröffentlicht: (2026)
von: Chen, Yongqiang, et al.
Veröffentlicht: (2026)
On the Thinking-Language Modeling Gap in Large Language Models
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
Is Gradient Ascent Really Necessary? Memorize to Forget for Machine Unlearning
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
Discovering and Reasoning of Causality in the Hidden World with Large Language Models
von: Liu, Chenxi, et al.
Veröffentlicht: (2024)
von: Liu, Chenxi, et al.
Veröffentlicht: (2024)
Can Large Language Models Help Experimental Design for Causal Discovery?
von: Li, Junyi, et al.
Veröffentlicht: (2025)
von: Li, Junyi, et al.
Veröffentlicht: (2025)
Does a Neural Network Really Encode Symbolic Concepts?
von: Li, Mingjie, et al.
Veröffentlicht: (2023)
von: Li, Mingjie, et al.
Veröffentlicht: (2023)
Does Alignment Tuning Really Break LLMs' Internal Confidence?
von: Oh, Hongseok, et al.
Veröffentlicht: (2024)
von: Oh, Hongseok, et al.
Veröffentlicht: (2024)
Noisy Test-Time Adaptation in Vision-Language Models
von: Cao, Chentao, et al.
Veröffentlicht: (2025)
von: Cao, Chentao, et al.
Veröffentlicht: (2025)
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
von: Wang, Xinyan, et al.
Veröffentlicht: (2026)
von: Wang, Xinyan, et al.
Veröffentlicht: (2026)
Identifiability and Asymptotics in Learning Homogeneous Linear ODE Systems from Discrete Observations
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2022)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2022)
Identifying Representations for Intervention Extrapolation
von: Saengkyongam, Sorawit, et al.
Veröffentlicht: (2023)
von: Saengkyongam, Sorawit, et al.
Veröffentlicht: (2023)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy Examples
von: Yuan, Suqin, et al.
Veröffentlicht: (2025)
von: Yuan, Suqin, et al.
Veröffentlicht: (2025)
On the Over-Memorization During Natural, Robust and Catastrophic Overfitting
von: Lin, Runqi, et al.
Veröffentlicht: (2023)
von: Lin, Runqi, et al.
Veröffentlicht: (2023)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
Robust Training of Federated Models with Extremely Label Deficiency
von: Zhang, Yonggang, et al.
Veröffentlicht: (2024)
von: Zhang, Yonggang, et al.
Veröffentlicht: (2024)
Enhancing One-Shot Federated Learning Through Data and Ensemble Co-Boosting
von: Dai, Rong, et al.
Veröffentlicht: (2024)
von: Dai, Rong, et al.
Veröffentlicht: (2024)
What If the Input is Expanded in OOD Detection?
von: Zhang, Boxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2024)
Envisioning Outlier Exposure by Large Language Models for Out-of-Distribution Detection
von: Cao, Chentao, et al.
Veröffentlicht: (2024)
von: Cao, Chentao, et al.
Veröffentlicht: (2024)
LLM Interpretability with Identifiable Temporal-Instantaneous Representation
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
Double Check My Desired Return: Transformer with Target Alignment for Offline Reinforcement Learning
von: Pei, Yue, et al.
Veröffentlicht: (2025)
von: Pei, Yue, et al.
Veröffentlicht: (2025)
Enhancing Evolving Domain Generalization through Dynamic Latent Representations
von: Xie, Binghui, et al.
Veröffentlicht: (2024)
von: Xie, Binghui, et al.
Veröffentlicht: (2024)
Instance-dependent Early Stopping
von: Yuan, Suqin, et al.
Veröffentlicht: (2025)
von: Yuan, Suqin, et al.
Veröffentlicht: (2025)
Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut Dependency
von: Lin, Runqi, et al.
Veröffentlicht: (2024)
von: Lin, Runqi, et al.
Veröffentlicht: (2024)
Understanding Robust Overfitting from the Feature Generalization Perspective
von: Yu, Chaojian, et al.
Veröffentlicht: (2023)
von: Yu, Chaojian, et al.
Veröffentlicht: (2023)
ERASE: Error-Resilient Representation Learning on Graphs for Label Noise Tolerance
von: Chen, Ling-Hao, et al.
Veröffentlicht: (2023)
von: Chen, Ling-Hao, et al.
Veröffentlicht: (2023)
Towards Effective Evaluations and Comparisons for LLM Unlearning Methods
von: Wang, Qizhou, et al.
Veröffentlicht: (2024)
von: Wang, Qizhou, et al.
Veröffentlicht: (2024)
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
von: Yang, Puning, et al.
Veröffentlicht: (2025)
von: Yang, Puning, et al.
Veröffentlicht: (2025)
BrokenBind: Universal Modality Exploration beyond Dataset Boundaries
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
von: Huang, Zhuo, et al.
Veröffentlicht: (2026)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
How Interpretable Are Interpretable Graph Neural Networks?
von: Chen, Yongqiang, et al.
Veröffentlicht: (2024)
von: Chen, Yongqiang, et al.
Veröffentlicht: (2024)
Towards Identifiability of Hierarchical Temporal Causal Representation Learning
von: Li, Zijian, et al.
Veröffentlicht: (2025)
von: Li, Zijian, et al.
Veröffentlicht: (2025)
Label Distribution Learning with Biased Annotations by Learning Multi-Label Representation
von: Kou, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Kou, Zhiqiang, et al.
Veröffentlicht: (2025)
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
von: Zhang, Zizhuo, et al.
Veröffentlicht: (2025)
von: Zhang, Zizhuo, et al.
Veröffentlicht: (2025)
Representational Transfer Learning for Matrix Completion
von: He, Yong, et al.
Veröffentlicht: (2024)
von: He, Yong, et al.
Veröffentlicht: (2024)
Enhancing Neural Subset Selection: Integrating Background Information into Set Representations
von: Xie, Binghui, et al.
Veröffentlicht: (2024)
von: Xie, Binghui, et al.
Veröffentlicht: (2024)
Generative Model Inversion Through the Lens of the Manifold Hypothesis
von: Peng, Xiong, et al.
Veröffentlicht: (2025)
von: Peng, Xiong, et al.
Veröffentlicht: (2025)
Advancing Counterfactual Inference through Nonlinear Quantile Regression
von: Xie, Shaoan, et al.
Veröffentlicht: (2023)
von: Xie, Shaoan, et al.
Veröffentlicht: (2023)
When Does Closeness in Distribution Imply Representational Similarity? An Identifiability Perspective
von: Nielsen, Beatrix M. G., et al.
Veröffentlicht: (2025)
von: Nielsen, Beatrix M. G., et al.
Veröffentlicht: (2025)
OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad
von: Tang, Luyao, et al.
Veröffentlicht: (2025)
von: Tang, Luyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CausalEvolve: Towards Open-Ended Discovery with Causal Scratchpad
von: Chen, Yongqiang, et al.
Veröffentlicht: (2026) -
On the Thinking-Language Modeling Gap in Large Language Models
von: Liu, Chenxi, et al.
Veröffentlicht: (2025) -
Is Gradient Ascent Really Necessary? Memorize to Forget for Machine Unlearning
von: Huang, Zhuo, et al.
Veröffentlicht: (2026) -
Discovering and Reasoning of Causality in the Hidden World with Large Language Models
von: Liu, Chenxi, et al.
Veröffentlicht: (2024) -
Can Large Language Models Help Experimental Design for Causal Discovery?
von: Li, Junyi, et al.
Veröffentlicht: (2025)