Latent Introspection: Models Can Detect Prior Concept Injections
Fuente:
arXiv
Saved in:
| Main Authors: | Pearson-Vogel, Theia, Vanek, Martin, Douglas, Raymond, Kulveit, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Artificial Self: Characterising the landscape of AI identity
by: Douglas, Raymond, et al.
Published: (2026)
by: Douglas, Raymond, et al.
Published: (2026)
Bayesian Concept Bottleneck Models with LLM Priors
by: Feng, Jean, et al.
Published: (2024)
by: Feng, Jean, et al.
Published: (2024)
Exploration Through Introspection: A Self-Aware Reward Model
by: Petrowski, Michael, et al.
Published: (2026)
by: Petrowski, Michael, et al.
Published: (2026)
Learning Discrete Concepts in Latent Hierarchical Models
by: Kong, Lingjing, et al.
Published: (2024)
by: Kong, Lingjing, et al.
Published: (2024)
Incoherence in Goal-Conditioned Autoregressive Models
by: Karwowski, Jacek, et al.
Published: (2025)
by: Karwowski, Jacek, et al.
Published: (2025)
SCALEX: Scalable Concept and Latent Exploration for Diffusion Models
by: Zeng, E. Zhixuan, et al.
Published: (2025)
by: Zeng, E. Zhixuan, et al.
Published: (2025)
Generating Counterfactual Trajectories with Latent Diffusion Models for Concept Discovery
by: Varshney, Payal, et al.
Published: (2024)
by: Varshney, Payal, et al.
Published: (2024)
Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
by: Qu, Xingwei, et al.
Published: (2025)
by: Qu, Xingwei, et al.
Published: (2025)
Nonparametric Identification of Latent Concepts
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
Distilling Symbolic Priors for Concept Learning into Neural Networks
by: Marinescu, Ioana, et al.
Published: (2024)
by: Marinescu, Ioana, et al.
Published: (2024)
AI-AI Bias: large language models favor communications generated by large language models
by: Laurito, Walter, et al.
Published: (2024)
by: Laurito, Walter, et al.
Published: (2024)
Sculpting Latent Spaces With MMD: Disentanglement With Programmable Priors
by: Fruytier, Quentin, et al.
Published: (2025)
by: Fruytier, Quentin, et al.
Published: (2025)
Analyzing Latent Concepts in Code Language Models
by: Sharma, Arushi, et al.
Published: (2025)
by: Sharma, Arushi, et al.
Published: (2025)
Latent Noise Injection for Private and Statistically Aligned Synthetic Data Generation
by: Shen, Rex, et al.
Published: (2025)
by: Shen, Rex, et al.
Published: (2025)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)
by: Qu, Yuxiao, et al.
Published: (2024)
Not All Latent Spaces Are Flat: Hyperbolic Concept Control
by: Briglia, Maria Rosaria, et al.
Published: (2026)
by: Briglia, Maria Rosaria, et al.
Published: (2026)
Latent Concept Disentanglement in Transformer-based Language Models
by: Hong, Guan Zhe, et al.
Published: (2025)
by: Hong, Guan Zhe, et al.
Published: (2025)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
Zero-Overhead Introspection for Adaptive Test-Time Compute
by: Manvi, Rohin, et al.
Published: (2025)
by: Manvi, Rohin, et al.
Published: (2025)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
by: Cui, Brandon, et al.
Published: (2026)
by: Cui, Brandon, et al.
Published: (2026)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
by: Liang, Kaiqu, et al.
Published: (2024)
by: Liang, Kaiqu, et al.
Published: (2024)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
by: Yang, Diji, et al.
Published: (2024)
by: Yang, Diji, et al.
Published: (2024)
Counterfactual Concept Bottleneck Models
by: Dominici, Gabriele, et al.
Published: (2024)
by: Dominici, Gabriele, et al.
Published: (2024)
Improving Value-based Process Verifier via Structural Prior Injection
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
Causal Concept Graphs in LLM Latent Space for Stepwise Reasoning
by: Meherab, Md Muntaqim, et al.
Published: (2026)
by: Meherab, Md Muntaqim, et al.
Published: (2026)
CoLa-DCE -- Concept-guided Latent Diffusion Counterfactual Explanations
by: Motzkus, Franz, et al.
Published: (2024)
by: Motzkus, Franz, et al.
Published: (2024)
Fair In-Context Learning via Latent Concept Variables
by: Bhaila, Karuna, et al.
Published: (2024)
by: Bhaila, Karuna, et al.
Published: (2024)
Residual Prior Diffusion: A Probabilistic Framework Integrating Coarse Latent Priors with Diffusion Models
by: Kutsuna, Takuro
Published: (2025)
by: Kutsuna, Takuro
Published: (2025)
PriorZero: Bridging Language Priors and World Models for Decision Making
by: Xiong, Junyu, et al.
Published: (2026)
by: Xiong, Junyu, et al.
Published: (2026)
Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
by: Röder, Frank, et al.
Published: (2025)
by: Röder, Frank, et al.
Published: (2025)
Online Drift Detection with Maximum Concept Discrepancy
by: Wan, Ke, et al.
Published: (2024)
by: Wan, Ke, et al.
Published: (2024)
The VampPrior Mixture Model
by: Stirn, Andrew A., et al.
Published: (2024)
by: Stirn, Andrew A., et al.
Published: (2024)
How Not to Detect Prompt Injections with an LLM
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
Behavior Injection: Preparing Language Models for Reinforcement Learning
by: Cen, Zhepeng, et al.
Published: (2025)
by: Cen, Zhepeng, et al.
Published: (2025)
Online Detection of Water Contamination Under Concept Drift
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
by: Li, YuQian, et al.
Published: (2025)
by: Li, YuQian, et al.
Published: (2025)
Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization
by: Penaloza, Emiliano, et al.
Published: (2025)
by: Penaloza, Emiliano, et al.
Published: (2025)
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
by: Zhan, Xiaohua, et al.
Published: (2026)
by: Zhan, Xiaohua, et al.
Published: (2026)
Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
by: Li, Xiaodan, et al.
Published: (2025)
by: Li, Xiaodan, et al.
Published: (2025)
Similar Items
-
The Artificial Self: Characterising the landscape of AI identity
by: Douglas, Raymond, et al.
Published: (2026) -
Bayesian Concept Bottleneck Models with LLM Priors
by: Feng, Jean, et al.
Published: (2024) -
Exploration Through Introspection: A Self-Aware Reward Model
by: Petrowski, Michael, et al.
Published: (2026) -
Learning Discrete Concepts in Latent Hierarchical Models
by: Kong, Lingjing, et al.
Published: (2024) -
Incoherence in Goal-Conditioned Autoregressive Models
by: Karwowski, Jacek, et al.
Published: (2025)