Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zur, Amir, Geiger, Atticus, Lubana, Ekdeep Singh, Bigelow, Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
von: Bigelow, Eric, et al.
Veröffentlicht: (2026)
von: Bigelow, Eric, et al.
Veröffentlicht: (2026)
In-Context Learning Dynamics with Random Binary Sequences
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
Emergence of Hierarchical Emotion Organization in Large Language Models
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
von: Pres, Itamar, et al.
Veröffentlicht: (2024)
von: Pres, Itamar, et al.
Veröffentlicht: (2024)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
von: Sarfati, Raphaël, et al.
Veröffentlicht: (2026)
von: Sarfati, Raphaël, et al.
Veröffentlicht: (2026)
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
von: Boppana, Siddharth, et al.
Veröffentlicht: (2026)
von: Boppana, Siddharth, et al.
Veröffentlicht: (2026)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
von: Feucht, Sheridan, et al.
Veröffentlicht: (2026)
von: Feucht, Sheridan, et al.
Veröffentlicht: (2026)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
von: Puyin, Li, et al.
Veröffentlicht: (2026)
von: Puyin, Li, et al.
Veröffentlicht: (2026)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Auditing language models for hidden objectives
von: Marks, Samuel, et al.
Veröffentlicht: (2025)
von: Marks, Samuel, et al.
Veröffentlicht: (2025)
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
ICLR: In-Context Learning of Representations
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems
von: Kutschka, Lorenz, et al.
Veröffentlicht: (2026)
von: Kutschka, Lorenz, et al.
Veröffentlicht: (2026)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
von: Ilia, Evgenia, et al.
Veröffentlicht: (2024)
von: Ilia, Evgenia, et al.
Veröffentlicht: (2024)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
Tokens, the oft-overlooked appetizer: Large language models, the distributional hypothesis, and meaning
von: Zimmerman, Julia Witte, et al.
Veröffentlicht: (2024)
von: Zimmerman, Julia Witte, et al.
Veröffentlicht: (2024)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
Forking Paths in Neural Text Generation
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
von: Park, Ji Won, et al.
Veröffentlicht: (2025)
von: Park, Ji Won, et al.
Veröffentlicht: (2025)
Token Statistics Reveal Conversational Drift in Multi-turn LLM Interaction
von: Hafez, Wael, et al.
Veröffentlicht: (2026)
von: Hafez, Wael, et al.
Veröffentlicht: (2026)
chDzDT: Word-level morphology-aware language model for Algerian social media text
von: Aries, Abdelkrime
Veröffentlicht: (2025)
von: Aries, Abdelkrime
Veröffentlicht: (2025)
Syllable-level lyrics generation from melody exploiting character-level language model
von: Zhang, Zhe, et al.
Veröffentlicht: (2023)
von: Zhang, Zhe, et al.
Veröffentlicht: (2023)
Dissociating language and thought in large language models
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
Entry-level guide to the use of large language models for medical research
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
Developmental trajectories of decision making and affective dynamics in large language models
von: Wang, Zhihao, et al.
Veröffentlicht: (2025)
von: Wang, Zhihao, et al.
Veröffentlicht: (2025)
Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
Extracting effective solutions hidden in large language models via generated comprehensive specialists: case studies in developing electronic devices
von: Tomita, Hikari, et al.
Veröffentlicht: (2024)
von: Tomita, Hikari, et al.
Veröffentlicht: (2024)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
LongTail-Swap: benchmarking language models' abilities on rare words
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
von: Bigelow, Eric, et al.
Veröffentlicht: (2026) -
In-Context Learning Dynamics with Random Binary Sequences
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023) -
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
von: Bigelow, Eric, et al.
Veröffentlicht: (2025) -
Emergence of Hierarchical Emotion Organization in Large Language Models
von: Zhao, Bo, et al.
Veröffentlicht: (2025) -
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)