Transformer Dynamics: A neuroscientific approach to interpretability of large language models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fernando, Jesseba, Guitchounts, Grigori |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
von: Fernando, Jesseba, et al.
Veröffentlicht: (2026)
von: Fernando, Jesseba, et al.
Veröffentlicht: (2026)
Ploutos: Towards interpretable stock movement prediction with financial large language model
von: Tong, Hanshuang, et al.
Veröffentlicht: (2024)
von: Tong, Hanshuang, et al.
Veröffentlicht: (2024)
CXR-LLAVA: a multimodal large language model for interpreting chest X-ray images
von: Lee, Seowoo, et al.
Veröffentlicht: (2023)
von: Lee, Seowoo, et al.
Veröffentlicht: (2023)
Georeferencing complex relative locality descriptions with large language models
von: Fernando, Aneesha, et al.
Veröffentlicht: (2025)
von: Fernando, Aneesha, et al.
Veröffentlicht: (2025)
The wall confronting large language models
von: Coveney, Peter V., et al.
Veröffentlicht: (2025)
von: Coveney, Peter V., et al.
Veröffentlicht: (2025)
Unsupervised decoding of encoded reasoning using language model interpretability
von: Fang, Ching, et al.
Veröffentlicht: (2025)
von: Fang, Ching, et al.
Veröffentlicht: (2025)
Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models
von: Peters, Uwe, et al.
Veröffentlicht: (2026)
von: Peters, Uwe, et al.
Veröffentlicht: (2026)
Dissociating language and thought in large language models
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
von: Cunha, Washington, et al.
Veröffentlicht: (2025)
von: Cunha, Washington, et al.
Veröffentlicht: (2025)
A critical review of methods and challenges in large language models
von: Moradi, Milad, et al.
Veröffentlicht: (2024)
von: Moradi, Milad, et al.
Veröffentlicht: (2024)
On the attribution of confidence to large language models
von: Keeling, Geoff, et al.
Veröffentlicht: (2024)
von: Keeling, Geoff, et al.
Veröffentlicht: (2024)
The challenge of uncertainty quantification of large language models in medicine
von: Atf, Zahra, et al.
Veröffentlicht: (2025)
von: Atf, Zahra, et al.
Veröffentlicht: (2025)
Domain adaptation of large language models for geotechnical applications
von: Fan, Lei, et al.
Veröffentlicht: (2025)
von: Fan, Lei, et al.
Veröffentlicht: (2025)
Chain-of-Authorization: Embedding authorization into large language models
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
Cognitive models can reveal interpretable value trade-offs in language models
von: Murthy, Sonia K., et al.
Veröffentlicht: (2025)
von: Murthy, Sonia K., et al.
Veröffentlicht: (2025)
Representation in large language models
von: Yetman, Cameron
Veröffentlicht: (2025)
von: Yetman, Cameron
Veröffentlicht: (2025)
EvoOpt-LLM: Evolving industrial optimization models with large language models
von: He, Yiliu, et al.
Veröffentlicht: (2026)
von: He, Yiliu, et al.
Veröffentlicht: (2026)
How large language models judge and influence human cooperation
von: Pires, Alexandre S., et al.
Veröffentlicht: (2025)
von: Pires, Alexandre S., et al.
Veröffentlicht: (2025)
Mechanistic interpretability of large language models with applications to the financial services industry
von: Golgoon, Ashkan, et al.
Veröffentlicht: (2024)
von: Golgoon, Ashkan, et al.
Veröffentlicht: (2024)
BadEdit: Backdooring large language models by model editing
von: Li, Yanzhou, et al.
Veröffentlicht: (2024)
von: Li, Yanzhou, et al.
Veröffentlicht: (2024)
A review on the use of large language models as virtual tutors
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
Optimizing watermarks for large language models
von: Wouters, Bram
Veröffentlicht: (2023)
von: Wouters, Bram
Veröffentlicht: (2023)
Alignment faking in large language models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
von: Diaz-Garcia, Jose A., et al.
Veröffentlicht: (2025)
von: Diaz-Garcia, Jose A., et al.
Veröffentlicht: (2025)
\(X\)-evolve: Solution space evolution powered by large language models
von: Zhai, Yi, et al.
Veröffentlicht: (2025)
von: Zhai, Yi, et al.
Veröffentlicht: (2025)
Benchmarking graph construction by large language models for coherence-driven inference
von: Huntsman, Steve, et al.
Veröffentlicht: (2025)
von: Huntsman, Steve, et al.
Veröffentlicht: (2025)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
MoleCode unlocks structural intelligence in large language models
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2026)
Quantifying non deterministic drift in large language models
von: Nicholson, Claire
Veröffentlicht: (2026)
von: Nicholson, Claire
Veröffentlicht: (2026)
Can large language models build causal graphs?
von: Long, Stephanie, et al.
Veröffentlicht: (2023)
von: Long, Stephanie, et al.
Veröffentlicht: (2023)
Quantifying construct validity in large language model evaluations
von: Kearns, Ryan Othniel
Veröffentlicht: (2026)
von: Kearns, Ryan Othniel
Veröffentlicht: (2026)
Multi-round jailbreak attack on large language models
von: Zhou, Yihua, et al.
Veröffentlicht: (2024)
von: Zhou, Yihua, et al.
Veröffentlicht: (2024)
Response: Emergent analogical reasoning in large language models
von: Hodel, Damian, et al.
Veröffentlicht: (2023)
von: Hodel, Damian, et al.
Veröffentlicht: (2023)
The 20 questions game to distinguish large language models
von: Richardeau, Gurvan, et al.
Veröffentlicht: (2024)
von: Richardeau, Gurvan, et al.
Veröffentlicht: (2024)
AI-AI Bias: large language models favor communications generated by large language models
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
Ranking protein-protein models with large language models and graph neural networks
von: Xu, Xiaotong, et al.
Veröffentlicht: (2024)
von: Xu, Xiaotong, et al.
Veröffentlicht: (2024)
Retrieval-augmented in-context learning for multimodal large language models in disease classification
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
Toward a unified framework for data-efficient evaluation of large language models
von: Liao, Lele, et al.
Veröffentlicht: (2025)
von: Liao, Lele, et al.
Veröffentlicht: (2025)
A blind spot for large language models: Supradiegetic linguistic information
von: Zimmerman, Julia Witte, et al.
Veröffentlicht: (2023)
von: Zimmerman, Julia Witte, et al.
Veröffentlicht: (2023)
Hyacinth6B: A large language model for Traditional Chinese
von: Song, Chih-Wei, et al.
Veröffentlicht: (2024)
von: Song, Chih-Wei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
von: Fernando, Jesseba, et al.
Veröffentlicht: (2026) -
Ploutos: Towards interpretable stock movement prediction with financial large language model
von: Tong, Hanshuang, et al.
Veröffentlicht: (2024) -
CXR-LLAVA: a multimodal large language model for interpreting chest X-ray images
von: Lee, Seowoo, et al.
Veröffentlicht: (2023) -
Georeferencing complex relative locality descriptions with large language models
von: Fernando, Aneesha, et al.
Veröffentlicht: (2025) -
The wall confronting large language models
von: Coveney, Peter V., et al.
Veröffentlicht: (2025)