Interpreting token compositionality in LLMs: A robustness analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Aljaafari, Nura, Carvalho, Danilo S., Freitas, André |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Mechanics of Conceptual Interpretation in GPT Models: Interpretative Insights
by: Aljaafari, Nura, et al.
Published: (2024)
by: Aljaafari, Nura, et al.
Published: (2024)
TRACE: Training and Inference-Time Interpretability Analysis for Language Models
by: Aljaafari, Nura, et al.
Published: (2025)
by: Aljaafari, Nura, et al.
Published: (2025)
Emergence and Localisation of Semantic Role Circuits in LLMs
by: Aljaafari, Nura, et al.
Published: (2025)
by: Aljaafari, Nura, et al.
Published: (2025)
CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment
by: Aljaafari, Nura, et al.
Published: (2025)
by: Aljaafari, Nura, et al.
Published: (2025)
TRACE for Tracking the Emergence of Semantic Representations in Transformers
by: Aljaafari, Nura, et al.
Published: (2025)
by: Aljaafari, Nura, et al.
Published: (2025)
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
by: Aljaafari, Nura, et al.
Published: (2026)
by: Aljaafari, Nura, et al.
Published: (2026)
From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach
by: Aljaafari, Nura, et al.
Published: (2026)
by: Aljaafari, Nura, et al.
Published: (2026)
Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder
by: Zhang, Yingji, et al.
Published: (2025)
by: Zhang, Yingji, et al.
Published: (2025)
Inductive Learning of Logical Theories with LLMs: An Expressivity-Graded Analysis
by: Gandarela, João Pedro, et al.
Published: (2024)
by: Gandarela, João Pedro, et al.
Published: (2024)
Quasi-symbolic Semantic Geometry over Transformer-based Variational AutoEncoder
by: Zhang, Yingji, et al.
Published: (2022)
by: Zhang, Yingji, et al.
Published: (2022)
Learning Disentangled Semantic Spaces of Explanations via Invertible Neural Networks
by: Zhang, Yingji, et al.
Published: (2023)
by: Zhang, Yingji, et al.
Published: (2023)
Multi-Relational Hyperbolic Word Embeddings from Natural Language Definitions
by: Valentino, Marco, et al.
Published: (2023)
by: Valentino, Marco, et al.
Published: (2023)
Learning to Disentangle Latent Reasoning Rules with Language VAEs: A Systematic Study
by: Zhang, Yingji, et al.
Published: (2025)
by: Zhang, Yingji, et al.
Published: (2025)
Why do LLMs attend to the first token?
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Towards Controllable Natural Language Inference through Lexical Inference Types
by: Zhang, Yingji, et al.
Published: (2023)
by: Zhang, Yingji, et al.
Published: (2023)
LangVAE and LangSpace: Building and Probing for Language Model VAEs
by: Carvalho, Danilo S., et al.
Published: (2025)
by: Carvalho, Danilo S., et al.
Published: (2025)
SylloBio-NLI: Evaluating Large Language Models on Biomedical Syllogistic Reasoning
by: Wysocka, Magdalena, et al.
Published: (2024)
by: Wysocka, Magdalena, et al.
Published: (2024)
Montague semantics and modifier consistency measurement in neural language models
by: Carvalho, Danilo S., et al.
Published: (2022)
by: Carvalho, Danilo S., et al.
Published: (2022)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
by: Witold, Waligóra
Published: (2024)
by: Witold, Waligóra
Published: (2024)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
Prediction hubs are context-informed frequent tokens in LLMs
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
Jacobian Scopes: token-level causal attributions in LLMs
by: Liu, Toni J. B., et al.
Published: (2026)
by: Liu, Toni J. B., et al.
Published: (2026)
Finetuning LLMs for EvaCun 2025 token prediction shared task
by: Jon, Josef, et al.
Published: (2025)
by: Jon, Josef, et al.
Published: (2025)
Comparative analysis of subword tokenization approaches for Indian languages
by: Das, Sudhansu Bala, et al.
Published: (2025)
by: Das, Sudhansu Bala, et al.
Published: (2025)
Improving Semantic Control in Discrete Latent Spaces with Transformer Quantized Variational Autoencoders
by: Zhang, Yingji, et al.
Published: (2024)
by: Zhang, Yingji, et al.
Published: (2024)
PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement
by: Quan, Xin, et al.
Published: (2025)
by: Quan, Xin, et al.
Published: (2025)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Is my model "mind blurting"? Interpreting the dynamics of reasoning tokens with Recurrence Quantification Analysis (RQA)
by: Pham, Quoc Tuan, et al.
Published: (2026)
by: Pham, Quoc Tuan, et al.
Published: (2026)
Accelerating Antibiotic Discovery with Large Language Models and Knowledge Graphs
by: Delmas, Maxime, et al.
Published: (2025)
by: Delmas, Maxime, et al.
Published: (2025)
Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
Contextual morphologically-guided tokenization for Latin encoder models
by: Hudspeth, Marisa, et al.
Published: (2025)
by: Hudspeth, Marisa, et al.
Published: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
by: Georgiou, Efthymios, et al.
Published: (2025)
by: Georgiou, Efthymios, et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
A Survey in Mathematical Language Processing
by: Meadows, Jordan, et al.
Published: (2022)
by: Meadows, Jordan, et al.
Published: (2022)
Similar Items
-
The Mechanics of Conceptual Interpretation in GPT Models: Interpretative Insights
by: Aljaafari, Nura, et al.
Published: (2024) -
TRACE: Training and Inference-Time Interpretability Analysis for Language Models
by: Aljaafari, Nura, et al.
Published: (2025) -
Emergence and Localisation of Semantic Role Circuits in LLMs
by: Aljaafari, Nura, et al.
Published: (2025) -
CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment
by: Aljaafari, Nura, et al.
Published: (2025) -
TRACE for Tracking the Emergence of Semantic Representations in Transformers
by: Aljaafari, Nura, et al.
Published: (2025)