Saved in:
| Main Authors: | Ebrahimi, M. Reza, Defferrard, Michaël, Panchal, Sunny, Memisevic, Roland |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.18333 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Your Context Is Not an Array: Unveiling Random Access Limitations in Transformers
by: Ebrahimi, MohammadReza, et al.
Published: (2024)
by: Ebrahimi, MohammadReza, et al.
Published: (2024)
Revisiting Bi-Linear State Transitions in Recurrent Neural Networks
by: Ebrahimi, M. Reza, et al.
Published: (2025)
by: Ebrahimi, M. Reza, et al.
Published: (2025)
Look, Remember and Reason: Grounded reasoning in videos with language models
by: Bhattacharyya, Apratim, et al.
Published: (2023)
by: Bhattacharyya, Apratim, et al.
Published: (2023)
Multi-Draft Speculative Sampling: Canonical Decomposition and Theoretical Limits
by: Khisti, Ashish, et al.
Published: (2024)
by: Khisti, Ashish, et al.
Published: (2024)
Enhancing Hallucination Detection through Noise Injection
by: Liu, Litian, et al.
Published: (2025)
by: Liu, Litian, et al.
Published: (2025)
Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
by: Phan, Buu, et al.
Published: (2025)
by: Phan, Buu, et al.
Published: (2025)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
by: Pourreza, Reza, et al.
Published: (2025)
by: Pourreza, Reza, et al.
Published: (2025)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
by: Badger, Benjamin L., et al.
Published: (2026)
by: Badger, Benjamin L., et al.
Published: (2026)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
by: Ruder, Sebastian, et al.
Published: (2018)
by: Ruder, Sebastian, et al.
Published: (2018)
Semi-Supervised Learning for Bilingual Lexicon Induction
by: Garnier, Paul, et al.
Published: (2024)
by: Garnier, Paul, et al.
Published: (2024)
Rethinking Associative Memory Mechanism in Induction Head
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
Large Language Models as Computable Approximations to Solomonoff Induction
by: Wan, Jun, et al.
Published: (2025)
by: Wan, Jun, et al.
Published: (2025)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
REQUAL-LM: Reliability and Equity through Aggregation in Large Language Models
by: Ebrahimi, Sana, et al.
Published: (2024)
by: Ebrahimi, Sana, et al.
Published: (2024)
Omitted Variable Bias in Language Models Under Distribution Shift
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Combining Induction and Transduction for Abstract Reasoning
by: Li, Wen-Ding, et al.
Published: (2024)
by: Li, Wen-Ding, et al.
Published: (2024)
Eliminating Position Bias of Language Models: A Mechanistic Approach
by: Wang, Ziqi, et al.
Published: (2024)
by: Wang, Ziqi, et al.
Published: (2024)
From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
by: Liu, Litian, et al.
Published: (2026)
by: Liu, Litian, et al.
Published: (2026)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
Universal Response and Emergence of Induction in LLMs
by: Luick, Niclas
Published: (2024)
by: Luick, Niclas
Published: (2024)
Race, Ethnicity and Their Implication on Bias in Large Language Models
by: Hu, Shiyue, et al.
Published: (2026)
by: Hu, Shiyue, et al.
Published: (2026)
Bias in Large Language Models: Origin, Evaluation, and Mitigation
by: Guo, Yufei, et al.
Published: (2024)
by: Guo, Yufei, et al.
Published: (2024)
How Quantization Shapes Bias in Large Language Models
by: Marcuzzi, Federico, et al.
Published: (2025)
by: Marcuzzi, Federico, et al.
Published: (2025)
An Analysis for Reasoning Bias of Language Models with Small Initialization
by: Yao, Junjie, et al.
Published: (2025)
by: Yao, Junjie, et al.
Published: (2025)
On Bilingual Lexicon Induction with Large Language Models
by: Li, Yaoyiran, et al.
Published: (2023)
by: Li, Yaoyiran, et al.
Published: (2023)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
by: Garg, Ankur, et al.
Published: (2025)
by: Garg, Ankur, et al.
Published: (2025)
Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
by: Yu, Xuemin, et al.
Published: (2026)
by: Yu, Xuemin, et al.
Published: (2026)
Sequence-to-Sequence Spanish Pre-trained Language Models
by: Araujo, Vladimir, et al.
Published: (2023)
by: Araujo, Vladimir, et al.
Published: (2023)
A Representation-Level Assessment of Bias Mitigation in Foundation Models
by: Nizhnichenkov, Svetoslav, et al.
Published: (2026)
by: Nizhnichenkov, Svetoslav, et al.
Published: (2026)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
Long-range Modeling and Processing of Multimodal Event Sequences
by: Li, Jichu, et al.
Published: (2026)
by: Li, Jichu, et al.
Published: (2026)
Bias after Prompting: Persistent Discrimination in Large Language Models
by: Sivakumar, Nivedha, et al.
Published: (2025)
by: Sivakumar, Nivedha, et al.
Published: (2025)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
Red-Teaming for Inducing Societal Bias in Large Language Models
by: Luo, Chu Fei, et al.
Published: (2024)
by: Luo, Chu Fei, et al.
Published: (2024)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
by: Reddy, Natesh, et al.
Published: (2025)
by: Reddy, Natesh, et al.
Published: (2025)
Temporal Tokenization Strategies for Event Sequence Modeling with Large Language Models
by: Liu, Zefang, et al.
Published: (2025)
by: Liu, Zefang, et al.
Published: (2025)
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
by: Shukla, Pushkar, et al.
Published: (2025)
by: Shukla, Pushkar, et al.
Published: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Neural Sequence-to-Sequence Modeling with Attention by Leveraging Deep Learning Architectures for Enhanced Contextual Understanding in Abstractive Text Summarization
by: Challagundla, Bhavith Chandra, et al.
Published: (2024)
by: Challagundla, Bhavith Chandra, et al.
Published: (2024)
Similar Items
-
Your Context Is Not an Array: Unveiling Random Access Limitations in Transformers
by: Ebrahimi, MohammadReza, et al.
Published: (2024) -
Revisiting Bi-Linear State Transitions in Recurrent Neural Networks
by: Ebrahimi, M. Reza, et al.
Published: (2025) -
Look, Remember and Reason: Grounded reasoning in videos with language models
by: Bhattacharyya, Apratim, et al.
Published: (2023) -
Multi-Draft Speculative Sampling: Canonical Decomposition and Theoretical Limits
by: Khisti, Ashish, et al.
Published: (2024) -
Enhancing Hallucination Detection through Noise Injection
by: Liu, Litian, et al.
Published: (2025)