Saved in:
| Main Authors: | Buonanno, Amedeo, Rivetti, Alessandro, Palmieri, Francesco A. N., Di Gennaro, Giovanni, Romano, Gianmarco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.15347 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A General Framework for Scalable UE-AP Association in User-Centric Cell-Free Massive MIMO based on Recurrent Neural Networks
by: Di Gennaro, Giovanni, et al.
Published: (2025)
by: Di Gennaro, Giovanni, et al.
Published: (2025)
A Deep Learning Approach for User-Centric Clustering in Cell-Free Massive MIMO Systems
by: Di Gennaro, Giovanni, et al.
Published: (2024)
by: Di Gennaro, Giovanni, et al.
Published: (2024)
Unsupervised Learning for Scalable Downlink Power Control in Cell-Free Massive MIMO
by: Di Gennaro, Giovanni, et al.
Published: (2026)
by: Di Gennaro, Giovanni, et al.
Published: (2026)
The Curved Spacetime of Transformer Architectures
by: Di Sipio, Riccardo, et al.
Published: (2025)
by: Di Sipio, Riccardo, et al.
Published: (2025)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
by: Basile, Lorenzo, et al.
Published: (2025)
by: Basile, Lorenzo, et al.
Published: (2025)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
by: Hoang, Nhat M., et al.
Published: (2025)
by: Hoang, Nhat M., et al.
Published: (2025)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024)
by: Kossen, Jannik, et al.
Published: (2024)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Bridging Emotions and Architecture: Sentiment Analysis in Modern Distributed Systems
by: Shah, Mahak, et al.
Published: (2025)
by: Shah, Mahak, et al.
Published: (2025)
Learning Novel Transformer Architecture for Time-series Forecasting
by: Zhang, Juyuan, et al.
Published: (2025)
by: Zhang, Juyuan, et al.
Published: (2025)
CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability
by: Capps, Chad A.
Published: (2026)
by: Capps, Chad A.
Published: (2026)
Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
by: Viswanathan, Karthik, et al.
Published: (2025)
by: Viswanathan, Karthik, et al.
Published: (2025)
Probing Cultural Signals in Large Language Models through Author Profiling
by: Lafargue, Valentin, et al.
Published: (2026)
by: Lafargue, Valentin, et al.
Published: (2026)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
by: Huang, Zeyi, et al.
Published: (2026)
by: Huang, Zeyi, et al.
Published: (2026)
Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
by: Bronzini, Marco, et al.
Published: (2025)
by: Bronzini, Marco, et al.
Published: (2025)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
by: Choi, Minsik, et al.
Published: (2025)
by: Choi, Minsik, et al.
Published: (2025)
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
by: Ringel, Liran, et al.
Published: (2025)
by: Ringel, Liran, et al.
Published: (2025)
Supernova: Achieving More with Less in Transformer Architectures
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
Residual Stream Duality in Modern Transformer Architectures
by: Zhang, Yifan
Published: (2026)
by: Zhang, Yifan
Published: (2026)
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026)
by: Park, Kiho, et al.
Published: (2026)
DisEmbed: Transforming Disease Understanding through Embeddings
by: Faroz, Salman
Published: (2024)
by: Faroz, Salman
Published: (2024)
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
by: Jelenić, Fran, et al.
Published: (2023)
by: Jelenić, Fran, et al.
Published: (2023)
Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
by: Omidi, Parsa, et al.
Published: (2025)
by: Omidi, Parsa, et al.
Published: (2025)
Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media
by: Hasan, Khalid, et al.
Published: (2025)
by: Hasan, Khalid, et al.
Published: (2025)
Towards Robust Few-Shot Text Classification Using Transformer Architectures and Dual Loss Strategies
by: Han, Xu, et al.
Published: (2025)
by: Han, Xu, et al.
Published: (2025)
Modeling Bilingual Sentence Processing: Evaluating RNN and Transformer Architectures for Cross-Language Structural Priming
by: Zhang, Demi, et al.
Published: (2024)
by: Zhang, Demi, et al.
Published: (2024)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025)
by: Kim, Jeonghoon, et al.
Published: (2025)
Teaching Transformers Causal Reasoning through Axiomatic Training
by: Vashishtha, Aniket, et al.
Published: (2024)
by: Vashishtha, Aniket, et al.
Published: (2024)
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets
by: Cheekati, Shravan
Published: (2024)
by: Cheekati, Shravan
Published: (2024)
Selective Attention: Enhancing Transformer through Principled Context Control
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Transformers through the lens of support-preserving maps between measures
by: Furuya, Takashi, et al.
Published: (2025)
by: Furuya, Takashi, et al.
Published: (2025)
Twitter Sentiment Analysis using Distributed Word and Sentence Representation
by: Reddy, Dwarampudi Mahidhar, et al.
Published: (2019)
by: Reddy, Dwarampudi Mahidhar, et al.
Published: (2019)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
by: Choi, Sehyun
Published: (2024)
by: Choi, Sehyun
Published: (2024)
The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling
by: Kerce, J. Clayton, et al.
Published: (2026)
by: Kerce, J. Clayton, et al.
Published: (2026)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
by: Song, Jiwon, et al.
Published: (2024)
by: Song, Jiwon, et al.
Published: (2024)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Metadata Matters for Time Series: Informative Forecasting with Transformers
by: Dong, Jiaxiang, et al.
Published: (2024)
by: Dong, Jiaxiang, et al.
Published: (2024)
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
by: Levinstein, B. A., et al.
Published: (2023)
by: Levinstein, B. A., et al.
Published: (2023)
Similar Items
-
A General Framework for Scalable UE-AP Association in User-Centric Cell-Free Massive MIMO based on Recurrent Neural Networks
by: Di Gennaro, Giovanni, et al.
Published: (2025) -
A Deep Learning Approach for User-Centric Clustering in Cell-Free Massive MIMO Systems
by: Di Gennaro, Giovanni, et al.
Published: (2024) -
Unsupervised Learning for Scalable Downlink Power Control in Cell-Free Massive MIMO
by: Di Gennaro, Giovanni, et al.
Published: (2026) -
The Curved Spacetime of Transformer Architectures
by: Di Sipio, Riccardo, et al.
Published: (2025) -
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
by: Basile, Lorenzo, et al.
Published: (2025)