The Depth Delusion: Why Transformers Should Be Wider, Not Deeper
Fuente:
arXiv
Saved in:
| Main Authors: | Fahim, Md Muhtasim Munif, Karim, Md Rezaul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pre-trained Encoders for Global Child Development: Transfer Learning Enables Deployment in Data-Scarce Settings
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
Child Mortality Prediction in Bangladesh: A Decade-Long Validation Study
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
The Dependency Divide: An Interpretable Machine Learning Framework for Profiling Student Digital Satisfaction in the Bangladesh Context
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
by: Su, Zelal, et al.
Published: (2026)
by: Su, Zelal, et al.
Published: (2026)
Distributed Causality in the SDG Network: Evidence from Panel VAR and Conditional Independence Analysis
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
Green-NAS: A Global-Scale Multi-Objective Neural Architecture Search for Robust and Efficient Edge-Native Weather Forecasting
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
Delving Deeper Into Astromorphic Transformers
by: Mia, Md Zesun Ahmed, et al.
Published: (2023)
by: Mia, Md Zesun Ahmed, et al.
Published: (2023)
LiPCoT: Linear Predictive Coding based Tokenizer for Self-supervised Learning of Time Series Data via Language Models
by: Anjum, Md Fahim
Published: (2024)
by: Anjum, Md Fahim
Published: (2024)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
by: Chen, Hung-Hsuan
Published: (2026)
by: Chen, Hung-Hsuan
Published: (2026)
Fast Think-on-Graph: Wider, Deeper and Faster Reasoning of Large Language Model on Knowledge Graph
by: Liang, Xujian, et al.
Published: (2025)
by: Liang, Xujian, et al.
Published: (2025)
Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data Free Distillation
by: Sikder, Md Fahim, et al.
Published: (2024)
by: Sikder, Md Fahim, et al.
Published: (2024)
Advancing Diffusion Models: Alias-Free Resampling and Enhanced Rotational Equivariance
by: Anjum, Md Fahim
Published: (2024)
by: Anjum, Md Fahim
Published: (2024)
Bt-GAN: Generating Fair Synthetic Healthdata via Bias-transforming Generative Adversarial Networks
by: Ramachandranpillai, Resmi, et al.
Published: (2024)
by: Ramachandranpillai, Resmi, et al.
Published: (2024)
FairX: A comprehensive benchmarking tool for model analysis using fairness, utility, and explainability
by: Sikder, Md Fahim, et al.
Published: (2024)
by: Sikder, Md Fahim, et al.
Published: (2024)
Generating Synthetic Fair Syntax-agnostic Data by Learning and Distilling Fair Representation
by: Sikder, Md Fahim, et al.
Published: (2024)
by: Sikder, Md Fahim, et al.
Published: (2024)
Forecasting NCAA Basketball Outcomes with Deep Learning: A Comparative Study of LSTM and Transformer Models
by: Habib, Md Imtiaz
Published: (2025)
by: Habib, Md Imtiaz
Published: (2025)
A Multi-Model Approach to English-Bangla Sentiment Classification of Government Mobile Banking App Reviews
by: Molla, Md. Naim, et al.
Published: (2026)
by: Molla, Md. Naim, et al.
Published: (2026)
Strategic Fusion Optimizes Transformer Compression
by: Rahman, Md Shoaibur
Published: (2025)
by: Rahman, Md Shoaibur
Published: (2025)
The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
by: Yao, Zexi, et al.
Published: (2025)
by: Yao, Zexi, et al.
Published: (2025)
ResNets Are Deeper Than You Think
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models
by: Yeung, Joshua Au, et al.
Published: (2025)
by: Yeung, Joshua Au, et al.
Published: (2025)
Graph Transformer-Based Pathway Embedding for Cancer Prognosis
by: Howlader, Koushik, et al.
Published: (2026)
by: Howlader, Koushik, et al.
Published: (2026)
Digging Deeper: Learning Multi-Level Concept Hierarchies
by: Hill, Oscar, et al.
Published: (2026)
by: Hill, Oscar, et al.
Published: (2026)
Activation Sparsity Opportunities for Compressing General Large Language Models
by: Dhar, Nobel, et al.
Published: (2024)
by: Dhar, Nobel, et al.
Published: (2024)
Interpretable Physics-Informed Load Forecasting for U.S. Grid Resilience: SHAP-Guided Ensemble Validation in Hybrid Deep Learning Under Extreme Weather
by: Abubakkar, Md, et al.
Published: (2026)
by: Abubakkar, Md, et al.
Published: (2026)
Explainable Machine Learning Framework for Cardiovascular Disease Diagnosis and Prognosis
by: Sourov, Md. Emon Akter, et al.
Published: (2025)
by: Sourov, Md. Emon Akter, et al.
Published: (2025)
Deeper Insights into Deep Graph Convolutional Networks: Stability and Generalization
by: Yang, Guangrui, et al.
Published: (2024)
by: Yang, Guangrui, et al.
Published: (2024)
Modeling Electric Vehicle Car-Following Behavior: Classical vs Machine Learning Approach
by: Uddin, Md. Shihab, et al.
Published: (2025)
by: Uddin, Md. Shihab, et al.
Published: (2025)
MCDFN: Supply Chain Demand Forecasting via an Explainable Multi-Channel Data Fusion Network Model
by: Jahin, Md Abrar, et al.
Published: (2024)
by: Jahin, Md Abrar, et al.
Published: (2024)
Energy-Efficient Deep Reinforcement Learning with Spiking Transformers
by: Uddin, Mohammad Irfan, et al.
Published: (2025)
by: Uddin, Mohammad Irfan, et al.
Published: (2025)
Unveiling Black-boxes: Explainable Deep Learning Models for Patent Classification
by: Shajalal, Md, et al.
Published: (2023)
by: Shajalal, Md, et al.
Published: (2023)
Secure Energy Transactions Using Blockchain Leveraging AI for Fraud Detection and Energy Market Stability
by: Khan, Md Asif Ul Hoq, et al.
Published: (2025)
by: Khan, Md Asif Ul Hoq, et al.
Published: (2025)
Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?
by: Omori, Yumi, et al.
Published: (2025)
by: Omori, Yumi, et al.
Published: (2025)
PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping
by: Huang, Nai-Chieh, et al.
Published: (2023)
by: Huang, Nai-Chieh, et al.
Published: (2023)
Learning To Play Atari Games Using Dueling Q-Learning and Hebbian Plasticity
by: Salehin, Md Ashfaq
Published: (2024)
by: Salehin, Md Ashfaq
Published: (2024)
Location Agnostic Adaptive Rain Precipitation Prediction using Deep Learning
by: Islam, Md Shazid, et al.
Published: (2024)
by: Islam, Md Shazid, et al.
Published: (2024)
N-ReLU: Zero-Mean Stochastic Extension of ReLU
by: Manik, Md Motaleb Hossen, et al.
Published: (2025)
by: Manik, Md Motaleb Hossen, et al.
Published: (2025)
Bayesian Regression for Predicting Subscription to Bank Term Deposits in Direct Marketing Campaigns
by: Tanvir, Muhammad Farhan, et al.
Published: (2024)
by: Tanvir, Muhammad Farhan, et al.
Published: (2024)
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
by: Shajalal, Md, et al.
Published: (2022)
by: Shajalal, Md, et al.
Published: (2022)
EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Similar Items
-
Pre-trained Encoders for Global Child Development: Transfer Learning Enables Deployment in Data-Scarce Settings
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026) -
Child Mortality Prediction in Bangladesh: A Decade-Long Validation Study
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026) -
The Dependency Divide: An Interpretable Machine Learning Framework for Profiling Student Digital Satisfaction in the Bangladesh Context
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026) -
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
by: Su, Zelal, et al.
Published: (2026) -
Distributed Causality in the SDG Network: Evidence from Panel VAR and Conditional Independence Analysis
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)