Rewiring the Transformer with Depth-Wise LSTMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Hongfei, Song, Yang, Liu, Qiuhui, van Genabith, Josef, Xiong, Deyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2020
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
von: Shi, Dan, et al.
Veröffentlicht: (2026)
von: Shi, Dan, et al.
Veröffentlicht: (2026)
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
von: Baeumel, Tanja, et al.
Veröffentlicht: (2026)
von: Baeumel, Tanja, et al.
Veröffentlicht: (2026)
Probing Context Localization of Polysemous Words in Pre-trained Language Model Sub-Layers
von: Vijayakumar, Soniya, et al.
Veröffentlicht: (2024)
von: Vijayakumar, Soniya, et al.
Veröffentlicht: (2024)
The Lookahead Limitation: Why Multi-Operand Addition is Hard for LLMs
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
Sign Language Translation with Sentence Embedding Supervision
von: Hamidullah, Yasser, et al.
Veröffentlicht: (2025)
von: Hamidullah, Yasser, et al.
Veröffentlicht: (2025)
Spatio-temporal Sign Language Representation and Translation
von: Hamidullah, Yasser, et al.
Veröffentlicht: (2025)
von: Hamidullah, Yasser, et al.
Veröffentlicht: (2025)
On Multilingual Encoder Language Model Compression for Low-Resource Languages
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification
von: Shcharbakova, Hanna, et al.
Veröffentlicht: (2025)
von: Shcharbakova, Hanna, et al.
Veröffentlicht: (2025)
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation
von: Tan, David, et al.
Veröffentlicht: (2026)
von: Tan, David, et al.
Veröffentlicht: (2026)
Seeing, Signing, and Saying: A Vision-Language Model-Assisted Pipeline for Sign Language Data Acquisition and Curation from Social Media
von: Yazdani, Shakib, et al.
Veröffentlicht: (2025)
von: Yazdani, Shakib, et al.
Veröffentlicht: (2025)
Measuring Spurious Correlation in Classification: 'Clever Hans' in Translationese
von: Borah, Angana, et al.
Veröffentlicht: (2023)
von: Borah, Angana, et al.
Veröffentlicht: (2023)
The Latin Substrate: How Language Models Represent and Mediate Script Choice
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2026)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2026)
When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
von: Bafna, Niyati, et al.
Veröffentlicht: (2023)
von: Bafna, Niyati, et al.
Veröffentlicht: (2023)
A Critical Study of Automatic Evaluation in Sign Language Translation
von: Yazdani, Shakib, et al.
Veröffentlicht: (2025)
von: Yazdani, Shakib, et al.
Veröffentlicht: (2025)
SONAR-SLT: Multilingual Sign Language Translation via Language-Agnostic Sentence Embedding Supervision
von: Hamidullah, Yasser, et al.
Veröffentlicht: (2025)
von: Hamidullah, Yasser, et al.
Veröffentlicht: (2025)
Multilingual Political Views of Large Language Models: Identification and Steering
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
PETra: A Multilingual Corpus of Pragmatic Explicitation in Translation
von: Osmelak, Doreen, et al.
Veröffentlicht: (2025)
von: Osmelak, Doreen, et al.
Veröffentlicht: (2025)
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
von: Song, Shixiang, et al.
Veröffentlicht: (2026)
von: Song, Shixiang, et al.
Veröffentlicht: (2026)
CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models
von: Shi, Ling, et al.
Veröffentlicht: (2024)
von: Shi, Ling, et al.
Veröffentlicht: (2024)
Modular Arithmetic: Language Models Solve Math Digit by Digit
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2026)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2026)
Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
von: Ghussin, Yusser Al, et al.
Veröffentlicht: (2026)
von: Ghussin, Yusser Al, et al.
Veröffentlicht: (2026)
Automatic Fact-checking in English and Telugu
von: Chikkala, Ravi Kiran, et al.
Veröffentlicht: (2025)
von: Chikkala, Ravi Kiran, et al.
Veröffentlicht: (2025)
DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge
von: Ghussin, Yusser Al, et al.
Veröffentlicht: (2026)
von: Ghussin, Yusser Al, et al.
Veröffentlicht: (2026)
An Empirical Study on the Robustness of Massively Multilingual Neural Machine Translation
von: Supryadi, et al.
Veröffentlicht: (2024)
von: Supryadi, et al.
Veröffentlicht: (2024)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
StockBot 2.0: Vanilla LSTMs Outperform Transformer-based Forecasting for Stock Prices
von: Mohanty, Shaswat
Veröffentlicht: (2026)
von: Mohanty, Shaswat
Veröffentlicht: (2026)
Towards Understanding Multi-Task Learning (Generalization) of LLMs via Detecting and Exploring Task-Specific Neurons
von: Leng, Yongqi, et al.
Veröffentlicht: (2024)
von: Leng, Yongqi, et al.
Veröffentlicht: (2024)
A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL
von: Yang, Lei, et al.
Veröffentlicht: (2026)
von: Yang, Lei, et al.
Veröffentlicht: (2026)
AutoPsyC: Automatic Recognition of Psychodynamic Conflicts from Semi-structured Interviews with Large Language Models
von: Hossain, Sayed Muddashir, et al.
Veröffentlicht: (2025)
von: Hossain, Sayed Muddashir, et al.
Veröffentlicht: (2025)
Self-Pluralising Culture Alignment for Large Language Models
von: Xu, Shaoyang, et al.
Veröffentlicht: (2024)
von: Xu, Shaoyang, et al.
Veröffentlicht: (2024)
DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Search
von: Yang, Lei, et al.
Veröffentlicht: (2024)
von: Yang, Lei, et al.
Veröffentlicht: (2024)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis
von: Lu, Junyu, et al.
Veröffentlicht: (2026)
von: Lu, Junyu, et al.
Veröffentlicht: (2026)
Evaluating Discourse Cohesion in Pre-trained Language Models
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
Reverse Probing: Evaluating Knowledge Transfer via Finetuned Task Embeddings for Coreference Resolution
von: Anikina, Tatiana, et al.
Veröffentlicht: (2025)
von: Anikina, Tatiana, et al.
Veröffentlicht: (2025)
LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language Models
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
von: Shi, Dan, et al.
Veröffentlicht: (2026) -
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
von: Baeumel, Tanja, et al.
Veröffentlicht: (2026) -
Probing Context Localization of Polysemous Words in Pre-trained Language Model Sub-Layers
von: Vijayakumar, Soniya, et al.
Veröffentlicht: (2024) -
The Lookahead Limitation: Why Multi-Operand Addition is Hard for LLMs
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025) -
Sign Language Translation with Sentence Embedding Supervision
von: Hamidullah, Yasser, et al.
Veröffentlicht: (2025)