MaBERT:A Padding Safe Interleaved Transformer Mamba Hybrid Encoder for Efficient Extended Context Masked Language Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Jinwoong, Park, Sangjin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReTAMamba: Reliability-Aware Temporal Aggregation with Mamba for Irregular Clinical Time Series Prediction
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
IKNet: Interpretable Stock Price Prediction via Keyword-Guided Integration of News and Technical Indicators
von: Kim, Jinwoong, et al.
Veröffentlicht: (2025)
von: Kim, Jinwoong, et al.
Veröffentlicht: (2025)
GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
GroupSHAP-Guided Integration of Financial News Keywords and Technical Indicators for Stock Price Prediction
von: Kim, Minjoo, et al.
Veröffentlicht: (2025)
von: Kim, Minjoo, et al.
Veröffentlicht: (2025)
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
Jamba: A Hybrid Transformer-Mamba Language Model
von: Lieber, Opher, et al.
Veröffentlicht: (2024)
von: Lieber, Opher, et al.
Veröffentlicht: (2024)
Safe-Embed: Unveiling the Safety-Critical Knowledge of Sentence Encoders
von: Kim, Jinseok, et al.
Veröffentlicht: (2024)
von: Kim, Jinseok, et al.
Veröffentlicht: (2024)
AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic
von: Elshehy, Omar, et al.
Veröffentlicht: (2026)
von: Elshehy, Omar, et al.
Veröffentlicht: (2026)
RexBERT: Context Specialized Bidirectional Encoders for E-commerce
von: Bajaj, Rahul, et al.
Veröffentlicht: (2026)
von: Bajaj, Rahul, et al.
Veröffentlicht: (2026)
GeoBuildBench: A Benchmark for Interactive and Executable Geometry Construction from Natural Language
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026)
EuroBERT: Scaling Multilingual Encoders for European Languages
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
Should We Still Pretrain Encoders with Masked Language Modeling?
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
von: Li, Yingyue, et al.
Veröffentlicht: (2025)
von: Li, Yingyue, et al.
Veröffentlicht: (2025)
LooComp: Leverage Leave-One-Out Strategy to Encoder-only Transformer for Efficient Query-aware Context Compression
von: Do, Thao, et al.
Veröffentlicht: (2026)
von: Do, Thao, et al.
Veröffentlicht: (2026)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining
von: Liang, Wen, et al.
Veröffentlicht: (2024)
von: Liang, Wen, et al.
Veröffentlicht: (2024)
IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method
von: Kim, Mihyeon, et al.
Veröffentlicht: (2025)
von: Kim, Mihyeon, et al.
Veröffentlicht: (2025)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long Documents
von: Czinczoll, Tamara, et al.
Veröffentlicht: (2024)
von: Czinczoll, Tamara, et al.
Veröffentlicht: (2024)
Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
von: Grosjean, Juri, et al.
Veröffentlicht: (2024)
von: Grosjean, Juri, et al.
Veröffentlicht: (2024)
Long-Context Encoder Models for Polish Language Understanding
von: Dadas, Sławomir, et al.
Veröffentlicht: (2026)
von: Dadas, Sławomir, et al.
Veröffentlicht: (2026)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP
von: Sounack, Thomas, et al.
Veröffentlicht: (2025)
von: Sounack, Thomas, et al.
Veröffentlicht: (2025)
mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
von: Marone, Marc, et al.
Veröffentlicht: (2025)
von: Marone, Marc, et al.
Veröffentlicht: (2025)
CultureBERT: Measuring Corporate Culture With Transformer-Based Language Models
von: Koch, Sebastian, et al.
Veröffentlicht: (2022)
von: Koch, Sebastian, et al.
Veröffentlicht: (2022)
Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models
von: Toker, Michael, et al.
Veröffentlicht: (2025)
von: Toker, Michael, et al.
Veröffentlicht: (2025)
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
von: Ling, Shaoshi, et al.
Veröffentlicht: (2023)
von: Ling, Shaoshi, et al.
Veröffentlicht: (2023)
MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining
von: Portes, Jacob, et al.
Veröffentlicht: (2023)
von: Portes, Jacob, et al.
Veröffentlicht: (2023)
Jamba-1.5: Hybrid Transformer-Mamba Models at Scale
von: Jamba Team, et al.
Veröffentlicht: (2024)
von: Jamba Team, et al.
Veröffentlicht: (2024)
m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2026)
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2026)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Symmetric Dot-Product Attention for Efficient Training of BERT Language Models
von: Courtois, Martin, et al.
Veröffentlicht: (2024)
von: Courtois, Martin, et al.
Veröffentlicht: (2024)
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
von: Chen, Wenchao, et al.
Veröffentlicht: (2024)
von: Chen, Wenchao, et al.
Veröffentlicht: (2024)
Characterizing Mamba's Selective Memory using Auto-Encoders
von: Hossain, Tamanna, et al.
Veröffentlicht: (2025)
von: Hossain, Tamanna, et al.
Veröffentlicht: (2025)
Detecting Redundant Health Survey Questions Using Language-agnostic BERT Sentence Embedding (LaBSE)
von: Kang, Sunghoon, et al.
Veröffentlicht: (2024)
von: Kang, Sunghoon, et al.
Veröffentlicht: (2024)
Chinese ModernBERT with Whole-Word Masking
von: Zhao, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhao, Zeyu, et al.
Veröffentlicht: (2025)
Extending Translate-Train for ColBERT-X to African Language CLIR
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
LakotaBERT: A Transformer-based Model for Low Resource Lakota Language
von: Parankusham, Kanishka, et al.
Veröffentlicht: (2025)
von: Parankusham, Kanishka, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ReTAMamba: Reliability-Aware Temporal Aggregation with Mamba for Irregular Clinical Time Series Prediction
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026) -
IKNet: Interpretable Stock Price Prediction via Keyword-Guided Integration of News and Technical Indicators
von: Kim, Jinwoong, et al.
Veröffentlicht: (2025) -
GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series
von: Kim, Jinwoong, et al.
Veröffentlicht: (2026) -
GroupSHAP-Guided Integration of Financial News Keywords and Technical Indicators for Stock Price Prediction
von: Kim, Minjoo, et al.
Veröffentlicht: (2025) -
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)