HalleluBERT: Let Every Token That Has Meaning Bear Its Weight
Fuente:
arXiv
Saved in:
| Main Author: | Schmitt, Raphael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SindBERT, the Sailor: Charting the Seas of Turkish NLP
by: Schmitt, Raphael, et al.
Published: (2025)
by: Schmitt, Raphael, et al.
Published: (2025)
GeistBERT: Breathing Life into German NLP
by: Scheible-Schmitt, Raphael, et al.
Published: (2025)
by: Scheible-Schmitt, Raphael, et al.
Published: (2025)
DunbaaBERT: From Sacrifice to Semantics
by: Maab, Iffat, et al.
Published: (2026)
by: Maab, Iffat, et al.
Published: (2026)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
by: Yu, Dian, et al.
Published: (2025)
by: Yu, Dian, et al.
Published: (2025)
BERT's Conceptual Cartography: Mapping the Landscapes of Meaning
by: Haket, Nina, et al.
Published: (2024)
by: Haket, Nina, et al.
Published: (2024)
Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature
by: Liu, Zheng, et al.
Published: (2025)
by: Liu, Zheng, et al.
Published: (2025)
MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation
by: Behrendt, Maike, et al.
Published: (2025)
by: Behrendt, Maike, et al.
Published: (2025)
KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis
by: Awlla, Kozhin muhealddin, et al.
Published: (2025)
by: Awlla, Kozhin muhealddin, et al.
Published: (2025)
TIDE: Every Layer Knows the Token Beneath the Context
by: Jaiswal, Ajay, et al.
Published: (2026)
by: Jaiswal, Ajay, et al.
Published: (2026)
Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding
by: Pan, Lehan, et al.
Published: (2026)
by: Pan, Lehan, et al.
Published: (2026)
An Encoder-Integrated PhoBERT with Graph Attention for Vietnamese Token-Level Classification
by: Nguyen, Ba-Quang
Published: (2025)
by: Nguyen, Ba-Quang
Published: (2025)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
by: Silva, Enzo S. N., et al.
Published: (2026)
by: Silva, Enzo S. N., et al.
Published: (2026)
Token-Sensitive Enclosure Semantics for Measurement-Bearing Expressions
by: Hulak, David B., et al.
Published: (2026)
by: Hulak, David B., et al.
Published: (2026)
Not Every Token Needs Forgetting: Selective Unlearning to Limit Change in Utility in Large Language Model Unlearning
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
by: Wang, Ante, et al.
Published: (2025)
by: Wang, Ante, et al.
Published: (2025)
Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens
by: Zhang, Yichi, et al.
Published: (2026)
by: Zhang, Yichi, et al.
Published: (2026)
PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation
by: Lim, Hyemin, et al.
Published: (2025)
by: Lim, Hyemin, et al.
Published: (2025)
FGR-ColBERT: Identifying Fine-Grained Relevance Tokens During Retrieval
by: Jarolím, Antonín, et al.
Published: (2026)
by: Jarolím, Antonín, et al.
Published: (2026)
Token Weighting for Long-Range Language Modeling
by: Helm, Falko, et al.
Published: (2025)
by: Helm, Falko, et al.
Published: (2025)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
by: Hu, Xiang, et al.
Published: (2025)
by: Hu, Xiang, et al.
Published: (2025)
Few Dimensions are Enough: Fine-tuning BERT with Selected Dimensions Revealed Its Redundant Nature
by: Fukuhata, Shion, et al.
Published: (2025)
by: Fukuhata, Shion, et al.
Published: (2025)
When Every Token Counts: Optimal Segmentation for Low-Resource Language Models
by: Raj, Bharath, et al.
Published: (2024)
by: Raj, Bharath, et al.
Published: (2024)
FaBERT: Pre-training BERT on Persian Blogs
by: Masumi, Mostafa, et al.
Published: (2024)
by: Masumi, Mostafa, et al.
Published: (2024)
NusaBERT: Teaching IndoBERT to be Multilingual and Multicultural
by: Wongso, Wilson, et al.
Published: (2024)
by: Wongso, Wilson, et al.
Published: (2024)
NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew
by: Shmidman, Shaltiel, et al.
Published: (2025)
by: Shmidman, Shaltiel, et al.
Published: (2025)
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
by: Scivetti, Wesley, et al.
Published: (2025)
by: Scivetti, Wesley, et al.
Published: (2025)
QiBERT -- Classifying Online Conversations Messages with BERT as a Feature
by: Ferreira-Saraiva, Bruno D., et al.
Published: (2024)
by: Ferreira-Saraiva, Bruno D., et al.
Published: (2024)
Pre-training technique to localize medical BERT and enhance biomedical BERT
by: Wada, Shoya, et al.
Published: (2020)
by: Wada, Shoya, et al.
Published: (2020)
Gazetteer-Enhanced Bangla Named Entity Recognition with BanglaBERT Semantic Embeddings K-Means-Infused CRF Model
by: Farhan, Niloy, et al.
Published: (2024)
by: Farhan, Niloy, et al.
Published: (2024)
GottBERT: a pure German Language Model
by: Scheible, Raphael, et al.
Published: (2020)
by: Scheible, Raphael, et al.
Published: (2020)
Weight Tying Biases Token Embeddings Towards the Output Space
by: Lopardo, Antonio, et al.
Published: (2026)
by: Lopardo, Antonio, et al.
Published: (2026)
NeoBERT: A Next-Generation BERT
by: Breton, Lola Le, et al.
Published: (2025)
by: Breton, Lola Le, et al.
Published: (2025)
SpikeBERT: A Language Spikformer Learned from BERT with Knowledge Distillation
by: Lv, Changze, et al.
Published: (2023)
by: Lv, Changze, et al.
Published: (2023)
Advancing Pancreatic Cancer Prediction with a Next Visit Token Prediction Head on top of Med-BERT
by: He, Jianping, et al.
Published: (2025)
by: He, Jianping, et al.
Published: (2025)
Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models
by: Xiang, Bajian, et al.
Published: (2026)
by: Xiang, Bajian, et al.
Published: (2026)
Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews
by: Christian, William, et al.
Published: (2025)
by: Christian, William, et al.
Published: (2025)
ACL: Aligned Contrastive Learning Improves BERT and Multi-exit BERT Fine-tuning
by: Li, Liz, et al.
Published: (2026)
by: Li, Liz, et al.
Published: (2026)
Similar Items
-
SindBERT, the Sailor: Charting the Seas of Turkish NLP
by: Schmitt, Raphael, et al.
Published: (2025) -
GeistBERT: Breathing Life into German NLP
by: Scheible-Schmitt, Raphael, et al.
Published: (2025) -
DunbaaBERT: From Sacrifice to Semantics
by: Maab, Iffat, et al.
Published: (2026) -
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
by: Yu, Dian, et al.
Published: (2025) -
BERT's Conceptual Cartography: Mapping the Landscapes of Meaning
by: Haket, Nina, et al.
Published: (2024)