Saved in:
| Main Authors: | Li, Zezheng, Yip, Kingston |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.08836 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
by: Reddy, K. Sahit, et al.
Published: (2025)
by: Reddy, K. Sahit, et al.
Published: (2025)
AI-Generated Text Detection and Classification Based on BERT Deep Learning Algorithm
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
BERT Learns (and Teaches) Chemistry
by: Payne, Josh, et al.
Published: (2020)
by: Payne, Josh, et al.
Published: (2020)
Injecting linguistic knowledge into BERT for Dialogue State Tracking
by: Feng, Xiaohan, et al.
Published: (2023)
by: Feng, Xiaohan, et al.
Published: (2023)
A Comprehensive Approach to Misspelling Correction with BERT and Levenshtein Distance
by: Naziri, Amirreza, et al.
Published: (2024)
by: Naziri, Amirreza, et al.
Published: (2024)
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction
by: Saoud, Abdulkader, et al.
Published: (2024)
by: Saoud, Abdulkader, et al.
Published: (2024)
BERTCaps: BERT Capsule for Persian Multi-Domain Sentiment Analysis
by: Memari, Mohammadali, et al.
Published: (2024)
by: Memari, Mohammadali, et al.
Published: (2024)
BERT-JEPA: Reorganizing CLS Embeddings for Language-Invariant Semantics
by: Gillin, Taj, et al.
Published: (2026)
by: Gillin, Taj, et al.
Published: (2026)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
by: Jung, Hee-Jun, et al.
Published: (2022)
by: Jung, Hee-Jun, et al.
Published: (2022)
Absolute convergence and error thresholds in non-active adaptive sampling
by: Ferro, Manuel Vilares, et al.
Published: (2024)
by: Ferro, Manuel Vilares, et al.
Published: (2024)
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
by: Zhao, Andrew, et al.
Published: (2025)
by: Zhao, Andrew, et al.
Published: (2025)
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
by: Lin, Zezheng, et al.
Published: (2026)
by: Lin, Zezheng, et al.
Published: (2026)
Seeing Through VisualBERT: A Causal Adventure on Memetic Landscapes
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2024)
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2024)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
by: Lee, Simon A., et al.
Published: (2025)
by: Lee, Simon A., et al.
Published: (2025)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
by: Liang, Chaoqi, et al.
Published: (2023)
by: Liang, Chaoqi, et al.
Published: (2023)
2-Tier SimCSE: Elevating BERT for Robust Sentence Embeddings
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Integrating LSTM and BERT for Long-Sequence Data Analysis in Intelligent Tutoring Systems
by: Li, Zhaoxing, et al.
Published: (2024)
by: Li, Zhaoxing, et al.
Published: (2024)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
by: Zain, Noor Ul, et al.
Published: (2025)
by: Zain, Noor Ul, et al.
Published: (2025)
BERT-ASC: Auxiliary-Sentence Construction for Implicit Aspect Learning in Sentiment Analysis
by: Ahmed, Murtadha, et al.
Published: (2022)
by: Ahmed, Murtadha, et al.
Published: (2022)
MrBERT: Modern Multilingual Encoders via Vocabulary, Domain, and Dimensional Adaptation
by: Tamayo, Daniel, et al.
Published: (2026)
by: Tamayo, Daniel, et al.
Published: (2026)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks
by: Lin, Zezheng, et al.
Published: (2026)
by: Lin, Zezheng, et al.
Published: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
by: Yang, Dongquan, et al.
Published: (2025)
by: Yang, Dongquan, et al.
Published: (2025)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Harnessing Large Language Models: Fine-tuned BERT for Detecting Charismatic Leadership Tactics in Natural Language
by: Saeid, Yasser, et al.
Published: (2024)
by: Saeid, Yasser, et al.
Published: (2024)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
by: Wu, Da, et al.
Published: (2023)
by: Wu, Da, et al.
Published: (2023)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
A Fusion of context-aware based BanglaBERT and Two-Layer Stacked LSTM Framework for Multi-Label Cyberbullying Detection
by: Raquib, Mirza, et al.
Published: (2026)
by: Raquib, Mirza, et al.
Published: (2026)
Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study
by: Hashtarkhani, Soheil, et al.
Published: (2025)
by: Hashtarkhani, Soheil, et al.
Published: (2025)
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
by: Ren, Weijieying, et al.
Published: (2024)
by: Ren, Weijieying, et al.
Published: (2024)
Improving VTE Identification through Language Models from Radiology Reports: A Comparative Study of Mamba, Phi-3 Mini, and BERT
by: Deng, Jamie, et al.
Published: (2024)
by: Deng, Jamie, et al.
Published: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Reducing the Scope of Language Models
by: Yunis, David, et al.
Published: (2024)
by: Yunis, David, et al.
Published: (2024)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
by: Zhao, Stephen, et al.
Published: (2025)
by: Zhao, Stephen, et al.
Published: (2025)
More Expressive Attention with Negative Weights
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
by: Sharma, Agniv, et al.
Published: (2024)
by: Sharma, Agniv, et al.
Published: (2024)
Similar Items
-
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
by: Reddy, K. Sahit, et al.
Published: (2025) -
AI-Generated Text Detection and Classification Based on BERT Deep Learning Algorithm
by: Wang, Hao, et al.
Published: (2024) -
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025) -
BERT Learns (and Teaches) Chemistry
by: Payne, Josh, et al.
Published: (2020) -
Injecting linguistic knowledge into BERT for Dialogue State Tracking
by: Feng, Xiaohan, et al.
Published: (2023)