Enregistré dans:
| Auteurs principaux: | Li, Zezheng, Yip, Kingston |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2404.08836 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
par: Reddy, K. Sahit, et autres
Publié: (2025)
par: Reddy, K. Sahit, et autres
Publié: (2025)
AI-Generated Text Detection and Classification Based on BERT Deep Learning Algorithm
par: Wang, Hao, et autres
Publié: (2024)
par: Wang, Hao, et autres
Publié: (2024)
Patent Language Model Pretraining with ModernBERT
par: Yousefiramandi, Amirhossein, et autres
Publié: (2025)
par: Yousefiramandi, Amirhossein, et autres
Publié: (2025)
BERT Learns (and Teaches) Chemistry
par: Payne, Josh, et autres
Publié: (2020)
par: Payne, Josh, et autres
Publié: (2020)
Injecting linguistic knowledge into BERT for Dialogue State Tracking
par: Feng, Xiaohan, et autres
Publié: (2023)
par: Feng, Xiaohan, et autres
Publié: (2023)
A Comprehensive Approach to Misspelling Correction with BERT and Levenshtein Distance
par: Naziri, Amirreza, et autres
Publié: (2024)
par: Naziri, Amirreza, et autres
Publié: (2024)
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction
par: Saoud, Abdulkader, et autres
Publié: (2024)
par: Saoud, Abdulkader, et autres
Publié: (2024)
BERTCaps: BERT Capsule for Persian Multi-Domain Sentiment Analysis
par: Memari, Mohammadali, et autres
Publié: (2024)
par: Memari, Mohammadali, et autres
Publié: (2024)
BERT-JEPA: Reorganizing CLS Embeddings for Language-Invariant Semantics
par: Gillin, Taj, et autres
Publié: (2026)
par: Gillin, Taj, et autres
Publié: (2026)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
par: Jung, Hee-Jun, et autres
Publié: (2022)
par: Jung, Hee-Jun, et autres
Publié: (2022)
Absolute convergence and error thresholds in non-active adaptive sampling
par: Ferro, Manuel Vilares, et autres
Publié: (2024)
par: Ferro, Manuel Vilares, et autres
Publié: (2024)
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
par: Zhao, Andrew, et autres
Publié: (2025)
par: Zhao, Andrew, et autres
Publié: (2025)
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
par: Lin, Zezheng, et autres
Publié: (2026)
par: Lin, Zezheng, et autres
Publié: (2026)
Seeing Through VisualBERT: A Causal Adventure on Memetic Landscapes
par: Bandyopadhyay, Dibyanayan, et autres
Publié: (2024)
par: Bandyopadhyay, Dibyanayan, et autres
Publié: (2024)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
par: Lee, Simon A., et autres
Publié: (2025)
par: Lee, Simon A., et autres
Publié: (2025)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
par: Liang, Chaoqi, et autres
Publié: (2023)
par: Liang, Chaoqi, et autres
Publié: (2023)
2-Tier SimCSE: Elevating BERT for Robust Sentence Embeddings
par: Wang, Yumeng, et autres
Publié: (2025)
par: Wang, Yumeng, et autres
Publié: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
par: S, Santhosh G, et autres
Publié: (2025)
par: S, Santhosh G, et autres
Publié: (2025)
Integrating LSTM and BERT for Long-Sequence Data Analysis in Intelligent Tutoring Systems
par: Li, Zhaoxing, et autres
Publié: (2024)
par: Li, Zhaoxing, et autres
Publié: (2024)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
par: Zain, Noor Ul, et autres
Publié: (2025)
par: Zain, Noor Ul, et autres
Publié: (2025)
BERT-ASC: Auxiliary-Sentence Construction for Implicit Aspect Learning in Sentiment Analysis
par: Ahmed, Murtadha, et autres
Publié: (2022)
par: Ahmed, Murtadha, et autres
Publié: (2022)
MrBERT: Modern Multilingual Encoders via Vocabulary, Domain, and Dimensional Adaptation
par: Tamayo, Daniel, et autres
Publié: (2026)
par: Tamayo, Daniel, et autres
Publié: (2026)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
par: Hallaç, İbrahim Rıza, et autres
Publié: (2026)
par: Hallaç, İbrahim Rıza, et autres
Publié: (2026)
The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks
par: Lin, Zezheng, et autres
Publié: (2026)
par: Lin, Zezheng, et autres
Publié: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
par: Yang, Dongquan, et autres
Publié: (2025)
par: Yang, Dongquan, et autres
Publié: (2025)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
par: S, Santhosh G, et autres
Publié: (2025)
par: S, Santhosh G, et autres
Publié: (2025)
Harnessing Large Language Models: Fine-tuned BERT for Detecting Charismatic Leadership Tactics in Natural Language
par: Saeid, Yasser, et autres
Publié: (2024)
par: Saeid, Yasser, et autres
Publié: (2024)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
par: Wu, Da, et autres
Publié: (2023)
par: Wu, Da, et autres
Publié: (2023)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
par: Huang, Yuxiang, et autres
Publié: (2026)
par: Huang, Yuxiang, et autres
Publié: (2026)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
par: Fu, Zichuan, et autres
Publié: (2026)
par: Fu, Zichuan, et autres
Publié: (2026)
A Fusion of context-aware based BanglaBERT and Two-Layer Stacked LSTM Framework for Multi-Label Cyberbullying Detection
par: Raquib, Mirza, et autres
Publié: (2026)
par: Raquib, Mirza, et autres
Publié: (2026)
Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study
par: Hashtarkhani, Soheil, et autres
Publié: (2025)
par: Hashtarkhani, Soheil, et autres
Publié: (2025)
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
par: Ren, Weijieying, et autres
Publié: (2024)
par: Ren, Weijieying, et autres
Publié: (2024)
Improving VTE Identification through Language Models from Radiology Reports: A Comparative Study of Mamba, Phi-3 Mini, and BERT
par: Deng, Jamie, et autres
Publié: (2024)
par: Deng, Jamie, et autres
Publié: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
par: Deng, Yichuan, et autres
Publié: (2024)
par: Deng, Yichuan, et autres
Publié: (2024)
Reducing the Scope of Language Models
par: Yunis, David, et autres
Publié: (2024)
par: Yunis, David, et autres
Publié: (2024)
What Matters in Transformers? Not All Attention is Needed
par: He, Shwai, et autres
Publié: (2024)
par: He, Shwai, et autres
Publié: (2024)
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
par: Zhao, Stephen, et autres
Publié: (2025)
par: Zhao, Stephen, et autres
Publié: (2025)
More Expressive Attention with Negative Weights
par: Lv, Ang, et autres
Publié: (2024)
par: Lv, Ang, et autres
Publié: (2024)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
par: Sharma, Agniv, et autres
Publié: (2024)
par: Sharma, Agniv, et autres
Publié: (2024)
Documents similaires
-
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
par: Reddy, K. Sahit, et autres
Publié: (2025) -
AI-Generated Text Detection and Classification Based on BERT Deep Learning Algorithm
par: Wang, Hao, et autres
Publié: (2024) -
Patent Language Model Pretraining with ModernBERT
par: Yousefiramandi, Amirhossein, et autres
Publié: (2025) -
BERT Learns (and Teaches) Chemistry
par: Payne, Josh, et autres
Publié: (2020) -
Injecting linguistic knowledge into BERT for Dialogue State Tracking
par: Feng, Xiaohan, et autres
Publié: (2023)