BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques
Fuente:
arXiv
Saved in:
| Main Authors: | Kabir, Muhammad Rafsan, Nabil, Md. Mohibur Rahman, Khan, Mohammad Ashrafuzzaman |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages
by: Gupta, Shivanshu, et al.
Published: (2023)
by: Gupta, Shivanshu, et al.
Published: (2023)
LLM Enhancer: Merged Approach using Vector Embedding for Reducing Large Language Model Hallucinations with External Knowledge
by: Rayhan, Naheed, et al.
Published: (2025)
by: Rayhan, Naheed, et al.
Published: (2025)
Linear Cross-Lingual Mapping of Sentence Embeddings
by: Vasilyev, Oleg, et al.
Published: (2023)
by: Vasilyev, Oleg, et al.
Published: (2023)
Cost-Efficient Cross-Lingual Retrieval-Augmented Generation for Low-Resource Languages: A Case Study in Bengali Agricultural Advisory
by: Hossain, Md. Asif, et al.
Published: (2026)
by: Hossain, Md. Asif, et al.
Published: (2026)
Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
by: Neogi, Dipta, et al.
Published: (2025)
by: Neogi, Dipta, et al.
Published: (2025)
BSpell: A CNN-Blended BERT Based Bangla Spell Checker
by: Rahman, Chowdhury Rafeed, et al.
Published: (2022)
by: Rahman, Chowdhury Rafeed, et al.
Published: (2022)
LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish Sentence Embeddings
by: Philippy, Fred, et al.
Published: (2024)
by: Philippy, Fred, et al.
Published: (2024)
Accelerating Bangla NLP Tasks with Automatic Mixed Precision: Resource-Efficient Training Preserving Model Efficacy
by: Opi, Md Mehrab Hossain, et al.
Published: (2025)
by: Opi, Md Mehrab Hossain, et al.
Published: (2025)
Semantically Enriched Cross-Lingual Sentence Embeddings for Crisis-related Social Media Texts
by: Lamsal, Rabindra, et al.
Published: (2024)
by: Lamsal, Rabindra, et al.
Published: (2024)
Simple Techniques for Enhancing Sentence Embeddings in Generative Language Models
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
by: Omnilingual SONAR Team, et al.
Published: (2026)
by: Omnilingual SONAR Team, et al.
Published: (2026)
Balancing Accuracy and Efficiency: CNN Fusion Models for Diabetic Retinopathy Screening
by: Islam, Md Rafid, et al.
Published: (2025)
by: Islam, Md Rafid, et al.
Published: (2025)
Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment
by: Miao, Zhongtao, et al.
Published: (2024)
by: Miao, Zhongtao, et al.
Published: (2024)
Uddessho: An Extensive Benchmark Dataset for Multimodal Author Intent Classification in Low-Resource Bangla Language
by: Faria, Fatema Tuj Johora, et al.
Published: (2024)
by: Faria, Fatema Tuj Johora, et al.
Published: (2024)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2025)
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2025)
Restoring Rhythm: Punctuation Restoration Using Transformer Models for Bangla, A Low-Resource Language
by: Mamun, Md Obyedullahil, et al.
Published: (2025)
by: Mamun, Md Obyedullahil, et al.
Published: (2025)
A Data Selection Approach for Enhancing Low Resource Machine Translation Using Cross-Lingual Sentence Representations
by: Kowtal, Nidhi, et al.
Published: (2024)
by: Kowtal, Nidhi, et al.
Published: (2024)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
by: Ahad, Jawad Ibn, et al.
Published: (2025)
by: Ahad, Jawad Ibn, et al.
Published: (2025)
Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
by: Bhuiyan, Samiul Basir, et al.
Published: (2025)
by: Bhuiyan, Samiul Basir, et al.
Published: (2025)
LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment
by: Kim, Jong Myoung, et al.
Published: (2025)
by: Kim, Jong Myoung, et al.
Published: (2025)
Using Multimodal and Language-Agnostic Sentence Embeddings for Abstractive Summarization
by: Chellaf, Chaimae, et al.
Published: (2026)
by: Chellaf, Chaimae, et al.
Published: (2026)
Refining Sentence Embedding Model through Ranking Sentences Generation with Large Language Models
by: He, Liyang, et al.
Published: (2025)
by: He, Liyang, et al.
Published: (2025)
Benchmarking Cross-Lingual Semantic Alignment in Multilingual Embeddings
by: Gong, Wen G.
Published: (2025)
by: Gong, Wen G.
Published: (2025)
Cross-Lingual Transfer for Low-Resource Natural Language Processing
by: García-Ferrero, Iker
Published: (2025)
by: García-Ferrero, Iker
Published: (2025)
LegalRAG: A Hybrid RAG System for Multilingual Legal Information Retrieval
by: Kabir, Muhammad Rafsan, et al.
Published: (2025)
by: Kabir, Muhammad Rafsan, et al.
Published: (2025)
Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
by: Bandarupalli, Srihari, et al.
Published: (2025)
by: Bandarupalli, Srihari, et al.
Published: (2025)
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi
by: Raihan, Md Nishat, et al.
Published: (2023)
by: Raihan, Md Nishat, et al.
Published: (2023)
Bootstrapping Embeddings for Low Resource Languages
by: Basoz, Merve, et al.
Published: (2026)
by: Basoz, Merve, et al.
Published: (2026)
Sign Language Translation with Sentence Embedding Supervision
by: Hamidullah, Yasser, et al.
Published: (2025)
by: Hamidullah, Yasser, et al.
Published: (2025)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
Zero-Shot Cross-Lingual NER Using Phonemic Representations for Low-Resource Languages
by: Sohn, Jimin, et al.
Published: (2024)
by: Sohn, Jimin, et al.
Published: (2024)
Franken-Adapter: Cross-Lingual Adaptation of LLMs by Embedding Surgery
by: Jiang, Fan, et al.
Published: (2025)
by: Jiang, Fan, et al.
Published: (2025)
Can Cross Encoders Produce Useful Sentence Embeddings?
by: Ananthakrishnan, Haritha, et al.
Published: (2025)
by: Ananthakrishnan, Haritha, et al.
Published: (2025)
Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction
by: Chavan, Ritesh Sunil, et al.
Published: (2025)
by: Chavan, Ritesh Sunil, et al.
Published: (2025)
Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual Alignment
by: Huang, Yongxin, et al.
Published: (2024)
by: Huang, Yongxin, et al.
Published: (2024)
Breaking the Fake News Barrier: Deep Learning Approaches in Bangla Language
by: Mondal, Pronoy Kumar, et al.
Published: (2025)
by: Mondal, Pronoy Kumar, et al.
Published: (2025)
Adaptative Bilingual Aligning Using Multilingual Sentence Embedding
by: Kraif, Olivier
Published: (2024)
by: Kraif, Olivier
Published: (2024)
BanglaForge: LLM Collaboration with Self-Refinement for Bangla Code Generation
by: Dihan, Mahir Labib, et al.
Published: (2025)
by: Dihan, Mahir Labib, et al.
Published: (2025)
Similar Items
-
Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages
by: Gupta, Shivanshu, et al.
Published: (2023) -
LLM Enhancer: Merged Approach using Vector Embedding for Reducing Large Language Model Hallucinations with External Knowledge
by: Rayhan, Naheed, et al.
Published: (2025) -
Linear Cross-Lingual Mapping of Sentence Embeddings
by: Vasilyev, Oleg, et al.
Published: (2023) -
Cost-Efficient Cross-Lingual Retrieval-Augmented Generation for Low-Resource Languages: A Case Study in Bengali Agricultural Advisory
by: Hossain, Md. Asif, et al.
Published: (2026) -
Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
by: Neogi, Dipta, et al.
Published: (2025)