VN-MTEB: Vietnamese Massive Text Embedding Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pham, Loc, Luu, Tung, Vo, Thu, Nguyen, Minh, Hoang, Viet |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning
von: Tung, Luu Quy, et al.
Veröffentlicht: (2025)
von: Tung, Luu Quy, et al.
Veröffentlicht: (2025)
ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair
von: Tran, Hong-Viet, et al.
Veröffentlicht: (2025)
von: Tran, Hong-Viet, et al.
Veröffentlicht: (2025)
A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
von: Hai, Toan Nguyen, et al.
Veröffentlicht: (2025)
von: Hai, Toan Nguyen, et al.
Veröffentlicht: (2025)
VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
von: Luu, Son T., et al.
Veröffentlicht: (2025)
von: Luu, Son T., et al.
Veröffentlicht: (2025)
New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis
von: Nguyen, Quy Hoang, et al.
Veröffentlicht: (2024)
von: Nguyen, Quy Hoang, et al.
Veröffentlicht: (2024)
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
von: Minh, Nguyen Huu Nhat, et al.
Veröffentlicht: (2025)
von: Minh, Nguyen Huu Nhat, et al.
Veröffentlicht: (2025)
PL-MTEB: Polish Massive Text Embedding Benchmark
von: Poświata, Rafał, et al.
Veröffentlicht: (2024)
von: Poświata, Rafał, et al.
Veröffentlicht: (2024)
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
von: Chung, Isaac, et al.
Veröffentlicht: (2025)
von: Chung, Isaac, et al.
Veröffentlicht: (2025)
Vietnamese AI Generated Text Detection
von: Tran, Quang-Dan, et al.
Veröffentlicht: (2024)
von: Tran, Quang-Dan, et al.
Veröffentlicht: (2024)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
ZeFaV: Boosting Large Language Models for Zero-shot Fact Verification
von: Luu, Son T., et al.
Veröffentlicht: (2024)
von: Luu, Son T., et al.
Veröffentlicht: (2024)
VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation
von: Phan, Hoang Hai, et al.
Veröffentlicht: (2025)
von: Phan, Hoang Hai, et al.
Veröffentlicht: (2025)
FinMTEB: Finance Massive Text Embedding Benchmark
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents
von: Pham, Khanh Trinh, et al.
Veröffentlicht: (2025)
von: Pham, Khanh Trinh, et al.
Veröffentlicht: (2025)
FaMTEB: Massive Text Embedding Benchmark in Persian Language
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
MMTEB: Massive Multilingual Text Embedding Benchmark
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2025)
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2025)
BERT-based model for Vietnamese Fact Verification Dataset
von: Tran, Bao, et al.
Veröffentlicht: (2025)
von: Tran, Bao, et al.
Veröffentlicht: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
von: Nguyen, Tan-Minh, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan-Minh, et al.
Veröffentlicht: (2025)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2026)
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2026)
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2025)
A Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media
von: Nguyen, Dung Ha, et al.
Veröffentlicht: (2024)
von: Nguyen, Dung Ha, et al.
Veröffentlicht: (2024)
Using Large Language Models for education managements in Vietnamese with low resources
von: Minh, Duc Do, et al.
Veröffentlicht: (2025)
von: Minh, Duc Do, et al.
Veröffentlicht: (2025)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
ViMMRC 2.0 -- Enhancing Machine Reading Comprehension on Vietnamese Literature Text
von: Luu, Son T., et al.
Veröffentlicht: (2023)
von: Luu, Son T., et al.
Veröffentlicht: (2023)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
Fact4ac at the Financial Misinformation Detection Challenge Task: Reference-Free Financial Misinformation Detection via Fine-Tuning and Few-Shot Prompting of Large Language Models
von: Hoang, Cuong, et al.
Veröffentlicht: (2026)
von: Hoang, Cuong, et al.
Veröffentlicht: (2026)
Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark
von: Cao, Hongliu
Veröffentlicht: (2024)
von: Cao, Hongliu
Veröffentlicht: (2024)
ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
von: Dang, Phuong-Nam, et al.
Veröffentlicht: (2025)
von: Dang, Phuong-Nam, et al.
Veröffentlicht: (2025)
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2024)
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2024)
Enhancing Retrieval Augmented Generation with Hierarchical Text Segmentation Chunking
von: Nguyen, Hai Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai Toan, et al.
Veröffentlicht: (2025)
NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
von: Nguyen, Hoang-Trung, et al.
Veröffentlicht: (2025)
von: Nguyen, Hoang-Trung, et al.
Veröffentlicht: (2025)
The Massive Legal Embedding Benchmark (MLEB)
von: Butler, Umar, et al.
Veröffentlicht: (2025)
von: Butler, Umar, et al.
Veröffentlicht: (2025)
The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design
von: Snegirev, Artem, et al.
Veröffentlicht: (2024)
von: Snegirev, Artem, et al.
Veröffentlicht: (2024)
Vietnamese Poem Generation & The Prospect Of Cross-Language Poem-To-Poem Translation
von: Huynh, Triet Minh, et al.
Veröffentlicht: (2024)
von: Huynh, Triet Minh, et al.
Veröffentlicht: (2024)
German Text Embedding Clustering Benchmark
von: Wehrli, Silvan, et al.
Veröffentlicht: (2024)
von: Wehrli, Silvan, et al.
Veröffentlicht: (2024)
FastDiSS: Few-step Match Many-step Diffusion Language Model on Sequence-to-Sequence Generation--Full Version
von: Nguyen-Cong, Dat, et al.
Veröffentlicht: (2026)
von: Nguyen-Cong, Dat, et al.
Veröffentlicht: (2026)
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
von: Tran, Hung Quang, et al.
Veröffentlicht: (2026)
von: Tran, Hung Quang, et al.
Veröffentlicht: (2026)
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning
von: Tung, Luu Quy, et al.
Veröffentlicht: (2025) -
ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair
von: Tran, Hong-Viet, et al.
Veröffentlicht: (2025) -
A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
von: Hai, Toan Nguyen, et al.
Veröffentlicht: (2025) -
VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
von: Luu, Son T., et al.
Veröffentlicht: (2025) -
New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis
von: Nguyen, Quy Hoang, et al.
Veröffentlicht: (2024)