Autoencoder-Based Framework to Capture Vocabulary Quality in NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Dang, Vu Minh Hoang, Verma, Rakesh M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
by: Verma, Rakesh M., et al.
Published: (2024)
by: Verma, Rakesh M., et al.
Published: (2024)
VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation
by: Phan, Hoang Hai, et al.
Published: (2025)
by: Phan, Hoang Hai, et al.
Published: (2025)
NLP Datasets for Idiom and Figurative Language Tasks
by: Matheny, Blake, et al.
Published: (2025)
by: Matheny, Blake, et al.
Published: (2025)
Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
by: Poświata, Rafał, et al.
Published: (2025)
by: Poświata, Rafał, et al.
Published: (2025)
Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models
by: Tuck, Bryan E., et al.
Published: (2025)
by: Tuck, Bryan E., et al.
Published: (2025)
Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated Tweets
by: Tuck, Bryan E., et al.
Published: (2024)
by: Tuck, Bryan E., et al.
Published: (2024)
Collaboration or Corporate Capture? Quantifying NLP's Reliance on Industry Artifacts and Contributions
by: Aitken, Will, et al.
Published: (2023)
by: Aitken, Will, et al.
Published: (2023)
DFKI-NLP at SemEval-2024 Task 2: Towards Robust LLMs Using Data Perturbations and MinMax Training
by: Verma, Bhuvanesh, et al.
Published: (2024)
by: Verma, Bhuvanesh, et al.
Published: (2024)
UETQuintet at BioCreative IX -- MedHopQA: Enhancing Biomedical QA with Selective Multi-hop Reasoning and Contextual Retrieval
by: Nguyen, Quoc-An, et al.
Published: (2026)
by: Nguyen, Quoc-An, et al.
Published: (2026)
Improving LLM Unlearning Robustness via Random Perturbations
by: Huu-Tien, Dang, et al.
Published: (2025)
by: Huu-Tien, Dang, et al.
Published: (2025)
The Nature of NLP: Analyzing Contributions in NLP Papers
by: Pramanick, Aniket, et al.
Published: (2024)
by: Pramanick, Aniket, et al.
Published: (2024)
Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
by: Remy, François, et al.
Published: (2024)
by: Remy, François, et al.
Published: (2024)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
by: Tuck, Bryan E., et al.
Published: (2025)
by: Tuck, Bryan E., et al.
Published: (2025)
Homograph Attacks on Maghreb Sentiment Analyzers
by: Qachfar, Fatima Zahra, et al.
Published: (2024)
by: Qachfar, Fatima Zahra, et al.
Published: (2024)
Effects of Soft-Domain Transfer and Named Entity Information on Deception Detection
by: Triplett, Steven, et al.
Published: (2024)
by: Triplett, Steven, et al.
Published: (2024)
VLSP 2023 -- LTER: A Summary of the Challenge on Legal Textual Entailment Recognition
by: Tran, Vu, et al.
Published: (2024)
by: Tran, Vu, et al.
Published: (2024)
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions
by: Hamed, Injy, et al.
Published: (2025)
by: Hamed, Injy, et al.
Published: (2025)
Enhancing Document Retrieval in COVID-19 Research: Leveraging Large Language Models for Hidden Relation Extraction
by: Trieu, Hoang-An, et al.
Published: (2025)
by: Trieu, Hoang-An, et al.
Published: (2025)
Capturing Opinion Shifts in Deliberative Discourse through Frequency-based Quantum deep learning methods
by: Thakur, Rakesh, et al.
Published: (2025)
by: Thakur, Rakesh, et al.
Published: (2025)
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?
by: Han, HyoJung, et al.
Published: (2024)
by: Han, HyoJung, et al.
Published: (2024)
TurkicNLP: An NLP Toolkit for Turkic Languages
by: Hakimov, Sherzod
Published: (2026)
by: Hakimov, Sherzod
Published: (2026)
ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models
by: Han, Simeng, et al.
Published: (2025)
by: Han, Simeng, et al.
Published: (2025)
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
by: Hoang, Nhat M., et al.
Published: (2024)
by: Hoang, Nhat M., et al.
Published: (2024)
Standardising the NLP Workflow: A Framework for Reproducible Linguistic Analysis
by: Pauli, Yves, et al.
Published: (2025)
by: Pauli, Yves, et al.
Published: (2025)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR
by: Kopparapu, Sunil Kumar
Published: (2026)
by: Kopparapu, Sunil Kumar
Published: (2026)
A conversational gesture synthesis system based on emotions and semantics
by: Hoang-Minh, Thanh
Published: (2025)
by: Hoang-Minh, Thanh
Published: (2025)
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
by: Tatariya, Kushal, et al.
Published: (2024)
by: Tatariya, Kushal, et al.
Published: (2024)
Fact4ac at the Financial Misinformation Detection Challenge Task: Reference-Free Financial Misinformation Detection via Fine-Tuning and Few-Shot Prompting of Large Language Models
by: Hoang, Cuong, et al.
Published: (2026)
by: Hoang, Cuong, et al.
Published: (2026)
SpecMind: Cognitively Inspired, Interactive Multi-Turn Framework for Postcondition Inference
by: Le, Cuong Chi, et al.
Published: (2026)
by: Le, Cuong Chi, et al.
Published: (2026)
NLP-AKG: Few-Shot Construction of NLP Academic Knowledge Graph Based on LLM
by: Lan, Jiayin, et al.
Published: (2025)
by: Lan, Jiayin, et al.
Published: (2025)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
by: Dong, Nguyen Tien, et al.
Published: (2025)
by: Dong, Nguyen Tien, et al.
Published: (2025)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Less Context, Same Performance: A RAG Framework for Resource-Efficient LLM-Based Clinical NLP
by: Cheetirala, Satya Narayana, et al.
Published: (2025)
by: Cheetirala, Satya Narayana, et al.
Published: (2025)
OWLViz: An Open-World Benchmark for Visual Question Answering
by: Nguyen, Thuy, et al.
Published: (2025)
by: Nguyen, Thuy, et al.
Published: (2025)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
by: Hasanaath, Ahmed, et al.
Published: (2025)
by: Hasanaath, Ahmed, et al.
Published: (2025)
Cross-Demographic Portability of Deep NLP-Based Depression Models
by: Rutowski, Tomek, et al.
Published: (2024)
by: Rutowski, Tomek, et al.
Published: (2024)
New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis
by: Nguyen, Quy Hoang, et al.
Published: (2024)
by: Nguyen, Quy Hoang, et al.
Published: (2024)
PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues
by: Zent, Matthew, et al.
Published: (2025)
by: Zent, Matthew, et al.
Published: (2025)
Few-Shot Optimized Framework for Hallucination Detection in Resource-Limited NLP Systems
by: Hikal, Baraa, et al.
Published: (2025)
by: Hikal, Baraa, et al.
Published: (2025)
Similar Items
-
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
by: Verma, Rakesh M., et al.
Published: (2024) -
VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation
by: Phan, Hoang Hai, et al.
Published: (2025) -
NLP Datasets for Idiom and Figurative Language Tasks
by: Matheny, Blake, et al.
Published: (2025) -
Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
by: Poświata, Rafał, et al.
Published: (2025) -
Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models
by: Tuck, Bryan E., et al.
Published: (2025)