Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
Fuente:
arXiv
Saved in:
| Main Authors: | Medvecki, Darija, Bašaragin, Bojana, Ljajić, Adela, Milošević, Nikola |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Verif.ai: Towards an Open-Source Scientific Generative Question-Answering System with Referenced and Verifiable Answers
by: Košprdić, Miloš, et al.
Published: (2024)
by: Košprdić, Miloš, et al.
Published: (2024)
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
by: Bašaragin, Bojana, et al.
Published: (2024)
by: Bašaragin, Bojana, et al.
Published: (2024)
VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
by: Košprdić, Miloš, et al.
Published: (2026)
by: Košprdić, Miloš, et al.
Published: (2026)
Scientific QA System with Verifiable Answers
by: Ljajić, Adela, et al.
Published: (2024)
by: Ljajić, Adela, et al.
Published: (2024)
From Zero to Hero: Harnessing Transformers for Biomedical Named Entity Recognition in Zero- and Few-shot Contexts
by: Košprdić, Miloš, et al.
Published: (2023)
by: Košprdić, Miloš, et al.
Published: (2023)
Improving Customer Service with Automatic Topic Detection in User Emails
by: Bašaragin, Bojana, et al.
Published: (2025)
by: Bašaragin, Bojana, et al.
Published: (2025)
De-identification of clinical free text using natural language processing: A systematic review of current approaches
by: Kovačević, Aleksandar, et al.
Published: (2023)
by: Kovačević, Aleksandar, et al.
Published: (2023)
Comparison of biomedical relationship extraction methods and models for knowledge graph creation
by: Milosevic, Nikola, et al.
Published: (2022)
by: Milosevic, Nikola, et al.
Published: (2022)
A comparative study of transformer-based embeddings for topic coherence
by: Ding, Alex, et al.
Published: (2026)
by: Ding, Alex, et al.
Published: (2026)
Abusive text transformation using LLMs
by: Chandra, Rohitash, et al.
Published: (2025)
by: Chandra, Rohitash, et al.
Published: (2025)
Personality testing of Large Language Models: Limited temporal stability, but highlighted prosociality
by: Bodroza, Bojana, et al.
Published: (2023)
by: Bodroza, Bojana, et al.
Published: (2023)
Med-gte-hybrid: A contextual embedding transformer model for extracting actionable information from clinical texts
by: Kumar, Aditya, et al.
Published: (2025)
by: Kumar, Aditya, et al.
Published: (2025)
Stylometry recognizes human and LLM-generated texts in short samples
by: Przystalski, Karol, et al.
Published: (2025)
by: Przystalski, Karol, et al.
Published: (2025)
Preference learning in shades of gray: Interpretable and bias-aware reward modeling for human preferences
by: Oprea, Simona-Vasilica, et al.
Published: (2026)
by: Oprea, Simona-Vasilica, et al.
Published: (2026)
A Human Word Association based model for topic detection in social networks
by: Khadivi, Mehrdad Ranjbar, et al.
Published: (2023)
by: Khadivi, Mehrdad Ranjbar, et al.
Published: (2023)
Generative AI for automatic topic labelling
by: Kozlowski, Diego, et al.
Published: (2024)
by: Kozlowski, Diego, et al.
Published: (2024)
Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents
by: Milosevic, Nikola
Published: (2026)
by: Milosevic, Nikola
Published: (2026)
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets
by: Münker, Simon, et al.
Published: (2024)
by: Münker, Simon, et al.
Published: (2024)
Large language models struggle with ethnographic text annotation
by: Goodall, Leonardo S., et al.
Published: (2026)
by: Goodall, Leonardo S., et al.
Published: (2026)
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
by: Gupta, Aryan, et al.
Published: (2025)
by: Gupta, Aryan, et al.
Published: (2025)
ARC-Encoder: learning compressed text representations for large language models
by: Pilchen, Hippolyte, et al.
Published: (2025)
by: Pilchen, Hippolyte, et al.
Published: (2025)
New Textual Corpora for Serbian Language Modeling
by: Škorić, Mihailo, et al.
Published: (2024)
by: Škorić, Mihailo, et al.
Published: (2024)
The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT
by: Xylogiannopoulos, Konstantinos, et al.
Published: (2025)
by: Xylogiannopoulos, Konstantinos, et al.
Published: (2025)
The study of short texts in digital politics: Document aggregation for topic modeling
by: Nakka, Nitheesha, et al.
Published: (2025)
by: Nakka, Nitheesha, et al.
Published: (2025)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
by: Cunha, Washington, et al.
Published: (2025)
by: Cunha, Washington, et al.
Published: (2025)
Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress
by: Odumakinde, Ayomide, et al.
Published: (2024)
by: Odumakinde, Ayomide, et al.
Published: (2024)
Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review
by: McGiff, Josh, et al.
Published: (2025)
by: McGiff, Josh, et al.
Published: (2025)
Performance Trade-offs of Optimizing Small Language Models for E-Commerce
by: Licardo, Josip Tomo, et al.
Published: (2025)
by: Licardo, Josip Tomo, et al.
Published: (2025)
Generalist embedding models are better at short-context clinical semantic search than specialized embedding models
by: Excoffier, Jean-Baptiste, et al.
Published: (2024)
by: Excoffier, Jean-Baptiste, et al.
Published: (2024)
Forma mentis networks predict creativity ratings of short texts via interpretable artificial intelligence in human and GPT-simulated raters
by: Haim, Edith, et al.
Published: (2024)
by: Haim, Edith, et al.
Published: (2024)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
by: Gupta, Ashray, et al.
Published: (2025)
by: Gupta, Ashray, et al.
Published: (2025)
Evaluating how LLM annotations represent diverse views on contentious topics
by: Brown, Megan A., et al.
Published: (2025)
by: Brown, Megan A., et al.
Published: (2025)
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
by: Gomes, Gonçalo, et al.
Published: (2025)
by: Gomes, Gonçalo, et al.
Published: (2025)
Benchmark of stylistic variation in LLM-generated texts
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Are generative AI text annotations systematically biased?
by: Stolwijk, Sjoerd B., et al.
Published: (2025)
by: Stolwijk, Sjoerd B., et al.
Published: (2025)
PAGE: Prompt Augmentation for text Generation Enhancement
by: Pacchiotti, Mauro Jose, et al.
Published: (2025)
by: Pacchiotti, Mauro Jose, et al.
Published: (2025)
Serialized EHR make for good text representations
by: Chou, Zhirong, et al.
Published: (2025)
by: Chou, Zhirong, et al.
Published: (2025)
Ensemble BERT: A student social network text sentiment classification model based on ensemble learning and BERT architecture
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
Multilingual Target-Stance Extraction
by: Mines, Ethan, et al.
Published: (2025)
by: Mines, Ethan, et al.
Published: (2025)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
by: Chen, Pinzhen, et al.
Published: (2024)
by: Chen, Pinzhen, et al.
Published: (2024)
Similar Items
-
Verif.ai: Towards an Open-Source Scientific Generative Question-Answering System with Referenced and Verifiable Answers
by: Košprdić, Miloš, et al.
Published: (2024) -
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
by: Bašaragin, Bojana, et al.
Published: (2024) -
VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
by: Košprdić, Miloš, et al.
Published: (2026) -
Scientific QA System with Verifiable Answers
by: Ljajić, Adela, et al.
Published: (2024) -
From Zero to Hero: Harnessing Transformers for Biomedical Named Entity Recognition in Zero- and Few-shot Contexts
by: Košprdić, Miloš, et al.
Published: (2023)