Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups
Fuente:
arXiv
Saved in:
| Main Authors: | Smădu, Răzvan-Alexandru, Ion, David-Gabriel, Cercel, Dumitru-Clementin, Pop, Florin, Cercel, Mihaela-Claudia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Cross-Lingual Meta-Learning Method Based on Domain Adaptation for Speech Emotion Recognition
by: Ion, David-Gabriel, et al.
Published: (2024)
by: Ion, David-Gabriel, et al.
Published: (2024)
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025)
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025)
MuSaRoNews: A Multidomain, Multimodal Satire Dataset from Romanian News Articles
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering
by: Crăciun, Cristian-George, et al.
Published: (2024)
by: Crăciun, Cristian-George, et al.
Published: (2024)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
by: Man, Andrei Vlad, et al.
Published: (2025)
by: Man, Andrei Vlad, et al.
Published: (2025)
Investigating the Impact of Semi-Supervised Methods with Data Augmentation on Offensive Language Detection in Romanian Language
by: Nicola, Elena-Beatrice, et al.
Published: (2024)
by: Nicola, Elena-Beatrice, et al.
Published: (2024)
RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
by: Timpuriu, Mircea, et al.
Published: (2026)
by: Timpuriu, Mircea, et al.
Published: (2026)
Enhancing Romanian Offensive Language Detection through Knowledge Distillation, Multi-Task Learning, and Data Augmentation
by: Matei, Vlad-Cristian, et al.
Published: (2024)
by: Matei, Vlad-Cristian, et al.
Published: (2024)
Air Pollution Forecasting in Bucharest
by: Şerban, Dragoş-Andrei, et al.
Published: (2025)
by: Şerban, Dragoş-Andrei, et al.
Published: (2025)
Multimodal Learning with Augmentation Techniques for Natural Disaster Assessment
by: Urse, Adrian-Dinu, et al.
Published: (2025)
by: Urse, Adrian-Dinu, et al.
Published: (2025)
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
by: Avram, Andrei-Marius, et al.
Published: (2025)
by: Avram, Andrei-Marius, et al.
Published: (2025)
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
by: Dima, George-Andrei, et al.
Published: (2025)
by: Dima, George-Andrei, et al.
Published: (2025)
RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
by: Avram, Andrei-Marius, et al.
Published: (2024)
by: Avram, Andrei-Marius, et al.
Published: (2024)
IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
by: Echim, Sebastian-Vasile, et al.
Published: (2025)
by: Echim, Sebastian-Vasile, et al.
Published: (2025)
Explainability-Driven Leaf Disease Classification Using Adversarial Training and Knowledge Distillation
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
RoQLlama: A Lightweight Romanian Adapted Language Model
by: Dima, George-Andrei, et al.
Published: (2024)
by: Dima, George-Andrei, et al.
Published: (2024)
Evaluating Data Augmentation Techniques for Coffee Leaf Disease Classification
by: Gheorghiu, Adrian, et al.
Published: (2024)
by: Gheorghiu, Adrian, et al.
Published: (2024)
Scaling Federated Learning Solutions with Kubernetes for Synthesizing Histopathology Images
by: Preda, Andrei-Alexandru, et al.
Published: (2025)
by: Preda, Andrei-Alexandru, et al.
Published: (2025)
RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
by: Avram, Andrei-Marius, et al.
Published: (2026)
by: Avram, Andrei-Marius, et al.
Published: (2026)
UniBERT: Adversarial Training for Language-Universal Representations
by: Avram, Andrei-Marius, et al.
Published: (2025)
by: Avram, Andrei-Marius, et al.
Published: (2025)
RoCoISLR: A Romanian Corpus for Isolated Sign Language Recognition
by: Rîpanu, Cătălin-Alexandru, et al.
Published: (2025)
by: Rîpanu, Cătălin-Alexandru, et al.
Published: (2025)
HistNERo: Historical Named Entity Recognition for the Romanian Language
by: Avram, Andrei-Marius, et al.
Published: (2024)
by: Avram, Andrei-Marius, et al.
Published: (2024)
Fine-tuning Large Language Models for Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection
by: Xiong, Feng, et al.
Published: (2024)
by: Xiong, Feng, et al.
Published: (2024)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
by: Dumitru, Alexandru, et al.
Published: (2025)
by: Dumitru, Alexandru, et al.
Published: (2025)
LLMic: Romanian Foundation Language Model
by: Bădoiu, Vlad-Andrei, et al.
Published: (2025)
by: Bădoiu, Vlad-Andrei, et al.
Published: (2025)
Multilingual Vision-Language Models, A Survey
by: Manea, Andrei-Alexandru, et al.
Published: (2025)
by: Manea, Andrei-Alexandru, et al.
Published: (2025)
ELLEN: Extremely Lightly Supervised Learning For Efficient Named Entity Recognition
by: Riaz, Haris, et al.
Published: (2024)
by: Riaz, Haris, et al.
Published: (2024)
FuLG: 150B Romanian Corpus for Language Model Pretraining
by: Bădoiu, Vlad-Andrei, et al.
Published: (2024)
by: Bădoiu, Vlad-Andrei, et al.
Published: (2024)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Multilingual Political Views of Large Language Models: Identification and Steering
by: Gurgurov, Daniil, et al.
Published: (2025)
by: Gurgurov, Daniil, et al.
Published: (2025)
SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
Multilingual Large Language Models and Curse of Multilinguality
by: Gurgurov, Daniil, et al.
Published: (2024)
by: Gurgurov, Daniil, et al.
Published: (2024)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
Pruning Multilingual Large Language Models for Multilingual Inference
by: Kim, Hwichan, et al.
Published: (2024)
by: Kim, Hwichan, et al.
Published: (2024)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
by: Lee, Lung-Hao, et al.
Published: (2026)
by: Lee, Lung-Hao, et al.
Published: (2026)
Language Surgery in Multilingual Large Language Models
by: Lopo, Joanito Agili, et al.
Published: (2025)
by: Lopo, Joanito Agili, et al.
Published: (2025)
Similar Items
-
A Cross-Lingual Meta-Learning Method Based on Domain Adaptation for Speech Emotion Recognition
by: Ion, David-Gabriel, et al.
Published: (2024) -
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025) -
MuSaRoNews: A Multidomain, Multimodal Satire Dataset from Romanian News Articles
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025) -
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering
by: Crăciun, Cristian-George, et al.
Published: (2024) -
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)