Self-Supervised Borrowing Detection on Multilingual Wordlists
Fuente:
arXiv
Guardado en:
| Autor principal: | Wientzek, Tim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More
por: Frydenlund, Arvid
Publicado: (2025)
por: Frydenlund, Arvid
Publicado: (2025)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
por: Borisov, Vadim
Publicado: (2026)
por: Borisov, Vadim
Publicado: (2026)
Next Token Prediction Is a Dead End for Creativity
por: Olatunji, Ibukun, et al.
Publicado: (2025)
por: Olatunji, Ibukun, et al.
Publicado: (2025)
What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
por: Wachowiak, Lennart, et al.
Publicado: (2025)
por: Wachowiak, Lennart, et al.
Publicado: (2025)
Transparent but Powerful: Explainability, Accuracy, and Generalizability in ADHD Detection from Social Media Data
por: Wiechmann, D., et al.
Publicado: (2024)
por: Wiechmann, D., et al.
Publicado: (2024)
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
por: Hua, Wenjie, et al.
Publicado: (2025)
por: Hua, Wenjie, et al.
Publicado: (2025)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
por: Banthia, Saumya, et al.
Publicado: (2020)
por: Banthia, Saumya, et al.
Publicado: (2020)
What is Wrong with Language Models that Can Not Tell a Story?
por: Yamshchikov, Ivan P., et al.
Publicado: (2022)
por: Yamshchikov, Ivan P., et al.
Publicado: (2022)
NAAQA: A Neural Architecture for Acoustic Question Answering
por: Abdelnour, Jerome, et al.
Publicado: (2021)
por: Abdelnour, Jerome, et al.
Publicado: (2021)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
por: Demir, M. Mikail, et al.
Publicado: (2025)
por: Demir, M. Mikail, et al.
Publicado: (2025)
Improving French Synthetic Speech Quality via SSML Prosody Control
por: Ouali, Nassima Ould, et al.
Publicado: (2025)
por: Ouali, Nassima Ould, et al.
Publicado: (2025)
LLMs Generate Kitsch
por: Klinge, Xenia, et al.
Publicado: (2026)
por: Klinge, Xenia, et al.
Publicado: (2026)
Comparative Study of Large Language Models on Chinese Film Script Continuation: An Empirical Analysis Based on GPT-5.2 and Qwen-Max
por: Cao, Yuxuan, et al.
Publicado: (2026)
por: Cao, Yuxuan, et al.
Publicado: (2026)
Beyond Subtokens: A Rich Character Embedding for Low-resource and Morphologically Complex Languages
por: Schneider, Felix, et al.
Publicado: (2026)
por: Schneider, Felix, et al.
Publicado: (2026)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
Adversarially Probing Cross-Family Sound Symbolism in 27 Languages
por: Sharma, Anika, et al.
Publicado: (2025)
por: Sharma, Anika, et al.
Publicado: (2025)
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
por: Wright, Devin R., et al.
Publicado: (2026)
por: Wright, Devin R., et al.
Publicado: (2026)
Towards Reliable Retrieval in RAG Systems for Large Legal Datasets
por: Reuter, Markus, et al.
Publicado: (2025)
por: Reuter, Markus, et al.
Publicado: (2025)
The Table of Media Bias Elements: A sentence-level taxonomy of media bias types and propaganda techniques
por: Menzner, Tim, et al.
Publicado: (2026)
por: Menzner, Tim, et al.
Publicado: (2026)
Probing neural audio codecs for distinctions among English nuclear tunes
por: Vigneaux, Juan Pablo, et al.
Publicado: (2026)
por: Vigneaux, Juan Pablo, et al.
Publicado: (2026)
ECLAIR: Enhanced Clarification for Interactive Responses in an Enterprise AI Assistant
por: Murzaku, John, et al.
Publicado: (2025)
por: Murzaku, John, et al.
Publicado: (2025)
Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages
por: de Zuazo, Xabier, et al.
Publicado: (2025)
por: de Zuazo, Xabier, et al.
Publicado: (2025)
Chronic pain patient narratives allow for the estimation of current pain intensity
por: Nunes, Diogo A. P., et al.
Publicado: (2022)
por: Nunes, Diogo A. P., et al.
Publicado: (2022)
DROID: Dual Representation for Out-of-Scope Intent Detection
por: Rashwan, Wael, et al.
Publicado: (2025)
por: Rashwan, Wael, et al.
Publicado: (2025)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
por: Salvi, Giampiero
Publicado: (2024)
por: Salvi, Giampiero
Publicado: (2024)
Where is my Glass Slipper? AI, Poetry and Art
por: Pagiaslis, Anastasios P.
Publicado: (2025)
por: Pagiaslis, Anastasios P.
Publicado: (2025)
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models
por: Cuscito, Miriam, et al.
Publicado: (2024)
por: Cuscito, Miriam, et al.
Publicado: (2024)
Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study
por: Baltaji, Razan, et al.
Publicado: (2023)
por: Baltaji, Razan, et al.
Publicado: (2023)
Reasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
por: Shih, Yu-Fei, et al.
Publicado: (2025)
por: Shih, Yu-Fei, et al.
Publicado: (2025)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
por: Kajare, Prajwal Vijay, et al.
Publicado: (2026)
por: Kajare, Prajwal Vijay, et al.
Publicado: (2026)
HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction
por: Arief, Hasan
Publicado: (2026)
por: Arief, Hasan
Publicado: (2026)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
por: Baranwal, Aaditya, et al.
Publicado: (2026)
por: Baranwal, Aaditya, et al.
Publicado: (2026)
MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
por: Wu, Yuexin, et al.
Publicado: (2025)
por: Wu, Yuexin, et al.
Publicado: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
por: Drchal, Jan, et al.
Publicado: (2023)
por: Drchal, Jan, et al.
Publicado: (2023)
Automating Clinical Information Retrieval from Finnish Electronic Health Records Using Large Language Models
por: Saukkoriipi, Mikko, et al.
Publicado: (2026)
por: Saukkoriipi, Mikko, et al.
Publicado: (2026)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English
por: Serditova, Dana, et al.
Publicado: (2026)
por: Serditova, Dana, et al.
Publicado: (2026)
Devanagari Handwritten Character Recognition using Convolutional Neural Network
por: Mehta, Diksha, et al.
Publicado: (2025)
por: Mehta, Diksha, et al.
Publicado: (2025)
Equivalence: An analysis of artists' roles with Image Generative AI from Conceptual Art perspective through an interactive installation design practice
por: Li, Yixuan, et al.
Publicado: (2024)
por: Li, Yixuan, et al.
Publicado: (2024)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
por: Xi, Wang, et al.
Publicado: (2025)
por: Xi, Wang, et al.
Publicado: (2025)
Ejemplares similares
-
Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More
por: Frydenlund, Arvid
Publicado: (2025) -
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
por: Borisov, Vadim
Publicado: (2026) -
Next Token Prediction Is a Dead End for Creativity
por: Olatunji, Ibukun, et al.
Publicado: (2025) -
What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
por: Wachowiak, Lennart, et al.
Publicado: (2025) -
Transparent but Powerful: Explainability, Accuracy, and Generalizability in ADHD Detection from Social Media Data
por: Wiechmann, D., et al.
Publicado: (2024)