RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saji, Alan, Husain, Jaavid Aktar, Jayakumar, Thanmay, Dabre, Raj, Kunchukuttan, Anoop, Puduppully, Ratish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Exploring the Maze of Multilingual Modeling
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2023)
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2023)
Towards Massive Multilingual Holistic Bias
von: Tan, Xiaoqing Ellen, et al.
Veröffentlicht: (2024)
von: Tan, Xiaoqing Ellen, et al.
Veröffentlicht: (2024)
What Drives Performance in Multilingual Language Models?
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2024)
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2024)
idT5: Indonesian Version of Multilingual T5 Transformer
von: Fuadi, Mukhlish, et al.
Veröffentlicht: (2023)
von: Fuadi, Mukhlish, et al.
Veröffentlicht: (2023)
DimStance: Multilingual Datasets for Dimensional Stance Analysis
von: Becker, Jonas, et al.
Veröffentlicht: (2026)
von: Becker, Jonas, et al.
Veröffentlicht: (2026)
Ensembling Multilingual Transformers for Robust Sentiment Analysis of Tweets
von: Bilehsavar, Meysam Shirdel, et al.
Veröffentlicht: (2025)
von: Bilehsavar, Meysam Shirdel, et al.
Veröffentlicht: (2025)
Socially Responsible Data for Large Multilingual Language Models
von: Smart, Andrew, et al.
Veröffentlicht: (2024)
von: Smart, Andrew, et al.
Veröffentlicht: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
von: Seki, Yohei, et al.
Veröffentlicht: (2024)
von: Seki, Yohei, et al.
Veröffentlicht: (2024)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
von: Kodali, Prashant, et al.
Veröffentlicht: (2025)
von: Kodali, Prashant, et al.
Veröffentlicht: (2025)
Graphemic Normalization of the Perso-Arabic Script
von: Doctor, Raiomond, et al.
Veröffentlicht: (2022)
von: Doctor, Raiomond, et al.
Veröffentlicht: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
von: Gutkin, Alexander, et al.
Veröffentlicht: (2023)
von: Gutkin, Alexander, et al.
Veröffentlicht: (2023)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
von: Lee, Lung-Hao, et al.
Veröffentlicht: (2026)
von: Lee, Lung-Hao, et al.
Veröffentlicht: (2026)
Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
von: Bilehsavar, Meysam Shirdel, et al.
Veröffentlicht: (2026)
von: Bilehsavar, Meysam Shirdel, et al.
Veröffentlicht: (2026)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
von: Ma, Longxuan, et al.
Veröffentlicht: (2023)
von: Ma, Longxuan, et al.
Veröffentlicht: (2023)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
von: Dang, Thao Anh, et al.
Veröffentlicht: (2024)
von: Dang, Thao Anh, et al.
Veröffentlicht: (2024)
LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models
von: Pitorro, Hugo, et al.
Veröffentlicht: (2025)
von: Pitorro, Hugo, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
LLMs and the Human Condition
von: Wallis, Peter
Veröffentlicht: (2024)
von: Wallis, Peter
Veröffentlicht: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
von: Cui, Hyang
Veröffentlicht: (2025)
von: Cui, Hyang
Veröffentlicht: (2025)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
von: Zheng, Qinyue, et al.
Veröffentlicht: (2025)
von: Zheng, Qinyue, et al.
Veröffentlicht: (2025)
Improving LLMs with a knowledge from databases
von: Máša, Petr
Veröffentlicht: (2025)
von: Máša, Petr
Veröffentlicht: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
von: Šindelář, Pavel, et al.
Veröffentlicht: (2025)
von: Šindelář, Pavel, et al.
Veröffentlicht: (2025)
Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
von: Ewais, Ahmed, et al.
Veröffentlicht: (2026)
von: Ewais, Ahmed, et al.
Veröffentlicht: (2026)
Personality, Role, and Expressive Style in Large Language Models: An Interactionist Analysis
von: Nagao, Moe, et al.
Veröffentlicht: (2026)
von: Nagao, Moe, et al.
Veröffentlicht: (2026)
MALT: Mechanistic Ablation of Lossy Translation in LLMs for a Low-Resource Language: Urdu
von: Bajwa, Taaha Saleem
Veröffentlicht: (2025)
von: Bajwa, Taaha Saleem
Veröffentlicht: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025) -
Exploring the Maze of Multilingual Modeling
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2023) -
Towards Massive Multilingual Holistic Bias
von: Tan, Xiaoqing Ellen, et al.
Veröffentlicht: (2024) -
What Drives Performance in Multilingual Language Models?
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2024)