HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-Training
Fuente:
arXiv
Guardado en:
| Autor principal: | Choi, Seungho |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Translating Hanja Historical Documents to Contemporary Korean and English
por: Son, Juhee, et al.
Publicado: (2022)
por: Son, Juhee, et al.
Publicado: (2022)
HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja
por: Song, Seyoung, et al.
Publicado: (2025)
por: Song, Seyoung, et al.
Publicado: (2025)
Small Language Model Helps Resolve Semantic Ambiguity of LLM Prompt
por: Huang, Zhenzhen, et al.
Publicado: (2026)
por: Huang, Zhenzhen, et al.
Publicado: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
por: Kim, Hyeonwoo, et al.
Publicado: (2024)
por: Kim, Hyeonwoo, et al.
Publicado: (2024)
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
por: Kim, SungHo, et al.
Publicado: (2026)
por: Kim, SungHo, et al.
Publicado: (2026)
Resolving Intent Ambiguities by Retrieving Discriminative Clarifying Questions
por: Dhole, Kaustubh D.
Publicado: (2020)
por: Dhole, Kaustubh D.
Publicado: (2020)
nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning
por: Luo, Tianqi, et al.
Publicado: (2025)
por: Luo, Tianqi, et al.
Publicado: (2025)
Resolving Transcription Ambiguity in Spanish: A Hybrid Acoustic-Lexical System for Punctuation Restoration
por: Zhu, Xiliang, et al.
Publicado: (2024)
por: Zhu, Xiliang, et al.
Publicado: (2024)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
por: Lee, Jiyoung, et al.
Publicado: (2025)
por: Lee, Jiyoung, et al.
Publicado: (2025)
Language Models Identify Ambiguities and Exploit Loopholes
por: Choi, Jio, et al.
Publicado: (2025)
por: Choi, Jio, et al.
Publicado: (2025)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
por: Zhang, Zhuoxuan, et al.
Publicado: (2025)
por: Zhang, Zhuoxuan, et al.
Publicado: (2025)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
por: Wu, Xinwei, et al.
Publicado: (2025)
por: Wu, Xinwei, et al.
Publicado: (2025)
Optimizing Language Augmentation for Multilingual Large Language Models: A Case Study on Korean
por: Choi, ChangSu, et al.
Publicado: (2024)
por: Choi, ChangSu, et al.
Publicado: (2024)
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
por: Jin, Kyohoon, et al.
Publicado: (2025)
por: Jin, Kyohoon, et al.
Publicado: (2025)
Cascaded Self-Evaluation Augmented Training for Lightweight Multimodal LLMs
por: Lv, Zheqi, et al.
Publicado: (2025)
por: Lv, Zheqi, et al.
Publicado: (2025)
Node Importance Estimation Leveraging LLMs for Semantic Augmentation in Knowledge Graphs
por: Lin, Xinyu, et al.
Publicado: (2024)
por: Lin, Xinyu, et al.
Publicado: (2024)
KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models
por: Kim, Taeeun, et al.
Publicado: (2025)
por: Kim, Taeeun, et al.
Publicado: (2025)
Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing
por: Saparina, Irina, et al.
Publicado: (2025)
por: Saparina, Irina, et al.
Publicado: (2025)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
por: Ceron, Tanise, et al.
Publicado: (2025)
por: Ceron, Tanise, et al.
Publicado: (2025)
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models
por: Kim, Yeeun, et al.
Publicado: (2024)
por: Kim, Yeeun, et al.
Publicado: (2024)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
por: Kumar, Rajeev, et al.
Publicado: (2025)
por: Kumar, Rajeev, et al.
Publicado: (2025)
Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs
por: Shore, Amber, et al.
Publicado: (2025)
por: Shore, Amber, et al.
Publicado: (2025)
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
por: Liu, Yaokun, et al.
Publicado: (2026)
por: Liu, Yaokun, et al.
Publicado: (2026)
From Ambiguity to Accuracy: The Transformative Effect of Coreference Resolution on Retrieval-Augmented Generation systems
por: Jang, Youngjoon, et al.
Publicado: (2025)
por: Jang, Youngjoon, et al.
Publicado: (2025)
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
por: Addepalli, Sravanti, et al.
Publicado: (2024)
por: Addepalli, Sravanti, et al.
Publicado: (2024)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
por: Zheng, Weihua, et al.
Publicado: (2026)
por: Zheng, Weihua, et al.
Publicado: (2026)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
por: He, Qianxi, et al.
Publicado: (2025)
por: He, Qianxi, et al.
Publicado: (2025)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
por: Tian, Junfeng, et al.
Publicado: (2024)
por: Tian, Junfeng, et al.
Publicado: (2024)
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
por: Kim, SungHo, et al.
Publicado: (2025)
por: Kim, SungHo, et al.
Publicado: (2025)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
por: Yao, Louie Hong, et al.
Publicado: (2025)
por: Yao, Louie Hong, et al.
Publicado: (2025)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
por: Keluskar, Aryan, et al.
Publicado: (2024)
por: Keluskar, Aryan, et al.
Publicado: (2024)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
por: Zhang, Yuhao, et al.
Publicado: (2025)
por: Zhang, Yuhao, et al.
Publicado: (2025)
Thoth: Mid-Training Bridges LLMs to Time Series Understanding
por: Lin, Jiafeng, et al.
Publicado: (2026)
por: Lin, Jiafeng, et al.
Publicado: (2026)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
por: Lim, Junghwan, et al.
Publicado: (2025)
por: Lim, Junghwan, et al.
Publicado: (2025)
ConflictRAG: Detecting and Resolving Knowledge Conflicts in Retrieval Augmented Generation
por: Wang, Chenyu, et al.
Publicado: (2026)
por: Wang, Chenyu, et al.
Publicado: (2026)
Rethinking Reflection in Pre-Training
por: AI, Essential, et al.
Publicado: (2025)
por: AI, Essential, et al.
Publicado: (2025)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
por: Park, Chanwoo, et al.
Publicado: (2025)
por: Park, Chanwoo, et al.
Publicado: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
por: Kang, Feiyang, et al.
Publicado: (2024)
por: Kang, Feiyang, et al.
Publicado: (2024)
Zero-shot Graph Reasoning via Retrieval Augmented Framework with LLMs
por: Li, Hanqing, et al.
Publicado: (2025)
por: Li, Hanqing, et al.
Publicado: (2025)
QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation
por: Li, Jiazheng, et al.
Publicado: (2025)
por: Li, Jiazheng, et al.
Publicado: (2025)
Ejemplares similares
-
Translating Hanja Historical Documents to Contemporary Korean and English
por: Son, Juhee, et al.
Publicado: (2022) -
HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja
por: Song, Seyoung, et al.
Publicado: (2025) -
Small Language Model Helps Resolve Semantic Ambiguity of LLM Prompt
por: Huang, Zhenzhen, et al.
Publicado: (2026) -
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
por: Kim, Hyeonwoo, et al.
Publicado: (2024) -
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
por: Kim, SungHo, et al.
Publicado: (2026)