Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Chae, Kyubyung, Kim, Gihoon, Lee, Gyuseong, Kim, Taesup, Lee, Jaejin, Kim, Heejin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
by: Hahm, Sungeun, et al.
Published: (2025)
by: Hahm, Sungeun, et al.
Published: (2025)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
by: Chae, Kyubyung, et al.
Published: (2024)
by: Chae, Kyubyung, et al.
Published: (2024)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)
by: Ko, Donghyeon, et al.
Published: (2025)
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
by: Jang, Jinkwan, et al.
Published: (2026)
by: Jang, Jinkwan, et al.
Published: (2026)
Model-based Preference Optimization in Abstractive Summarization without Human Feedback
by: Choi, Jaepill, et al.
Published: (2024)
by: Choi, Jaepill, et al.
Published: (2024)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
by: So, Yeonkyoung, et al.
Published: (2025)
by: So, Yeonkyoung, et al.
Published: (2025)
Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
by: Yang, Seoyun, et al.
Published: (2025)
by: Yang, Seoyun, et al.
Published: (2025)
Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
by: Kim, Gihoon, et al.
Published: (2025)
by: Kim, Gihoon, et al.
Published: (2025)
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
by: Kim, Seorin, et al.
Published: (2025)
by: Kim, Seorin, et al.
Published: (2025)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026)
by: Lee, Joosung, et al.
Published: (2026)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
by: Son, Youngjun, et al.
Published: (2025)
by: Son, Youngjun, et al.
Published: (2025)
When Vision Models Meet Parameter Efficient Look-Aside Adapters Without Large-Scale Audio Pretraining
by: Yeo, Juan, et al.
Published: (2024)
by: Yeo, Juan, et al.
Published: (2024)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025)
by: Kim, Jinpyo, et al.
Published: (2025)
Multimodal Cognitive Reframing Therapy via Multi-hop Psychotherapeutic Reasoning
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
Autoregressive Score Generation for Multi-trait Essay Scoring
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
by: Yeom, Jewon, et al.
Published: (2026)
by: Yeom, Jewon, et al.
Published: (2026)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
X-PEFT: eXtremely Parameter-Efficient Fine-Tuning for Extreme Multi-Profile Scenarios
by: Kwak, Namju, et al.
Published: (2024)
by: Kwak, Namju, et al.
Published: (2024)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)
by: Jung, Sungmok, et al.
Published: (2026)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
by: Han, Hojae, et al.
Published: (2026)
by: Han, Hojae, et al.
Published: (2026)
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
by: Yeom, Jewon, et al.
Published: (2026)
by: Yeom, Jewon, et al.
Published: (2026)
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
by: Kim, Sean, et al.
Published: (2025)
by: Kim, Sean, et al.
Published: (2025)
DNA 1.0 Technical Report
by: Lee, Jungyup, et al.
Published: (2025)
by: Lee, Jungyup, et al.
Published: (2025)
Raon-Speech Technical Report
by: Kim, Beomsoo, et al.
Published: (2026)
by: Kim, Beomsoo, et al.
Published: (2026)
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
by: Lee, Jiyoung, et al.
Published: (2024)
by: Lee, Jiyoung, et al.
Published: (2024)
Multi-Facet Blending for Faceted Query-by-Example Retrieval
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
K-EXAONE Technical Report
by: Choi, Eunbi, et al.
Published: (2026)
by: Choi, Eunbi, et al.
Published: (2026)
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
by: Park, Seongchan, et al.
Published: (2026)
by: Park, Seongchan, et al.
Published: (2026)
EXAONE 4.5 Technical Report
by: Choi, Eunbi, et al.
Published: (2026)
by: Choi, Eunbi, et al.
Published: (2026)
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
by: Lee, Seohyeong, et al.
Published: (2025)
by: Lee, Seohyeong, et al.
Published: (2025)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
by: Lee, Sangyub, et al.
Published: (2026)
by: Lee, Sangyub, et al.
Published: (2026)
PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation
by: Park, Junho, et al.
Published: (2026)
by: Park, Junho, et al.
Published: (2026)
RAISE: Enhancing Scientific Reasoning in LLMs via Step-by-Step Retrieval
by: Oh, Minhae, et al.
Published: (2025)
by: Oh, Minhae, et al.
Published: (2025)
Exploring Iterative Controllable Summarization with Large Language Models
by: Ryu, Sangwon, et al.
Published: (2024)
by: Ryu, Sangwon, et al.
Published: (2024)
How Language Directions Align with Token Geometry in Multilingual LLMs
by: Kim, JaeSeong, et al.
Published: (2025)
by: Kim, JaeSeong, et al.
Published: (2025)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
by: Kim, Kyuhee, et al.
Published: (2025)
by: Kim, Kyuhee, et al.
Published: (2025)
Similar Items
-
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025) -
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
by: Hahm, Sungeun, et al.
Published: (2025) -
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
by: Chae, Kyubyung, et al.
Published: (2024) -
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025) -
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
by: Jang, Jinkwan, et al.
Published: (2026)