COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Dasol, Lee, DongGeon, Kartono, Brigitta Jesica, Berndt, Helena, Kwon, Taeyoun, Jang, Joonwon, Park, Haon, Yu, Hwanjo, Kahng, Minsuk |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
von: Dingeto, Hiskias, et al.
Veröffentlicht: (2025)
von: Dingeto, Hiskias, et al.
Veröffentlicht: (2025)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark
von: Jeong, Jihae, et al.
Veröffentlicht: (2025)
von: Jeong, Jihae, et al.
Veröffentlicht: (2025)
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
Rectifying Demonstration Shortcut in In-Context Learning
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
von: Kim, Kihyun, et al.
Veröffentlicht: (2026)
von: Kim, Kihyun, et al.
Veröffentlicht: (2026)
Verbosity-Aware Rationale Reduction: Effective Reduction of Redundant Rationale via Principled Criteria
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2025)
Typed-RAG: Type-Aware Decomposition of Non-Factoid Questions for Retrieval-Augmented Generation
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
ToDi: Token-wise Distillation via Fine-Grained Divergence Control
von: Jung, Seongryong, et al.
Veröffentlicht: (2025)
von: Jung, Seongryong, et al.
Veröffentlicht: (2025)
Understanding the Dataset Practitioners Behind Large Language Model Development
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
VLSlice: Interactive Vision-and-Language Slice Discovery
von: Slyman, Eric, et al.
Veröffentlicht: (2023)
von: Slyman, Eric, et al.
Veröffentlicht: (2023)
Eliciting Instruction-tuned Code Language Models' Capabilities to Utilize Auxiliary Function for Code Generation
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2024)
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2024)
Theme-Explanation Structure for Table Summarization using Large Language Models: A Case Study on Korean Tabular Data
von: Kwack, TaeYoon, et al.
Veröffentlicht: (2025)
von: Kwack, TaeYoon, et al.
Veröffentlicht: (2025)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
von: Jung, Minji, et al.
Veröffentlicht: (2026)
von: Jung, Minji, et al.
Veröffentlicht: (2026)
Automatic Histograms: Leveraging Language Models for Text Dataset Exploration
von: Reif, Emily, et al.
Veröffentlicht: (2024)
von: Reif, Emily, et al.
Veröffentlicht: (2024)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
von: Lee, Kyungjae, et al.
Veröffentlicht: (2024)
von: Lee, Kyungjae, et al.
Veröffentlicht: (2024)
No Language Data Left Behind: A Comparative Study of CJK Language Datasets in the Hugging Face Ecosystem
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
From Perception to Decision: Assessing the Role of Chart Types Affordances in High-Level Decision Tasks
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
Mathematical Foundations of Polyphonic Music Generation via Structural Inductive Bias
von: Seo, Joonwon
Veröffentlicht: (2026)
von: Seo, Joonwon
Veröffentlicht: (2026)
SPARTA: Advancing Sparse Attention in Spiking Neural Networks via Spike-Timing-Based Prioritization
von: Jang, Minsuk, et al.
Veröffentlicht: (2025)
von: Jang, Minsuk, et al.
Veröffentlicht: (2025)
Mine-JEPA: In-Domain Self-Supervised Learning for Mine-Like Object Classification in Side-Scan Sonar
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2026)
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2026)
Heuristic Algorithm-based Action Masking Reinforcement Learning (HAAM-RL) with Ensemble Inference Method
von: Choi, Kyuwon, et al.
Veröffentlicht: (2024)
von: Choi, Kyuwon, et al.
Veröffentlicht: (2024)
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
von: Panpatil, Siddhant, et al.
Veröffentlicht: (2025)
von: Panpatil, Siddhant, et al.
Veröffentlicht: (2025)
El impacto de la variable de género en la migración Honduras-México: el caso de las Hondureñas en Frontera Comalapa
von: Nicanor Madueño Haon
Veröffentlicht: (2010)
von: Nicanor Madueño Haon
Veröffentlicht: (2010)
Distribution-Level Feature Distancing for Machine Unlearning: Towards a Better Trade-off Between Model Utility and Forgetting
von: Choi, Dasol, et al.
Veröffentlicht: (2024)
von: Choi, Dasol, et al.
Veröffentlicht: (2024)
Hybrid Synchronization with Continuous Varying Exponent in Decentralized Power Grid
von: Park, Jinha, et al.
Veröffentlicht: (2024)
von: Park, Jinha, et al.
Veröffentlicht: (2024)
Green functions of mixed boundary value problems for stationary Stokes systems in two dimensions
von: Choi, Jongkeun, et al.
Veröffentlicht: (2024)
von: Choi, Jongkeun, et al.
Veröffentlicht: (2024)
BrainDecoder: Style-Based Visual Decoding of EEG Signals
von: Choi, Minsuk, et al.
Veröffentlicht: (2024)
von: Choi, Minsuk, et al.
Veröffentlicht: (2024)
KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
von: Jang, Seongbo, et al.
Veröffentlicht: (2024)
von: Jang, Seongbo, et al.
Veröffentlicht: (2024)
Interactive Prompt Debugging with Sequence Salience
von: Tenney, Ian, et al.
Veröffentlicht: (2024)
von: Tenney, Ian, et al.
Veröffentlicht: (2024)
Conical Kähler-Einstein metrics on K-unstable del Pezzo surfaces
von: Jeong, Dasol, et al.
Veröffentlicht: (2025)
von: Jeong, Dasol, et al.
Veröffentlicht: (2025)
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
von: Ha, Junwoo, et al.
Veröffentlicht: (2025)
von: Ha, Junwoo, et al.
Veröffentlicht: (2025)
SELFI: Selective Fusion of Identity for Generalizable Deepfake Detection
von: Kim, Younghun, et al.
Veröffentlicht: (2025)
von: Kim, Younghun, et al.
Veröffentlicht: (2025)
Density Matrix RNN (DM-RNN): A Quantum Information Theoretic Framework for Modeling Musical Context and Polyphony
von: Seo, Joonwon, et al.
Veröffentlicht: (2026)
von: Seo, Joonwon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
von: Lee, DongGeon, et al.
Veröffentlicht: (2025) -
When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
von: Dingeto, Hiskias, et al.
Veröffentlicht: (2025) -
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
von: Lee, DongGeon, et al.
Veröffentlicht: (2025) -
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
von: Choi, Dasol, et al.
Veröffentlicht: (2025) -
Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark
von: Jeong, Jihae, et al.
Veröffentlicht: (2025)