ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lee, Michael S., Maurya, Yash, Rein, Drew, Herring, Bert, Nguyen, Jonathan, Song, Kyungho, Sehwag, Udari Madhushani, Cho, Jiyeon, Deshpande, Kaustubh, Jang, Yeongkyun, Joo, Jiyeon, Choi, Minn Seok, Fuelle, Evi, Knight, Christina Q, Brandifino, Joseph, Fenkell, Max |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
par: Sehwag, Udari Madhushani, et autres
Publié: (2026)
par: Sehwag, Udari Madhushani, et autres
Publié: (2026)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
par: Campbell, David, et autres
Publié: (2026)
par: Campbell, David, et autres
Publié: (2026)
In-Context Learning with Topological Information for Knowledge Graph Completion
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
par: Knight, Christina Q., et autres
Publié: (2025)
par: Knight, Christina Q., et autres
Publié: (2025)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
par: Pathmanathan, Pankayaraj, et autres
Publié: (2024)
par: Pathmanathan, Pankayaraj, et autres
Publié: (2024)
Can LLMs be Scammed? A Baseline Measurement Study
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
par: Sehwag, Udari Madhushani, et autres
Publié: (2024)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
par: Sehwag, Udari Madhushani, et autres
Publié: (2025)
par: Sehwag, Udari Madhushani, et autres
Publié: (2025)
LHAW: Controllable Underspecification for Long-Horizon Tasks
par: Pu, George, et autres
Publié: (2026)
par: Pu, George, et autres
Publié: (2026)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
par: Xu, Yuancheng, et autres
Publié: (2024)
par: Xu, Yuancheng, et autres
Publié: (2024)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
par: Cook, Thomas, et autres
Publié: (2025)
par: Cook, Thomas, et autres
Publié: (2025)
Influence of social pension on well‐being and health of the rural elderly: the case of South Korea
par: Jiyeon An, et autres
Publié: (2024)
par: Jiyeon An, et autres
Publié: (2024)
Environmental, Social, and Governance ( ESG ) Research: A Systematic Review of Recent Trends (2020–2024)
par: Jiyeon Kim, et autres
Publié: (2025)
par: Jiyeon Kim, et autres
Publié: (2025)
A Multi-Level Visual Analytics Approach to Artist-Era Alignment in Popular Music
par: Bae, Jiyeon, et autres
Publié: (2026)
par: Bae, Jiyeon, et autres
Publié: (2026)
Attachment Between Nurses and Patients in Hospital Settings: Concept Analysis Using Walker and Avant's Method
par: Jiyeon Lee, et autres
Publié: (2026)
par: Jiyeon Lee, et autres
Publié: (2026)
Assessing differential impacts of a trade agreement using a quantile regression approach
par: Jiyeon Kim, et autres
Publié: (2025)
par: Jiyeon Kim, et autres
Publié: (2025)
ESG Performance Evolution in Retail: A Systematic Review and Meta‐Analysis
par: Jiyeon Kim, et autres
Publié: (2025)
par: Jiyeon Kim, et autres
Publié: (2025)
AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
par: Karthikeyan, Harish, et autres
Publié: (2025)
par: Karthikeyan, Harish, et autres
Publié: (2025)
Global Comparison of Codes of Ethics for Nurses: A Mixed‐Method Collective Case Study Differentiating Aspirational and Mandatory Ethics
par: Min Ji Kim, et autres
Publié: (2024)
par: Min Ji Kim, et autres
Publié: (2024)
Metric Design != Metric Behavior: Improving Metric Selection for the Unbiased Evaluation of Dimensionality Reduction
par: Bae, Jiyeon, et autres
Publié: (2025)
par: Bae, Jiyeon, et autres
Publié: (2025)
On the topology of real Lagrangians in toric symplectic manifolds
par: Brendel, Joé, et autres
Publié: (2019)
par: Brendel, Joé, et autres
Publié: (2019)
Comparative Analysis of Deep Learning Techniques for Load Forecasting in Power Systems Using Single‐Layer and Hybrid Models
par: Jiyeon Jang, et autres
Publié: (2024)
par: Jiyeon Jang, et autres
Publié: (2024)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
par: Chakraborty, Souradip, et autres
Publié: (2025)
par: Chakraborty, Souradip, et autres
Publié: (2025)
Teaching Molecular Dynamics to a Non-Autoregressive Ionic Transport Predictor
par: Kim, Jiyeon, et autres
Publié: (2026)
par: Kim, Jiyeon, et autres
Publié: (2026)
Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding
par: Kim, Suyoung, et autres
Publié: (2024)
par: Kim, Suyoung, et autres
Publié: (2024)
Semi-Supervised Neural Super-Resolution for Mesh-Based Simulations
par: Kim, Jiyeon, et autres
Publié: (2026)
par: Kim, Jiyeon, et autres
Publié: (2026)
NMR spectroscopic investigations of transition metal complexes in organometallic and bioinorganic chemistry
par: Jeongcheol Shin, et autres
Publié: (2024)
par: Jeongcheol Shin, et autres
Publié: (2024)
Gender Disparities in Interventional Pain Medicine: Representation, Leadership, and Compensation
par: Marissa Catalanotto, et autres
Publié: (2026)
par: Marissa Catalanotto, et autres
Publié: (2026)
Towards Automatic Evaluation for Image Transcreation
par: Khanuja, Simran, et autres
Publié: (2024)
par: Khanuja, Simran, et autres
Publié: (2024)
The Formation of Japan-ROK Security Relations
par: Choi, Kyungwon
Publié: (2025)
par: Choi, Kyungwon
Publié: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
par: Chiu, Yu Ying, et autres
Publié: (2025)
par: Chiu, Yu Ying, et autres
Publié: (2025)
First-principles study on Small Polaron and Li diffusion in layered LiCoO2
par: Ahn, Seryung, et autres
Publié: (2022)
par: Ahn, Seryung, et autres
Publié: (2022)
Diverse Rare Sample Generation with Pretrained GANs
par: Lee, Subeen, et autres
Publié: (2024)
par: Lee, Subeen, et autres
Publié: (2024)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
par: Manglik, Akshay, et autres
Publié: (2026)
par: Manglik, Akshay, et autres
Publié: (2026)
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
par: Xiao, Yuchen, et autres
Publié: (2023)
par: Xiao, Yuchen, et autres
Publié: (2023)
OAM spatial demultiplexing by diffraction-based noiseless mode conversion with axicon
par: Kim, Junsu, et autres
Publié: (2025)
par: Kim, Junsu, et autres
Publié: (2025)
CREward: A Type-Specific Creativity Reward Model
par: Han, Jiyeon, et autres
Publié: (2025)
par: Han, Jiyeon, et autres
Publié: (2025)
PyGRF: An improved Python Geographical Random Forest model and case studies in public health and natural disasters
par: Sun, Kai, et autres
Publié: (2024)
par: Sun, Kai, et autres
Publié: (2024)
Synergistic Effect of Prolonged Oxygenation and Reactive Oxygen Species Scavenging on Diabetic Wound Healing Using an Injectable Thermoresponsive Hydrogel
par: Jiyeon Lee, et autres
Publié: (2025)
par: Jiyeon Lee, et autres
Publié: (2025)
Randomised Controlled Trial: Influence of Subconjunctival Anaesthesia Duration on Pain Perception During Intravitreal Injections: Response
par: Jiyeon Kim, et autres
Publié: (2026)
par: Jiyeon Kim, et autres
Publié: (2026)
PyGRF: An Improved Python Geographical Random Forest Model and Case Studies in Public Health and Natural Disasters
par: Kai Sun, et autres
Publié: (2024)
par: Kai Sun, et autres
Publié: (2024)
Documents similaires
-
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
par: Sehwag, Udari Madhushani, et autres
Publié: (2026) -
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
par: Campbell, David, et autres
Publié: (2026) -
In-Context Learning with Topological Information for Knowledge Graph Completion
par: Sehwag, Udari Madhushani, et autres
Publié: (2024) -
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
par: Knight, Christina Q., et autres
Publié: (2025) -
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
par: Pathmanathan, Pankayaraj, et autres
Publié: (2024)