Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Rui, Quan, Yihao, Shi, Zeru, Wang, Zhenting, Li, Yanshu, Tang, Ruixiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Auto-Prompt Generation is Not Robust: Prompt Optimization Driven by Pseudo Gradient
por: Shi, Zeru, et al.
Publicado: (2024)
por: Shi, Zeru, et al.
Publicado: (2024)
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
por: Li, Jiaqian, et al.
Publicado: (2026)
por: Li, Jiaqian, et al.
Publicado: (2026)
A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models
por: Shi, Zeru, et al.
Publicado: (2026)
por: Shi, Zeru, et al.
Publicado: (2026)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
por: Wu, Rui, et al.
Publicado: (2026)
por: Wu, Rui, et al.
Publicado: (2026)
Meaningless Tokens, Meaningful Gains: How Activation Shifts Enhance LLM Reasoning
por: Shi, Zeru, et al.
Publicado: (2025)
por: Shi, Zeru, et al.
Publicado: (2025)
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering
por: Wang, Yanshu, et al.
Publicado: (2024)
por: Wang, Yanshu, et al.
Publicado: (2024)
Reinforcing Consistency in Video MLLMs with Structured Rewards
por: Quan, Yihao, et al.
Publicado: (2026)
por: Quan, Yihao, et al.
Publicado: (2026)
Script Gap: Evaluating LLM Triage on Indian Languages in Native vs Romanized Scripts in a Real World Setting
por: Khullar, Manurag, et al.
Publicado: (2025)
por: Khullar, Manurag, et al.
Publicado: (2025)
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
por: Chen, Yuefei, et al.
Publicado: (2026)
por: Chen, Yuefei, et al.
Publicado: (2026)
Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
por: Wang, Zhenting, et al.
Publicado: (2026)
por: Wang, Zhenting, et al.
Publicado: (2026)
Token-Budget-Aware LLM Reasoning
por: Han, Tingxu, et al.
Publicado: (2024)
por: Han, Tingxu, et al.
Publicado: (2024)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
por: Wang, Yanshu, et al.
Publicado: (2024)
por: Wang, Yanshu, et al.
Publicado: (2024)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
por: Zhou, Yuqing, et al.
Publicado: (2024)
por: Zhou, Yuqing, et al.
Publicado: (2024)
DBR: Divergence-Based Regularization for Debiasing Natural Language Understanding Models
por: Li, Zihao, et al.
Publicado: (2025)
por: Li, Zihao, et al.
Publicado: (2025)
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
por: Wang, Zhenting, et al.
Publicado: (2025)
por: Wang, Zhenting, et al.
Publicado: (2025)
Prompting with Phonemes: Enhancing LLMs' Multilinguality for Non-Latin Script Languages
por: Nguyen, Hoang H, et al.
Publicado: (2024)
por: Nguyen, Hoang H, et al.
Publicado: (2024)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
por: Chen, Yen-Shan, et al.
Publicado: (2026)
por: Chen, Yen-Shan, et al.
Publicado: (2026)
Secure LLM Fine-Tuning via Safety-Aware Probing
por: Wu, Chengcan, et al.
Publicado: (2025)
por: Wu, Chengcan, et al.
Publicado: (2025)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
por: Wei, Zeming, et al.
Publicado: (2026)
por: Wei, Zeming, et al.
Publicado: (2026)
LongSafety: Enhance Safety for Long-Context LLMs
por: Huang, Mianqiu, et al.
Publicado: (2024)
por: Huang, Mianqiu, et al.
Publicado: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
por: Zhou, Yujun, et al.
Publicado: (2024)
por: Zhou, Yujun, et al.
Publicado: (2024)
Advancing LLM Safe Alignment with Safety Representation Ranking
por: Du, Tianqi, et al.
Publicado: (2025)
por: Du, Tianqi, et al.
Publicado: (2025)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
por: Wang, Guoli, et al.
Publicado: (2026)
por: Wang, Guoli, et al.
Publicado: (2026)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
por: Yang, Xikang, et al.
Publicado: (2025)
por: Yang, Xikang, et al.
Publicado: (2025)
Bypassing Safety Guardrails in LLMs Using Humor
por: Cisneros-Velarde, Pedro
Publicado: (2025)
por: Cisneros-Velarde, Pedro
Publicado: (2025)
Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety
por: Shen, Qian, et al.
Publicado: (2026)
por: Shen, Qian, et al.
Publicado: (2026)
Continuous Approximations for Improving Quantization Aware Training of LLMs
por: Li, He, et al.
Publicado: (2024)
por: Li, He, et al.
Publicado: (2024)
Towards Generalizable Implicit In-Context Learning with Attention Routing
por: Li, Jiaqian, et al.
Publicado: (2025)
por: Li, Jiaqian, et al.
Publicado: (2025)
Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety
por: Mao, Junyu, et al.
Publicado: (2025)
por: Mao, Junyu, et al.
Publicado: (2025)
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs
por: Liu, Zhihao, et al.
Publicado: (2024)
por: Liu, Zhihao, et al.
Publicado: (2024)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
por: Hu, Zhiyuan, et al.
Publicado: (2026)
por: Hu, Zhiyuan, et al.
Publicado: (2026)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
por: Qiu, Quantong, et al.
Publicado: (2026)
por: Qiu, Quantong, et al.
Publicado: (2026)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
por: Wei, Rongzhe, et al.
Publicado: (2025)
por: Wei, Rongzhe, et al.
Publicado: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
por: Lu, Wenquan, et al.
Publicado: (2025)
por: Lu, Wenquan, et al.
Publicado: (2025)
Multilingual Language Models Encode Script Over Linguistic Structure
por: Verma, Aastha A K, et al.
Publicado: (2026)
por: Verma, Aastha A K, et al.
Publicado: (2026)
Text-Based Detection of On-Hold Scripts in Contact Center Calls
por: Galimzianov, Dmitrii, et al.
Publicado: (2024)
por: Galimzianov, Dmitrii, et al.
Publicado: (2024)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
por: Wang, Kai, et al.
Publicado: (2025)
por: Wang, Kai, et al.
Publicado: (2025)
Machine-Assisted Script Curation
por: Ciosici, Manuel R., et al.
Publicado: (2021)
por: Ciosici, Manuel R., et al.
Publicado: (2021)
Ejemplares similares
-
Auto-Prompt Generation is Not Robust: Prompt Optimization Driven by Pseudo Gradient
por: Shi, Zeru, et al.
Publicado: (2024) -
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
por: Li, Jiaqian, et al.
Publicado: (2026) -
A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models
por: Shi, Zeru, et al.
Publicado: (2026) -
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
por: Wu, Rui, et al.
Publicado: (2026) -
Meaningless Tokens, Meaningful Gains: How Activation Shifts Enhance LLM Reasoning
por: Shi, Zeru, et al.
Publicado: (2025)