Saved in:
| Main Authors: | Belo, Ruben, Guimaraes, Marta, Soares, Claudia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.12672 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025)
by: Mendu, Sai Krishna, et al.
Published: (2025)
Probability of Collision of satellites and space debris for short-term encounters: Rederivation and fast-to-compute upper and lower bounds
by: Ferreira, Ricardo, et al.
Published: (2023)
by: Ferreira, Ricardo, et al.
Published: (2023)
Generalizing Trilateration: Approximate Maximum Likelihood Estimator for Initial Orbit Determination in Low-Earth Orbit
by: Ferreira, Ricardo, et al.
Published: (2024)
by: Ferreira, Ricardo, et al.
Published: (2024)
Harmful Visual Content Manipulation Matters in Misinformation Detection Under Multimedia Scenarios
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
If Concept Bottlenecks are the Question, are Foundation Models the Answer?
by: Debole, Nicola, et al.
Published: (2025)
by: Debole, Nicola, et al.
Published: (2025)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
LookAhead Tuning: Safer Language Models via Partial Answer Previews
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Towards Understanding and Avoiding Limitations of Convolutions on Graphs
by: Roth, Andreas
Published: (2026)
by: Roth, Andreas
Published: (2026)
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
by: Yamabe, Shojiro, et al.
Published: (2025)
by: Yamabe, Shojiro, et al.
Published: (2025)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Safety Representations for Safer Policy Learning
by: Mani, Kaustubh, et al.
Published: (2025)
by: Mani, Kaustubh, et al.
Published: (2025)
Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Two Calm Ends and the Wild Middle: A Geometric Picture of Memorization in Diffusion Models
by: Dodson, Nick, et al.
Published: (2026)
by: Dodson, Nick, et al.
Published: (2026)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
by: K, Swaminathan S, et al.
Published: (2026)
by: K, Swaminathan S, et al.
Published: (2026)
Generalized Distributional Alignment Games for Unbiased Answer-Level Fine-Tuning
by: Mohri, Mehryar, et al.
Published: (2026)
by: Mohri, Mehryar, et al.
Published: (2026)
Concept Alignment
by: Rane, Sunayana, et al.
Published: (2024)
by: Rane, Sunayana, et al.
Published: (2024)
Uncertainty-aware Latent Safety Filters for Avoiding Out-of-Distribution Failures
by: Seo, Junwon, et al.
Published: (2025)
by: Seo, Junwon, et al.
Published: (2025)
Toward Real-World IoT Security: Concept Drift-Resilient IoT Botnet Detection via Latent Space Representation Learning and Alignment
by: Wasswa, Hassan, et al.
Published: (2025)
by: Wasswa, Hassan, et al.
Published: (2025)
Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach
by: Chai, Yidong, et al.
Published: (2025)
by: Chai, Yidong, et al.
Published: (2025)
Nonparametric Identification of Latent Concepts
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
Distributional Alignment Games for Answer-Level Fine-Tuning
by: Mohri, Mehryar, et al.
Published: (2026)
by: Mohri, Mehryar, et al.
Published: (2026)
Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
by: Zarlenga, Mateo Espinosa, et al.
Published: (2025)
by: Zarlenga, Mateo Espinosa, et al.
Published: (2025)
Consensus Sampling for Safer Generative AI
by: Kalai, Adam Tauman, et al.
Published: (2025)
by: Kalai, Adam Tauman, et al.
Published: (2025)
Model Stitching by Functional Latent Alignment
by: Athanasiadis, Ioannis, et al.
Published: (2025)
by: Athanasiadis, Ioannis, et al.
Published: (2025)
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
by: Cao, Chentao, et al.
Published: (2025)
by: Cao, Chentao, et al.
Published: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
The Cost of Avoiding Backpropagation
by: Panchal, Kunjal, et al.
Published: (2025)
by: Panchal, Kunjal, et al.
Published: (2025)
Latent Space Translation via Semantic Alignment
by: Maiorca, Valentino, et al.
Published: (2023)
by: Maiorca, Valentino, et al.
Published: (2023)
Safer Policy Compliance with Dynamic Epistemic Fallback
by: Imperial, Joseph Marvin, et al.
Published: (2026)
by: Imperial, Joseph Marvin, et al.
Published: (2026)
Learning Discrete Concepts in Latent Hierarchical Models
by: Kong, Lingjing, et al.
Published: (2024)
by: Kong, Lingjing, et al.
Published: (2024)
Simple Mechanisms for Representing, Indexing and Manipulating Concepts
by: Li, Yuanzhi, et al.
Published: (2023)
by: Li, Yuanzhi, et al.
Published: (2023)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
by: Aakanksha, et al.
Published: (2024)
by: Aakanksha, et al.
Published: (2024)
Exchangeable Sequence Models Quantify Uncertainty Over Latent Concepts
by: Ye, Naimeng, et al.
Published: (2024)
by: Ye, Naimeng, et al.
Published: (2024)
Latent Adaptive Planner for Dynamic Manipulation
by: Noh, Donghun, et al.
Published: (2025)
by: Noh, Donghun, et al.
Published: (2025)
Toward Faithful and Complete Answer Construction from a Single Document
by: Chen, Zhaoyang, et al.
Published: (2026)
by: Chen, Zhaoyang, et al.
Published: (2026)
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
by: Burnat, Florian A. D., et al.
Published: (2026)
by: Burnat, Florian A. D., et al.
Published: (2026)
Prototype-Grounded Concept Models for Verifiable Concept Alignment
by: Colamonaco, Stefano, et al.
Published: (2026)
by: Colamonaco, Stefano, et al.
Published: (2026)
Similar Items
-
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025) -
Probability of Collision of satellites and space debris for short-term encounters: Rederivation and fast-to-compute upper and lower bounds
by: Ferreira, Ricardo, et al.
Published: (2023) -
Generalizing Trilateration: Approximate Maximum Likelihood Estimator for Initial Orbit Determination in Low-Earth Orbit
by: Ferreira, Ricardo, et al.
Published: (2024) -
Harmful Visual Content Manipulation Matters in Misinformation Detection Under Multimedia Scenarios
by: Wang, Bing, et al.
Published: (2026) -
If Concept Bottlenecks are the Question, are Foundation Models the Answer?
by: Debole, Nicola, et al.
Published: (2025)