MetaSC: Test-Time Safety Specification Optimization for Language Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Gallego, Víctor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
von: Gallego, Víctor
Veröffentlicht: (2025)
von: Gallego, Víctor
Veröffentlicht: (2025)
Discovering Agentic Safety Specifications from 1-Bit Danger Signals
von: Gallego, Víctor
Veröffentlicht: (2026)
von: Gallego, Víctor
Veröffentlicht: (2026)
Configurable Safety Tuning of Language Models with Synthetic Preference Data
von: Gallego, Victor
Veröffentlicht: (2024)
von: Gallego, Victor
Veröffentlicht: (2024)
Configurable Preference Tuning with Rubric-Guided Synthetic Data
von: Gallego, Víctor
Veröffentlicht: (2025)
von: Gallego, Víctor
Veröffentlicht: (2025)
Merging Improves Self-Critique Against Jailbreak Attacks
von: Gallego, Victor
Veröffentlicht: (2024)
von: Gallego, Victor
Veröffentlicht: (2024)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
von: Banerjee, Somnath, et al.
Veröffentlicht: (2025)
von: Banerjee, Somnath, et al.
Veröffentlicht: (2025)
Test-Time Safety Alignment
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
Precision Knowledge Editing: Enhancing Safety in Large Language Models
von: Li, Xuying, et al.
Veröffentlicht: (2024)
von: Li, Xuying, et al.
Veröffentlicht: (2024)
Test-Time Fairness and Robustness in Large Language Models
von: Cotta, Leonardo, et al.
Veröffentlicht: (2024)
von: Cotta, Leonardo, et al.
Veröffentlicht: (2024)
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
von: Diao, Muxi, et al.
Veröffentlicht: (2024)
von: Diao, Muxi, et al.
Veröffentlicht: (2024)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
von: Röttger, Paul, et al.
Veröffentlicht: (2023)
von: Röttger, Paul, et al.
Veröffentlicht: (2023)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
SafeLLM: Domain-Specific Safety Monitoring for Large Language Models: A Case Study of Offshore Wind Maintenance
von: Walker, Connor, et al.
Veröffentlicht: (2024)
von: Walker, Connor, et al.
Veröffentlicht: (2024)
Training-Free Test-Time Contrastive Learning for Large Language Models
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2026)
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2026)
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
von: Miles, Roy, et al.
Veröffentlicht: (2026)
von: Miles, Roy, et al.
Veröffentlicht: (2026)
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)
von: Lu, Yida, et al.
Veröffentlicht: (2025)
Test-Time Learning for Large Language Models
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
All Languages Matter: On the Multilingual Safety of Large Language Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
RAFT: Adapting Language Model to Domain Specific RAG
von: Zhang, Tianjun, et al.
Veröffentlicht: (2024)
von: Zhang, Tianjun, et al.
Veröffentlicht: (2024)
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
von: Kour, George, et al.
Veröffentlicht: (2025)
von: Kour, George, et al.
Veröffentlicht: (2025)
MedAdapter: Efficient Test-Time Adaptation of Large Language Models towards Medical Reasoning
von: Shi, Wenqi, et al.
Veröffentlicht: (2024)
von: Shi, Wenqi, et al.
Veröffentlicht: (2024)
CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models
von: Zhang, Wenjing, et al.
Veröffentlicht: (2024)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2024)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties
von: Ong, Jasmine Chiat Ling, et al.
Veröffentlicht: (2024)
von: Ong, Jasmine Chiat Ling, et al.
Veröffentlicht: (2024)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Fine-Tuned Language Models for Domain-Specific Summarization and Tagging
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
SC-Phi2: A Fine-tuned Small Language Model for StarCraft II Macromanagement Tasks
von: Khan, Muhammad Junaid, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Junaid, et al.
Veröffentlicht: (2024)
Specialty-Specific Medical Language Model for Immune-Mediated Diseases
von: Kocaman, Veysel, et al.
Veröffentlicht: (2026)
von: Kocaman, Veysel, et al.
Veröffentlicht: (2026)
Safety-Aware Fine-Tuning of Large Language Models
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024)
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024)
Large Language Model Safety: A Holistic Survey
von: Shi, Dan, et al.
Veröffentlicht: (2024)
von: Shi, Dan, et al.
Veröffentlicht: (2024)
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
von: Tan, Yingshui, et al.
Veröffentlicht: (2024)
von: Tan, Yingshui, et al.
Veröffentlicht: (2024)
A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
SLOT: Sample-specific Language Model Optimization at Test-time
von: Hu, Yang, et al.
Veröffentlicht: (2025)
von: Hu, Yang, et al.
Veröffentlicht: (2025)
CRANE: Causal Relevance Analysis of Language-Specific Neurons in Multilingual Large Language Models
von: Le, Yifan, et al.
Veröffentlicht: (2026)
von: Le, Yifan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
von: Gallego, Víctor
Veröffentlicht: (2025) -
Discovering Agentic Safety Specifications from 1-Bit Danger Signals
von: Gallego, Víctor
Veröffentlicht: (2026) -
Configurable Safety Tuning of Language Models with Synthetic Preference Data
von: Gallego, Victor
Veröffentlicht: (2024) -
Configurable Preference Tuning with Rubric-Guided Synthetic Data
von: Gallego, Víctor
Veröffentlicht: (2025) -
Merging Improves Self-Critique Against Jailbreak Attacks
von: Gallego, Victor
Veröffentlicht: (2024)