SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Qin, Wang, Fei, Xiao, Chaowei, Chen, Muhao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
di: Liu, Qin, et al.
Pubblicazione: (2023)
di: Liu, Qin, et al.
Pubblicazione: (2023)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
di: Liu, Xiaogeng, et al.
Pubblicazione: (2023)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2023)
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
di: Liu, Qin, et al.
Pubblicazione: (2025)
di: Liu, Qin, et al.
Pubblicazione: (2025)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
di: Wang, Jiongxiao, et al.
Pubblicazione: (2023)
di: Wang, Jiongxiao, et al.
Pubblicazione: (2023)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
di: Liu, Qin, et al.
Pubblicazione: (2024)
di: Liu, Qin, et al.
Pubblicazione: (2024)
Instructional Fingerprinting of Large Language Models
di: Xu, Jiashu, et al.
Pubblicazione: (2024)
di: Xu, Jiashu, et al.
Pubblicazione: (2024)
Two Heads Are Better Than One: Integrating Knowledge from Knowledge Graphs and Large Language Models for Entity Alignment
di: Yang, Linyao, et al.
Pubblicazione: (2024)
di: Yang, Linyao, et al.
Pubblicazione: (2024)
DeepEdit: Knowledge Editing as Decoding with Constraints
di: Wang, Yiwei, et al.
Pubblicazione: (2024)
di: Wang, Yiwei, et al.
Pubblicazione: (2024)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
di: Wang, Cheng, et al.
Pubblicazione: (2026)
di: Wang, Cheng, et al.
Pubblicazione: (2026)
CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
di: Jiang, Yi, et al.
Pubblicazione: (2025)
di: Jiang, Yi, et al.
Pubblicazione: (2025)
KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning
di: Yu, Peng, et al.
Pubblicazione: (2024)
di: Yu, Peng, et al.
Pubblicazione: (2024)
KG-BiLM: Knowledge Graph Embedding via Bidirectional Language Models
di: Chen, Zirui, et al.
Pubblicazione: (2025)
di: Chen, Zirui, et al.
Pubblicazione: (2025)
ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models
di: Guo, Wenbin, et al.
Pubblicazione: (2025)
di: Guo, Wenbin, et al.
Pubblicazione: (2025)
HumanLM: Simulating Users with State Alignment Beats Response Imitation
di: Wu, Shirley, et al.
Pubblicazione: (2026)
di: Wu, Shirley, et al.
Pubblicazione: (2026)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
di: Tong, Terry, et al.
Pubblicazione: (2025)
di: Tong, Terry, et al.
Pubblicazione: (2025)
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
di: Fang, Tianqing, et al.
Pubblicazione: (2023)
di: Fang, Tianqing, et al.
Pubblicazione: (2023)
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
di: Sun, Zetian, et al.
Pubblicazione: (2025)
di: Sun, Zetian, et al.
Pubblicazione: (2025)
Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation
di: Zhou, Jiang, et al.
Pubblicazione: (2026)
di: Zhou, Jiang, et al.
Pubblicazione: (2026)
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
di: Wang, Xinyan, et al.
Pubblicazione: (2026)
di: Wang, Xinyan, et al.
Pubblicazione: (2026)
FamiCom: Further Demystifying Prompts for Language Models with Task-Agnostic Performance Estimation
di: Li, Bangzheng, et al.
Pubblicazione: (2024)
di: Li, Bangzheng, et al.
Pubblicazione: (2024)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
di: Li, Bangzheng, et al.
Pubblicazione: (2023)
di: Li, Bangzheng, et al.
Pubblicazione: (2023)
Xmodel-LM Technical Report
di: Wang, Yichuan, et al.
Pubblicazione: (2024)
di: Wang, Yichuan, et al.
Pubblicazione: (2024)
From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering
di: Li, Lei, et al.
Pubblicazione: (2025)
di: Li, Lei, et al.
Pubblicazione: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
CrossIn: An Efficient Instruction Tuning Approach for Cross-Lingual Knowledge Alignment
di: Lin, Geyu, et al.
Pubblicazione: (2024)
di: Lin, Geyu, et al.
Pubblicazione: (2024)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
di: Askari, Hadi, et al.
Pubblicazione: (2025)
di: Askari, Hadi, et al.
Pubblicazione: (2025)
Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement
di: Lin, Chenyu, et al.
Pubblicazione: (2025)
di: Lin, Chenyu, et al.
Pubblicazione: (2025)
DERA: Dense Entity Retrieval for Entity Alignment in Knowledge Graphs
di: Wang, Zhichun, et al.
Pubblicazione: (2024)
di: Wang, Zhichun, et al.
Pubblicazione: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
InternLM2 Technical Report
di: Cai, Zheng, et al.
Pubblicazione: (2024)
di: Cai, Zheng, et al.
Pubblicazione: (2024)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
Extracting and Understanding the Superficial Knowledge in Alignment
di: Chen, Runjin, et al.
Pubblicazione: (2025)
di: Chen, Runjin, et al.
Pubblicazione: (2025)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
di: Tong, Terry, et al.
Pubblicazione: (2024)
di: Tong, Terry, et al.
Pubblicazione: (2024)
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
di: Pham, Hung Manh, et al.
Pubblicazione: (2026)
di: Pham, Hung Manh, et al.
Pubblicazione: (2026)
Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
di: Liu, Genglin, et al.
Pubblicazione: (2023)
di: Liu, Genglin, et al.
Pubblicazione: (2023)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
di: Liu, Qin, et al.
Pubblicazione: (2025)
di: Liu, Qin, et al.
Pubblicazione: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
di: Huang, James Y., et al.
Pubblicazione: (2025)
di: Huang, James Y., et al.
Pubblicazione: (2025)
Dynamic Parametric Retrieval Augmented Generation for Test-time Knowledge Enhancement
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
di: Liu, Qin, et al.
Pubblicazione: (2023) -
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
di: Liu, Xiaogeng, et al.
Pubblicazione: (2023) -
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
di: Liu, Qin, et al.
Pubblicazione: (2025) -
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
di: Wang, Jiongxiao, et al.
Pubblicazione: (2023) -
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
di: Xu, Jiashu, et al.
Pubblicazione: (2023)