HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
Fuente:
arXiv
Saved in:
| Main Authors: | Mei, Lingrui, Liu, Shenghua, Wang, Yiwei, Bi, Baolong, Yuan, Ruibin, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLANG: New Concept Comprehension of Large Language Models
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
LPNL: Scalable Link Prediction with Large Language Models
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
by: Yao, Jiayu, et al.
Published: (2025)
by: Yao, Jiayu, et al.
Published: (2025)
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
Decoding by Contrasting Knowledge: Enhancing LLMs' Confidence on Edited Facts
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Adaptive Token Biaser: Knowledge Editing via Biasing Key Entities
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
StruEdit: Structured Outputs Enable the Fast and Accurate Knowledge Editing for Large Language Models
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
a1: Steep Test-time Scaling Law via Environment Augmented Generation
by: Mei, Lingrui, et al.
Published: (2025)
by: Mei, Lingrui, et al.
Published: (2025)
Not in Sync: Unveiling Temporal Bias in Audio Chat Models
by: Yao, Jiayu, et al.
Published: (2025)
by: Yao, Jiayu, et al.
Published: (2025)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
by: Ge, Yuyao, et al.
Published: (2026)
by: Ge, Yuyao, et al.
Published: (2026)
Gated Differentiable Working Memory for Long-Context Language Modeling
by: Mei, Lingrui, et al.
Published: (2026)
by: Mei, Lingrui, et al.
Published: (2026)
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Rethinking All Evidence: Enhancing Trustworthy Retrieval-Augmented Generation via Conflict-Driven Summarization
by: Chen, Juan, et al.
Published: (2025)
by: Chen, Juan, et al.
Published: (2025)
ALiiCE: Evaluating Positional Fine-grained Citation Generation
by: Xu, Yilong, et al.
Published: (2024)
by: Xu, Yilong, et al.
Published: (2024)
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Can Graph Descriptive Order Affect Solving Graph Problems with LLMs?
by: Ge, Yuyao, et al.
Published: (2024)
by: Ge, Yuyao, et al.
Published: (2024)
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
A Survey of Context Engineering for Large Language Models
by: Mei, Lingrui, et al.
Published: (2025)
by: Mei, Lingrui, et al.
Published: (2025)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
You Know What I'm Saying: Jailbreak Attack via Implicit Reference
by: Wu, Tianyu, et al.
Published: (2024)
by: Wu, Tianyu, et al.
Published: (2024)
LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?
by: Cheng, Xueqi, et al.
Published: (2026)
by: Cheng, Xueqi, et al.
Published: (2026)
Context-DPO: Aligning Language Models for Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models
by: Xu, Yilong, et al.
Published: (2025)
by: Xu, Yilong, et al.
Published: (2025)
RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering
by: Yuan, Bo, et al.
Published: (2026)
by: Yuan, Bo, et al.
Published: (2026)
Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generation
by: Gao, Linfeng, et al.
Published: (2025)
by: Gao, Linfeng, et al.
Published: (2025)
How to Make Large Language Models Generate 100% Valid Molecules?
by: Tao, Wen, et al.
Published: (2025)
by: Tao, Wen, et al.
Published: (2025)
CP-Router: An Uncertainty-Aware Router Between LLM and LRM
by: Su, Jiayuan, et al.
Published: (2025)
by: Su, Jiayuan, et al.
Published: (2025)
MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning
by: Cui, Wanqing, et al.
Published: (2024)
by: Cui, Wanqing, et al.
Published: (2024)
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation
by: Wang, Xinglin, et al.
Published: (2024)
by: Wang, Xinglin, et al.
Published: (2024)
Large Language Models as Computable Approximations to Solomonoff Induction
by: Wan, Jun, et al.
Published: (2025)
by: Wan, Jun, et al.
Published: (2025)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
by: Ni, Shiyu, et al.
Published: (2024)
by: Ni, Shiyu, et al.
Published: (2024)
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
WebRouter: Query-specific Router via Variational Information Bottleneck for Cost-sensitive Web Agent
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
by: Jin, Weiyang, et al.
Published: (2025)
by: Jin, Weiyang, et al.
Published: (2025)
Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text Detection
by: Pu, Xiao, et al.
Published: (2026)
by: Pu, Xiao, et al.
Published: (2026)
ExpGuard: LLM Content Moderation in Specialized Domains
by: Choi, Minseok, et al.
Published: (2026)
by: Choi, Minseok, et al.
Published: (2026)
Similar Items
-
SLANG: New Concept Comprehension of Large Language Models
by: Mei, Lingrui, et al.
Published: (2024) -
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
by: Bi, Baolong, et al.
Published: (2025) -
LPNL: Scalable Link Prediction with Large Language Models
by: Bi, Baolong, et al.
Published: (2024) -
Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
by: Yao, Jiayu, et al.
Published: (2025) -
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
by: Mei, Lingrui, et al.
Published: (2024)