Saved in:
| Main Authors: | Wu, Tianyu, Mei, Lingrui, Yuan, Ruibin, Li, Lujun, Xue, Wei, Guo, Yike |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.03857 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
by: Gaznepoglu, Ünal Ege, et al.
Published: (2025)
by: Gaznepoglu, Ünal Ege, et al.
Published: (2025)
HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026)
by: Lee, Joosung, et al.
Published: (2026)
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
by: Tan, Chenchen, et al.
Published: (2025)
by: Tan, Chenchen, et al.
Published: (2025)
Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles
by: Wang, Kuang, et al.
Published: (2025)
by: Wang, Kuang, et al.
Published: (2025)
Large Language Models Know What To Say But Not When To Speak
by: Umair, Muhammad, et al.
Published: (2024)
by: Umair, Muhammad, et al.
Published: (2024)
RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation
by: Chan, Chi-Min, et al.
Published: (2024)
by: Chan, Chi-Min, et al.
Published: (2024)
ImF: Implicit Fingerprint for Large Language Models
by: Wu, Jiaxuan, et al.
Published: (2025)
by: Wu, Jiaxuan, et al.
Published: (2025)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
It Is Not About What You Say, It Is About How You Say It: A Surprisingly Simple Approach for Improving Reading Comprehension
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
by: Chen, Xinxi, et al.
Published: (2024)
by: Chen, Xinxi, et al.
Published: (2024)
Teacher: Can You See What I'm Saying? A Research Experience with Deaf Learners
by: Olga Lucía Ávila Caica
Published: (2011)
by: Olga Lucía Ávila Caica
Published: (2011)
Say What You Mean: Natural Language Access Control with Large Language Models for Internet of Things
by: Cheng, Ye, et al.
Published: (2025)
by: Cheng, Ye, et al.
Published: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024)
by: Zhou, Yukai, et al.
Published: (2024)
Large Language Models Might Not Care What You Are Saying: Prompt Format Beats Descriptions
by: Tang, Chenming, et al.
Published: (2024)
by: Tang, Chenming, et al.
Published: (2024)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
by: Zhang, Hanning, et al.
Published: (2023)
by: Zhang, Hanning, et al.
Published: (2023)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
by: Zhou, Ziya, et al.
Published: (2024)
by: Zhou, Ziya, et al.
Published: (2024)
Say It Differently: Linguistic Styles as Jailbreak Vectors
by: Panda, Srikant, et al.
Published: (2025)
by: Panda, Srikant, et al.
Published: (2025)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
by: Rezaeimanesh, Sara, et al.
Published: (2026)
by: Rezaeimanesh, Sara, et al.
Published: (2026)
LLMs Know More About Numbers than They Can Say
by: Yuchi, Fengting, et al.
Published: (2026)
by: Yuchi, Fengting, et al.
Published: (2026)
Large Language Models as Computable Approximations to Solomonoff Induction
by: Wan, Jun, et al.
Published: (2025)
by: Wan, Jun, et al.
Published: (2025)
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
by: Yeom, Jewon, et al.
Published: (2026)
by: Yeom, Jewon, et al.
Published: (2026)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
by: Prakash, Nirmalendu, et al.
Published: (2025)
by: Prakash, Nirmalendu, et al.
Published: (2025)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
SnapKV: LLM Knows What You are Looking for Before Generation
by: Li, Yuhong, et al.
Published: (2024)
by: Li, Yuhong, et al.
Published: (2024)
Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"
by: Madhwal, Dhruv, et al.
Published: (2026)
by: Madhwal, Dhruv, et al.
Published: (2026)
Dialogue Injection Attack: Jailbreaking LLMs through Context Manipulation
by: Meng, Wenlong, et al.
Published: (2025)
by: Meng, Wenlong, et al.
Published: (2025)
We Know I Know You Know; Choreographic Programming With Multicast and Multiply Located Values
by: Bates, Mako, et al.
Published: (2024)
by: Bates, Mako, et al.
Published: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
by: Li, Qizhang, et al.
Published: (2024)
by: Li, Qizhang, et al.
Published: (2024)
Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction
by: Huang, Yuting, et al.
Published: (2025)
by: Huang, Yuting, et al.
Published: (2025)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
by: Wang, Zijun, et al.
Published: (2024)
by: Wang, Zijun, et al.
Published: (2024)
SafeInt: Shielding Large Language Models from Jailbreak Attacks via Safety-Aware Representation Intervention
by: Wu, Jiaqi, et al.
Published: (2025)
by: Wu, Jiaqi, et al.
Published: (2025)
All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
by: Takemoto, Kazuhiro
Published: (2024)
by: Takemoto, Kazuhiro
Published: (2024)
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented Generation
by: Li, Zhuohang, et al.
Published: (2024)
by: Li, Zhuohang, et al.
Published: (2024)
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
by: Ding, Peng, et al.
Published: (2025)
by: Ding, Peng, et al.
Published: (2025)
Similar Items
-
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
by: Luo, Wen, et al.
Published: (2026) -
You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
by: Gaznepoglu, Ünal Ege, et al.
Published: (2025) -
HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
by: Mei, Lingrui, et al.
Published: (2024) -
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026) -
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
by: Tan, Chenchen, et al.
Published: (2025)