Saved in:
| Main Authors: | Levinstein, B. A., Herrmann, Daniel A. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2307.00175 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing the Limits of the Lie Detector Approach to LLM Deception
by: Berger, Tom-Felix
Published: (2026)
by: Berger, Tom-Felix
Published: (2026)
Language Models Optimized to Fool Detectors Still Have a Distinct Style (And How to Change It)
by: Soto, Rafael Rivera, et al.
Published: (2025)
by: Soto, Rafael Rivera, et al.
Published: (2025)
Erasing Conceptual Knowledge from Language Models
by: Gandikota, Rohit, et al.
Published: (2024)
by: Gandikota, Rohit, et al.
Published: (2024)
Still "Talking About Large Language Models": Some Clarifications
by: Shanahan, Murray
Published: (2024)
by: Shanahan, Murray
Published: (2024)
An Empirical Study of Mamba-based Language Models
by: Waleffe, Roger, et al.
Published: (2024)
by: Waleffe, Roger, et al.
Published: (2024)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
by: To, Bang Trinh Tran, et al.
Published: (2025)
by: To, Bang Trinh Tran, et al.
Published: (2025)
Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry
by: Mathewson, Kyle Elliott
Published: (2026)
by: Mathewson, Kyle Elliott
Published: (2026)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models
by: Mundra, Nandini, et al.
Published: (2024)
by: Mundra, Nandini, et al.
Published: (2024)
Assessing Adversarial Robustness of Large Language Models: An Empirical Study
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
Temporally Consistent Factuality Probing for Large Language Models
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
by: Wang, Qianli, et al.
Published: (2026)
by: Wang, Qianli, et al.
Published: (2026)
Comparing Template-based and Template-free Language Model Probing
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
Monitoring Latent World States in Language Models with Propositional Probes
by: Feng, Jiahai, et al.
Published: (2024)
by: Feng, Jiahai, et al.
Published: (2024)
An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models
by: Luo, Hanrui, et al.
Published: (2026)
by: Luo, Hanrui, et al.
Published: (2026)
Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey
by: Lee, Kunil, et al.
Published: (2026)
by: Lee, Kunil, et al.
Published: (2026)
Cross-Lingual Empirical Evaluation of Large Language Models for Arabic Medical Tasks
by: Abouzahir, Chaimae, et al.
Published: (2026)
by: Abouzahir, Chaimae, et al.
Published: (2026)
Editing Conceptual Knowledge for Large Language Models
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
Probing Cultural Signals in Large Language Models through Author Profiling
by: Lafargue, Valentin, et al.
Published: (2026)
by: Lafargue, Valentin, et al.
Published: (2026)
Detecting Conceptual Abstraction in LLMs
by: Regneri, Michaela, et al.
Published: (2024)
by: Regneri, Michaela, et al.
Published: (2024)
Small Models Are (Still) Effective Cross-Domain Argument Extractors
by: Gantt, William, et al.
Published: (2024)
by: Gantt, William, et al.
Published: (2024)
From Imitation to Introspection: Probing Self-Consciousness in Language Models
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Empirical Analysis of Efficient Fine-Tuning Methods for Large Pre-Trained Language Models
by: Doering, Nigel, et al.
Published: (2024)
by: Doering, Nigel, et al.
Published: (2024)
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
by: Sharma, Kartik, et al.
Published: (2025)
by: Sharma, Kartik, et al.
Published: (2025)
Pretraining Language Models with Subword Regularization: An Empirical Study of BPE Dropout in Low-Resource NLP
by: Visser, Ruan, et al.
Published: (2026)
by: Visser, Ruan, et al.
Published: (2026)
Reliability Under Randomness: An Empirical Analysis of Sparse and Dense Language Models Across Decoding Temperatures
by: Grover, Kabir
Published: (2026)
by: Grover, Kabir
Published: (2026)
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
by: Kocak, Aysenur, et al.
Published: (2025)
by: Kocak, Aysenur, et al.
Published: (2025)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
by: Fu, Wenjie, et al.
Published: (2024)
by: Fu, Wenjie, et al.
Published: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
Diagnosing Robotics Systems Issues with Large Language Models
by: Herrmann, Jordis Emilia, et al.
Published: (2024)
by: Herrmann, Jordis Emilia, et al.
Published: (2024)
Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector
by: Zhang, Andi, et al.
Published: (2024)
by: Zhang, Andi, et al.
Published: (2024)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
by: Li, Daoyang, et al.
Published: (2024)
by: Li, Daoyang, et al.
Published: (2024)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
by: Noels, Sander, et al.
Published: (2025)
by: Noels, Sander, et al.
Published: (2025)
Science is Exploration: Computational Frontiers for Conceptual Metaphor Theory
by: Hicke, Rebecca M. M., et al.
Published: (2024)
by: Hicke, Rebecca M. M., et al.
Published: (2024)
Dynamic Evaluation of Large Language Models by Meta Probing Agents
by: Zhu, Kaijie, et al.
Published: (2024)
by: Zhu, Kaijie, et al.
Published: (2024)
Probing the Decision Boundaries of In-context Learning in Large Language Models
by: Zhao, Siyan, et al.
Published: (2024)
by: Zhao, Siyan, et al.
Published: (2024)
The Zero Body Problem: Probing LLM Use of Sensory Language
by: Hicke, Rebecca M. M., et al.
Published: (2025)
by: Hicke, Rebecca M. M., et al.
Published: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025)
by: Malek, Alan, et al.
Published: (2025)
FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models
by: Wang, Yanling, et al.
Published: (2024)
by: Wang, Yanling, et al.
Published: (2024)
A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models
by: Sahu, Utkarsh, et al.
Published: (2025)
by: Sahu, Utkarsh, et al.
Published: (2025)
Similar Items
-
Probing the Limits of the Lie Detector Approach to LLM Deception
by: Berger, Tom-Felix
Published: (2026) -
Language Models Optimized to Fool Detectors Still Have a Distinct Style (And How to Change It)
by: Soto, Rafael Rivera, et al.
Published: (2025) -
Erasing Conceptual Knowledge from Language Models
by: Gandikota, Rohit, et al.
Published: (2024) -
Still "Talking About Large Language Models": Some Clarifications
by: Shanahan, Murray
Published: (2024) -
An Empirical Study of Mamba-based Language Models
by: Waleffe, Roger, et al.
Published: (2024)