Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
Fuente:
arXiv
Saved in:
| Main Authors: | Boxo, Gerard, Neelappa, Aman, Raval, Shivam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
by: Dawes, Cutter, et al.
Published: (2026)
by: Dawes, Cutter, et al.
Published: (2026)
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
by: Goel, Aman, et al.
Published: (2025)
by: Goel, Aman, et al.
Published: (2025)
Caught in the Act: a mechanistic approach to detecting deception
by: Boxo, Gerard, et al.
Published: (2025)
by: Boxo, Gerard, et al.
Published: (2025)
Enriching language models with graph-based context information to better understand textual data
by: Roethel, Albert, et al.
Published: (2023)
by: Roethel, Albert, et al.
Published: (2023)
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)
by: Lampinen, Andrew K., et al.
Published: (2025)
The language of time: a language model perspective on time-series foundation models
by: Xie, Yi, et al.
Published: (2025)
by: Xie, Yi, et al.
Published: (2025)
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
Representation in large language models
by: Yetman, Cameron
Published: (2025)
by: Yetman, Cameron
Published: (2025)
Auditing language models for hidden objectives
by: Marks, Samuel, et al.
Published: (2025)
by: Marks, Samuel, et al.
Published: (2025)
Alignment faking in large language models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
Lightweight reranking for language model generations
by: Jain, Siddhartha, et al.
Published: (2023)
by: Jain, Siddhartha, et al.
Published: (2023)
Long-form factuality in large language models
by: Wei, Jerry, et al.
Published: (2024)
by: Wei, Jerry, et al.
Published: (2024)
Can large language models explore in-context?
by: Krishnamurthy, Akshay, et al.
Published: (2024)
by: Krishnamurthy, Akshay, et al.
Published: (2024)
Extracting effective solutions hidden in large language models via generated comprehensive specialists: case studies in developing electronic devices
by: Tomita, Hikari, et al.
Published: (2024)
by: Tomita, Hikari, et al.
Published: (2024)
AI-AI Bias: large language models favor communications generated by large language models
by: Laurito, Walter, et al.
Published: (2024)
by: Laurito, Walter, et al.
Published: (2024)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
Leveraging Large Language Models for Web Scraping
by: Ahluwalia, Aman, et al.
Published: (2024)
by: Ahluwalia, Aman, et al.
Published: (2024)
STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models
by: Basavatia, Shreyas, et al.
Published: (2024)
by: Basavatia, Shreyas, et al.
Published: (2024)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
Fresh in memory: Training-order recency is linearly encoded in language model activations
by: Krasheninnikov, Dmitrii, et al.
Published: (2025)
by: Krasheninnikov, Dmitrii, et al.
Published: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
by: Park, Ji Won, et al.
Published: (2025)
by: Park, Ji Won, et al.
Published: (2025)
Safety and accuracy follow different scaling laws in clinical large language models
by: Wind, Sebastian, et al.
Published: (2026)
by: Wind, Sebastian, et al.
Published: (2026)
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
Vocabulary shapes cross-lingual variation of word-order learnability in language models
by: Martins, Jonas Mayer, et al.
Published: (2026)
by: Martins, Jonas Mayer, et al.
Published: (2026)
CogBench: a large language model walks into a psychology lab
by: Coda-Forno, Julian, et al.
Published: (2024)
by: Coda-Forno, Julian, et al.
Published: (2024)
Perturbation: A simple and efficient adversarial tracer for representation learning in language models
by: Rozner, Joshua, et al.
Published: (2026)
by: Rozner, Joshua, et al.
Published: (2026)
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
by: Khraisha, Qusai, et al.
Published: (2023)
by: Khraisha, Qusai, et al.
Published: (2023)
Multi-step retrieval and reasoning improves radiology question answering with large language models
by: Wind, Sebastian, et al.
Published: (2025)
by: Wind, Sebastian, et al.
Published: (2025)
MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
by: Yoshida, Davis, et al.
Published: (2023)
by: Yoshida, Davis, et al.
Published: (2023)
SteuerLLM: Local specialized large language model for German tax law analysis
by: Wind, Sebastian, et al.
Published: (2026)
by: Wind, Sebastian, et al.
Published: (2026)
Multimodal large language model for wheat breeding: a new exploration of smart breeding
by: Yang, Guofeng, et al.
Published: (2024)
by: Yang, Guofeng, et al.
Published: (2024)
Can Large Language Models Infer Causal Relationships from Real-World Text?
by: Saklad, Ryan, et al.
Published: (2025)
by: Saklad, Ryan, et al.
Published: (2025)
IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts
by: Basak, Udvas, et al.
Published: (2024)
by: Basak, Udvas, et al.
Published: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
An artificial intelligence framework for end-to-end rare disease phenotyping from clinical notes using large language models
by: Shyr, Cathy, et al.
Published: (2026)
by: Shyr, Cathy, et al.
Published: (2026)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
Large language models can learn and generalize steganographic chain-of-thought under process supervision
by: Skaf, Joey, et al.
Published: (2025)
by: Skaf, Joey, et al.
Published: (2025)
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
by: Reddy, K. Sahit, et al.
Published: (2025)
by: Reddy, K. Sahit, et al.
Published: (2025)
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
by: Sharma, Aryan, et al.
Published: (2026)
by: Sharma, Aryan, et al.
Published: (2026)
Zero-shot data citation function classification using transformer-based large language models (LLMs)
by: Byers, Neil, et al.
Published: (2025)
by: Byers, Neil, et al.
Published: (2025)
Similar Items
-
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
by: Dawes, Cutter, et al.
Published: (2026) -
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
by: Goel, Aman, et al.
Published: (2025) -
Caught in the Act: a mechanistic approach to detecting deception
by: Boxo, Gerard, et al.
Published: (2025) -
Enriching language models with graph-based context information to better understand textual data
by: Roethel, Albert, et al.
Published: (2023) -
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)