Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Ghandeharioun, Asma, Caciularu, Avi, Pearce, Adam, Dixon, Lucas, Geva, Mor |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Can Transformers Count to n?
di: Yehudai, Gilad, et al.
Pubblicazione: (2024)
di: Yehudai, Gilad, et al.
Pubblicazione: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
di: Katz, Shahar, et al.
Pubblicazione: (2024)
di: Katz, Shahar, et al.
Pubblicazione: (2024)
Think Before You Lie: How Reasoning Leads to Honesty
di: Yuan, Ann, et al.
Pubblicazione: (2026)
di: Yuan, Ann, et al.
Pubblicazione: (2026)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
di: Goldman, Omer, et al.
Pubblicazione: (2024)
di: Goldman, Omer, et al.
Pubblicazione: (2024)
Interpretability Illusions in the Generalization of Simplified Models
di: Friedman, Dan, et al.
Pubblicazione: (2023)
di: Friedman, Dan, et al.
Pubblicazione: (2023)
Latent Reasoning with Supervised Thinking States
di: Amos, Ido, et al.
Pubblicazione: (2026)
di: Amos, Ido, et al.
Pubblicazione: (2026)
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
di: Blum, Carter, et al.
Pubblicazione: (2025)
di: Blum, Carter, et al.
Pubblicazione: (2025)
Layer by Layer: Uncovering Hidden Representations in Language Models
di: Skean, Oscar, et al.
Pubblicazione: (2025)
di: Skean, Oscar, et al.
Pubblicazione: (2025)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
Who's asking? User personas and the mechanics of latent misalignment
di: Ghandeharioun, Asma, et al.
Pubblicazione: (2024)
di: Ghandeharioun, Asma, et al.
Pubblicazione: (2024)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
di: Cattan, Arie, et al.
Pubblicazione: (2025)
di: Cattan, Arie, et al.
Pubblicazione: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
di: Huang, Jing, et al.
Pubblicazione: (2024)
di: Huang, Jing, et al.
Pubblicazione: (2024)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
di: Chen, Mayee F., et al.
Pubblicazione: (2024)
di: Chen, Mayee F., et al.
Pubblicazione: (2024)
UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation
di: Guo, Shuhan, et al.
Pubblicazione: (2024)
di: Guo, Shuhan, et al.
Pubblicazione: (2024)
A Unified Framework for Model Editing
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
Decoding-time Realignment of Language Models
di: Liu, Tianlin, et al.
Pubblicazione: (2024)
di: Liu, Tianlin, et al.
Pubblicazione: (2024)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
di: Yona, Itay, et al.
Pubblicazione: (2026)
di: Yona, Itay, et al.
Pubblicazione: (2026)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
di: Ivgi, Maor, et al.
Pubblicazione: (2024)
di: Ivgi, Maor, et al.
Pubblicazione: (2024)
Correlated Errors in Large Language Models
di: Kim, Elliot, et al.
Pubblicazione: (2025)
di: Kim, Elliot, et al.
Pubblicazione: (2025)
A Study on Hidden Layer Distillation for Large Language Model Pre-Training
di: Guigon, Maxime, et al.
Pubblicazione: (2026)
di: Guigon, Maxime, et al.
Pubblicazione: (2026)
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
di: Jin, Yiqiao, et al.
Pubblicazione: (2026)
di: Jin, Yiqiao, et al.
Pubblicazione: (2026)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
di: Weng, Zixuan, et al.
Pubblicazione: (2026)
di: Weng, Zixuan, et al.
Pubblicazione: (2026)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
di: Flamant, Cedric, et al.
Pubblicazione: (2026)
di: Flamant, Cedric, et al.
Pubblicazione: (2026)
Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models
di: Wong, Wai Tuck, et al.
Pubblicazione: (2026)
di: Wong, Wai Tuck, et al.
Pubblicazione: (2026)
Introducing Background Temperature to Characterise Hidden Randomness in Large Language Models
di: Messina, Alberto, et al.
Pubblicazione: (2026)
di: Messina, Alberto, et al.
Pubblicazione: (2026)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
di: Sun, Shaoning, et al.
Pubblicazione: (2026)
di: Sun, Shaoning, et al.
Pubblicazione: (2026)
Hidden State Poisoning Attacks against Mamba-based Language Models
di: Mercier, Alexandre Le, et al.
Pubblicazione: (2026)
di: Mercier, Alexandre Le, et al.
Pubblicazione: (2026)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
di: Cheng, Jiale, et al.
Pubblicazione: (2024)
di: Cheng, Jiale, et al.
Pubblicazione: (2024)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
di: Li, Melody Zixuan, et al.
Pubblicazione: (2025)
di: Li, Melody Zixuan, et al.
Pubblicazione: (2025)
Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling
di: Amin, Adil
Pubblicazione: (2026)
di: Amin, Adil
Pubblicazione: (2026)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
di: Yang, Daniel, et al.
Pubblicazione: (2026)
di: Yang, Daniel, et al.
Pubblicazione: (2026)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
di: Shafran, Or, et al.
Pubblicazione: (2025)
di: Shafran, Or, et al.
Pubblicazione: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
PromptBench: A Unified Library for Evaluation of Large Language Models
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
di: Ye, Tian, et al.
Pubblicazione: (2024)
di: Ye, Tian, et al.
Pubblicazione: (2024)
Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image Models
di: Ventura, Mor, et al.
Pubblicazione: (2023)
di: Ventura, Mor, et al.
Pubblicazione: (2023)
ReFT: Representation Finetuning for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
Affinity and Diversity: A Unified Metric for Demonstration Selection via Internal Representations
di: Kato, Mariko, et al.
Pubblicazione: (2025)
di: Kato, Mariko, et al.
Pubblicazione: (2025)
Documenti analoghi
-
When Can Transformers Count to n?
di: Yehudai, Gilad, et al.
Pubblicazione: (2024) -
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
di: Katz, Shahar, et al.
Pubblicazione: (2024) -
Think Before You Lie: How Reasoning Leads to Honesty
di: Yuan, Ann, et al.
Pubblicazione: (2026) -
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
di: Goldman, Omer, et al.
Pubblicazione: (2024) -
Interpretability Illusions in the Generalization of Simplified Models
di: Friedman, Dan, et al.
Pubblicazione: (2023)