Log-linear Guardedness and its Implications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ravfogel, Shauli, Goldberg, Yoav, Cotterell, Ryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
A Practical Method for Generating String Counterfactuals
von: Avitan, Matan, et al.
Veröffentlicht: (2024)
von: Avitan, Matan, et al.
Veröffentlicht: (2024)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
von: Ben-Zaken, Elad, et al.
Veröffentlicht: (2021)
von: Ben-Zaken, Elad, et al.
Veröffentlicht: (2021)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
von: Cohen, Amir DN, et al.
Veröffentlicht: (2024)
von: Cohen, Amir DN, et al.
Veröffentlicht: (2024)
Gumbel Counterfactual Generation From Language Models
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
LEACE: Perfect linear concept erasure in closed form
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
Description-Based Text Similarity
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2023)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2023)
Representation Surgery: Theory and Practice of Affine Steering
von: Singh, Shashwat, et al.
Veröffentlicht: (2024)
von: Singh, Shashwat, et al.
Veröffentlicht: (2024)
State over Tokens: Characterizing the Role of Reasoning Tokens
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
On Affine Homotopy between Language Encoders
von: Chan, Robin SM, et al.
Veröffentlicht: (2024)
von: Chan, Robin SM, et al.
Veröffentlicht: (2024)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
PRISM: PRIor from corpus Statistics for topic Modeling
von: Ishon, Tal, et al.
Veröffentlicht: (2026)
von: Ishon, Tal, et al.
Veröffentlicht: (2026)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
Preserving Task-Relevant Information Under Linear Concept Removal
von: Holstege, Floris, et al.
Veröffentlicht: (2025)
von: Holstege, Floris, et al.
Veröffentlicht: (2025)
Transformers Can Represent $n$-gram Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Compared to What? Baselines and Metrics for Counterfactual Prompting
von: Yang, Zihao, et al.
Veröffentlicht: (2026)
von: Yang, Zihao, et al.
Veröffentlicht: (2026)
An Algebraic View of the Expressivity of Recurrent Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2026)
von: Nowak, Franz, et al.
Veröffentlicht: (2026)
On the Representational Capacity of Recurrent Neural Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
von: Amini, Afra, et al.
Veröffentlicht: (2025)
von: Amini, Afra, et al.
Veröffentlicht: (2025)
Direct Preference Optimization with an Offset
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
On the Role of Context in Reading Time Prediction
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
On Efficiently Representing Regular Languages as RNNs
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
GRADE: Quantifying Sample Diversity in Text-to-Image Models
von: Rassin, Royi, et al.
Veröffentlicht: (2024)
von: Rassin, Royi, et al.
Veröffentlicht: (2024)
Can Language Models Learn Typologically Implausible Languages?
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
Geometric Factual Recall in Transformers
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2026)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2026)
Unique Hard Attention: A Tale of Two Sides
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
Syntactic Control of Language Models by Posterior Inference
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
Variational Best-of-N Alignment
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
von: Trainin, Itamar, et al.
Veröffentlicht: (2026)
von: Trainin, Itamar, et al.
Veröffentlicht: (2026)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
von: Svete, Anej, et al.
Veröffentlicht: (2026)
von: Svete, Anej, et al.
Veröffentlicht: (2026)
Training Neural Networks as Recognizers of Formal Languages
von: Butoi, Alexandra, et al.
Veröffentlicht: (2024)
von: Butoi, Alexandra, et al.
Veröffentlicht: (2024)
Context-Free Recognition with Transformers
von: Jerad, Selim, et al.
Veröffentlicht: (2026)
von: Jerad, Selim, et al.
Veröffentlicht: (2026)
The Truthfulness Spectrum Hypothesis
von: Ying, Zhuofan Josh, et al.
Veröffentlicht: (2026)
von: Ying, Zhuofan Josh, et al.
Veröffentlicht: (2026)
Transformers are Inherently Succinct
von: Bergsträßer, Pascal, et al.
Veröffentlicht: (2025)
von: Bergsträßer, Pascal, et al.
Veröffentlicht: (2025)
A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior
von: Re, Francesco Ignazio, et al.
Veröffentlicht: (2025)
von: Re, Francesco Ignazio, et al.
Veröffentlicht: (2025)
Emergence of Linear Truth Encodings in Language Models
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2025)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022) -
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022) -
A Practical Method for Generating String Counterfactuals
von: Avitan, Matan, et al.
Veröffentlicht: (2024) -
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
von: Ben-Zaken, Elad, et al.
Veröffentlicht: (2021) -
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
von: Cohen, Amir DN, et al.
Veröffentlicht: (2024)