A Practical Method for Generating String Counterfactuals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Avitan, Matan, Cotterell, Ryan, Goldberg, Yoav, Ravfogel, Shauli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Log-linear Guardedness and its Implications
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Representation Surgery: Theory and Practice of Affine Steering
von: Singh, Shashwat, et al.
Veröffentlicht: (2024)
von: Singh, Shashwat, et al.
Veröffentlicht: (2024)
Gumbel Counterfactual Generation From Language Models
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
von: Ben-Zaken, Elad, et al.
Veröffentlicht: (2021)
von: Ben-Zaken, Elad, et al.
Veröffentlicht: (2021)
LEACE: Perfect linear concept erasure in closed form
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
von: Cohen, Amir DN, et al.
Veröffentlicht: (2024)
von: Cohen, Amir DN, et al.
Veröffentlicht: (2024)
Description-Based Text Similarity
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2023)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2023)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
State over Tokens: Characterizing the Role of Reasoning Tokens
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
Compared to What? Baselines and Metrics for Counterfactual Prompting
von: Yang, Zihao, et al.
Veröffentlicht: (2026)
von: Yang, Zihao, et al.
Veröffentlicht: (2026)
On Affine Homotopy between Language Encoders
von: Chan, Robin SM, et al.
Veröffentlicht: (2024)
von: Chan, Robin SM, et al.
Veröffentlicht: (2024)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
PRISM: PRIor from corpus Statistics for topic Modeling
von: Ishon, Tal, et al.
Veröffentlicht: (2026)
von: Ishon, Tal, et al.
Veröffentlicht: (2026)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
von: Hamman, Faisal, et al.
Veröffentlicht: (2025)
von: Hamman, Faisal, et al.
Veröffentlicht: (2025)
Efficient Decoding Methods for Language Models on Encrypted Data
von: Avitan, Matan, et al.
Veröffentlicht: (2025)
von: Avitan, Matan, et al.
Veröffentlicht: (2025)
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
A Regularized LSTM Method for Detecting Fake News Articles
von: Camelia, Tanjina Sultana, et al.
Veröffentlicht: (2024)
von: Camelia, Tanjina Sultana, et al.
Veröffentlicht: (2024)
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
von: Christian, Brian, et al.
Veröffentlicht: (2026)
von: Christian, Brian, et al.
Veröffentlicht: (2026)
Unintended Impacts of LLM Alignment on Global Representation
von: Ryan, Michael J., et al.
Veröffentlicht: (2024)
von: Ryan, Michael J., et al.
Veröffentlicht: (2024)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
von: Noels, Sander, et al.
Veröffentlicht: (2025)
von: Noels, Sander, et al.
Veröffentlicht: (2025)
Transformers Can Represent $n$-gram Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
von: Yang, Hongkun, et al.
Veröffentlicht: (2026)
von: Yang, Hongkun, et al.
Veröffentlicht: (2026)
Reasoning Models Generate Societies of Thought
von: Kim, Junsol, et al.
Veröffentlicht: (2026)
von: Kim, Junsol, et al.
Veröffentlicht: (2026)
KPoEM: A Human-Annotated Dataset for Emotion Classification and RAG-Based Poetry Generation in Korean Modern Poetry
von: Lim, Iro, et al.
Veröffentlicht: (2025)
von: Lim, Iro, et al.
Veröffentlicht: (2025)
On Efficiently Representing Regular Languages as RNNs
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
On the Representational Capacity of Recurrent Neural Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
Improving Socratic Question Generation using Data Augmentation and Preference Optimization
von: Kumar, Nischal Ashok, et al.
Veröffentlicht: (2024)
von: Kumar, Nischal Ashok, et al.
Veröffentlicht: (2024)
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
von: Friedrich, Felix, et al.
Veröffentlicht: (2024)
von: Friedrich, Felix, et al.
Veröffentlicht: (2024)
Reflecting in the Reflection: Integrating a Socratic Questioning Framework into Automated AI-Based Question Generation
von: Holub, Ondřej, et al.
Veröffentlicht: (2026)
von: Holub, Ondřej, et al.
Veröffentlicht: (2026)
Unique Hard Attention: A Tale of Two Sides
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment
von: Borchers, Conrad, et al.
Veröffentlicht: (2025)
von: Borchers, Conrad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Log-linear Guardedness and its Implications
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022) -
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022) -
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022) -
Representation Surgery: Theory and Practice of Affine Steering
von: Singh, Shashwat, et al.
Veröffentlicht: (2024) -
Gumbel Counterfactual Generation From Language Models
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)