Aligned Probing: Relating Toxic Behavior and Model Internals
Fuente:
arXiv
Saved in:
| Main Authors: | Waldis, Andreas, Gautam, Vagrant, Lauscher, Anne, Klakow, Dietrich, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
by: Mewes, Fabian, et al.
Published: (2026)
by: Mewes, Fabian, et al.
Published: (2026)
How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
by: Waldis, Andreas, et al.
Published: (2023)
by: Waldis, Andreas, et al.
Published: (2023)
Overview of PerpectiveArg2024: The First Shared Task on Perspective Argument Retrieval
by: Falk, Neele, et al.
Published: (2024)
by: Falk, Neele, et al.
Published: (2024)
Understanding "Democratization" in NLP and ML Research
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
by: García-de-Herreros, Paloma, et al.
Published: (2025)
by: García-de-Herreros, Paloma, et al.
Published: (2025)
Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets
by: Schiller, Benjamin, et al.
Published: (2022)
by: Schiller, Benjamin, et al.
Published: (2022)
Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
by: Purkayastha, Sukannya, et al.
Published: (2025)
by: Purkayastha, Sukannya, et al.
Published: (2025)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
by: Mosbach, Marius, et al.
Published: (2024)
by: Mosbach, Marius, et al.
Published: (2024)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
A Pipeline to Assess Merging Methods via Behavior and Internals
by: Sigrist, Yutaro, et al.
Published: (2025)
by: Sigrist, Yutaro, et al.
Published: (2025)
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
by: Beck, Tilman, et al.
Published: (2023)
by: Beck, Tilman, et al.
Published: (2023)
LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
by: Purkayastha, Sukannya, et al.
Published: (2025)
by: Purkayastha, Sukannya, et al.
Published: (2025)
Teaching and Critiquing Conceptualization and Operationalization in NLP
by: Gautam, Vagrant
Published: (2025)
by: Gautam, Vagrant
Published: (2025)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
by: Purkayastha, Sukannya, et al.
Published: (2026)
by: Purkayastha, Sukannya, et al.
Published: (2026)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
The Impact of Demonstrations on Multilingual In-Context Learning: A Multidimensional Analysis
by: Zhang, Miaoran, et al.
Published: (2024)
by: Zhang, Miaoran, et al.
Published: (2024)
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
by: Sachdeva, Rachneet, et al.
Published: (2025)
by: Sachdeva, Rachneet, et al.
Published: (2025)
Expert Preference-based Evaluation of Automated Related Work Generation
by: Şahinuç, Furkan, et al.
Published: (2025)
by: Şahinuç, Furkan, et al.
Published: (2025)
Your Multimodal Speech Model Says I Have a Face for Radio
by: Nachesa, Maya K., et al.
Published: (2026)
by: Nachesa, Maya K., et al.
Published: (2026)
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
by: Holtermann, Carolin, et al.
Published: (2025)
by: Holtermann, Carolin, et al.
Published: (2025)
Utilizing Multimodal Data for Edge Case Robust Call-sign Recognition and Understanding
by: Blatt, Alexander, et al.
Published: (2024)
by: Blatt, Alexander, et al.
Published: (2024)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
by: Dycke, Nils, et al.
Published: (2025)
by: Dycke, Nils, et al.
Published: (2025)
Citation Failure: Definition, Analysis and Efficient Mitigation
by: Buchmann, Jan, et al.
Published: (2025)
by: Buchmann, Jan, et al.
Published: (2025)
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification
by: Bates, Luke, et al.
Published: (2023)
by: Bates, Luke, et al.
Published: (2023)
Reward Modeling for Scientific Writing Evaluation
by: Şahinuç, Furkan, et al.
Published: (2026)
by: Şahinuç, Furkan, et al.
Published: (2026)
Token Weighting for Long-Range Language Modeling
by: Helm, Falko, et al.
Published: (2025)
by: Helm, Falko, et al.
Published: (2025)
Attribute or Abstain: Large Language Models as Long Document Assistants
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
by: Paul, Indraneil, et al.
Published: (2024)
by: Paul, Indraneil, et al.
Published: (2024)
IGC: Integrating a Gated Calculator into an LLM to Solve Arithmetic Tasks Reliably and Efficiently
by: Dietz, Florian, et al.
Published: (2025)
by: Dietz, Florian, et al.
Published: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
by: Baumgärtner, Tim, et al.
Published: (2026)
by: Baumgärtner, Tim, et al.
Published: (2026)
Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
Similar Items
-
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
by: Waldis, Andreas, et al.
Published: (2024) -
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024) -
WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case
by: Gautam, Vagrant, et al.
Published: (2024) -
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024) -
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
by: Mewes, Fabian, et al.
Published: (2026)