Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
Fuente:
arXiv
Guardado en:
| Autores principales: | Kowal, Matthew, Paulo, Goncalo, Jaburi, Louis, Tseng, Tom, McKinney, Lev E, Heimersheim, Stefan, Tucker, Aaron David, Gleave, Adam, Pelrine, Kellin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
por: Struppek, Lukas, et al.
Publicado: (2026)
por: Struppek, Lukas, et al.
Publicado: (2026)
Can Go AIs be adversarially robust?
por: Tseng, Tom, et al.
Publicado: (2024)
por: Tseng, Tom, et al.
Publicado: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
por: Murphy, Brendan, et al.
Publicado: (2025)
por: Murphy, Brendan, et al.
Publicado: (2025)
Scaling Trends for Data Poisoning in LLMs
por: Bowen, Dillon, et al.
Publicado: (2024)
por: Bowen, Dillon, et al.
Publicado: (2024)
Exploiting Novel GPT-4 APIs
por: Pelrine, Kellin, et al.
Publicado: (2023)
por: Pelrine, Kellin, et al.
Publicado: (2023)
Large language models can effectively convince people to believe conspiracies
por: Costello, Thomas H., et al.
Publicado: (2026)
por: Costello, Thomas H., et al.
Publicado: (2026)
It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
por: Kowal, Matthew, et al.
Publicado: (2025)
por: Kowal, Matthew, et al.
Publicado: (2025)
GULPS: Two-Qubit Gate Synthesis via Linear Programming for Heterogeneous Instruction Sets
por: McKinney, Evan, et al.
Publicado: (2025)
por: McKinney, Evan, et al.
Publicado: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
por: Taufeeque, Mohammad, et al.
Publicado: (2026)
por: Taufeeque, Mohammad, et al.
Publicado: (2026)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
por: Ayonrinde, Kola, et al.
Publicado: (2025)
por: Ayonrinde, Kola, et al.
Publicado: (2025)
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
por: Ayonrinde, Kola, et al.
Publicado: (2025)
por: Ayonrinde, Kola, et al.
Publicado: (2025)
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
por: Chang, Vincent, et al.
Publicado: (2025)
por: Chang, Vincent, et al.
Publicado: (2025)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
por: Hossain, Saad, et al.
Publicado: (2026)
por: Hossain, Saad, et al.
Publicado: (2026)
Leakage Safe Graph Features for Interpretable Fraud Detection in Temporal Transaction Networks
por: Khaleghpour, Hamideh, et al.
Publicado: (2026)
por: Khaleghpour, Hamideh, et al.
Publicado: (2026)
Student Conceptions of Group Work: Visual Research into LIS Student Group Work Using the Draw-and-Write Technique
por: McKinney, Pamela, et al.
Publicado: (2018)
por: McKinney, Pamela, et al.
Publicado: (2018)
Human Frailties: Springboard to Increased Systems Engineering Influence
por: Eileen Patrice Arnold, et al.
Publicado: (2024)
por: Eileen Patrice Arnold, et al.
Publicado: (2024)
Online Influence Campaigns: Strategies and Vulnerabilities
por: Musulan, Andreea, et al.
Publicado: (2024)
por: Musulan, Andreea, et al.
Publicado: (2024)
Postcolonialism and Migration in French Comics
por: McKinney, Mark
Publicado: (2025)
por: McKinney, Mark
Publicado: (2025)
Grandmothering While Black: A Twenty‐First‐Century Story of Love, Coercion, and Survival. By Lashawnda L.Pittman. University of California Press, Oakland, California, 2023. 336 pp. $92.04 (hardcover). ISBN: 978‐0‐52‐038995‐3; $29.95 (paperback). ISBN: 978‐0‐52‐038996‐0; $29.95 (ebook). ISBN: 978‐0‐52‐038997‐7
por: Elliana McKinney
Publicado: (2025)
por: Elliana McKinney
Publicado: (2025)
Schools Inquiring About Seven-Day School Rerecording of Public and Instructional Television Programs.
por: McKinney, Eleanor
Publicado: (1975)
por: McKinney, Eleanor
Publicado: (1975)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
por: Rivera, Mauricio, et al.
Publicado: (2024)
por: Rivera, Mauricio, et al.
Publicado: (2024)
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
por: Vergho, Tyler, et al.
Publicado: (2024)
por: Vergho, Tyler, et al.
Publicado: (2024)
Scaling Trends in Language Model Robustness
por: Howe, Nikolaus, et al.
Publicado: (2024)
por: Howe, Nikolaus, et al.
Publicado: (2024)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
por: Taufeeque, Mohammad, et al.
Publicado: (2025)
por: Taufeeque, Mohammad, et al.
Publicado: (2025)
Water conservation surveys of New South Wales
por: McKinney, Hugh Giffen
Publicado: (1896)
por: McKinney, Hugh Giffen
Publicado: (1896)
Evolution of erect marine bryozoan faunas : repeated succes of unilaminate species
por: McKinney, F.K
Publicado: (1986)
por: McKinney, F.K
Publicado: (1986)
Created from nafta : the structure, function, and significance of the treatys related institutions / Joseph A. McKinney
por: McKinney, Joseph A
por: McKinney, Joseph A
Media Utilization in the Classroom.
por: Bowie, Melvin McKinney
Publicado: (1985)
por: Bowie, Melvin McKinney
Publicado: (1985)
Conceptual and Practical Matters: The Challenges and Benefits of Conducting Educational Research Using Historical Data. Sage Research Methods Cases Part 2
por: Stephen J. McKinney
Publicado: (2017)
por: Stephen J. McKinney
Publicado: (2017)
The Contribution of Iona and Peter Opie to Children's Literature.
por: McKinney, Barbara J.
Publicado: (1996)
por: McKinney, Barbara J.
Publicado: (1996)
Another Degree? What For?
por: McKinney, Eleanor R.
Publicado: (1969)
por: McKinney, Eleanor R.
Publicado: (1969)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
por: Braun, Dan, et al.
Publicado: (2025)
por: Braun, Dan, et al.
Publicado: (2025)
Negation Neglect: When models fail to learn negations in training
por: Mayne, Harry, et al.
Publicado: (2026)
por: Mayne, Harry, et al.
Publicado: (2026)
STACK: Adversarial Attacks on LLM Safeguard Pipelines
por: McKenzie, Ian R., et al.
Publicado: (2025)
por: McKenzie, Ian R., et al.
Publicado: (2025)
Optimizing Neuro-Fuzzy and Colonial Competition Algorithms for Skin Cancer Diagnosis in Dermatoscopic Images
por: Khaleghpour, Hamideh, et al.
Publicado: (2025)
por: Khaleghpour, Hamideh, et al.
Publicado: (2025)
Unified AI for Accurate Audio Anomaly Detection
por: Khaleghpour, Hamideh, et al.
Publicado: (2025)
por: Khaleghpour, Hamideh, et al.
Publicado: (2025)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
por: Cundy, Chris, et al.
Publicado: (2025)
por: Cundy, Chris, et al.
Publicado: (2025)
You can remove GPT2's LayerNorm by fine-tuning
por: Heimersheim, Stefan
Publicado: (2024)
por: Heimersheim, Stefan
Publicado: (2024)
Uncertainty Resolution in Misinformation Detection
por: Orlovskiy, Yury, et al.
Publicado: (2024)
por: Orlovskiy, Yury, et al.
Publicado: (2024)
Ejemplares similares
-
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
por: Struppek, Lukas, et al.
Publicado: (2026) -
Can Go AIs be adversarially robust?
por: Tseng, Tom, et al.
Publicado: (2024) -
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
por: Murphy, Brendan, et al.
Publicado: (2025) -
Scaling Trends for Data Poisoning in LLMs
por: Bowen, Dillon, et al.
Publicado: (2024) -
Exploiting Novel GPT-4 APIs
por: Pelrine, Kellin, et al.
Publicado: (2023)