Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Frank, Aranguri, Santiago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inference-Time Toxicity Mitigation in Protein Language Models
von: Burda, Manuel Fernández, et al.
Veröffentlicht: (2026)
von: Burda, Manuel Fernández, et al.
Veröffentlicht: (2026)
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
von: Tu, Zhenfeng, et al.
Veröffentlicht: (2024)
von: Tu, Zhenfeng, et al.
Veröffentlicht: (2024)
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
von: Balappanawar, Ishwar, et al.
Veröffentlicht: (2025)
von: Balappanawar, Ishwar, et al.
Veröffentlicht: (2025)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
Phase-aware Training Schedule Simplifies Learning in Flow-Based Generative Models
von: Aranguri, Santiago, et al.
Veröffentlicht: (2024)
von: Aranguri, Santiago, et al.
Veröffentlicht: (2024)
AIM: Attributing, Interpreting, Mitigating Data Unfairness
von: Liu, Zhining, et al.
Veröffentlicht: (2024)
von: Liu, Zhining, et al.
Veröffentlicht: (2024)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
Discovering Behavioral Predispositions in Data to Improve Human Activity Recognition
von: Popko, Maximilian, et al.
Veröffentlicht: (2022)
von: Popko, Maximilian, et al.
Veröffentlicht: (2022)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
von: Hoang, Huy, et al.
Veröffentlicht: (2024)
von: Hoang, Huy, et al.
Veröffentlicht: (2024)
Efficient Ensembles Improve Training Data Attribution
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
von: Chen, Jianhui, et al.
Veröffentlicht: (2026)
von: Chen, Jianhui, et al.
Veröffentlicht: (2026)
Automatic Configuration of LLM Post-Training Pipelines
von: Chwa, Channe, et al.
Veröffentlicht: (2026)
von: Chwa, Channe, et al.
Veröffentlicht: (2026)
Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
von: Nam, Yunhun, et al.
Veröffentlicht: (2025)
von: Nam, Yunhun, et al.
Veröffentlicht: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026)
von: Rank, Ben, et al.
Veröffentlicht: (2026)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
von: Wu, Jiayun, et al.
Veröffentlicht: (2025)
von: Wu, Jiayun, et al.
Veröffentlicht: (2025)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
von: Hoang, Huy, et al.
Veröffentlicht: (2025)
von: Hoang, Huy, et al.
Veröffentlicht: (2025)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Models
von: Lin, Jinxu, et al.
Veröffentlicht: (2024)
von: Lin, Jinxu, et al.
Veröffentlicht: (2024)
Distributional Training Data Attribution: What do Influence Functions Sample?
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
Correlation-Aware Feature Attribution Based Explainable AI
von: Sengupta, Poushali, et al.
Veröffentlicht: (2025)
von: Sengupta, Poushali, et al.
Veröffentlicht: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
PT$^2$-LLM: Post-Training Ternarization for Large Language Models
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
von: Zhao, Stephen, et al.
Veröffentlicht: (2025)
von: Zhao, Stephen, et al.
Veröffentlicht: (2025)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
von: Han, Zhenyu, et al.
Veröffentlicht: (2025)
von: Han, Zhenyu, et al.
Veröffentlicht: (2025)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
von: Murata, Naoki, et al.
Veröffentlicht: (2026)
von: Murata, Naoki, et al.
Veröffentlicht: (2026)
Measuring and Mitigating Bias for Tabular Datasets with Multiple Protected Attributes
von: Duong, Manh Khoi, et al.
Veröffentlicht: (2024)
von: Duong, Manh Khoi, et al.
Veröffentlicht: (2024)
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
von: Lai, Song, et al.
Veröffentlicht: (2025)
von: Lai, Song, et al.
Veröffentlicht: (2025)
Explanova: Automatically Discover Data Insights in N \times M Table via XAI Combined LLM Workflow
von: Huang, Yiming
Veröffentlicht: (2026)
von: Huang, Yiming
Veröffentlicht: (2026)
TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models
von: Tang, Yuchi, et al.
Veröffentlicht: (2025)
von: Tang, Yuchi, et al.
Veröffentlicht: (2025)
Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking
von: Vardasbi, Ali, et al.
Veröffentlicht: (2025)
von: Vardasbi, Ali, et al.
Veröffentlicht: (2025)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Inference-Time Toxicity Mitigation in Protein Language Models
von: Burda, Manuel Fernández, et al.
Veröffentlicht: (2026) -
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
von: Tu, Zhenfeng, et al.
Veröffentlicht: (2024) -
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
von: Li, Zhe, et al.
Veröffentlicht: (2025) -
Who's the Evil Twin? Differential Auditing for Undesired Behavior
von: Balappanawar, Ishwar, et al.
Veröffentlicht: (2025) -
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)