ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhixue, Shan, Boxuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023)
by: Zhao, Zhixue, et al.
Published: (2023)
RePo: Language Models with Context Re-Positioning
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
by: Kurz, Simon, et al.
Published: (2024)
by: Kurz, Simon, et al.
Published: (2024)
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
by: Ramesh, Samarth N, et al.
Published: (2024)
by: Ramesh, Samarth N, et al.
Published: (2024)
Reliable, Adaptable, and Attributable Language Models with Retrieval
by: Asai, Akari, et al.
Published: (2024)
by: Asai, Akari, et al.
Published: (2024)
ReALM: Reference Resolution As Language Modeling
by: Moniz, Joel Ruben Antony, et al.
Published: (2024)
by: Moniz, Joel Ruben Antony, et al.
Published: (2024)
ReFT: Representation Finetuning for Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
by: Shan, Zikang, et al.
Published: (2026)
by: Shan, Zikang, et al.
Published: (2026)
Multi-Attribute Steering of Language Models via Targeted Intervention
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
Large Language Models as Topological Structure Enhancers for Text-Attributed Graphs
by: Sun, Shengyin, et al.
Published: (2023)
by: Sun, Shengyin, et al.
Published: (2023)
ELF-Gym: Evaluating Large Language Models Generated Features for Tabular Prediction
by: Zhang, Yanlin, et al.
Published: (2024)
by: Zhang, Yanlin, et al.
Published: (2024)
FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
by: Li, Zhuoyun, et al.
Published: (2026)
by: Li, Zhuoyun, et al.
Published: (2026)
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
by: Qi, Jirui, et al.
Published: (2024)
by: Qi, Jirui, et al.
Published: (2024)
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
by: Niu, Peizhi, et al.
Published: (2025)
by: Niu, Peizhi, et al.
Published: (2025)
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
by: Chuang, Yung-Sung, et al.
Published: (2025)
by: Chuang, Yung-Sung, et al.
Published: (2025)
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
by: Deng, Xinle, et al.
Published: (2026)
by: Deng, Xinle, et al.
Published: (2026)
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
by: Li, Jia-Nan, et al.
Published: (2025)
by: Li, Jia-Nan, et al.
Published: (2025)
Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
by: Su, Guangxin, et al.
Published: (2025)
by: Su, Guangxin, et al.
Published: (2025)
Deep Learning Detection Method for Large Language Models-Generated Scientific Content
by: Alhijawi, Bushra, et al.
Published: (2024)
by: Alhijawi, Bushra, et al.
Published: (2024)
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
by: Zhao, Siyan, et al.
Published: (2024)
by: Zhao, Siyan, et al.
Published: (2024)
Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
by: Zhang, Rongzhi, et al.
Published: (2025)
by: Zhang, Rongzhi, et al.
Published: (2025)
LLM-Select: Feature Selection with Large Language Models
by: Jeong, Daniel P., et al.
Published: (2024)
by: Jeong, Daniel P., et al.
Published: (2024)
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks
by: Zhu, Yinghao, et al.
Published: (2024)
by: Zhu, Yinghao, et al.
Published: (2024)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
by: Cai, Hengrui, et al.
Published: (2023)
by: Cai, Hengrui, et al.
Published: (2023)
Re-Tuning: Overcoming the Compositionality Limits of Large Language Models with Recursive Tuning
by: Pasewark, Eric, et al.
Published: (2024)
by: Pasewark, Eric, et al.
Published: (2024)
RePrompt: Planning by Automatic Prompt Engineering for Large Language Models Agents
by: Chen, Weizhe, et al.
Published: (2024)
by: Chen, Weizhe, et al.
Published: (2024)
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
by: Gao, Bo, et al.
Published: (2025)
by: Gao, Bo, et al.
Published: (2025)
CEM: A Data-Efficient Method for Large Language Models to Continue Evolving From Mistakes
by: Zhao, Haokun, et al.
Published: (2024)
by: Zhao, Haokun, et al.
Published: (2024)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
by: Jacobi, Jonathan, et al.
Published: (2025)
by: Jacobi, Jonathan, et al.
Published: (2025)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
by: Lin, Xuan, et al.
Published: (2025)
by: Lin, Xuan, et al.
Published: (2025)
Efficient Estimation of Kernel Surrogate Models for Task Attribution
by: Zhang, Zhenshuo, et al.
Published: (2026)
by: Zhang, Zhenshuo, et al.
Published: (2026)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
by: Soo, Samuel, et al.
Published: (2025)
by: Soo, Samuel, et al.
Published: (2025)
RevOrder: A Novel Method for Enhanced Arithmetic in Language Models
by: Shen, Si, et al.
Published: (2024)
by: Shen, Si, et al.
Published: (2024)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition
by: Uddin, Mohammed Mudassir, et al.
Published: (2026)
by: Uddin, Mohammed Mudassir, et al.
Published: (2026)
Merino: Entropy-driven Design for Generative Language Models on IoT Devices
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Similar Items
-
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023) -
RePo: Language Models with Context Re-Positioning
by: Li, Huayang, et al.
Published: (2025) -
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
by: Kurz, Simon, et al.
Published: (2024) -
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
by: Ramesh, Samarth N, et al.
Published: (2024) -
Reliable, Adaptable, and Attributable Language Models with Retrieval
by: Asai, Akari, et al.
Published: (2024)