Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Xuansheng, Zhao, Haiyan, Zhu, Yaochen, Shi, Yucheng, Yang, Fan, Hu, Lijie, Liu, Tianming, Zhai, Xiaoming, Yao, Wenlin, Li, Jundong, Du, Mengnan, Liu, Ninghao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
di: Wang, Yun, et al.
Pubblicazione: (2026)
di: Wang, Yun, et al.
Pubblicazione: (2026)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
di: Yang, Tianze, et al.
Pubblicazione: (2025)
di: Yang, Tianze, et al.
Pubblicazione: (2025)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
di: Wang, Yun, et al.
Pubblicazione: (2026)
di: Wang, Yun, et al.
Pubblicazione: (2026)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
di: Wang, Yun, et al.
Pubblicazione: (2025)
di: Wang, Yun, et al.
Pubblicazione: (2025)
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
di: Li, Zihao, et al.
Pubblicazione: (2024)
di: Li, Zihao, et al.
Pubblicazione: (2024)
Towards Uncovering How Large Language Model Works: An Explainability Perspective
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Soundness-Aware Level: A Microscopic Signature that Predicts LLM Reasoning Potential
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Foundation Models for Low-Resource Language Education (Vision Paper)
di: Ding, Zhaojun, et al.
Pubblicazione: (2024)
di: Ding, Zhaojun, et al.
Pubblicazione: (2024)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering?
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
di: Guo, Shuchen, et al.
Pubblicazione: (2025)
di: Guo, Shuchen, et al.
Pubblicazione: (2025)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
Strategic Demonstration Selection for Improved Fairness in LLM In-Context Learning
di: Hu, Jingyu, et al.
Pubblicazione: (2024)
di: Hu, Jingyu, et al.
Pubblicazione: (2024)
MGH Radiology Llama: A Llama 3 70B Model for Radiology
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
MKRAG: Medical Knowledge Retrieval Augmented Generation for Medical Question Answering
di: Shi, Yucheng, et al.
Pubblicazione: (2023)
di: Shi, Yucheng, et al.
Pubblicazione: (2023)
Transforming Teachers' Roles and Agencies in the Era of Generative AI: Perceptions, Acceptance, Knowledge, and Practices
di: Zhai, Xiaoming
Pubblicazione: (2024)
di: Zhai, Xiaoming
Pubblicazione: (2024)
Investigating CoT Monitorability in Large Reasoning Models
di: Yang, Shu, et al.
Pubblicazione: (2025)
di: Yang, Shu, et al.
Pubblicazione: (2025)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
Knowledge Graph-Enhanced Large Language Models via Path Selection
di: Liu, Haochen, et al.
Pubblicazione: (2024)
di: Liu, Haochen, et al.
Pubblicazione: (2024)
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
di: Liu, Weiqi, et al.
Pubblicazione: (2026)
di: Liu, Weiqi, et al.
Pubblicazione: (2026)
Explainable AI in Usable Privacy and Security: Challenges and Opportunities
di: Freiberger, Vincent, et al.
Pubblicazione: (2025)
di: Freiberger, Vincent, et al.
Pubblicazione: (2025)
Transforming Science Learning Materials in the Era of Artificial Intelligence
di: Zhai, Xiaoming, et al.
Pubblicazione: (2026)
di: Zhai, Xiaoming, et al.
Pubblicazione: (2026)
AGI: Artificial General Intelligence for Education
di: Latif, Ehsan, et al.
Pubblicazione: (2023)
di: Latif, Ehsan, et al.
Pubblicazione: (2023)
Towards Trustworthy GUI Agents: A Survey
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
di: Shi, Yucheng, et al.
Pubblicazione: (2025)
Explainable AI Reloaded: Challenging the XAI Status Quo in the Era of Large Language Models
di: Ehsan, Upol, et al.
Pubblicazione: (2024)
di: Ehsan, Upol, et al.
Pubblicazione: (2024)
Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models
di: Lu, Jinliang, et al.
Pubblicazione: (2024)
di: Lu, Jinliang, et al.
Pubblicazione: (2024)
Knowledge Editing for Large Language Models: A Survey
di: Wang, Song, et al.
Pubblicazione: (2023)
di: Wang, Song, et al.
Pubblicazione: (2023)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
di: Shu, Dong, et al.
Pubblicazione: (2024)
di: Shu, Dong, et al.
Pubblicazione: (2024)
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
di: Afroogh, Saleh, et al.
Pubblicazione: (2026)
di: Afroogh, Saleh, et al.
Pubblicazione: (2026)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
di: Li, Daoyang, et al.
Pubblicazione: (2024)
di: Li, Daoyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
di: Wu, Xuansheng, et al.
Pubblicazione: (2025) -
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
di: Wu, Xuansheng, et al.
Pubblicazione: (2025) -
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025) -
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025) -
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)