Quantifying True Robustness: Synonymity-Weighted Similarity for Trustworthy XAI Evaluation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Burger, Christopher |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Individual Facts: Investigating Categorical Knowledge Locality of Taxonomy and Meronomy Concepts in GPT Models
von: Burger, Christopher, et al.
Veröffentlicht: (2024)
von: Burger, Christopher, et al.
Veröffentlicht: (2024)
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
von: Jan, Essa, et al.
Veröffentlicht: (2025)
von: Jan, Essa, et al.
Veröffentlicht: (2025)
A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025)
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025)
Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025)
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
von: Mo, Yichuan, et al.
Veröffentlicht: (2026)
von: Mo, Yichuan, et al.
Veröffentlicht: (2026)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
von: Wang, Dongwei, et al.
Veröffentlicht: (2024)
von: Wang, Dongwei, et al.
Veröffentlicht: (2024)
Rethinking Suicidal Ideation Detection: A Trustworthy Annotation Framework and Cross-Lingual Model Evaluation
von: Dzafic, Amina, et al.
Veröffentlicht: (2025)
von: Dzafic, Amina, et al.
Veröffentlicht: (2025)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
von: Tan, Weihao, et al.
Veröffentlicht: (2024)
von: Tan, Weihao, et al.
Veröffentlicht: (2024)
XAI4LLM. Let Machine Learning Models and LLMs Collaborate for Enhanced In-Context Learning in Healthcare
von: Nazary, Fatemeh, et al.
Veröffentlicht: (2024)
von: Nazary, Fatemeh, et al.
Veröffentlicht: (2024)
LLMs for XAI: Future Directions for Explaining Explanations
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
Harmonic LLMs are Trustworthy
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2024)
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2024)
Trustworthy Summarization via Uncertainty Quantification and Risk Awareness in Large Language Models
von: Pan, Shuaidong, et al.
Veröffentlicht: (2025)
von: Pan, Shuaidong, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
Evaluating the Robustness of Analogical Reasoning in Large Language Models
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
von: Joshi, Abhinav, et al.
Veröffentlicht: (2025)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2025)
Quantifying the Capabilities of LLMs across Scale and Precision
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
von: Sato, Ryoma
Veröffentlicht: (2026)
von: Sato, Ryoma
Veröffentlicht: (2026)
Quantifying Modality Contributions via Disentangling Multimodal Representations
von: Amit, Padegal, et al.
Veröffentlicht: (2025)
von: Amit, Padegal, et al.
Veröffentlicht: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
von: Punjwani, Saif, et al.
Veröffentlicht: (2025)
von: Punjwani, Saif, et al.
Veröffentlicht: (2025)
Mitigate Negative Transfer with Similarity Heuristic Lifelong Prompt Tuning
von: Wu, Chenyuan, et al.
Veröffentlicht: (2024)
von: Wu, Chenyuan, et al.
Veröffentlicht: (2024)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
Optimizing Decoding Paths in Masked Diffusion Models by Quantifying Uncertainty
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
von: Yang, Chao, et al.
Veröffentlicht: (2024)
von: Yang, Chao, et al.
Veröffentlicht: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
von: Chawla, Krrish, et al.
Veröffentlicht: (2025)
von: Chawla, Krrish, et al.
Veröffentlicht: (2025)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
Scalability Matters: Overcoming Challenges in InstructGLM with Similarity-Degree-Based Sampling
von: Lee, Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Hyun, et al.
Veröffentlicht: (2025)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
von: Jeong, Hyejun, et al.
Veröffentlicht: (2024)
von: Jeong, Hyejun, et al.
Veröffentlicht: (2024)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity
von: Burnat, Florian A. D., et al.
Veröffentlicht: (2026)
von: Burnat, Florian A. D., et al.
Veröffentlicht: (2026)
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
House of Cards: Massive Weights in LLMs
von: Oh, Jaehoon, et al.
Veröffentlicht: (2024)
von: Oh, Jaehoon, et al.
Veröffentlicht: (2024)
More Expressive Attention with Negative Weights
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models
von: Min, Zeping, et al.
Veröffentlicht: (2025)
von: Min, Zeping, et al.
Veröffentlicht: (2025)
GCC-Spam: Spam Detection via GAN, Contrastive Learning, and Character Similarity Networks
von: Wang, Zhijie, et al.
Veröffentlicht: (2025)
von: Wang, Zhijie, et al.
Veröffentlicht: (2025)
Task-Aware LoRA Adapter Composition via Similarity Retrieval in Vector Databases
von: Adsul, Riya, et al.
Veröffentlicht: (2026)
von: Adsul, Riya, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond Individual Facts: Investigating Categorical Knowledge Locality of Taxonomy and Meronomy Concepts in GPT Models
von: Burger, Christopher, et al.
Veröffentlicht: (2024) -
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
von: Jan, Essa, et al.
Veröffentlicht: (2025) -
A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025) -
Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2025) -
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
von: Yang, Shiping, et al.
Veröffentlicht: (2025)