Multi-Level Explanations for Generative Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Paes, Lucas Monteiro, Wei, Dennis, Do, Hyo Jin, Strobelt, Hendrik, Luss, Ronny, Dhurandhar, Amit, Nagireddy, Manish, Ramamurthy, Karthikeyan Natesan, Sattigeri, Prasanna, Geyer, Werner, Ghosh, Soumya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Stability meets Sufficiency: Informative Explanations that do not Overwhelm
di: Luss, Ronny, et al.
Pubblicazione: (2021)
di: Luss, Ronny, et al.
Pubblicazione: (2021)
Trust Regions for Explanations via Black-Box Probabilistic Certification
di: Dhurandhar, Amit, et al.
Pubblicazione: (2024)
di: Dhurandhar, Amit, et al.
Pubblicazione: (2024)
ICX360: In-Context eXplainability 360 Toolkit
di: Wei, Dennis, et al.
Pubblicazione: (2025)
di: Wei, Dennis, et al.
Pubblicazione: (2025)
Value Alignment from Unstructured Text
di: Padhi, Inkit, et al.
Pubblicazione: (2024)
di: Padhi, Inkit, et al.
Pubblicazione: (2024)
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
di: Nagireddy, Manish, et al.
Pubblicazione: (2024)
di: Nagireddy, Manish, et al.
Pubblicazione: (2024)
CELL your Model: Contrastive Explanations for Large Language Models
di: Luss, Ronny, et al.
Pubblicazione: (2024)
di: Luss, Ronny, et al.
Pubblicazione: (2024)
Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods
di: Wei, Dennis, et al.
Pubblicazione: (2024)
di: Wei, Dennis, et al.
Pubblicazione: (2024)
Identifying Sub-networks in Neural Networks via Functionally Similar Representations
di: Gao, Tian, et al.
Pubblicazione: (2024)
di: Gao, Tian, et al.
Pubblicazione: (2024)
Programming Refusal with Conditional Activation Steering
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
Large Language Model Confidence Estimation via Black-Box Access
di: Pedapati, Tejaswini, et al.
Pubblicazione: (2024)
di: Pedapati, Tejaswini, et al.
Pubblicazione: (2024)
LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems
di: Asif, Sadia, et al.
Pubblicazione: (2026)
di: Asif, Sadia, et al.
Pubblicazione: (2026)
Ranking Large Language Models without Ground Truth
di: Dhurandhar, Amit, et al.
Pubblicazione: (2024)
di: Dhurandhar, Amit, et al.
Pubblicazione: (2024)
Sparsity May Be All You Need: Sparse Random Parameter Adaptation
di: Rios, Jesus, et al.
Pubblicazione: (2025)
di: Rios, Jesus, et al.
Pubblicazione: (2025)
Agentic AI Needs a Systems Theory
di: Miehling, Erik, et al.
Pubblicazione: (2025)
di: Miehling, Erik, et al.
Pubblicazione: (2025)
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
di: Gourabathina, Abinitha, et al.
Pubblicazione: (2026)
di: Gourabathina, Abinitha, et al.
Pubblicazione: (2026)
Contextual Moral Value Alignment Through Context-Based Aggregation
di: Dognin, Pierre, et al.
Pubblicazione: (2024)
di: Dognin, Pierre, et al.
Pubblicazione: (2024)
AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows
di: Ngong, Ivoline C., et al.
Pubblicazione: (2026)
di: Ngong, Ivoline C., et al.
Pubblicazione: (2026)
Mitigating Misalignment Contagion by Steering with Implicit Traits
di: Chang, Maria, et al.
Pubblicazione: (2026)
di: Chang, Maria, et al.
Pubblicazione: (2026)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
di: Ngong, Ivoline, et al.
Pubblicazione: (2025)
di: Ngong, Ivoline, et al.
Pubblicazione: (2025)
Facilitating Human-LLM Collaboration through Factuality Scores and Source Attributions
di: Do, Hyo Jin, et al.
Pubblicazione: (2024)
di: Do, Hyo Jin, et al.
Pubblicazione: (2024)
Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
di: Do, Hyo Jin, et al.
Pubblicazione: (2025)
di: Do, Hyo Jin, et al.
Pubblicazione: (2025)
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
di: Villa, Danielle, et al.
Pubblicazione: (2025)
di: Villa, Danielle, et al.
Pubblicazione: (2025)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
di: Miehling, Erik, et al.
Pubblicazione: (2024)
di: Miehling, Erik, et al.
Pubblicazione: (2024)
NeuroPrune: A Neuro-inspired Topological Sparse Training Algorithm for Large Language Models
di: Dhurandhar, Amit, et al.
Pubblicazione: (2024)
di: Dhurandhar, Amit, et al.
Pubblicazione: (2024)
Selective Explanations
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2024)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2024)
Reasoning about concepts with LLMs: Inconsistencies abound
di: Uceda-Sosa, Rosario, et al.
Pubblicazione: (2024)
di: Uceda-Sosa, Rosario, et al.
Pubblicazione: (2024)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
di: Mozannar, Hussein, et al.
Pubblicazione: (2024)
di: Mozannar, Hussein, et al.
Pubblicazione: (2024)
Thermometer: Towards Universal Calibration for Large Language Models
di: Shen, Maohao, et al.
Pubblicazione: (2024)
di: Shen, Maohao, et al.
Pubblicazione: (2024)
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
di: Achintalwar, Swapnaja, et al.
Pubblicazione: (2024)
di: Achintalwar, Swapnaja, et al.
Pubblicazione: (2024)
Active Sequential Two-Sample Testing
di: Li, Weizhi, et al.
Pubblicazione: (2023)
di: Li, Weizhi, et al.
Pubblicazione: (2023)
Attribute Graphs Underlying Molecular Generative Models: Path to Learning with Limited Data
di: Hoffman, Samuel C., et al.
Pubblicazione: (2022)
di: Hoffman, Samuel C., et al.
Pubblicazione: (2022)
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
di: Do, Hyo Jin, et al.
Pubblicazione: (2025)
di: Do, Hyo Jin, et al.
Pubblicazione: (2025)
Language Models Coupled with Metacognition Can Outperform Reasoning Models
di: Khandelwal, Vedant, et al.
Pubblicazione: (2025)
di: Khandelwal, Vedant, et al.
Pubblicazione: (2025)
Are Uncertainty Quantification Capabilities of Evidential Deep Learning a Mirage?
di: Shen, Maohao, et al.
Pubblicazione: (2024)
di: Shen, Maohao, et al.
Pubblicazione: (2024)
Causal Bandits with General Causal Models and Interventions
di: Yan, Zirui, et al.
Pubblicazione: (2024)
di: Yan, Zirui, et al.
Pubblicazione: (2024)
Interventional Causal Discovery in a Mixture of DAGs
di: Varıcı, Burak, et al.
Pubblicazione: (2024)
di: Varıcı, Burak, et al.
Pubblicazione: (2024)
Evaluating the Prompt Steerability of Large Language Models
di: Miehling, Erik, et al.
Pubblicazione: (2024)
di: Miehling, Erik, et al.
Pubblicazione: (2024)
Improved binary black hole searches through better discrimination against noise transients
di: Choudhary, Sunil, et al.
Pubblicazione: (2022)
di: Choudhary, Sunil, et al.
Pubblicazione: (2022)
Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
di: Boggust, Angie, et al.
Pubblicazione: (2024)
di: Boggust, Angie, et al.
Pubblicazione: (2024)
GPT-2 Through the Lens of Vector Symbolic Architectures
di: Knittel, Johannes, et al.
Pubblicazione: (2024)
di: Knittel, Johannes, et al.
Pubblicazione: (2024)
Documenti analoghi
-
When Stability meets Sufficiency: Informative Explanations that do not Overwhelm
di: Luss, Ronny, et al.
Pubblicazione: (2021) -
Trust Regions for Explanations via Black-Box Probabilistic Certification
di: Dhurandhar, Amit, et al.
Pubblicazione: (2024) -
ICX360: In-Context eXplainability 360 Toolkit
di: Wei, Dennis, et al.
Pubblicazione: (2025) -
Value Alignment from Unstructured Text
di: Padhi, Inkit, et al.
Pubblicazione: (2024) -
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
di: Nagireddy, Manish, et al.
Pubblicazione: (2024)