Dissenting Explanations: Leveraging Disagreement to Reduce Model Overreliance
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Reingold, Omer, Shen, Judy Hanwen, Talati, Aditi |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
par: Sela, Omer
Publié: (2026)
par: Sela, Omer
Publié: (2026)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
par: Khan, Omer Jauhar
Publié: (2025)
par: Khan, Omer Jauhar
Publié: (2025)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
par: Da, Longchao, et autres
Publié: (2025)
par: Da, Longchao, et autres
Publié: (2025)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
par: Mathew, Aby Mammen
Publié: (2026)
par: Mathew, Aby Mammen
Publié: (2026)
A Collaborative Content Moderation Framework for Toxicity Detection based on Conformalized Estimates of Annotation Disagreement
par: Villate-Castillo, Guillermo, et autres
Publié: (2024)
par: Villate-Castillo, Guillermo, et autres
Publié: (2024)
Semantic Modeling for World-Centered Architectures
par: Mantsivoda, Andrei, et autres
Publié: (2026)
par: Mantsivoda, Andrei, et autres
Publié: (2026)
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
par: Wang, Yongjie, et autres
Publié: (2025)
par: Wang, Yongjie, et autres
Publié: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
par: Fan, Jingxing, et autres
Publié: (2025)
par: Fan, Jingxing, et autres
Publié: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
par: Mahale, Ajay Pravin
Publié: (2026)
par: Mahale, Ajay Pravin
Publié: (2026)
Formal Abductive Latent Explanations for Prototype-Based Networks
par: Soria, Jules, et autres
Publié: (2025)
par: Soria, Jules, et autres
Publié: (2025)
STACHE: Local Black-Box Explanations for Reinforcement Learning Policies
par: Elashkin, Andrew, et autres
Publié: (2025)
par: Elashkin, Andrew, et autres
Publié: (2025)
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
par: Teklehaymanot, Hailay Kidu, et autres
Publié: (2025)
par: Teklehaymanot, Hailay Kidu, et autres
Publié: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
par: Jia, Xiao
Publié: (2026)
par: Jia, Xiao
Publié: (2026)
Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior
par: Kim, Dongmin, et autres
Publié: (2025)
par: Kim, Dongmin, et autres
Publié: (2025)
Communication-Efficient and Accurate Approach for Aggregation in Federated Low-Rank Adaptation
par: Nguyen, Le-Tuan, et autres
Publié: (2025)
par: Nguyen, Le-Tuan, et autres
Publié: (2025)
Automatic selection of primary studies in systematic reviews with evolutionary rule-based classification
par: de la Torre-López, José, et autres
Publié: (2025)
par: de la Torre-López, José, et autres
Publié: (2025)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
par: Chahine, Makram, et autres
Publié: (2024)
par: Chahine, Makram, et autres
Publié: (2024)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
par: Parcalabescu, Letitia, et autres
Publié: (2023)
par: Parcalabescu, Letitia, et autres
Publié: (2023)
Harnessing non-adversarial robustness in large language models
par: Zhou, Qinghua, et autres
Publié: (2026)
par: Zhou, Qinghua, et autres
Publié: (2026)
ATEX-CF: Attack-Informed Counterfactual Explanations for Graph Neural Networks
par: Zhang, Yu, et autres
Publié: (2026)
par: Zhang, Yu, et autres
Publié: (2026)
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
par: Wang, Yongjie, et autres
Publié: (2025)
par: Wang, Yongjie, et autres
Publié: (2025)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
par: Consoli, Sergio, et autres
Publié: (2025)
par: Consoli, Sergio, et autres
Publié: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
par: Pepe, Alberto, et autres
Publié: (2026)
par: Pepe, Alberto, et autres
Publié: (2026)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
par: Wang, Yihao, et autres
Publié: (2026)
par: Wang, Yihao, et autres
Publié: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
par: Kim, Heejun, et autres
Publié: (2026)
par: Kim, Heejun, et autres
Publié: (2026)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
par: Sáez, Arnau Igualde, et autres
Publié: (2025)
par: Sáez, Arnau Igualde, et autres
Publié: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
par: Asperti, Andrea, et autres
Publié: (2025)
par: Asperti, Andrea, et autres
Publié: (2025)
AI Agents: Evolution, Architecture, and Real-World Applications
par: Krishnan, Naveen
Publié: (2025)
par: Krishnan, Naveen
Publié: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
par: Zhang, Xue
Publié: (2025)
par: Zhang, Xue
Publié: (2025)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
par: Shekar, Pavan C, et autres
Publié: (2025)
par: Shekar, Pavan C, et autres
Publié: (2025)
Reducing Instability in Synthetic Data Evaluation with a Super-Metric in MalDataGen
par: da Silva, Anna Luiza Gomes, et autres
Publié: (2025)
par: da Silva, Anna Luiza Gomes, et autres
Publié: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
par: Yang, Yibo
Publié: (2025)
par: Yang, Yibo
Publié: (2025)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
par: Alpay, Faruk, et autres
Publié: (2025)
par: Alpay, Faruk, et autres
Publié: (2025)
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery
par: Su, Jiarui, et autres
Publié: (2026)
par: Su, Jiarui, et autres
Publié: (2026)
Scalable Heterogeneous Graph Foundation Models for Data-Driven Optimal Power Flow in Smart Grids
par: Pasini, Massimiliano Lupo, et autres
Publié: (2026)
par: Pasini, Massimiliano Lupo, et autres
Publié: (2026)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
par: Viveiros, André G., et autres
Publié: (2025)
par: Viveiros, André G., et autres
Publié: (2025)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
par: Badshah, Sher, et autres
Publié: (2024)
par: Badshah, Sher, et autres
Publié: (2024)
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
par: Koubaa, Anis, et autres
Publié: (2025)
par: Koubaa, Anis, et autres
Publié: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
par: Tang, Zhengzheng
Publié: (2026)
par: Tang, Zhengzheng
Publié: (2026)
Documents similaires
-
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
par: Sela, Omer
Publié: (2026) -
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
par: Khan, Omer Jauhar
Publié: (2025) -
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
par: Da, Longchao, et autres
Publié: (2025) -
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
par: Mathew, Aby Mammen
Publié: (2026) -
A Collaborative Content Moderation Framework for Toxicity Detection based on Conformalized Estimates of Annotation Disagreement
par: Villate-Castillo, Guillermo, et autres
Publié: (2024)