Interpretability-Guided Bi-objective Optimization: Aligning Accuracy and Explainability
Fuente:
arXiv
Salvato in:
| Autori principali: | Fouladi, Kasra, Rahmani, Hamta |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Axiomatic Approach to General Intelligence: SANC(E3) -- Self-organizing Active Network of Concepts with Energy E3
di: Kwon, Daesuk, et al.
Pubblicazione: (2026)
di: Kwon, Daesuk, et al.
Pubblicazione: (2026)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
di: Singha, Disha
Pubblicazione: (2026)
di: Singha, Disha
Pubblicazione: (2026)
Graceful task adaptation with a bi-hemispheric RL agent
di: Nicholas, Grant, et al.
Pubblicazione: (2024)
di: Nicholas, Grant, et al.
Pubblicazione: (2024)
Simulation-Driven Railway Delay Prediction: An Imitation Learning Approach
di: Elliker, Clément, et al.
Pubblicazione: (2025)
di: Elliker, Clément, et al.
Pubblicazione: (2025)
The ARC of Progress towards AGI: A Living Survey of Abstraction and Reasoning
di: Vahdati, Sahar, et al.
Pubblicazione: (2026)
di: Vahdati, Sahar, et al.
Pubblicazione: (2026)
Deep Reinforcement Learning for Adverse Garage Scenario Generation
di: Li, Kai
Pubblicazione: (2024)
di: Li, Kai
Pubblicazione: (2024)
Intervention Complexity as a Canonical Reward and a Measure of Intelligence
di: McCane, Brendan
Pubblicazione: (2026)
di: McCane, Brendan
Pubblicazione: (2026)
LaPro-DTA: Latent Dual-View Drug Representations and Salient Protein Feature Extraction for Generalizable Drug--Target Affinity Prediction
di: Dun, Zihan, et al.
Pubblicazione: (2026)
di: Dun, Zihan, et al.
Pubblicazione: (2026)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
di: Kim, Sejin, et al.
Pubblicazione: (2025)
di: Kim, Sejin, et al.
Pubblicazione: (2025)
Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents
di: Minai, Ali A.
Pubblicazione: (2025)
di: Minai, Ali A.
Pubblicazione: (2025)
Beyond Mimicry: Preference Coherence in LLMs
di: Mikaelson, Luhan, et al.
Pubblicazione: (2025)
di: Mikaelson, Luhan, et al.
Pubblicazione: (2025)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
di: Ged, François, et al.
Pubblicazione: (2023)
di: Ged, François, et al.
Pubblicazione: (2023)
Training Language Models to Win Debates with Self-Play Improves Judge Accuracy
di: Arnesen, Samuel, et al.
Pubblicazione: (2024)
di: Arnesen, Samuel, et al.
Pubblicazione: (2024)
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
di: Lin, Zezheng, et al.
Pubblicazione: (2026)
di: Lin, Zezheng, et al.
Pubblicazione: (2026)
Learning Can Converge Stably to the Wrong Belief under Latent Reliability
di: Zhang, Zhipeng, et al.
Pubblicazione: (2026)
di: Zhang, Zhipeng, et al.
Pubblicazione: (2026)
Feel-Good Thompson Sampling for Contextual Bandits: a Markov Chain Monte Carlo Showdown
di: Anand, Emile, et al.
Pubblicazione: (2025)
di: Anand, Emile, et al.
Pubblicazione: (2025)
Implicit Counterfactual Data Augmentation for Robust Learning
di: Zhou, Xiaoling, et al.
Pubblicazione: (2023)
di: Zhou, Xiaoling, et al.
Pubblicazione: (2023)
Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates
di: Kaplanski, Pawel
Pubblicazione: (2026)
di: Kaplanski, Pawel
Pubblicazione: (2026)
Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
di: Ackermann, Richard, et al.
Pubblicazione: (2025)
di: Ackermann, Richard, et al.
Pubblicazione: (2025)
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
di: Vegner, Ivan, et al.
Pubblicazione: (2025)
di: Vegner, Ivan, et al.
Pubblicazione: (2025)
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
di: Baxi, Rahul
Pubblicazione: (2025)
di: Baxi, Rahul
Pubblicazione: (2025)
Choosing DAG Models Using Markov and Minimal Edge Count in the Absence of Ground Truth
di: Ramsey, Joseph D., et al.
Pubblicazione: (2024)
di: Ramsey, Joseph D., et al.
Pubblicazione: (2024)
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
di: Wang, Xiaoxiao, et al.
Pubblicazione: (2026)
di: Wang, Xiaoxiao, et al.
Pubblicazione: (2026)
FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
di: Ma, Jingxiao, et al.
Pubblicazione: (2025)
di: Ma, Jingxiao, et al.
Pubblicazione: (2025)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
di: Hooshyar, Danial, et al.
Pubblicazione: (2023)
di: Hooshyar, Danial, et al.
Pubblicazione: (2023)
Study Design and Demystification of Physics Informed Neural Networks for Power Flow Simulation
di: Leyli-abadi, Milad, et al.
Pubblicazione: (2025)
di: Leyli-abadi, Milad, et al.
Pubblicazione: (2025)
Prompt Readiness Levels (PRL): a maturity scale and scoring framework for production grade prompt assets
di: Guinard, Sebastien
Pubblicazione: (2026)
di: Guinard, Sebastien
Pubblicazione: (2026)
Improving Fairness with Ensemble Combination: Margin-Dependent Bounds
di: Bian, Yijun
Pubblicazione: (2023)
di: Bian, Yijun
Pubblicazione: (2023)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
di: Shafieinejad, Masoumeh, et al.
Pubblicazione: (2026)
di: Shafieinejad, Masoumeh, et al.
Pubblicazione: (2026)
Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence
di: Ren, Wanying, et al.
Pubblicazione: (2026)
di: Ren, Wanying, et al.
Pubblicazione: (2026)
The Parameters of Educability
di: Valiant, Leslie G.
Pubblicazione: (2024)
di: Valiant, Leslie G.
Pubblicazione: (2024)
Reasoning Beyond the Obvious: Evaluating Divergent and Convergent Thinking in LLMs for Financial Scenarios
di: Bok, Zhuang Qiang, et al.
Pubblicazione: (2025)
di: Bok, Zhuang Qiang, et al.
Pubblicazione: (2025)
Tensor Generalized Approximate Message Passing
di: Li, Yinchuan, et al.
Pubblicazione: (2025)
di: Li, Yinchuan, et al.
Pubblicazione: (2025)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
di: Sarkar, Nilesh, et al.
Pubblicazione: (2026)
di: Sarkar, Nilesh, et al.
Pubblicazione: (2026)
The Efficiency Attenuation Phenomenon: A Computational Challenge to the Language of Thought Hypothesis
di: Zhang, Di
Pubblicazione: (2026)
di: Zhang, Di
Pubblicazione: (2026)
What Teaches Robots to Walk, Teaches Them to Trade too -- Regime Adaptive Execution using Informed Data and LLMs
di: Saqur, Raeid
Pubblicazione: (2024)
di: Saqur, Raeid
Pubblicazione: (2024)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
di: Khandelwal, Vedant, et al.
Pubblicazione: (2024)
di: Khandelwal, Vedant, et al.
Pubblicazione: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
di: Hedar, Abdel-Rahman, et al.
Pubblicazione: (2024)
di: Hedar, Abdel-Rahman, et al.
Pubblicazione: (2024)
A survey of air combat behavior modeling using machine learning
di: Gorton, Patrick Ribu, et al.
Pubblicazione: (2024)
di: Gorton, Patrick Ribu, et al.
Pubblicazione: (2024)
Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
di: Nayak, Nikhil, et al.
Pubblicazione: (2026)
di: Nayak, Nikhil, et al.
Pubblicazione: (2026)
Documenti analoghi
-
An Axiomatic Approach to General Intelligence: SANC(E3) -- Self-organizing Active Network of Concepts with Energy E3
di: Kwon, Daesuk, et al.
Pubblicazione: (2026) -
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
di: Singha, Disha
Pubblicazione: (2026) -
Graceful task adaptation with a bi-hemispheric RL agent
di: Nicholas, Grant, et al.
Pubblicazione: (2024) -
Simulation-Driven Railway Delay Prediction: An Imitation Learning Approach
di: Elliker, Clément, et al.
Pubblicazione: (2025) -
The ARC of Progress towards AGI: A Living Survey of Abstraction and Reasoning
di: Vahdati, Sahar, et al.
Pubblicazione: (2026)