Demystifying the Accuracy-Interpretability Trade-Off: A Case Study of Inferring Ratings from Reviews
Fuente:
arXiv
Guardado en:
| Autores principales: | Atrey, Pranjal, Brundage, Michael P., Wu, Min, Dutta, Sanghamitra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
por: Dissanayake, Pasan, et al.
Publicado: (2025)
por: Dissanayake, Pasan, et al.
Publicado: (2025)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
por: Hamman, Faisal, et al.
Publicado: (2025)
por: Hamman, Faisal, et al.
Publicado: (2025)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
por: Ye, Qinyuan, et al.
Publicado: (2025)
por: Ye, Qinyuan, et al.
Publicado: (2025)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
por: Yao, Chaorui, et al.
Publicado: (2025)
por: Yao, Chaorui, et al.
Publicado: (2025)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
por: Kang, Feiyang, et al.
Publicado: (2025)
por: Kang, Feiyang, et al.
Publicado: (2025)
VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
por: Egea, David, et al.
Publicado: (2025)
por: Egea, David, et al.
Publicado: (2025)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
por: Aggarwal, Pranjal, et al.
Publicado: (2025)
por: Aggarwal, Pranjal, et al.
Publicado: (2025)
Demystifying Chains, Trees, and Graphs of Thoughts
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Synthetic Data for Veterinary EHR De-identification: Benefits, Limits, and Safety Trade-offs Under Fixed Compute
por: Brundage, David
Publicado: (2026)
por: Brundage, David
Publicado: (2026)
Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
por: Hamman, Faisal, et al.
Publicado: (2025)
por: Hamman, Faisal, et al.
Publicado: (2025)
Quantifying the Accuracy-Interpretability Trade-Off in Concept-Based Sidechannel Models
por: Debot, David, et al.
Publicado: (2025)
por: Debot, David, et al.
Publicado: (2025)
Exploring Accuracy-Fairness Trade-off in Large Language Models
por: Zhang, Qingquan, et al.
Publicado: (2024)
por: Zhang, Qingquan, et al.
Publicado: (2024)
Demystifying Local and Global Fairness Trade-offs in Federated Learning Using Partial Information Decomposition
por: Hamman, Faisal, et al.
Publicado: (2023)
por: Hamman, Faisal, et al.
Publicado: (2023)
Demystifying the Slash Pattern in Attention: The Role of RoPE
por: Cheng, Yuan, et al.
Publicado: (2026)
por: Cheng, Yuan, et al.
Publicado: (2026)
Demystifying Embedding Spaces using Large Language Models
por: Tennenholtz, Guy, et al.
Publicado: (2023)
por: Tennenholtz, Guy, et al.
Publicado: (2023)
Towards Interpreting Language Models: A Case Study in Multi-Hop Reasoning
por: Sakarvadia, Mansi
Publicado: (2024)
por: Sakarvadia, Mansi
Publicado: (2024)
Infer Human's Intentions Before Following Natural Language Instructions
por: Wan, Yanming, et al.
Publicado: (2024)
por: Wan, Yanming, et al.
Publicado: (2024)
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
por: Habibi, Reza, et al.
Publicado: (2026)
por: Habibi, Reza, et al.
Publicado: (2026)
Demystifying the unreasonable effectiveness of online alignment methods
por: Kang, Enoch Hyunwook
Publicado: (2026)
por: Kang, Enoch Hyunwook
Publicado: (2026)
Agentic-R1: Distilled Dual-Strategy Reasoning
por: Du, Weihua, et al.
Publicado: (2025)
por: Du, Weihua, et al.
Publicado: (2025)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
por: Wang, Shouren, et al.
Publicado: (2025)
por: Wang, Shouren, et al.
Publicado: (2025)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
por: Tian, Yuandong, et al.
Publicado: (2023)
por: Tian, Yuandong, et al.
Publicado: (2023)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
por: Li, Pingzhi, et al.
Publicado: (2023)
por: Li, Pingzhi, et al.
Publicado: (2023)
Can Large Language Models Infer Causation from Correlation?
por: Jin, Zhijing, et al.
Publicado: (2023)
por: Jin, Zhijing, et al.
Publicado: (2023)
A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
por: Chen, Michael K.
Publicado: (2025)
por: Chen, Michael K.
Publicado: (2025)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
por: Khan, Zaid, et al.
Publicado: (2025)
por: Khan, Zaid, et al.
Publicado: (2025)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
por: Huang, Yuzhen, et al.
Publicado: (2025)
por: Huang, Yuzhen, et al.
Publicado: (2025)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
por: Noukhovitch, Michael, et al.
Publicado: (2024)
por: Noukhovitch, Michael, et al.
Publicado: (2024)
Agentless: Demystifying LLM-based Software Engineering Agents
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
por: Cho, Hanseul, et al.
Publicado: (2024)
por: Cho, Hanseul, et al.
Publicado: (2024)
Prompting Strategies for Enabling Large Language Models to Infer Causation from Correlation
por: Sgouritsa, Eleni, et al.
Publicado: (2024)
por: Sgouritsa, Eleni, et al.
Publicado: (2024)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
por: Huang, Xinting, et al.
Publicado: (2025)
por: Huang, Xinting, et al.
Publicado: (2025)
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
por: Setlur, Amrith, et al.
Publicado: (2026)
por: Setlur, Amrith, et al.
Publicado: (2026)
Learning to Reason under Off-Policy Guidance
por: Yan, Jianhao, et al.
Publicado: (2025)
por: Yan, Jianhao, et al.
Publicado: (2025)
Can Large Language Models Infer Causal Relationships from Real-World Text?
por: Saklad, Ryan, et al.
Publicado: (2025)
por: Saklad, Ryan, et al.
Publicado: (2025)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
por: Chen, Justin Chih-Yao, et al.
Publicado: (2025)
por: Chen, Justin Chih-Yao, et al.
Publicado: (2025)
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection
por: Ma, Mingyu Derek, et al.
Publicado: (2025)
por: Ma, Mingyu Derek, et al.
Publicado: (2025)
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics
por: Nair, Jishnu Sethumadhavan, et al.
Publicado: (2026)
por: Nair, Jishnu Sethumadhavan, et al.
Publicado: (2026)
Interpretable Predictability-Based AI Text Detection: A Replication Study
por: Skurla, Adam, et al.
Publicado: (2026)
por: Skurla, Adam, et al.
Publicado: (2026)
Chain-of-Thought Unfaithfulness as Disguised Accuracy
por: Bentham, Oliver, et al.
Publicado: (2024)
por: Bentham, Oliver, et al.
Publicado: (2024)
Ejemplares similares
-
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
por: Dissanayake, Pasan, et al.
Publicado: (2025) -
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
por: Hamman, Faisal, et al.
Publicado: (2025) -
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
por: Ye, Qinyuan, et al.
Publicado: (2025) -
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
por: Yao, Chaorui, et al.
Publicado: (2025) -
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
por: Kang, Feiyang, et al.
Publicado: (2025)