RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Zhengyang, Dickens, Charles, Pham, Derek, Dsouza, Amanda, Parchami, Armin, Sala, Frederic, Varma, Paroma |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
by: Bauer, Justin, et al.
Published: (2026)
by: Bauer, Justin, et al.
Published: (2026)
Automating Benchmark Design
by: Dsouza, Amanda, et al.
Published: (2025)
by: Dsouza, Amanda, et al.
Published: (2025)
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
by: Dsouza, Amanda, et al.
Published: (2024)
by: Dsouza, Amanda, et al.
Published: (2024)
Benchmarking Agents in Insurance Underwriting Environments
by: Dsouza, Amanda, et al.
Published: (2026)
by: Dsouza, Amanda, et al.
Published: (2026)
Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
by: Vinay, Vaishali
Published: (2025)
by: Vinay, Vaishali
Published: (2025)
TrajPRed: Trajectory Prediction with Region-based Relation Learning
by: Zhou, Chen, et al.
Published: (2024)
by: Zhou, Chen, et al.
Published: (2024)
Automated Computation of Therapies Using Failure Mode and Effects Analysis in the Medical Domain
by: Luttermann, Malte, et al.
Published: (2024)
by: Luttermann, Malte, et al.
Published: (2024)
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
by: Liu, Zehua, et al.
Published: (2026)
by: Liu, Zehua, et al.
Published: (2026)
RIFT: Reordered Instruction Following Testbed To Evaluate Instruction Following in Singular Multistep Prompt Structures
by: Jaffe, Andrew, et al.
Published: (2026)
by: Jaffe, Andrew, et al.
Published: (2026)
Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025
by: Ansari, Samar
Published: (2026)
by: Ansari, Samar
Published: (2026)
RIFT: A Scalable Methodology for LLM Accelerator Fault Assessment using Reinforcement Learning
by: Khalil, Khurram, et al.
Published: (2025)
by: Khalil, Khurram, et al.
Published: (2025)
The Bucolic Mode in Byzantine Art
by: Chatterjee, Paroma
Published: (2026)
by: Chatterjee, Paroma
Published: (2026)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
by: Huang, Tzu-Heng, et al.
Published: (2024)
by: Huang, Tzu-Heng, et al.
Published: (2024)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
by: Wang, Jiayu, et al.
Published: (2025)
by: Wang, Jiayu, et al.
Published: (2025)
Weak-to-Strong Generalization Through the Data-Centric Lens
by: Shin, Changho, et al.
Published: (2024)
by: Shin, Changho, et al.
Published: (2024)
Entropy Collapse: A Universal Failure Mode of Intelligent Systems
by: Khanh, Truong Xuan, et al.
Published: (2025)
by: Khanh, Truong Xuan, et al.
Published: (2025)
Revealing Interpretable Failure Modes of VLMs
by: Chaudhary, Isha, et al.
Published: (2026)
by: Chaudhary, Isha, et al.
Published: (2026)
Automated Bug Report Prioritization in Large Open-Source Projects
by: Pierson, Riley, et al.
Published: (2025)
by: Pierson, Riley, et al.
Published: (2025)
A Nascent Taxonomy of Machine Learning in Intelligent Robotic Process Automation
by: Laakmann, Lukas, et al.
Published: (2025)
by: Laakmann, Lukas, et al.
Published: (2025)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Automated Analysis of Global AI Safety Initiatives: A Taxonomy-Driven LLM Approach
by: Semitsu, Takayuki, et al.
Published: (2026)
by: Semitsu, Takayuki, et al.
Published: (2026)
TrojFlow: Flow Models are Natural Targets for Trojan Attacks
by: Qi, Zhengyang, et al.
Published: (2024)
by: Qi, Zhengyang, et al.
Published: (2024)
Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset
by: Anastasopoulou, Panagiota, et al.
Published: (2024)
by: Anastasopoulou, Panagiota, et al.
Published: (2024)
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
by: Chen, Si, et al.
Published: (2026)
by: Chen, Si, et al.
Published: (2026)
Automated Analysis of Learning Outcomes and Exam Questions Based on Bloom's Taxonomy
by: Kumar, Ramya, et al.
Published: (2025)
by: Kumar, Ramya, et al.
Published: (2025)
Degradation Modeling and Prognostic Analysis Under Unknown Failure Modes
by: Fu, Ying, et al.
Published: (2024)
by: Fu, Ying, et al.
Published: (2024)
From Documents to Database: Failure Modes for Industrial Assets
by: Kabakci-Zorlu, Duygu, et al.
Published: (2025)
by: Kabakci-Zorlu, Duygu, et al.
Published: (2025)
Failure Modes for Deep Learning-Based Online Mapping: How to Measure and Address Them
by: Hubbertz, Michael, et al.
Published: (2026)
by: Hubbertz, Michael, et al.
Published: (2026)
Zero-Shot Robustification of Zero-Shot Models
by: Adila, Dyah, et al.
Published: (2023)
by: Adila, Dyah, et al.
Published: (2023)
Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation
by: Zhao, Jitian, et al.
Published: (2025)
by: Zhao, Jitian, et al.
Published: (2025)
Quantifying Automation Risk in High-Automation AI Systems: A Bayesian Framework for Failure Propagation and Optimal Oversight
by: Srivastava, Vishal, et al.
Published: (2026)
by: Srivastava, Vishal, et al.
Published: (2026)
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
by: Pandey, Mukund
Published: (2026)
by: Pandey, Mukund
Published: (2026)
Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models
by: Kim, Dueun, et al.
Published: (2026)
by: Kim, Dueun, et al.
Published: (2026)
From Failure Modes to Reliability Awareness in Generative and Agentic AI System
by: Janet, et al.
Published: (2025)
by: Janet, et al.
Published: (2025)
Solving Dual Sourcing Problems with Supply Mode Dependent Failure Rates
by: Akkerman, Fabian, et al.
Published: (2024)
by: Akkerman, Fabian, et al.
Published: (2024)
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
by: Feuer, Benjamin, et al.
Published: (2024)
by: Feuer, Benjamin, et al.
Published: (2024)
Studying How to Efficiently and Effectively Guide Models with Explanations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
by: Parchami-Araghi, Amin, et al.
Published: (2024)
by: Parchami-Araghi, Amin, et al.
Published: (2024)
Similar Items
-
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
by: Bauer, Justin, et al.
Published: (2026) -
Automating Benchmark Design
by: Dsouza, Amanda, et al.
Published: (2025) -
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
by: Dsouza, Amanda, et al.
Published: (2024) -
Benchmarking Agents in Insurance Underwriting Environments
by: Dsouza, Amanda, et al.
Published: (2026) -
Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
by: Vinay, Vaishali
Published: (2025)