STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Bonagiri, Akash, Anderias, Gerard Janno, Patil, Saee, Lai, Angelina, Borkar, Devang, Kang, Gezheng, Gandhi, Ishant, Rafatirad, Setareh, Homayoun, Houman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
di: Bonagiri, Akash, et al.
Pubblicazione: (2026)
di: Bonagiri, Akash, et al.
Pubblicazione: (2026)
Human-AI Interaction: Evaluating LLM Reasoning on Digital Logic Circuit included Graph Problems, in terms of creativity in design and analysis
di: Thota, Yogeswar Reddy, et al.
Pubblicazione: (2026)
di: Thota, Yogeswar Reddy, et al.
Pubblicazione: (2026)
When Models Ignore Definitions: Measuring Semantic Override Hallucinations in LLM Reasoning
di: Thota, Yogeswar Reddy, et al.
Pubblicazione: (2026)
di: Thota, Yogeswar Reddy, et al.
Pubblicazione: (2026)
Kumo: A Security-Focused Serverless Cloud Simulator
di: Shao, Wei, et al.
Pubblicazione: (2026)
di: Shao, Wei, et al.
Pubblicazione: (2026)
Self-Supervised and Topological Signal-Quality Assessment for Any PPG Device
di: Shao, Wei, et al.
Pubblicazione: (2025)
di: Shao, Wei, et al.
Pubblicazione: (2025)
Lightweight Cross-Device Sleep Tracking on the WeBe Wearable Platform
di: Shao, Wei, et al.
Pubblicazione: (2026)
di: Shao, Wei, et al.
Pubblicazione: (2026)
Rapid Adaptation of SpO2 Estimation to Wearable Devices via Transfer Learning on Low-Sampling-Rate PPG
di: Liang, Zequan, et al.
Pubblicazione: (2025)
di: Liang, Zequan, et al.
Pubblicazione: (2025)
APEX: Attention on Personality based Emotion ReXgnition Framework
di: Fang, Ruijie, et al.
Pubblicazione: (2024)
di: Fang, Ruijie, et al.
Pubblicazione: (2024)
Bit of a Close Talker: A Practical Guide to Serverless Cloud Co-Location Attacks
di: Shao, Wei, et al.
Pubblicazione: (2025)
di: Shao, Wei, et al.
Pubblicazione: (2025)
Generalizable Blood Pressure Estimation from Multi-Wavelength PPG Using Curriculum-Adversarial Learning
di: Liang, Zequan, et al.
Pubblicazione: (2025)
di: Liang, Zequan, et al.
Pubblicazione: (2025)
Know Me by My Pulse: Toward Practical Continuous Authentication on Wearable Devices via Wrist-Worn PPG
di: Shao, Wei, et al.
Pubblicazione: (2025)
di: Shao, Wei, et al.
Pubblicazione: (2025)
A Comprehensive Study of Implicit and Explicit Biases in Large Language Models
di: Kazi, Fatima, et al.
Pubblicazione: (2025)
di: Kazi, Fatima, et al.
Pubblicazione: (2025)
Uniqueness of simultaneous reconstruction of general space- and time-dependent sources and initial states in fractional diffusion equations and systems from single boundary measurements
di: Janno, Jaan
Pubblicazione: (2026)
di: Janno, Jaan
Pubblicazione: (2026)
Inverse problems for a generalized fractional diffusion equation with unknown history
di: Janno, Jaan
Pubblicazione: (2024)
di: Janno, Jaan
Pubblicazione: (2024)
FaRAccel: FPGA-Accelerated Defense Architecture for Efficient Bit-Flip Attack Resilience in Transformer Models
di: Nazari, Najmeh, et al.
Pubblicazione: (2025)
di: Nazari, Najmeh, et al.
Pubblicazione: (2025)
The AI Companion in Education: Analyzing the Pedagogical Potential of ChatGPT in Computer Science and Engineering
di: He, Zhangying, et al.
Pubblicazione: (2024)
di: He, Zhangying, et al.
Pubblicazione: (2024)
FFCL: Forward-Forward Net with Cortical Loops, Training and Inference on Edge Without Backpropagation
di: Karkehabadi, Ali, et al.
Pubblicazione: (2024)
di: Karkehabadi, Ali, et al.
Pubblicazione: (2024)
Engineering Practical Succinct Bit Vectors: A Space-Time Pareto Analysis on Apple Silicon ARM64 Cores
di: Garg, Ishant
Pubblicazione: (2026)
di: Garg, Ishant
Pubblicazione: (2026)
Inverse problem to determine simultaneously several scalar parameters and a time-dependent source term in a superdiffusion equation involving a multiterm fractional Laplacian
di: Gerges, Hany, et al.
Pubblicazione: (2025)
di: Gerges, Hany, et al.
Pubblicazione: (2025)
SaliencyDecor: Enhancing Neural Network Interpretability through Feature Decorrelation
di: Karkehabadi, Ali, et al.
Pubblicazione: (2026)
di: Karkehabadi, Ali, et al.
Pubblicazione: (2026)
Controllable Hand Grasp Generation for HOI and Efficient Evaluation Methods
di: Ishant, et al.
Pubblicazione: (2025)
di: Ishant, et al.
Pubblicazione: (2025)
Role of Intermetallics in Creep Behavior of Squeezed Cast Ca‐ and Sr‐Modified AZ91 Magnesium Alloy
di: Hitesh Patil, et al.
Pubblicazione: (2024)
di: Hitesh Patil, et al.
Pubblicazione: (2024)
The Value of Disagreement in AI Design, Evaluation, and Alignment
di: Fazelpour, Sina, et al.
Pubblicazione: (2025)
di: Fazelpour, Sina, et al.
Pubblicazione: (2025)
HW-V2W-Map: Hardware Vulnerability to Weakness Mapping Framework for Root Cause Analysis with GPT-assisted Mitigation Suggestion
di: Lin, Yu-Zheng, et al.
Pubblicazione: (2023)
di: Lin, Yu-Zheng, et al.
Pubblicazione: (2023)
DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge
di: Mohanty, Asmita, et al.
Pubblicazione: (2025)
di: Mohanty, Asmita, et al.
Pubblicazione: (2025)
A Benchmark for Procedural Memory Retrieval in Language Agents
di: Kohar, Ishant, et al.
Pubblicazione: (2025)
di: Kohar, Ishant, et al.
Pubblicazione: (2025)
Generative AI-Based Effective Malware Detection for Embedded Computing Systems
di: Kasarapu, Sreenitha, et al.
Pubblicazione: (2024)
di: Kasarapu, Sreenitha, et al.
Pubblicazione: (2024)
NLP4Gov: A Comprehensive Library for Computational Policy Analysis
di: Chakraborti, Mahasweta, et al.
Pubblicazione: (2024)
di: Chakraborti, Mahasweta, et al.
Pubblicazione: (2024)
Towards unbiased recovery of cosmic filament properties: the role of spine curvature and optimized smoothing
di: Dhawalikar, Saee, et al.
Pubblicazione: (2024)
di: Dhawalikar, Saee, et al.
Pubblicazione: (2024)
Evaluating Male Fetal External Genital Morphology
di: Nikit Venishetty, et al.
Pubblicazione: (2025)
di: Nikit Venishetty, et al.
Pubblicazione: (2025)
Impact of LLMs on Team Collaboration in Software Development
di: Dhanuka, Devang
Pubblicazione: (2025)
di: Dhanuka, Devang
Pubblicazione: (2025)
Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging
di: Latibari, Banafsheh Saber, et al.
Pubblicazione: (2025)
di: Latibari, Banafsheh Saber, et al.
Pubblicazione: (2025)
Fuzzing BusyBox: Leveraging LLM and Crash Reuse for Embedded Bug Unearthing
di: Asmita, et al.
Pubblicazione: (2024)
di: Asmita, et al.
Pubblicazione: (2024)
AI INTEGRATED E LEARNING PLATFORM FOR PERSONALIZED LEARNING, ADAPTIVE ASSESSMENT, REAL-TIME ANALYTICS
di: Ishant Chauhan, Gaurav Kumar, Kalpendra Kumar and Dheeraj Kumar Singh
Pubblicazione: (2026)
di: Ishant Chauhan, Gaurav Kumar, Kalpendra Kumar and Dheeraj Kumar Singh
Pubblicazione: (2026)
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
di: Bonagiri, Akash, et al.
Pubblicazione: (2025)
di: Bonagiri, Akash, et al.
Pubblicazione: (2025)
Advanced Energy-Efficient System for Precision Electrodermal Activity Monitoring in Stress Detection
di: Zhang, Ruoyu, et al.
Pubblicazione: (2024)
di: Zhang, Ruoyu, et al.
Pubblicazione: (2024)
DualBind: A Dual-Loss Framework for Protein-Ligand Binding Affinity Prediction
di: Liu, Meng, et al.
Pubblicazione: (2024)
di: Liu, Meng, et al.
Pubblicazione: (2024)
Mechanical and Durability Performance of Recycled Coarse Aggregate Concrete Modified with Class-F Fly Ash for M50-Grade Pavement Applications
di: Suresh M. Patil, Anushka R. Borkar, Devanshu N. Choudhary
Pubblicazione: (2026)
di: Suresh M. Patil, Anushka R. Borkar, Devanshu N. Choudhary
Pubblicazione: (2026)
SoK: The Attack Surface of Agentic AI -- Tools, and Autonomy
di: Dehghantanha, Ali, et al.
Pubblicazione: (2026)
di: Dehghantanha, Ali, et al.
Pubblicazione: (2026)
Stochastic Approximation with Two Time Scales: The General Case
di: Borkar, Vivek S
Pubblicazione: (2024)
di: Borkar, Vivek S
Pubblicazione: (2024)
Documenti analoghi
-
CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
di: Bonagiri, Akash, et al.
Pubblicazione: (2026) -
Human-AI Interaction: Evaluating LLM Reasoning on Digital Logic Circuit included Graph Problems, in terms of creativity in design and analysis
di: Thota, Yogeswar Reddy, et al.
Pubblicazione: (2026) -
When Models Ignore Definitions: Measuring Semantic Override Hallucinations in LLM Reasoning
di: Thota, Yogeswar Reddy, et al.
Pubblicazione: (2026) -
Kumo: A Security-Focused Serverless Cloud Simulator
di: Shao, Wei, et al.
Pubblicazione: (2026) -
Self-Supervised and Topological Signal-Quality Assessment for Any PPG Device
di: Shao, Wei, et al.
Pubblicazione: (2025)