Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Shahnazari, Kourosh, Ayyoubzadeh, Seyed Moein, Keshtparvar, Mohammadali, Ghaffari, Pegah |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Dynamic Atlas of Persian Poetic Symbolism: Families, Fields, and the Historical Rewiring of Meaning
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
Eigenmood Space: Uncertainty-Aware Spectral Graph Analysis of Psychological Patterns in Classical Persian Poetry
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
NAZM: Network Analysis of Zonal Metrics in Persian Poetic Tradition
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
BASIL: Best-Action Symbolic Interpretable Learning for Evolving Compact RL Policies
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
PARSI: Persian Authorship Recognition via Stylometric Integration
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
Between Century and Poet: Graph-Based Lexical Semantic Change in Persian Poetry
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
A Novel Nearest Neighbors Algorithm Based on Power Muirhead Mean
por: Shahnazari, Kourosh, et al.
Publicado: (2022)
por: Shahnazari, Kourosh, et al.
Publicado: (2022)
Echoes Across Centuries: Phonetic Signatures of Persian Poets
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
por: Shahnazari, Kourosh, et al.
Publicado: (2026)
Who Are You Behind the Screen? Implicit MBTI and Gender Detection Using Artificial Intelligence
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
Word Sense Disambiguation in Persian: Can AI Finally Get It Right?
por: Ayyoubzadeh, Seyed Moein, et al.
Publicado: (2024)
por: Ayyoubzadeh, Seyed Moein, et al.
Publicado: (2024)
Designing DSIC Mechanisms for Data Sharing in the Era of Large Language Models
por: Ayyoubzadeh, Seyed Moein, et al.
Publicado: (2025)
por: Ayyoubzadeh, Seyed Moein, et al.
Publicado: (2025)
Can Invisible Psychological Traits Organize Visible Network Structure? A Complex Network Analysis of Myers-Briggs Type Indicator-Based Interaction Patterns in Anonymous Social Networks
por: Ayyoubzadeh, Seyed Moein, et al.
Publicado: (2025)
por: Ayyoubzadeh, Seyed Moein, et al.
Publicado: (2025)
Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback
por: Hosseini, Seyed Amir, et al.
Publicado: (2026)
por: Hosseini, Seyed Amir, et al.
Publicado: (2026)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
por: Zhou, Jin Peng, et al.
Publicado: (2024)
por: Zhou, Jin Peng, et al.
Publicado: (2024)
Benchmarking is Broken -- Don't Let AI be its Own Judge
por: Cheng, Zerui, et al.
Publicado: (2025)
por: Cheng, Zerui, et al.
Publicado: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
por: Bajpai, Divya Jyoti, et al.
Publicado: (2025)
por: Bajpai, Divya Jyoti, et al.
Publicado: (2025)
Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
por: Lade, Ankit Hemant, et al.
Publicado: (2026)
por: Lade, Ankit Hemant, et al.
Publicado: (2026)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
por: Zhang, Jiefu, et al.
Publicado: (2026)
por: Zhang, Jiefu, et al.
Publicado: (2026)
Intelligent Reservoir Decision Support: An Integrated Framework Combining Large Language Models, Advanced Prompt Engineering, and Multimodal Data Fusion for Real-Time Petroleum Operations
por: Mahjour, Seyed Kourosh, et al.
Publicado: (2025)
por: Mahjour, Seyed Kourosh, et al.
Publicado: (2025)
Transformers Don't In-Context Learn Least Squares Regression
por: Hill, Joshua, et al.
Publicado: (2025)
por: Hill, Joshua, et al.
Publicado: (2025)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
por: Park, Young-Jin, et al.
Publicado: (2025)
por: Park, Young-Jin, et al.
Publicado: (2025)
Don't stop me now: Rethinking Validation Criteria for Model Parameter Selection
por: Apicella, Andrea, et al.
Publicado: (2026)
por: Apicella, Andrea, et al.
Publicado: (2026)
Reasoning Models Don't Always Say What They Think
por: Chen, Yanda, et al.
Publicado: (2025)
por: Chen, Yanda, et al.
Publicado: (2025)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
por: Khoriaty, Matthew, et al.
Publicado: (2025)
por: Khoriaty, Matthew, et al.
Publicado: (2025)
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
por: Kim, Hyunseung, et al.
Publicado: (2024)
por: Kim, Hyunseung, et al.
Publicado: (2024)
Don't Waste Your Time: Early Stopping Cross-Validation
por: Bergman, Edward, et al.
Publicado: (2024)
por: Bergman, Edward, et al.
Publicado: (2024)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
por: Wang, Qinsi, et al.
Publicado: (2025)
por: Wang, Qinsi, et al.
Publicado: (2025)
Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
por: Kazoom, Roie, et al.
Publicado: (2025)
por: Kazoom, Roie, et al.
Publicado: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
por: Dey, Nolan, et al.
Publicado: (2025)
por: Dey, Nolan, et al.
Publicado: (2025)
xAI-Drop: Don't Use What You Cannot Explain
por: De Luca, Vincenzo Marco, et al.
Publicado: (2024)
por: De Luca, Vincenzo Marco, et al.
Publicado: (2024)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
por: Hariri, Mohsen, et al.
Publicado: (2025)
por: Hariri, Mohsen, et al.
Publicado: (2025)
LLM Cyber Evaluations Don't Capture Real-World Risk
por: Lukošiūtė, Kamilė, et al.
Publicado: (2025)
por: Lukošiūtė, Kamilė, et al.
Publicado: (2025)
Large Language Models Must Be Taught to Know What They Don't Know
por: Kapoor, Sanyam, et al.
Publicado: (2024)
por: Kapoor, Sanyam, et al.
Publicado: (2024)
Don't throw the baby out with the bathwater: How and why deep learning for ARC
por: Cole, Jack, et al.
Publicado: (2025)
por: Cole, Jack, et al.
Publicado: (2025)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
por: Yang, Yongyi, et al.
Publicado: (2026)
por: Yang, Yongyi, et al.
Publicado: (2026)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
por: Plyusov, Daniil, et al.
Publicado: (2026)
por: Plyusov, Daniil, et al.
Publicado: (2026)
Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models
por: Zeng, Qingyuan, et al.
Publicado: (2026)
por: Zeng, Qingyuan, et al.
Publicado: (2026)
Don't Play Favorites: Minority Guidance for Diffusion Models
por: Um, Soobin, et al.
Publicado: (2023)
por: Um, Soobin, et al.
Publicado: (2023)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
por: Leemann, Tobias, et al.
Publicado: (2024)
por: Leemann, Tobias, et al.
Publicado: (2024)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
por: Kachaev, Nikita, et al.
Publicado: (2025)
por: Kachaev, Nikita, et al.
Publicado: (2025)
Ejemplares similares
-
A Dynamic Atlas of Persian Poetic Symbolism: Families, Fields, and the Historical Rewiring of Meaning
por: Shahnazari, Kourosh, et al.
Publicado: (2026) -
Eigenmood Space: Uncertainty-Aware Spectral Graph Analysis of Psychological Patterns in Classical Persian Poetry
por: Shahnazari, Kourosh, et al.
Publicado: (2026) -
NAZM: Network Analysis of Zonal Metrics in Persian Poetic Tradition
por: Shahnazari, Kourosh, et al.
Publicado: (2025) -
BASIL: Best-Action Symbolic Interpretable Learning for Evolving Compact RL Policies
por: Shahnazari, Kourosh, et al.
Publicado: (2025) -
PARSI: Persian Authorship Recognition via Stylometric Integration
por: Shahnazari, Kourosh, et al.
Publicado: (2025)