AI Transparency Atlas: Framework, Scoring, and Real-Time Model Card Evaluation Pipeline
Fuente:
arXiv
Saved in:
| Main Authors: | Mamirov, Akhmadillo, Azmain, Faiaz, Wang, Hanyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling
by: Mamirov, Akhmadillo
Published: (2025)
by: Mamirov, Akhmadillo
Published: (2025)
P4OMP: Retrieval-Augmented Prompting for OpenMP Parallelism in Serial Code
by: Abdullah, Wali Mohammad, et al.
Published: (2025)
by: Abdullah, Wali Mohammad, et al.
Published: (2025)
Sentiment Analysis in Software Engineering: Evaluating Generative Pre-trained Transformers
by: Saifullah, KM Khalid, et al.
Published: (2025)
by: Saifullah, KM Khalid, et al.
Published: (2025)
LoCoML: A Framework for Real-World ML Inference Pipelines
by: Maddireddy, Kritin, et al.
Published: (2025)
by: Maddireddy, Kritin, et al.
Published: (2025)
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs
by: Kabir, Azmain, et al.
Published: (2024)
by: Kabir, Azmain, et al.
Published: (2024)
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
by: Schwartz, Reva, et al.
Published: (2026)
by: Schwartz, Reva, et al.
Published: (2026)
What's documented in AI? Systematic Analysis of 32K AI Model Cards
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
Future of Code with Generative AI: Transparency and Safety in the Era of AI Generated Software
by: Hanson, David
Published: (2025)
by: Hanson, David
Published: (2025)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026)
by: Bandi, Chaithanya, et al.
Published: (2026)
AI-Augmented CI/CD Pipelines: From Code Commit to Production with Autonomous Decisions
by: Baqar, Mohammad, et al.
Published: (2025)
by: Baqar, Mohammad, et al.
Published: (2025)
Digital Twins & ZeroConf AI: Structuring Automated Intelligent Pipelines for Industrial Applications
by: Picone, Marco, et al.
Published: (2026)
by: Picone, Marco, et al.
Published: (2026)
EGI: A Multimodal Emotional AI Framework for Enhancing Scrum Master Real-time Self-Awareness
by: Huang, Jingni, et al.
Published: (2026)
by: Huang, Jingni, et al.
Published: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
MORTAR: A Model-based Runtime Action Repair Framework for AI-enabled Cyber-Physical Systems
by: Wang, Renzhi, et al.
Published: (2024)
by: Wang, Renzhi, et al.
Published: (2024)
From Queries to Insights: Agentic LLM Pipelines for Spatio-Temporal Text-to-SQL
by: Redd, Manu, et al.
Published: (2025)
by: Redd, Manu, et al.
Published: (2025)
AIBoMGen: Generating an AI Bill of Materials for Secure, Transparent, and Compliant Model Training
by: Vandendriessche, Wiebe, et al.
Published: (2026)
by: Vandendriessche, Wiebe, et al.
Published: (2026)
ICE-Score: Instructing Large Language Models to Evaluate Code
by: Zhuo, Terry Yue
Published: (2023)
by: Zhuo, Terry Yue
Published: (2023)
DREAM: Debugging and Repairing AutoML Pipelines
by: Zhang, Xiaoyu, et al.
Published: (2023)
by: Zhang, Xiaoyu, et al.
Published: (2023)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
EvoClaw: Evaluating AI Agents on Continuous Software Evolution
by: Deng, Gangda, et al.
Published: (2026)
by: Deng, Gangda, et al.
Published: (2026)
SmartMLOps Studio: Design of an LLM-Integrated IDE with Automated MLOps Pipelines for Model Development and Monitoring
by: Jin, Jiawei, et al.
Published: (2025)
by: Jin, Jiawei, et al.
Published: (2025)
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
A Defect Classification Framework for AI-Based Software Systems (AI-ODC)
by: Alannsary, Mohammed O.
Published: (2025)
by: Alannsary, Mohammed O.
Published: (2025)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
by: Nunes, Henrique, et al.
Published: (2025)
by: Nunes, Henrique, et al.
Published: (2025)
Traceability and Accountability in Role-Specialized Multi-Agent LLM Pipelines
by: Barrak, Amine
Published: (2025)
by: Barrak, Amine
Published: (2025)
VulnLLMEval: A Framework for Evaluating Large Language Models in Software Vulnerability Detection and Patching
by: Zibaeirad, Arastoo, et al.
Published: (2024)
by: Zibaeirad, Arastoo, et al.
Published: (2024)
Opus: A Quantitative Framework for Workflow Evaluation
by: Seroul, Alan, et al.
Published: (2025)
by: Seroul, Alan, et al.
Published: (2025)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
by: Wu, Qinyun, et al.
Published: (2024)
by: Wu, Qinyun, et al.
Published: (2024)
Enhancing Debugging Skills with AI-Powered Assistance: A Real-Time Tool for Debugging Support
by: Artser, Elizaveta, et al.
Published: (2026)
by: Artser, Elizaveta, et al.
Published: (2026)
An Empirical Study on Compliance with Ranking Transparency in the Software Documentation of EU Online Platforms
by: Sovrano, Francesco, et al.
Published: (2023)
by: Sovrano, Francesco, et al.
Published: (2023)
ReusStdFlow: A Standardized Reusability Framework for Dynamic Workflow Construction in Agentic AI
by: Zhang, Gaoyang, et al.
Published: (2026)
by: Zhang, Gaoyang, et al.
Published: (2026)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
Enhancing Deployment-Time Predictive Model Robustness for Code Analysis and Optimization
by: Wang, Huanting, et al.
Published: (2024)
by: Wang, Huanting, et al.
Published: (2024)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
by: Li, Jingyue, et al.
Published: (2026)
by: Li, Jingyue, et al.
Published: (2026)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
by: Lian, Keke, et al.
Published: (2025)
by: Lian, Keke, et al.
Published: (2025)
Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production
by: Fehlis, Yao, et al.
Published: (2026)
by: Fehlis, Yao, et al.
Published: (2026)
An Empirical Framework for Evaluating Semantic Preservation Using Hugging Face
by: Jia, Nan, et al.
Published: (2025)
by: Jia, Nan, et al.
Published: (2025)
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
by: Li, Hanyu, et al.
Published: (2026)
by: Li, Hanyu, et al.
Published: (2026)
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
by: Pandey, Ruchika, et al.
Published: (2024)
by: Pandey, Ruchika, et al.
Published: (2024)
Similar Items
-
Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling
by: Mamirov, Akhmadillo
Published: (2025) -
P4OMP: Retrieval-Augmented Prompting for OpenMP Parallelism in Serial Code
by: Abdullah, Wali Mohammad, et al.
Published: (2025) -
Sentiment Analysis in Software Engineering: Evaluating Generative Pre-trained Transformers
by: Saifullah, KM Khalid, et al.
Published: (2025) -
LoCoML: A Framework for Real-World ML Inference Pipelines
by: Maddireddy, Kritin, et al.
Published: (2025) -
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs
by: Kabir, Azmain, et al.
Published: (2024)