AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Qiuhai, Jin, Claire, Wang, Xinyue, Zheng, Yuhan, Li, Qunhua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
Model-in-the-Loop (MILO): Accelerating Multimodal AI Data Annotation with LLMs
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
by: Soni, Nikita, et al.
Published: (2025)
by: Soni, Nikita, et al.
Published: (2025)
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
by: Sun, Yuxuan, et al.
Published: (2025)
by: Sun, Yuxuan, et al.
Published: (2025)
Harmonic LLMs are Trustworthy
by: Kersting, Nicholas S., et al.
Published: (2024)
by: Kersting, Nicholas S., et al.
Published: (2024)
AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
by: Ou, Yixin, et al.
Published: (2025)
by: Ou, Yixin, et al.
Published: (2025)
Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications
by: Sakib, Abu Noman Md, et al.
Published: (2026)
by: Sakib, Abu Noman Md, et al.
Published: (2026)
Interaction Dynamics as a Reward Signal for LLMs
by: Gooding, Sian, et al.
Published: (2025)
by: Gooding, Sian, et al.
Published: (2025)
LLMs for XAI: Future Directions for Explaining Explanations
by: Zytek, Alexandra, et al.
Published: (2024)
by: Zytek, Alexandra, et al.
Published: (2024)
Generative UI: LLMs are Effective UI Generators
by: Leviathan, Yaniv, et al.
Published: (2026)
by: Leviathan, Yaniv, et al.
Published: (2026)
Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
by: Ma, Cheng Charles, et al.
Published: (2024)
by: Ma, Cheng Charles, et al.
Published: (2024)
What Large Language Models Know and What People Think They Know
by: Steyvers, Mark, et al.
Published: (2024)
by: Steyvers, Mark, et al.
Published: (2024)
How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
by: Xu, Ziwen, et al.
Published: (2026)
by: Xu, Ziwen, et al.
Published: (2026)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
Are You Being Tracked? Discover the Power of Zero-Shot Trajectory Tracing with LLMs!
by: Yang, Huanqi, et al.
Published: (2024)
by: Yang, Huanqi, et al.
Published: (2024)
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
by: Garbacea, Cristina, et al.
Published: (2026)
by: Garbacea, Cristina, et al.
Published: (2026)
Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
by: Liu, Jonathan, et al.
Published: (2025)
by: Liu, Jonathan, et al.
Published: (2025)
MotionTeller: Multi-modal Integration of Wearable Time-Series with LLMs for Health and Behavioral Understanding
by: Zhang, Aiwei, et al.
Published: (2025)
by: Zhang, Aiwei, et al.
Published: (2025)
Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
by: Zhang, Ziyuan, et al.
Published: (2025)
by: Zhang, Ziyuan, et al.
Published: (2025)
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
Teaching According to Students' Aptitude: Personalized Mathematics Tutoring via Persona-, Memory-, and Forgetting-Aware LLMs
by: Wu, Yang, et al.
Published: (2025)
by: Wu, Yang, et al.
Published: (2025)
Heterogeneous Value Alignment Evaluation for Large Language Models
by: Zhang, Zhaowei, et al.
Published: (2023)
by: Zhang, Zhaowei, et al.
Published: (2023)
"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
by: Wu, Yang, et al.
Published: (2026)
by: Wu, Yang, et al.
Published: (2026)
Evaluating Large Language Models for Health-related Queries with Presuppositions
by: Kaur, Navreet, et al.
Published: (2023)
by: Kaur, Navreet, et al.
Published: (2023)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
by: Ngueajio, Mikel K., et al.
Published: (2025)
by: Ngueajio, Mikel K., et al.
Published: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
by: Kahng, Minsuk, et al.
Published: (2024)
by: Kahng, Minsuk, et al.
Published: (2024)
If in a Crowdsourced Data Annotation Pipeline, a GPT-4
by: He, Zeyu, et al.
Published: (2024)
by: He, Zeyu, et al.
Published: (2024)
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
by: Ben-Zion, Ziv, et al.
Published: (2025)
by: Ben-Zion, Ziv, et al.
Published: (2025)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024)
by: Park, Jung In, et al.
Published: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
by: Baidya, Avinash, et al.
Published: (2025)
by: Baidya, Avinash, et al.
Published: (2025)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI
by: Saha, Agnik, et al.
Published: (2025)
by: Saha, Agnik, et al.
Published: (2025)
PRECISE Framework: GPT-based Text For Improved Readability, Reliability, and Understandability of Radiology Reports For Patient-Centered Care
by: Tripathi, Satvik, et al.
Published: (2024)
by: Tripathi, Satvik, et al.
Published: (2024)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
by: Rajagopal, Shreya, et al.
Published: (2024)
by: Rajagopal, Shreya, et al.
Published: (2024)
UniAutoML: A Human-Centered Framework for Unified Discriminative and Generative AutoML with Large Language Models
by: Guo, Jiayi, et al.
Published: (2024)
by: Guo, Jiayi, et al.
Published: (2024)
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
by: He, Zeyu, et al.
Published: (2025)
by: He, Zeyu, et al.
Published: (2025)
AutoGLM: Autonomous Foundation Agents for GUIs
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
LLMs as Writing Assistants: Exploring Perspectives on Sense of Ownership and Reasoning
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
Similar Items
-
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
by: Ayonrinde, Kola, et al.
Published: (2025) -
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025) -
Model-in-the-Loop (MILO): Accelerating Multimodal AI Data Annotation with LLMs
by: Wang, Yifan, et al.
Published: (2024) -
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
by: Soni, Nikita, et al.
Published: (2025) -
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
by: Sun, Yuxuan, et al.
Published: (2025)