ConDABench: Interactive Evaluation of Language Models for Data Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Dutta, Avik, Gupta, Priyanshu, Hasanbeig, Hosein, Singh, Rahul Pratap, Nigam, Harshit, Gulwani, Sumit, Radhakrishna, Arjun, Soares, Gustavo, Tiwari, Ashish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Investigation of Robustness in Large Language Models under Tabular Distortions
by: Dutta, Avik, et al.
Published: (2026)
by: Dutta, Avik, et al.
Published: (2026)
Soft-Label Training Preserves Epistemic Uncertainty
by: Singh, Agamdeep, et al.
Published: (2025)
by: Singh, Agamdeep, et al.
Published: (2025)
Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
MetaReflection: Learning Instructions for Language Agents using Past Reflections
by: Gupta, Priyanshu, et al.
Published: (2024)
by: Gupta, Priyanshu, et al.
Published: (2024)
TEN: Table Explicitization, Neurosymbolically
by: Mehrotra, Nikita, et al.
Published: (2025)
by: Mehrotra, Nikita, et al.
Published: (2025)
LLM-Guided Compositional Program Synthesis
by: Khan, Ruhma, et al.
Published: (2025)
by: Khan, Ruhma, et al.
Published: (2025)
Do Code Models Suffer from the Dunning-Kruger Effect?
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Exploring Interaction Patterns for Debugging: Enhancing Conversational Capabilities of AI-assistants
by: Chopra, Bhavya, et al.
Published: (2024)
by: Chopra, Bhavya, et al.
Published: (2024)
STACKFEED: Structured Textual Actor-Critic Knowledge Base Editing with FeedBack
by: Kirtania, Shashank, et al.
Published: (2024)
by: Kirtania, Shashank, et al.
Published: (2024)
Ordered Semantically Diverse Sampling for Textual Data
by: Tiwari, Ashish, et al.
Published: (2025)
by: Tiwari, Ashish, et al.
Published: (2025)
Improving Language Agents through BREW
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
TableTalk: Scaffolding Spreadsheet Development with a Language Agent
by: Liang, Jenny T., et al.
Published: (2025)
by: Liang, Jenny T., et al.
Published: (2025)
IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
by: Biyani, Param, et al.
Published: (2025)
by: Biyani, Param, et al.
Published: (2025)
InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
by: Hu, Xueyu, et al.
Published: (2024)
by: Hu, Xueyu, et al.
Published: (2024)
Effect of PCM and nano‐embedded PCM on the solar pond performance as sensible heat storage: An experimental approach
by: Ajay Pratap Singh, et al.
Published: (2024)
by: Ajay Pratap Singh, et al.
Published: (2024)
LOGIC-LM++: Multi-Step Refinement for Symbolic Formulations
by: Kirtania, Shashank, et al.
Published: (2024)
by: Kirtania, Shashank, et al.
Published: (2024)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
by: Mhatre, Sanket, et al.
Published: (2025)
by: Mhatre, Sanket, et al.
Published: (2025)
Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
by: Kumar, Aayush, et al.
Published: (2025)
by: Kumar, Aayush, et al.
Published: (2025)
Bridging Gaps Between Student and Expert Evaluations of AI-Generated Programming Hints
by: Phung, Tung, et al.
Published: (2025)
by: Phung, Tung, et al.
Published: (2025)
Mission-driven Exploration for Accelerated Deep Reinforcement Learning with Temporal Logic Task Specifications
by: Wang, Jun, et al.
Published: (2023)
by: Wang, Jun, et al.
Published: (2023)
Decoding In-Context Learning: Neuroscience-inspired Analysis of Representations in Large Language Models
by: Yousefi, Safoora, et al.
Published: (2023)
by: Yousefi, Safoora, et al.
Published: (2023)
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Diffusion is a code repair operator and generator
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Entanglement Entropy for Screened Interactions via Dimensional Mapping to Harmonic Oscillators
by: Kulkarni, Akshay, et al.
Published: (2026)
by: Kulkarni, Akshay, et al.
Published: (2026)
OriCon3D: Effective 3D Object Detection using Orientation and Confidence
by: Rajani, Dhyey Manish, et al.
Published: (2023)
by: Rajani, Dhyey Manish, et al.
Published: (2023)
Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
by: Omi, Nabil, et al.
Published: (2024)
by: Omi, Nabil, et al.
Published: (2024)
Enhancing Creativity in Large Language Models through Associative Thinking Strategies
by: Mehrotra, Pronita, et al.
Published: (2024)
by: Mehrotra, Pronita, et al.
Published: (2024)
Asian option valuation under price impact
by: Tiwari, Priyanshu, et al.
Published: (2025)
by: Tiwari, Priyanshu, et al.
Published: (2025)
Minimization of the circuit components with modified cascaded multilevel inverter topology
by: Singh, Raghvendra Pratap, et al.
Published: (2025)
by: Singh, Raghvendra Pratap, et al.
Published: (2025)
Improved Twin Actor Twin Delayed Deep Deterministic Policy Gradient With Spatial Occupancy for Traffic Signal Control System
by: Nikhil Nigam, et al.
Published: (2025)
by: Nikhil Nigam, et al.
Published: (2025)
Digital Twin-assisted belief-state reinforcement learning for latency-robust ISAC in 6G networks
by: Tiwari, Himanshu, et al.
Published: (2026)
by: Tiwari, Himanshu, et al.
Published: (2026)
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
Automating Human Tutor-Style Programming Feedback: Leveraging GPT-4 Tutor Model for Hint Generation and GPT-3.5 Student Model for Hint Validation
by: Phung, Tung, et al.
Published: (2023)
by: Phung, Tung, et al.
Published: (2023)
Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Tabularis Formatus: Predictive Formatting for Tables
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
An Empirical Study of Validating Synthetic Data for Formula Generation
by: Singh, Usneek, et al.
Published: (2024)
by: Singh, Usneek, et al.
Published: (2024)
Virtual Reality in Social Media: A New Era of Immersive Social Interactions
by: Chaubey, Priyanshu
Published: (2025)
by: Chaubey, Priyanshu
Published: (2025)
PU-Lie: Lightweight Deception Detection in Imbalanced Diplomatic Dialogues via Positive-Unlabeled Learning
by: Kuwar, Bhavinkumar Vinodbhai, et al.
Published: (2025)
by: Kuwar, Bhavinkumar Vinodbhai, et al.
Published: (2025)
Similar Items
-
An Empirical Investigation of Robustness in Large Language Models under Tabular Distortions
by: Dutta, Avik, et al.
Published: (2026) -
Soft-Label Training Preserves Epistemic Uncertainty
by: Singh, Agamdeep, et al.
Published: (2025) -
Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory
by: Singh, Mukul, et al.
Published: (2025) -
MetaReflection: Learning Instructions for Language Agents using Past Reflections
by: Gupta, Priyanshu, et al.
Published: (2024) -
TEN: Table Explicitization, Neurosymbolically
by: Mehrotra, Nikita, et al.
Published: (2025)