Benchmarking Text-to-Python against Text-to-SQL: The Impact of Explicit Logic and Ambiguity
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Hangle, Hou, Chenyu, Cao, Bin, Li, Ruizhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
R$^3$-SQL: Ranking Reward and Resampling for Text-to-SQL
di: Han, Hojae, et al.
Pubblicazione: (2026)
di: Han, Hojae, et al.
Pubblicazione: (2026)
Enhancing LLM Fine-tuning for Text-to-SQLs by SQL Quality Measurement
di: Sarker, Shouvon, et al.
Pubblicazione: (2024)
di: Sarker, Shouvon, et al.
Pubblicazione: (2024)
From Queries to Insights: Agentic LLM Pipelines for Spatio-Temporal Text-to-SQL
di: Redd, Manu, et al.
Pubblicazione: (2025)
di: Redd, Manu, et al.
Pubblicazione: (2025)
AutoPLC: Generating Vendor-Aware Structured Text for Programmable Logic Controllers
di: Yang, Donghao, et al.
Pubblicazione: (2024)
di: Yang, Donghao, et al.
Pubblicazione: (2024)
A Study of In-Context-Learning-Based Text-to-SQL Errors
di: Shen, Jiawei, et al.
Pubblicazione: (2025)
di: Shen, Jiawei, et al.
Pubblicazione: (2025)
Dialect2SQL: A Novel Text-to-SQL Dataset for Arabic Dialects with a Focus on Moroccan Darija
di: Chafik, Salmane, et al.
Pubblicazione: (2025)
di: Chafik, Salmane, et al.
Pubblicazione: (2025)
LeGo-Code: Can Modular Curriculum Learning Advance Complex Code Generation? Insights from Text-to-SQL
di: Chafik, Salmane, et al.
Pubblicazione: (2026)
di: Chafik, Salmane, et al.
Pubblicazione: (2026)
LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning
di: Hou, Bo, et al.
Pubblicazione: (2025)
di: Hou, Bo, et al.
Pubblicazione: (2025)
From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python
di: Wang, Jinhua, et al.
Pubblicazione: (2026)
di: Wang, Jinhua, et al.
Pubblicazione: (2026)
In-Context Code-Text Learning for Bimodal Software Engineering
di: Tang, Xunzhu, et al.
Pubblicazione: (2024)
di: Tang, Xunzhu, et al.
Pubblicazione: (2024)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
di: Dong, Zeming, et al.
Pubblicazione: (2023)
di: Dong, Zeming, et al.
Pubblicazione: (2023)
Ambiguity Detection and Elimination in Automated Executable Process Modeling
di: Matei, Ion, et al.
Pubblicazione: (2026)
di: Matei, Ion, et al.
Pubblicazione: (2026)
Ambiguity Resolution with Human Feedback for Code Writing Tasks
di: Nandan, Aditey, et al.
Pubblicazione: (2025)
di: Nandan, Aditey, et al.
Pubblicazione: (2025)
Modeling Code: Is Text All You Need?
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
Text Tells the Cost: Predicting and Analyzing Repayment Effort of Self-Admitted Technical Debt
di: Li, Yikun, et al.
Pubblicazione: (2023)
di: Li, Yikun, et al.
Pubblicazione: (2023)
Lost in Transcription: How Speech-to-Text Errors Derail Code Understanding
di: Havare, Jayant, et al.
Pubblicazione: (2026)
di: Havare, Jayant, et al.
Pubblicazione: (2026)
Better Python Programming for all: With the focus on Maintainability
di: Shivashankar, Karthik, et al.
Pubblicazione: (2024)
di: Shivashankar, Karthik, et al.
Pubblicazione: (2024)
Accurate and Consistent Graph Model Generation from Text with Large Language Models
di: Chen, Boqi, et al.
Pubblicazione: (2025)
di: Chen, Boqi, et al.
Pubblicazione: (2025)
GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries
di: Hou, Shuyang, et al.
Pubblicazione: (2025)
di: Hou, Shuyang, et al.
Pubblicazione: (2025)
Copilot-in-the-Loop: Fixing Code Smells in Copilot-Generated Python Code using Copilot
di: Zhang, Beiqi, et al.
Pubblicazione: (2024)
di: Zhang, Beiqi, et al.
Pubblicazione: (2024)
Text2Scenario: Text-Driven Scenario Generation for Autonomous Driving Test
di: Cai, Xuan, et al.
Pubblicazione: (2025)
di: Cai, Xuan, et al.
Pubblicazione: (2025)
The Last Dependency Crusade: Solving Python Dependency Conflicts with LLMs
di: Bartlett, Antony, et al.
Pubblicazione: (2025)
di: Bartlett, Antony, et al.
Pubblicazione: (2025)
Machine Learning Techniques for Python Source Code Vulnerability Detection
di: Farasat, Talaya, et al.
Pubblicazione: (2024)
di: Farasat, Talaya, et al.
Pubblicazione: (2024)
Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
di: Almukhtar, Mohamed, et al.
Pubblicazione: (2026)
di: Almukhtar, Mohamed, et al.
Pubblicazione: (2026)
Adaptive Hierarchical Evaluation of LLMs and SAST tools for CWE Prediction in Python
di: Adnan, Muntasir, et al.
Pubblicazione: (2026)
di: Adnan, Muntasir, et al.
Pubblicazione: (2026)
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
di: Maaz, Muhammad, et al.
Pubblicazione: (2025)
di: Maaz, Muhammad, et al.
Pubblicazione: (2025)
KGCompiler: Deep Learning Compilation Optimization for Knowledge Graph Complex Logical Query Answering
di: Lin, Hongyu, et al.
Pubblicazione: (2025)
di: Lin, Hongyu, et al.
Pubblicazione: (2025)
PyGen: A Collaborative Human-AI Approach to Python Package Creation
di: Barua, Saikat, et al.
Pubblicazione: (2024)
di: Barua, Saikat, et al.
Pubblicazione: (2024)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
di: Li, Yikun, et al.
Pubblicazione: (2026)
di: Li, Yikun, et al.
Pubblicazione: (2026)
MEMRES: A Memory-Augmented Resolver with Confidence Cascade for Agentic Python Dependency Resolution
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
di: Zhu, Qiming, et al.
Pubblicazione: (2024)
di: Zhu, Qiming, et al.
Pubblicazione: (2024)
Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
di: Diehl, Patrick, et al.
Pubblicazione: (2024)
di: Diehl, Patrick, et al.
Pubblicazione: (2024)
PyResBugs: A Dataset of Residual Python Bugs for Natural Language-Driven Fault Injection
di: Cotroneo, Domenico, et al.
Pubblicazione: (2025)
di: Cotroneo, Domenico, et al.
Pubblicazione: (2025)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
di: Wang, Kaixin, et al.
Pubblicazione: (2025)
di: Wang, Kaixin, et al.
Pubblicazione: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
di: Zhang, Wentao, et al.
Pubblicazione: (2026)
di: Zhang, Wentao, et al.
Pubblicazione: (2026)
An Empirical Study of Vulnerabilities in Python Packages and Their Detection
di: Quan, Haowei, et al.
Pubblicazione: (2025)
di: Quan, Haowei, et al.
Pubblicazione: (2025)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
di: Yang, Zhonghao, et al.
Pubblicazione: (2026)
di: Yang, Zhonghao, et al.
Pubblicazione: (2026)
EmbedAgent: Benchmarking Large Language Models in Embedded System Development
di: Xu, Ruiyang, et al.
Pubblicazione: (2025)
di: Xu, Ruiyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
R$^3$-SQL: Ranking Reward and Resampling for Text-to-SQL
di: Han, Hojae, et al.
Pubblicazione: (2026) -
Enhancing LLM Fine-tuning for Text-to-SQLs by SQL Quality Measurement
di: Sarker, Shouvon, et al.
Pubblicazione: (2024) -
From Queries to Insights: Agentic LLM Pipelines for Spatio-Temporal Text-to-SQL
di: Redd, Manu, et al.
Pubblicazione: (2025) -
AutoPLC: Generating Vendor-Aware Structured Text for Programmable Logic Controllers
di: Yang, Donghao, et al.
Pubblicazione: (2024) -
A Study of In-Context-Learning-Based Text-to-SQL Errors
di: Shen, Jiawei, et al.
Pubblicazione: (2025)