Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Jinyang, Huo, Nan, Gao, Yan, Shi, Jiayi, Zhao, Yingxiu, Qu, Ge, Wu, Yurong, Ma, Chenhao, Lou, Jian-Guang, Cheng, Reynold |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation
por: Qu, Ge, et al.
Publicado: (2024)
por: Qu, Ge, et al.
Publicado: (2024)
SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL
por: Qu, Ge, et al.
Publicado: (2025)
por: Qu, Ge, et al.
Publicado: (2025)
Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning
por: Huo, Nan, et al.
Publicado: (2025)
por: Huo, Nan, et al.
Publicado: (2025)
Automatic Instruction Evolving for Large Language Models
por: Zeng, Weihao, et al.
Publicado: (2024)
por: Zeng, Weihao, et al.
Publicado: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
por: Wang, Xu, et al.
Publicado: (2025)
por: Wang, Xu, et al.
Publicado: (2025)
Debiasing Recommendation with Personal Popularity
por: Ning, Wentao, et al.
Publicado: (2024)
por: Ning, Wentao, et al.
Publicado: (2024)
Finding Locally Densest Subgraphs: Convex Programming with Edge and Triangle Density
por: Yang, Yi, et al.
Publicado: (2025)
por: Yang, Yi, et al.
Publicado: (2025)
BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions
por: Huo, Nan, et al.
Publicado: (2025)
por: Huo, Nan, et al.
Publicado: (2025)
Exploring the Technical Knowledge Interaction of Global Digital Humanities: Three-decade Evidence from Bibliometric-based perspectives
por: Li, Jiayi, et al.
Publicado: (2025)
por: Li, Jiayi, et al.
Publicado: (2025)
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
por: Garbacea, Cristina, et al.
Publicado: (2026)
por: Garbacea, Cristina, et al.
Publicado: (2026)
SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications
por: Li, Jinyang, et al.
Publicado: (2025)
por: Li, Jinyang, et al.
Publicado: (2025)
A Remark on the Set of Exactly Approximable Vectors in the Simultaneous Case
por: Fregoli, Reynold
Publicado: (2023)
por: Fregoli, Reynold
Publicado: (2023)
Sums of Reciprocals of Fractional Parts II
por: Fregoli, Reynold
Publicado: (2023)
por: Fregoli, Reynold
Publicado: (2023)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent
por: Cheng, Yuhao, et al.
Publicado: (2025)
por: Cheng, Yuhao, et al.
Publicado: (2025)
Divergent responses of the diatom Thalassiosira weissflogii to ocean acidification during light and dark periods
por: Guang Gao, et al.
Publicado: (2026)
por: Guang Gao, et al.
Publicado: (2026)
R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
por: Chen, Jiayi, et al.
Publicado: (2025)
por: Chen, Jiayi, et al.
Publicado: (2025)
Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
por: Chong, Yee Hin, et al.
Publicado: (2026)
por: Chong, Yee Hin, et al.
Publicado: (2026)
AgentEvolver: Towards Efficient Self-Evolving Agent System
por: Zhai, Yunpeng, et al.
Publicado: (2025)
por: Zhai, Yunpeng, et al.
Publicado: (2025)
StraGo: Harnessing Strategic Guidance for Prompt Optimization
por: Wu, Yurong, et al.
Publicado: (2024)
por: Wu, Yurong, et al.
Publicado: (2024)
DeMarking: A Defense for Network Flow Watermarking in Real-Time
por: Yuan, Yali, et al.
Publicado: (2024)
por: Yuan, Yali, et al.
Publicado: (2024)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
por: Li, Chance Jiajie, et al.
Publicado: (2025)
por: Li, Chance Jiajie, et al.
Publicado: (2025)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
por: Ji, Haonian, et al.
Publicado: (2026)
por: Ji, Haonian, et al.
Publicado: (2026)
Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution
por: Wei, Zixin, et al.
Publicado: (2025)
por: Wei, Zixin, et al.
Publicado: (2025)
Alita-G: Self-Evolving Generative Agent for Agent Generation
por: Qiu, Jiahao, et al.
Publicado: (2025)
por: Qiu, Jiahao, et al.
Publicado: (2025)
Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment
por: Qiu, Pengcheng, et al.
Publicado: (2025)
por: Qiu, Pengcheng, et al.
Publicado: (2025)
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
por: Dai, Yuhong, et al.
Publicado: (2026)
por: Dai, Yuhong, et al.
Publicado: (2026)
BEACON: A Benchmark for Efficient and Accurate Counting of Subgraphs
por: Najafi, Mohammad Matin, et al.
Publicado: (2025)
por: Najafi, Mohammad Matin, et al.
Publicado: (2025)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
por: Gu, Jihao, et al.
Publicado: (2025)
por: Gu, Jihao, et al.
Publicado: (2025)
What Types of Start-ups Receive Funding from the Small Business Innovation Research (SBIR) Program? Evidence from the Kauffman Firm Survey
por: Reynold V. Galope
Publicado: (2014)
por: Reynold V. Galope
Publicado: (2014)
Self-Evolving LLMs via Continual Instruction Tuning
por: Kang, Jiazheng, et al.
Publicado: (2025)
por: Kang, Jiazheng, et al.
Publicado: (2025)
Higher-Dimensional Moving Averages and Submanifold Genericity
por: Cheng, Jiajun, et al.
Publicado: (2025)
por: Cheng, Jiajun, et al.
Publicado: (2025)
Step-Opt: Boosting Optimization Modeling in LLMs through Iterative Data Synthesis and Structured Validation
por: Wu, Yang, et al.
Publicado: (2025)
por: Wu, Yang, et al.
Publicado: (2025)
As Content and Layout Co-Evolve: TangibleSite for Scaffolding Blind People's Webpage Design through Multimodal Interaction
por: Li, Jiasheng, et al.
Publicado: (2026)
por: Li, Jiasheng, et al.
Publicado: (2026)
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
por: Yildiz, Alperen, et al.
Publicado: (2025)
por: Yildiz, Alperen, et al.
Publicado: (2025)
Evolving Deception: When Agents Evolve, Deception Wins
por: Ying, Zonghao, et al.
Publicado: (2026)
por: Ying, Zonghao, et al.
Publicado: (2026)
Nitrogen/Oxygen Co‐Doped Carbon Quantum Dots with Efficient High Color‐Purity Red Emission for Bright Electroluminescent LEDs
por: Chenhao Li, et al.
Publicado: (2025)
por: Chenhao Li, et al.
Publicado: (2025)
Nitrogen/Oxygen Co‐Doped Carbon Quantum Dots with Efficient High Color‐Purity Red Emission for Bright Electroluminescent LEDs
por: Chenhao Li, et al.
Publicado: (2025)
por: Chenhao Li, et al.
Publicado: (2025)
Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
por: Cheng, Zihao, et al.
Publicado: (2026)
por: Cheng, Zihao, et al.
Publicado: (2026)
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
por: Qu, Heng, et al.
Publicado: (2026)
por: Qu, Heng, et al.
Publicado: (2026)
Ejemplares similares
-
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation
por: Qu, Ge, et al.
Publicado: (2024) -
SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL
por: Qu, Ge, et al.
Publicado: (2025) -
Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning
por: Huo, Nan, et al.
Publicado: (2025) -
Automatic Instruction Evolving for Large Language Models
por: Zeng, Weihao, et al.
Publicado: (2024) -
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
por: Wang, Xu, et al.
Publicado: (2025)