GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
Fuente:
arXiv
Saved in:
| Main Authors: | Fachada, Nuno, Fernandes, Daniel, Fernandes, Carlos M., Ferreira-Saraiva, Bruno D., Matos-Carvalho, João P. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks
by: Fernandes, Daniel, et al.
Published: (2025)
by: Fernandes, Daniel, et al.
Published: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
by: Haque, Md. Asraful, et al.
Published: (2026)
by: Haque, Md. Asraful, et al.
Published: (2026)
Can Large Language Models Implement Agent-Based Models? An ODD-based Replication Study
by: Fachada, Nuno, et al.
Published: (2026)
by: Fachada, Nuno, et al.
Published: (2026)
Text Clustering with Large Language Model Embeddings
by: Petukhova, Alina, et al.
Published: (2024)
by: Petukhova, Alina, et al.
Published: (2024)
Advancing Explainability in Neural Machine Translation: Analytical Metrics for Attention and Alignment Consistency
by: Mishra, Anurag
Published: (2024)
by: Mishra, Anurag
Published: (2024)
React-ing to Grace Hopper 200: Five Open-Weights Coding Models, One React Native App, One GH200, One Weekend
by: Potanin, Alex
Published: (2026)
by: Potanin, Alex
Published: (2026)
Rango: Adaptive Retrieval-Augmented Proving for Automated Software Verification
by: Thompson, Kyle, et al.
Published: (2024)
by: Thompson, Kyle, et al.
Published: (2024)
AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents
by: Dasgupta, Sudip, et al.
Published: (2025)
by: Dasgupta, Sudip, et al.
Published: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
by: Xia, Bowei, et al.
Published: (2026)
by: Xia, Bowei, et al.
Published: (2026)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
by: Badshah, Sher, et al.
Published: (2025)
by: Badshah, Sher, et al.
Published: (2025)
Text-to-SQL based on Large Language Models and Database Keyword Search
by: Nascimento, Eduardo R., et al.
Published: (2025)
by: Nascimento, Eduardo R., et al.
Published: (2025)
Knowledge-Guided Multi-Agent Framework for Automated Requirements Development: A Vision
by: Huang, Jiangping, et al.
Published: (2025)
by: Huang, Jiangping, et al.
Published: (2025)
Optimizing Large Language Models for OpenAPI Code Completion
by: Petryshyn, Bohdan, et al.
Published: (2024)
by: Petryshyn, Bohdan, et al.
Published: (2024)
Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model
by: Qiu, Tianhao, et al.
Published: (2026)
by: Qiu, Tianhao, et al.
Published: (2026)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
by: Bayram, M. Ali, et al.
Published: (2024)
by: Bayram, M. Ali, et al.
Published: (2024)
Efficient Telecom Specific LLM: TSLAM-Mini with QLoRA and Digital Twin Data
by: Ethiraj, Vignesh, et al.
Published: (2025)
by: Ethiraj, Vignesh, et al.
Published: (2025)
Is It Time To Treat Prompts As Code? A Multi-Use Case Study For Prompt Optimization Using DSPy
by: Lemos, Francisca, et al.
Published: (2025)
by: Lemos, Francisca, et al.
Published: (2025)
The Transformative Influence of LLMs on Software Development & Developer Productivity
by: Jalil, Sajed
Published: (2023)
by: Jalil, Sajed
Published: (2023)
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
by: Peyronnet, Antoine, et al.
Published: (2026)
by: Peyronnet, Antoine, et al.
Published: (2026)
Collaborative LLM Agents for C4 Software Architecture Design Automation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
by: Berman, Shmuel, et al.
Published: (2024)
by: Berman, Shmuel, et al.
Published: (2024)
PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning
by: Qiu, Xiaoqi, et al.
Published: (2024)
by: Qiu, Xiaoqi, et al.
Published: (2024)
TCProF: Time-Complexity Prediction SSL Framework
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
MicroRemed: Benchmarking LLMs in Microservices Remediation
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
by: Görge, Rebekka, et al.
Published: (2025)
by: Görge, Rebekka, et al.
Published: (2025)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
by: Plevris, Vagelis, et al.
Published: (2023)
by: Plevris, Vagelis, et al.
Published: (2023)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
by: Zhang, Xue
Published: (2025)
by: Zhang, Xue
Published: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
by: Basit, Abdul, et al.
Published: (2025)
by: Basit, Abdul, et al.
Published: (2025)
ESALE: Enhancing Code-Summary Alignment Learning for Source Code Summarization
by: Fang, Chunrong, et al.
Published: (2024)
by: Fang, Chunrong, et al.
Published: (2024)
Commenting Higher-level Code Unit: Full Code, Reduced Code, or Hierarchical Code Summarization
by: Sun, Weisong, et al.
Published: (2025)
by: Sun, Weisong, et al.
Published: (2025)
Source Code Summarization in the Era of Large Language Models
by: Sun, Weisong, et al.
Published: (2024)
by: Sun, Weisong, et al.
Published: (2024)
GEML: A Grammar-based Evolutionary Machine Learning Approach for Design-Pattern Detection
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
Evaluating Pixel Language Models on Non-Standardized Languages
by: Muñoz-Ortiz, Alberto, et al.
Published: (2024)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2024)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
Enhancing Large Language Models through Neuro-Symbolic Integration and Ontological Reasoning
by: Vsevolodovna, Ruslan Idelfonso Magana, et al.
Published: (2025)
by: Vsevolodovna, Ruslan Idelfonso Magana, et al.
Published: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Similar Items
-
DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks
by: Fernandes, Daniel, et al.
Published: (2025) -
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
by: Haque, Md. Asraful, et al.
Published: (2026) -
Can Large Language Models Implement Agent-Based Models? An ODD-based Replication Study
by: Fachada, Nuno, et al.
Published: (2026) -
Text Clustering with Large Language Model Embeddings
by: Petukhova, Alina, et al.
Published: (2024) -
Advancing Explainability in Neural Machine Translation: Analytical Metrics for Attention and Alignment Consistency
by: Mishra, Anurag
Published: (2024)