CVEvolve: Autonomous Algorithm Discovery for Unstructured Scientific Data Processing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918534888554496 |
|---|---|
| author | Du, Ming Yin, Xiangyu Luo, Yanqi Beniwal, Dishant Tang, Songyuan Sharma, Hemant Cherukara, Mathew J. |
| author_facet | Du, Ming Yin, Xiangyu Luo, Yanqi Beniwal, Dishant Tang, Songyuan Sharma, Hemant Cherukara, Mathew J. |
| contents | Scientific data processing often requires task-specific algorithms or AI models, creating a barrier for domain scientists who need to analyze their data but may not have extensive computing or image-processing expertise. This barrier is especially pronounced when data are noisy, have a high dynamic range, are sparsely labeled, or are only loosely specified. We introduce CVEvolve, an autonomous agentic harness with a zero-code interface for scientific data-processing algorithm discovery. CVEvolve combines a multi-round search strategy with tools for code execution, evaluation implementation, history management, holdout testing, and optional inspection of scientific data and visual outputs. The search alternates between discovery and improvement actions, and uses lineage-aware stochastic candidate sampling to balance exploration and exploitation. We demonstrate CVEvolve on X-ray fluorescence microscopy image registration, Bragg peak detection, high-energy diffraction microscopy image segmentation, and hybrid analytical-learning-based affine registration. Across these tasks, CVEvolve discovers algorithms that improve over baseline methods, while holdout test tracking helps identify candidates that generalize better than later over-optimized alternatives. These results show that zero-code, autonomous LLM-powered algorithm development can help domain scientists turn unstructured scientific image data into practical algorithms and downstream scientific discoveries. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11359 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CVEvolve: Autonomous Algorithm Discovery for Unstructured Scientific Data Processing Du, Ming Yin, Xiangyu Luo, Yanqi Beniwal, Dishant Tang, Songyuan Sharma, Hemant Cherukara, Mathew J. Artificial Intelligence Data Analysis, Statistics and Probability 68T42 I.2.2 Scientific data processing often requires task-specific algorithms or AI models, creating a barrier for domain scientists who need to analyze their data but may not have extensive computing or image-processing expertise. This barrier is especially pronounced when data are noisy, have a high dynamic range, are sparsely labeled, or are only loosely specified. We introduce CVEvolve, an autonomous agentic harness with a zero-code interface for scientific data-processing algorithm discovery. CVEvolve combines a multi-round search strategy with tools for code execution, evaluation implementation, history management, holdout testing, and optional inspection of scientific data and visual outputs. The search alternates between discovery and improvement actions, and uses lineage-aware stochastic candidate sampling to balance exploration and exploitation. We demonstrate CVEvolve on X-ray fluorescence microscopy image registration, Bragg peak detection, high-energy diffraction microscopy image segmentation, and hybrid analytical-learning-based affine registration. Across these tasks, CVEvolve discovers algorithms that improve over baseline methods, while holdout test tracking helps identify candidates that generalize better than later over-optimized alternatives. These results show that zero-code, autonomous LLM-powered algorithm development can help domain scientists turn unstructured scientific image data into practical algorithms and downstream scientific discoveries. |
| title | CVEvolve: Autonomous Algorithm Discovery for Unstructured Scientific Data Processing |
| topic | Artificial Intelligence Data Analysis, Statistics and Probability 68T42 I.2.2 |
| url | https://arxiv.org/abs/2605.11359 |