Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Tingting, Lin, Beibei, Anumasa, Srinivas, Shah, Vedant, Yuan, Zifeng, Zou, Qiran, Goyal, Anirudh, Liu, Dianbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
von: Anumasa, Srinivas, et al.
Veröffentlicht: (2025)
von: Anumasa, Srinivas, et al.
Veröffentlicht: (2025)
HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
von: Chen, Tingting, et al.
Veröffentlicht: (2025)
von: Chen, Tingting, et al.
Veröffentlicht: (2025)
Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via inference-time Multi-Views Brainstorming
von: Lagzian, Arash, et al.
Veröffentlicht: (2025)
von: Lagzian, Arash, et al.
Veröffentlicht: (2025)
Laplacian Score Sharpening for Mitigating Hallucination in Diffusion Models
von: C, Barath Chandran., et al.
Veröffentlicht: (2025)
von: C, Barath Chandran., et al.
Veröffentlicht: (2025)
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
von: Zou, Qiran, et al.
Veröffentlicht: (2026)
von: Zou, Qiran, et al.
Veröffentlicht: (2026)
Representation Collapsing Problems in Vector Quantization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024)
RGB-to-Polarization Estimation: A New Task and Benchmark Study
von: Lin, Beibei, et al.
Veröffentlicht: (2025)
von: Lin, Beibei, et al.
Veröffentlicht: (2025)
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
Early Quantization Shrinks Codebook: A Simple Fix for Diversity-Preserving Tokenization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
Physical Reasoning and Object Planning for Household Embodied Agents
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
Masked Generative Priors Improve World Models Sequence Modelling Capabilities
von: Meo, Cristian, et al.
Veröffentlicht: (2024)
von: Meo, Cristian, et al.
Veröffentlicht: (2024)
Efficient Causal Graph Discovery Using Large Language Models
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2024)
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2024)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
von: Du, Wanjun, et al.
Veröffentlicht: (2026)
von: Du, Wanjun, et al.
Veröffentlicht: (2026)
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
RaindropGS: A Benchmark for 3D Gaussian Splatting under Raindrop Conditions
von: Teng, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Teng, Zhiqiang, et al.
Veröffentlicht: (2025)
Unsupervised Concept Discovery Mitigates Spurious Correlations
von: Arefin, Md Rifat, et al.
Veröffentlicht: (2024)
von: Arefin, Md Rifat, et al.
Veröffentlicht: (2024)
Automated Discovery of Test Oracles for Database Management Systems Using LLMs
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
DeepFDR: A Deep Learning-based False Discovery Rate Control Method for Neuroimaging Data
von: Kim, Taehyo, et al.
Veröffentlicht: (2023)
von: Kim, Taehyo, et al.
Veröffentlicht: (2023)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
Unlearning via Sparse Representations
von: Shah, Vedant, et al.
Veröffentlicht: (2023)
von: Shah, Vedant, et al.
Veröffentlicht: (2023)
A False Discovery Rate Control Method Using a Fully Connected Hidden Markov Random Field for Neuroimaging Data
von: Kim, Taehyo, et al.
Veröffentlicht: (2025)
von: Kim, Taehyo, et al.
Veröffentlicht: (2025)
SciPaths: Forecasting Pathways to Scientific Discovery
von: Chamoun, Eric, et al.
Veröffentlicht: (2026)
von: Chamoun, Eric, et al.
Veröffentlicht: (2026)
AutoDataset: A Lightweight System for Continuous Dataset Discovery and Search
von: Yang, Junzhe, et al.
Veröffentlicht: (2026)
von: Yang, Junzhe, et al.
Veröffentlicht: (2026)
GeoComplete: Geometry-Aware Diffusion for Reference-Driven Image Completion
von: Lin, Beibei, et al.
Veröffentlicht: (2025)
von: Lin, Beibei, et al.
Veröffentlicht: (2025)
Improving Sparse Memory Finetuning
von: Goyal, Satyam, et al.
Veröffentlicht: (2026)
von: Goyal, Satyam, et al.
Veröffentlicht: (2026)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
von: Gupta, Prakhar, et al.
Veröffentlicht: (2026)
von: Gupta, Prakhar, et al.
Veröffentlicht: (2026)
Semiparametric Causal Discovery and Inference with Invalid Instruments
von: Zou, Jing, et al.
Veröffentlicht: (2025)
von: Zou, Jing, et al.
Veröffentlicht: (2025)
Evolution Guided Generative Flow Networks
von: Ikram, Zarif, et al.
Veröffentlicht: (2024)
von: Ikram, Zarif, et al.
Veröffentlicht: (2024)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
von: Padwal, Vedant
Veröffentlicht: (2026)
von: Padwal, Vedant
Veröffentlicht: (2026)
The Biased Oracle: Assessing LLMs' Understandability and Empathy in Medical Diagnoses
von: Yao, Jianzhou, et al.
Veröffentlicht: (2025)
von: Yao, Jianzhou, et al.
Veröffentlicht: (2025)
Safeguarding the Truth of High-Value Price Oracle Task: A Dynamically Adjusted Truth Discovery Method
von: Xian, Youquan, et al.
Veröffentlicht: (2024)
von: Xian, Youquan, et al.
Veröffentlicht: (2024)
Adapting to Heterophilic Graph Data with Structure-Guided Neighbor Discovery
von: Tenorio, Victor M., et al.
Veröffentlicht: (2025)
von: Tenorio, Victor M., et al.
Veröffentlicht: (2025)
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
Diversed Model Discovery via Structured Table Discovery
von: Dong, Zhengyuan, et al.
Veröffentlicht: (2026)
von: Dong, Zhengyuan, et al.
Veröffentlicht: (2026)
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
von: Luo, Yulin, et al.
Veröffentlicht: (2024)
von: Luo, Yulin, et al.
Veröffentlicht: (2024)
Language Guided Skill Discovery
von: Rho, Seungeun, et al.
Veröffentlicht: (2024)
von: Rho, Seungeun, et al.
Veröffentlicht: (2024)
Neural-Guided Equation Discovery
von: Brugger, Jannis, et al.
Veröffentlicht: (2025)
von: Brugger, Jannis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
von: Anumasa, Srinivas, et al.
Veröffentlicht: (2025) -
HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
von: Chen, Tingting, et al.
Veröffentlicht: (2025) -
Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via inference-time Multi-Views Brainstorming
von: Lagzian, Arash, et al.
Veröffentlicht: (2025) -
Laplacian Score Sharpening for Mitigating Hallucination in Diffusion Models
von: C, Barath Chandran., et al.
Veröffentlicht: (2025) -
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
von: Zou, Qiran, et al.
Veröffentlicht: (2026)