Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Attrach, Rafi Al, Fani, Rajna, Lobentanzer, Sebastian, Giner-Miguelez, Joan, Das, Debanshu, K., Varuni H., Sarwar, Nobin, Ghosh, Rajat, Archit, Anwai, Motghare, Surbhi, Parry, Christina Conrad, Oala, Luis, Grosso, Lara, Vanschoren, Joaquin, Vogler, Steffen, Goswami, Sujata, Rosenthal, Eric S., Ghassemi, Marzyeh, McDermott, Matthew, Pollard, Tom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations
von: Benjelloun, Omar, et al.
Veröffentlicht: (2026)
von: Benjelloun, Omar, et al.
Veröffentlicht: (2026)
FedMentalCare: Towards Privacy-Preserving Fine-Tuned LLMs to Analyze Mental Health Status Using Federated Learning Framework
von: Sarwar, Nobin
Veröffentlicht: (2025)
von: Sarwar, Nobin
Veröffentlicht: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
von: Sarwar, Nobin
Veröffentlicht: (2025)
von: Sarwar, Nobin
Veröffentlicht: (2025)
Croissant: A Metadata Format for ML-Ready Datasets
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2024)
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2024)
Rethinking Tokenization for Clinical Time Series: When Less is More
von: Attrach, Rafi Al, et al.
Veröffentlicht: (2025)
von: Attrach, Rafi Al, et al.
Veröffentlicht: (2025)
ViM-UNet: Vision Mamba for Biomedical Segmentation
von: Archit, Anwai, et al.
Veröffentlicht: (2024)
von: Archit, Anwai, et al.
Veröffentlicht: (2024)
Probabilistic Domain Adaptation for Biomedical Image Segmentation
von: Archit, Anwai, et al.
Veröffentlicht: (2023)
von: Archit, Anwai, et al.
Veröffentlicht: (2023)
Revisiting foundation models for cell instance segmentation
von: Archit, Anwai, et al.
Veröffentlicht: (2026)
von: Archit, Anwai, et al.
Veröffentlicht: (2026)
M3: Conversational LLMs Simplify Secure Clinical Data Access, Understanding, and Analysis
von: Attrach, Rafi Al, et al.
Veröffentlicht: (2025)
von: Attrach, Rafi Al, et al.
Veröffentlicht: (2025)
Coefficient of Variation Masking: A Volatility-Aware Strategy for EHR Foundation Models
von: Fani, Rajna, et al.
Veröffentlicht: (2025)
von: Fani, Rajna, et al.
Veröffentlicht: (2025)
FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
von: Sarwar, Nobin, et al.
Veröffentlicht: (2025)
von: Sarwar, Nobin, et al.
Veröffentlicht: (2025)
MedicoSAM: Robust Improvement of SAM for Medical Imaging
von: Archit, Anwai, et al.
Veröffentlicht: (2025)
von: Archit, Anwai, et al.
Veröffentlicht: (2025)
Parameter Efficient Fine-Tuning of Segment Anything Model for Biomedical Imaging
von: Teuber, Carolin, et al.
Veröffentlicht: (2025)
von: Teuber, Carolin, et al.
Veröffentlicht: (2025)
Segment Anything for Histopathology
von: Griebel, Titus, et al.
Veröffentlicht: (2025)
von: Griebel, Titus, et al.
Veröffentlicht: (2025)
A Standardized Machine-readable Dataset Documentation Format for Responsible AI
von: Jain, Nitisha, et al.
Veröffentlicht: (2024)
von: Jain, Nitisha, et al.
Veröffentlicht: (2024)
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
von: Afzal, Anum, et al.
Veröffentlicht: (2024)
von: Afzal, Anum, et al.
Veröffentlicht: (2024)
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
von: Jin, Qixuan, et al.
Veröffentlicht: (2024)
von: Jin, Qixuan, et al.
Veröffentlicht: (2024)
Measuring Stochastic Data Complexity with Boltzmann Influence Functions
von: Ng, Nathan, et al.
Veröffentlicht: (2024)
von: Ng, Nathan, et al.
Veröffentlicht: (2024)
SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
von: Ranjbar, Sepehr Kazemi, et al.
Veröffentlicht: (2025)
von: Ranjbar, Sepehr Kazemi, et al.
Veröffentlicht: (2025)
BioimageAIpub: a toolbox for AI-ready bioimaging data publishing
von: Dvoretskii, Stefan, et al.
Veröffentlicht: (2025)
von: Dvoretskii, Stefan, et al.
Veröffentlicht: (2025)
Quantifying the Expectation-Realisation Gap for Agentic AI Systems
von: Lobentanzer, Sebastian
Veröffentlicht: (2026)
von: Lobentanzer, Sebastian
Veröffentlicht: (2026)
The Agentic Automation Canvas: a structured framework for agentic AI project design
von: Lobentanzer, Sebastian
Veröffentlicht: (2026)
von: Lobentanzer, Sebastian
Veröffentlicht: (2026)
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
von: Gourabathina, Abinitha, et al.
Veröffentlicht: (2025)
von: Gourabathina, Abinitha, et al.
Veröffentlicht: (2025)
Views Can Be Deceiving: Improved SSL Through Feature Space Augmentation
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
Enabling Content Management Systems as an Information Source in Model-driven Projects
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2025)
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2025)
On the Readiness of Scientific Data for a Fair and Transparent Use in Machine Learning
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2024)
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2024)
Using Large Language Models to Enrich the Documentation of Datasets for Machine Learning
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2024)
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2024)
A domain-specific language for describing machine learning datasets
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2022)
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2022)
Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy
von: Teuber, Carolin, et al.
Veröffentlicht: (2026)
von: Teuber, Carolin, et al.
Veröffentlicht: (2026)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
Tiling artifacts and trade-offs of feature normalization in the segmentation of large biological images
von: Buglakova, Elena, et al.
Veröffentlicht: (2025)
von: Buglakova, Elena, et al.
Veröffentlicht: (2025)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
An Investigation of Memorization Risk in Healthcare Foundation Models
von: Tonekaboni, Sana, et al.
Veröffentlicht: (2025)
von: Tonekaboni, Sana, et al.
Veröffentlicht: (2025)
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2026)
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2026)
Can AI Relate: Testing Large Language Model Response for Mental Health Support
von: Gabriel, Saadia, et al.
Veröffentlicht: (2024)
von: Gabriel, Saadia, et al.
Veröffentlicht: (2024)
Robustness Beyond Known Groups with Low-rank Adaptation
von: Gourabathina, Abinitha, et al.
Veröffentlicht: (2026)
von: Gourabathina, Abinitha, et al.
Veröffentlicht: (2026)
What's in a Query: Polarity-Aware Distribution-Based Fair Ranking
von: Balagopalan, Aparna, et al.
Veröffentlicht: (2025)
von: Balagopalan, Aparna, et al.
Veröffentlicht: (2025)
Identifying Implicit Social Biases in Vision-Language Models
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations
von: Benjelloun, Omar, et al.
Veröffentlicht: (2026) -
FedMentalCare: Towards Privacy-Preserving Fine-Tuned LLMs to Analyze Mental Health Status Using Federated Learning Framework
von: Sarwar, Nobin
Veröffentlicht: (2025) -
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
von: Sarwar, Nobin
Veröffentlicht: (2025) -
Croissant: A Metadata Format for ML-Ready Datasets
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2024) -
Rethinking Tokenization for Clinical Time Series: When Less is More
von: Attrach, Rafi Al, et al.
Veröffentlicht: (2025)