Skip to content
Go back

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility

Published:  at  02:00 AM

Publication Details

Status: Preprint on ArXiv

Authors: Sasi Kiran Gaddipati, Diyana Muhammed, Farhana Keya, Gollam Rabby, Sören Auer

ArXiv: 2605.16616

Abstract

Autonomous research systems capable of generating complete scientific manuscripts have advanced rapidly, yet robust and realistic evaluation frameworks have failed to keep pace. To bridge this gap, we introduce MLReplicate, an end-to-end benchmark evaluating autonomous research systems on machine learning reproducibility.

The benchmark was constructed from ICML 2025 outstanding papers reformulated into standardized input specifications and evaluated across 6 state-of-the-art research systems: AI Scientist-V1, AI Scientist-V2, Agent Laboratory, CycleResearcher, AI Researcher, and Tiny Scientist, yielding 45 generated manuscripts, with 3 failed experiments.

Outputs are assessed using a dual-protocol approach that combines automated conference-style review and structured expert human evaluation, while tracking computational cost, runtime, and the amount of required human intervention.

Key Findings

MLReplicate exposes a substantial gap between current autonomous research systems and genuine scientific rigor, and establishes a practical, extensible evaluation framework for systematic progress toward trustworthy AI-driven scientific discovery.

Access



Next Post
AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science