ML experiment agents
agents that run training experiments, tune and evolve models
SakanaAI/AI-Scientist SakanaAI/AI-Scientist is the de facto leader by adoption and offers a broad, end-to-end autonomous scientific-discovery stack; the single most useful thing to know is to pick based on deployment and audit needs (self-hosted vs cloud, reproducibility, and long-running state), not star count alone.
Which one matches your setup?
Answer any of the questions — the shortlist updates as you go. Recommendations come from the capability passports below, nothing else.
Comparison matrix
Axes are extracted from each project's docs by our review pipeline; the maturity score is computed from stars, growth and commit activity — not an opinion. Click a column to sort.
| Runs in | Models | Context | Cost to run | |||||
|---|---|---|---|---|---|---|---|---|
⭐ 13.8k | CLICoding-agent plugin | BYOKOpenAIAnthropic | Diff + related files | Severity gating & audits | ●●●●● | Cloud APIs (your key) | ●●●●● | Your API key or model subscription |
⭐ 6.8k +59/7d | CLI | BYOKOpenAIAnthropicGeminiOllama / local | File + related files | Cascade & novelty filters | ●●●●● | Run fully local | ●●●●● | Your API key; per-iteration LLM costs (local models nearly free) |
⭐ 4.3k | CLIWeb app | BYOKOpenAIAnthropicFixed provider | Diff + related files | Basic filtering & gating | ●●●●● | Self-hostable | ●●●●● | Requires your cloud model API keys (per-use billing) |
⭐ 3.2k | CLIWeb appCoding-agent pluginIDE | OpenAIAnthropicGeminiOllama / localFixed provider | Whole-repo analysis | Human-in-the-loop | ●●●●● | Self-hostable | ●●●●● | Depends on chosen runner — uses your model provider (API) or local models |
⭐ 522 | CLICoding-agent pluginIDE | OpenAIAnthropicGeminiFixed provider | Manifest + on-demand | Verification & provenance | ●●●●● | Cloud APIs (user keys) | ●●●●● | Uses your agent/provider API (may be paid). |
⭐ 14.3k | CLI | BYOKOpenAIAnthropicGeminiFixed provider | Template-level view | Ensemble reviews | ●●●●● | Self-hostable (local GPU) | ●●●●● | Requires your API keys for cloud models; local GPU needed for open-weight runs. |
⭐ 6.9k | CLI | OpenAIAnthropicGemini | Related files | Config-based controls | ●●●●● | Cloud APIs (your keys) | ●●●●● | Your API key, per-run (README cites ≈$15–$20 for experiments + ≈$5 for writing with default models) |
⭐ 1.5k | CLIWeb app | BYOK | Sandbox workspace | None mentioned | ●●●●● | Cloud via API key | ●●●●● | Your API key; cloud model usage charges. |
⭐ 1.4k | CLIWeb app | BYOKOpenAIAnthropicGeminiOllama / local | Workspace-wide | Metric-guided pruning | ●●●●● | Fully local option | ●●●●● | Requires your API key (OpenAI/Anthropic/etc.), or can run with local LLMs. |
⭐ 1.4k +9/7d | CLI | OpenAIAnthropicBYOK | Whole-repo access | Configurable filters | ●●●●● | Cloud API keys | ●●●●● | Your API keys (OpenAI / Anthropic / OpenRouter) |
⭐ 1.2k | CLI | BYOKOpenAIAnthropic | Whole-repo analysis | Phase gates & rate-limit | ●●●●● | API-key cloud | ●●●●● | Your API key; low LLM cost (README reports ~$0.08/day). |
⭐ 249 | CLICoding-agent pluginIDEWeb app | AnthropicOpenAIFixed provider | Diff + related files | Convergence guards | ●●●●● | Agent-dependent (cloud APIs) | ●●●●● | Runs via host agent (e.g., Claude, Copilot) — may require subscription or API key |
Capability profiles
Six axes, 0–5 each. The shape tells you the strategy: a wide hexagon is a generalist, a spike is a specialist. Showing the 8 most established — the rest are in the full catalog.
Provides a lightweight, skill-based autonomous research workflow with built-in cross-model reviewer loops and deterministic integrity audits (Anti-Autoresearch), targeted at reproducible ML research rather than generic agent tasks.
Combines MAP-Elites, island-based evolution and LLM ensembles to autonomously discover novel, hardware-optimized algorithms with reproducible scientific pipelines.
Auto-distills a growing memory/knowledge graph into reusable AutoSkills and a multi-agent research loop, enabling long-running self-evolving AI scientists.
A local-first, long-horizon autonomous research studio that preserves full experiment state, branches, artifacts and lets humans inspect and take over—unlike one-shot cloud chat agents.
Provides an agent-native, structured research artifact format that enforces verifiable, auditable experiments and preserves dead-end knowledge, unlike general-purpose agent toolkits that produce ephemeral, untraceable outputs.
Automates end-to-end scientific discovery by generating hypotheses, running experiments, and producing full LaTeX papers from templates.
End-to-end autonomous scientific discovery: it uses progressive agentic tree search to ideate, run experiments, analyze results, and draft papers without human-authored templates.
Provides a ready sandbox that combines Scientific Agent Skills with the Claude/ADK stack to run interactive, agentic ML experiments locally.
All repositories (16)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑🔬
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model re…
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
Open-source implementation of AlphaEvolve
Now, Stronger AI Pushes Frontiers, Stronger Our Shared Future.
AIDE: AI-Driven Exploration in the Space of Code. The machine Learning engineering agent that automates AI R&D…
InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery
🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring,…
Kosmos: An AI Scientist for Autonomous Discovery - An implementation and adaptation to be driven by Claude Cod…
Research Ecosystem for Rigorous and Trustworthy AI Scientists
An autonomous AI scientist: a multi-agent loop over literature, experiments, self-critique and write-up, with…
Code associated with the paper An AI system to help scientists write expert-level empirical software
One file. Your AI coding agent becomes a scientist. 30+ experiments while you sleep.
Multi-agent AI scientist that turns experimental data into > publication-ready research papers.