Toolbox/Applied agents

ML experiment agents

agents that run training experiments, tune and evolve models

16 repositories, most starred first
State of the category

SakanaAI/AI-Scientist SakanaAI/AI-Scientist is the de facto leader by adoption and offers a broad, end-to-end autonomous scientific-discovery stack; the single most useful thing to know is to pick based on deployment and audit needs (self-hosted vs cloud, reproducibility, and long-running state), not star count alone.

Reproducibility & Auditing
Teams running experiments need verifiable, auditable outputs and deterministic guards so results are trustworthy and debuggable, otherwise agent-generated pipelines are hard to validate or reuse.
Local / Self-hosting
Privacy, regulatory compliance, and control over cost/performance push teams to prefer projects that can run fully on-prem or on local GPUs instead of cloud-only offerings.
Long-horizon Statefulness
Real research workflows need persistent memory, branching experiments and resumability so knowledge and artifacts accumulate across long-running projects rather than being one-shot outputs.
Cost-efficient Runtime
Automating continuous or long-running experiments requires techniques to minimize LLM calls and optimize compute so teams can run 24/7 without prohibitive cloud bills.
Specialized Workflows & Integrations
Different teams need tailored workflows (ML code optimization, publication pipelines, IDE/agent plugin support, or specific search algorithms) rather than a one-size-fits-all agent.
Find your fit

Which one matches your setup?

Answer any of the questions — the shortlist updates as you go. Recommendations come from the capability passports below, nothing else.

Where should reviews happen?
Can code leave your infrastructure?
What matters most?
Model access?
Pick at least one answer to get a shortlist.
Side by side

Comparison matrix

Axes are extracted from each project's docs by our review pipeline; the maturity score is computed from stars, growth and commit activity — not an opinion. Click a column to sort.

Runs inModelsContextCost to run
13.8k
CLICoding-agent pluginBYOKOpenAIAnthropicDiff + related filesSeverity gating & auditsCloud APIs (your key)Your API key or model subscription
6.8k +59/7d
CLIBYOKOpenAIAnthropicGeminiOllama / localFile + related filesCascade & novelty filtersRun fully localYour API key; per-iteration LLM costs (local models nearly free)
4.3k
CLIWeb appBYOKOpenAIAnthropicFixed providerDiff + related filesBasic filtering & gatingSelf-hostableRequires your cloud model API keys (per-use billing)
3.2k
CLIWeb appCoding-agent pluginIDEOpenAIAnthropicGeminiOllama / localFixed providerWhole-repo analysisHuman-in-the-loopSelf-hostableDepends on chosen runner — uses your model provider (API) or local models
522
CLICoding-agent pluginIDEOpenAIAnthropicGeminiFixed providerManifest + on-demandVerification & provenanceCloud APIs (user keys)Uses your agent/provider API (may be paid).
14.3k
CLIBYOKOpenAIAnthropicGeminiFixed providerTemplate-level viewEnsemble reviewsSelf-hostable (local GPU)Requires your API keys for cloud models; local GPU needed for open-weight runs.
6.9k
CLIOpenAIAnthropicGeminiRelated filesConfig-based controlsCloud APIs (your keys)Your API key, per-run (README cites ≈$15–$20 for experiments + ≈$5 for writing with default models)
1.5k
CLIWeb appBYOKSandbox workspaceNone mentionedCloud via API keyYour API key; cloud model usage charges.
1.4k
CLIWeb appBYOKOpenAIAnthropicGeminiOllama / localWorkspace-wideMetric-guided pruningFully local optionRequires your API key (OpenAI/Anthropic/etc.), or can run with local LLMs.
1.4k +9/7d
CLIOpenAIAnthropicBYOKWhole-repo accessConfigurable filtersCloud API keysYour API keys (OpenAI / Anthropic / OpenRouter)
1.2k
CLIBYOKOpenAIAnthropicWhole-repo analysisPhase gates & rate-limitAPI-key cloudYour API key; low LLM cost (README reports ~$0.08/day).
249
CLICoding-agent pluginIDEWeb appAnthropicOpenAIFixed providerDiff + related filesConvergence guardsAgent-dependent (cloud APIs)Runs via host agent (e.g., Claude, Copilot) — may require subscription or API key
Ranked by maturity — 4 more in the full catalog below.
At a glance

Capability profiles

Six axes, 0–5 each. The shape tells you the strategy: a wide hexagon is a generalist, a spike is a specialist. Showing the 8 most established — the rest are in the full catalog.

ContextNoiseCustomPrivacyModelsMaturity

Provides a lightweight, skill-based autonomous research workflow with built-in cross-model reviewer loops and deterministic integrity audits (Anti-Autoresearch), targeted at reproducible ML research rather than generic agent tasks.

ContextNoiseCustomPrivacyModelsMaturity

Combines MAP-Elites, island-based evolution and LLM ensembles to autonomously discover novel, hardware-optimized algorithms with reproducible scientific pipelines.

ContextNoiseCustomPrivacyModelsMaturity

Auto-distills a growing memory/knowledge graph into reusable AutoSkills and a multi-agent research loop, enabling long-running self-evolving AI scientists.

ContextNoiseCustomPrivacyModelsMaturity

A local-first, long-horizon autonomous research studio that preserves full experiment state, branches, artifacts and lets humans inspect and take over—unlike one-shot cloud chat agents.

ContextNoiseCustomPrivacyModelsMaturity

Provides an agent-native, structured research artifact format that enforces verifiable, auditable experiments and preserves dead-end knowledge, unlike general-purpose agent toolkits that produce ephemeral, untraceable outputs.

ContextNoiseCustomPrivacyModelsMaturity

Automates end-to-end scientific discovery by generating hypotheses, running experiments, and producing full LaTeX papers from templates.

ContextNoiseCustomPrivacyModelsMaturity

End-to-end autonomous scientific discovery: it uses progressive agentic tree search to ideate, run experiments, analyze results, and draft papers without human-authored templates.

ContextNoiseCustomPrivacyModelsMaturity

Provides a ready sandbox that combines Scientific Agent Skills with the Claude/ADK stack to run interactive, agentic ML experiments locally.

All repositories (16)

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑‍🔬

14.3k
Jupyter Notebook
1 yr

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model re…

13.8k
Python
4 mo

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

6.9k
Python
1 yr

Open-source implementation of AlphaEvolve

6.8k
+597d
Python
1 yr

🔬 Harness Vibe Research with Self-evolving AI Scientists

4.3k
Python
6 mo

Now, Stronger AI Pushes Frontiers, Stronger Our Shared Future.

3.2k
TypeScript
10 mo

An agentic Machine Learning Engineer

1.5k
Python
8 mo

AIDE: AI-Driven Exploration in the Space of Code. The machine Learning engineering agent that automates AI R&D…

1.4k
Python
2 yrs

InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery

1.4k
+97d
Python
1 yr

🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring,…

1.2k
-247d
Python
3 mo

Kosmos: An AI Scientist for Autonomous Discovery - An implementation and adaptation to be driven by Claude Cod…

556
Python
8 mo

Research Ecosystem for Rigorous and Trustworthy AI Scientists

522
HTML
3 mo

An autonomous AI scientist: a multi-agent loop over literature, experiments, self-critique and write-up, with…

463
Python
1 mo

Code associated with the paper An AI system to help scientists write expert-level empirical software

298
Jupyter Notebook
10 mo

One file. Your AI coding agent becomes a scientist. 30+ experiments while you sleep.

249
Python
4 mo

Multi-agent AI scientist that turns experimental data into > publication-ready research papers.

164
Python
2 mo

More in Applied agents