Toolbox/Agent capabilities

Computer use

agents that operate a desktop or GUI

25 repositories, most starred first
State of the category

NanmiCoder/cc-haha NanmiCoder/cc-haha is the de-facto leader by adoption and breadth, offering a desktop-first, local-capable workstation that bundles code/desktop control and integrations. The single most useful decision for teams is whether they need a fully local/deterministic agent or are comfortable relying on cloud models and specialized device integrations — that choice narrows the useful projects quickly.

Local determinism
Teams that require offline, auditable or deterministic automation (no model calls on healthy runs) need projects that prioritize fully local execution and deterministic workflows rather than cloud-driven guessing.
App & device control
Controlling mobile apps, simulators and specific platform GUIs requires specialized tooling for taps/swipes, accessibility outlines and device streaming that general-purpose agents do not provide.
Sandboxing & audit
Enterprise and security-conscious teams need per-agent isolation, policy enforcement and auditable action trails to safely run autonomous agents in production.
Domain-specific workflows
Many teams work in vertical workflows (office docs, CAD/EDA, Blender, bioinformatics, system ops) and need agents that integrate domain libraries, typed actions and reproducible session logs.
Adaptive & research
Researchers and power users are building self-improving skill trees, research-backed agent frameworks and centralized session/workspace tooling that focus on learning, evaluation and reusable Skills rather than one-off automation.
Find your fit

Which one matches your setup?

Answer any of the questions — the shortlist updates as you go. Recommendations come from the capability passports below, nothing else.

Where should reviews happen?
Can code leave your infrastructure?
What matters most?
Model access?
Pick at least one answer to get a shortlist.
Side by side

Comparison matrix

Axes are extracted from each project's docs by our review pipeline; the maturity score is computed from stars, growth and commit activity — not an opinion. Click a column to sort.

Runs inModelsContextCost to run
14.3k +38/7d
CLIWeb appAnthropicOpenAIBYOKOllama / localDiff + WorktreeApproval gatingSelf-hostableRequires configuring your model provider API key (third‑party billing applies)
14.1k +48/7d
CLIWeb appBYOKGeminiWhole-repo analysisHuman-in-the-loopCloud API keyYour LLM API key; provider charges apply
3.8k +225/7d
CLIFixed providerRecording onlyManual review/editCopilot cloud requiredRequires GitHub Copilot access
1.9k +69/7d
Web appCLIBYOKOpenAIAnthropicGeminiDiff + related filesNone mentionedSelf-hostableUse your API key or third‑party relays for external models; desktop app is local-first
1.9k +94/7d
CLICoding-agent pluginFixed providerLocal UI onlyNone mentionedFully localFree, local
1.7k +84/7d
CLIWeb appCoding-agent pluginBYOKAnthropicOpenAIGeminiFixed providerDevice-only contextReview & retryBYO API keys (cloud)Requires your model API keys for model usage; local backend and Android client run from the repo.
1.3k +35/7d
CLICICoding-agent pluginFixed providerDiff onlyBasic filtersFully localFree, local
750 +51/7d
CLIWeb appBYOKWorkspace-awarePlan-mode gatingSelf-hostableRequires pi and your provider API keys (configured via pi)
575 +141/7d
CLIBYOKProject + related filesRevision & transaction checksCloud APIs (your key)Your model provider API key (provider charges may apply)
12.2k +22/7d
CLIWeb appBYOKOpenAIAnthropicGeminiFixed providerSingle-screen contextBasic filtersCloud APIs (your key)Your API keys (provider-dependent)
5.0k
Web appBYOKDocument-levelSnapshots & diffsSelf-hostable (BYOK)Free app; uses Genspark proxy by default or your API key for model costs.
4.2k +955/7d
Web appCLICIBYOKOpenAIAnthropicGeminiWorkspace accessPolicy and auditCloud APIs w/ keyRequires CopilotKit license and your model API key.
Ranked by maturity — 13 more in the full catalog below.
At a glance

Capability profiles

Six axes, 0–5 each. The shape tells you the strategy: a wide hexagon is a generalist, a spike is a specialist. Showing the 8 most established — the rest are in the full catalog.

ContextNoiseCustomPrivacyModelsMaturity

A desktop-first, local-first Claude Code workstation that combines worktree-aware code diffs, Computer Use desktop control, H5 remote access and IM integrations in one app.

ContextNoiseCustomPrivacyModelsMaturity

Self-evolving agent that crystallizes each solved task into a persistent, reusable personal skill tree from a minimal (~3K-line) seed.

ContextNoiseCustomPrivacyModelsMaturity

Converts a single on-screen recording (clicks, window switches, URLs, and optional narration) into a generalized, reusable Skill or scheduled Automation using the GitHub Copilot CLI.

ContextNoiseCustomPrivacyModelsMaturity

A local‑first desktop AI agent that actually performs file-system operations, runs bash/managed processes, and extends via an MCP/Skills ecosystem with an optional remote Gateway.

ContextNoiseCustomPrivacyModelsMaturity

Open-source, self-hosted MCP implementation of Computer Use that provides a local alternative to Codex Computer Use and integrates with multiple agent clients.

ContextNoiseCustomPrivacyModelsMaturity

Built as a backend-managed mobile operator system with explicit orchestration, standby device dispatch, and separable planning/VLM roles for long-running phone workflows.

ContextNoiseCustomPrivacyModelsMaturity

Provides a token-efficient, agent-native observe→act CLI that lets agents read compact accessibility outlines and perform alias-cached taps on both iOS Simulators and Android devices.

ContextNoiseCustomPrivacyModelsMaturity

Provides a desktop workspace that runs multiple local pi RPC agent processes to manage multi-project sessions, history, Git integration and import of Codex/Claude local sessions in one unified GUI.

All repositories (25)

Local-first cross-platform desktop workspace for Claude Code / agents: multi-agent, Git worktrees, code diffs,…

14.3k
+387d
TypeScript
5 mo

Self-evolving agent: grows skill tree from 3.3K-line seed, achieving full system control with 6x less token co…

14.1k
+487d
Python
7 mo

Agent S: an open agentic framework that uses computers like a human

12.2k
+227d
Python
1 yr

AppAgent: Multimodal Agents as Smartphone Users, an LLM-based multimodal agent framework designed to operate s…

6.9k
+87d
Python
2 yrs

Free, open-source alternative to Microsoft Office with built-in AI agents — Word (.docx), Excel (.xlsx), Power…

5.0k
TypeScript
1 mo

Open-source AI coworkers that each get a computer of their own: a browser, files and tools, with every action…

4.2k
+9557d
TypeScript
2 wk

Desktop app that records your on-screen work session and uses the GitHub Copilot CLI to reconstruct it as an i…

3.8k
+2257d
TypeScript
1 mo

Open-source AI agent desktop app for Windows & macOS. One-click install Claude Code, MCP tools, and Skills — w…

2.1k
+247d
TypeScript
7 mo

A fully functional AI Agent desktop client that supports Webui access and can be creatively customized and exp…

1.9k
+697d
TypeScript
3 mo

👾 Open Computer Use – Open-Source Alternative to Codex Computer Use

1.9k
+947d
Swift
4 mo

Compiles a demonstrated GUI task into a program that reports VERIFIED only if an independent check agrees. pip…

1.7k
+117d
Python
3 yrs

OpenGUI is an Android GUI agent framework for phone-use AI that can see, plan, and operate real mobile apps th…

1.7k
+847d
TypeScript
4 mo

The first open-source Artificial Narrow Intelligence generalist agentic framework Computer-Using-Agent that fu…

1.3k
Python
2 yrs

Give your AI agent eyes and hands on iOS Simulator and Android emulator/devices.

1.3k
+357d
Swift
2 mo

PiDeck 是一个开源的桌面工作台,用于在本地项目目录中统一管理 pi Agent 会话,并支持导入 Codex、Claude 本地会话以便统一浏览和恢复。支持多项目工作区、会话历史、Git 集成、内置终端、模型配置和…

750
+517d
TypeScript
3 mo

One chat. Every agent. A universal interface for AI agents running on your computer.

704
+37d
TypeScript
5 mo

Autonomous self-evolving agents. Vision-grounded layered memory and self-written skills for LLM agents that op…

692
-57d
Python
4 mo

Open source previs software in the browser: block a scene, pose characters, author camera moves and cuts, then…

575
+1417d
JavaScript
1 mo

嘉立创EDA专业版(EasyEDA Pro)自动化:给 AI harness 装上画板的「手」—— 一套 typed 原理图/PCB 动作,CLI / Agent Skill / stdio MCP 三形态融合接入。承接…

349
+407d
Go
2 mo

DeepSeek Harness (DSH) plugin: a live iOS Simulator — and a USB-connected iPhone — inside the conversation. 22…

275
+127d
TypeScript
2 wk

A macOS Dynamic Island for AI coding agents: Claude Code, Codex, Gemini CLI, Cursor, hooks, skills, and local…

212
Rust
4 mo

One macOS app for Claude Code, Codex, and every agent runtime you use — scheduled runs, global hotkey launcher…

142
+37d
TypeScript
2 mo

基于多模态视觉感知与 LLM Agent 的 macOS 微信自动化框架 | Visual RPA for WeChat

93
+67d
Python
4 mo

A menu-bar macOS agent loop: every N minutes it asks Claude/Codex what's overloading your Mac and suggests one…

22
Swift
2 mo

The bioinformatics agent desktop for clinicians and wet-lab scientists — built on DeepSeek Harness. One-click…

20
Python
5 d
Don't want to self-host?

Hosted alternatives

If running your own reviewer is more ops than you want, these managed services cover the same job.

Claude Computer Use

Anthropic's hosted computer-use capability: Claude drives a virtual desktop through the API, so you get screen-grounded control without building the perceive-act loop or running a VM.

Try Claude Computer Use
OpenAI Operator

A managed agent that operates a browser and desktop tasks on OpenAI's infrastructure, the hosted counterpart to running a local computer-use harness yourself.

Try OpenAI Operator
Scrapybara

Rents ready-made cloud desktops and browsers for agents to control, so you point an existing agent at a managed machine instead of provisioning and securing one.

Try Scrapybara

More in Agent capabilities