better-harness vs deep-swe
The two are close in size: 2.2k stars for better-harness, 1.6k for deep-swe. Over the days we have tracked them better-harness moved +177.5% and deep-swe +130.1%, so better-harness is growing faster right now.
Neither one leads on the six capability axes, so the choice comes down to which of them fits the way you already work.
Stars and commit dates come from our own daily tracking. The six axes are read off each project's documentation by our review pipeline, so they describe what a project says about itself, not what we measured in its code.
Where they stand today
Reviews AI coding workflows end-to-end and produces evidence-bounded, prioritized findings across five Agent Work Loop dimensions rather than only analyzing final diffs.
- Stars
- 2.2k
- Tracked growth
- +177.5%
- Maturity
- ●●●●●
- Last commit
- 1d ago
- Language
- JavaScript
- License
- MIT
- Cost to run
- Depends on host and agent (may use paid services)
Provides long-horizon, behaviorally-graded software-engineering tasks with isolated sandbox execution and separate verifier environments (via Pier).
- Stars
- 1.6k
- Tracked growth
- +130.1%
- Maturity
- ●●●●●
- Last commit
- 9d ago
- Language
- Python
- License
- Apache-2.0
- Cost to run
- Your API key (OpenAI/Anthropic), pay-per-run model costs
Six axes, head to head
Each axis runs 0 to 5. The label under a score is what that project's own docs claim, not a category average.
| Axis | better-harness | deep-swe |
|---|---|---|
Context depth How much of your codebase it sees before it answers: the open diff, the diff plus related files, or the whole repository. | ●●●●● Whole-repo analysis | ●●●●● Whole-repo analysis |
Noise control How it keeps output volume down — severity thresholds, deduplication, incremental runs over new commits only. | ●●●●● Evidence-backed findings | ●●●●● Behavioral verification |
Customization How far it bends to your team: custom rules, prompts, style guides, per-path config. | ●●●●● Extensible configs & plugins | ●●●●● Config + prompts |
Privacy Whether your code stays on your own infrastructure: fully local, self-hostable, or cloud API only. | ●●●●● Self-hostable | ●●●●● Cloud via API key |
Model freedom Whether you can point it at any provider, or it is wired to one. | ●●●●● Multiple host providers | ●●●●● Multiple providers |
Setup ease What it takes to get a first useful run out of it. | ●●●●● Host-specific quick start | ●●●●● CLI + API key |
Which one to pick
Pick better-harness if…
Workflow-focused — pick Better Harness when you want honest, evidence-backed reviews of your agents' task understanding, execution, validation, delivery, and learning capture.
Pick deep-swe if…
Long-horizon evaluation — choose DeepSWE when you need realistic, multi-step engineering tasks with programmatic verifiers and sandboxed grading to measure end-to-end agent behavior.
What people want from each one
QoderAI/better-harness
datacurve-ai/deep-swe
Hacker News: DeepSWE results are unreliable – 3/3 DSv4 "failed" tasks solved with same model drew 3 points and 0 comments.
Questions people ask
Is better-harness better than deep-swe?
Neither one leads on the six capability axes, so the choice comes down to which of them fits the way you already work. better-harness is worth picking when workflow-focused — pick Better Harness when you want honest, evidence-backed reviews of your agents' task understanding, execution, validation, delivery, and learning capture.
Which of better-harness and deep-swe keeps my code private?
better-harness: Self-hostable (4/5). deep-swe: Cloud via API key (3/5).
What does each one cost to run?
better-harness: Depends on host and agent (may use paid services). deep-swe: Your API key (OpenAI/Anthropic), pay-per-run model costs.
Full profiles: QoderAI/better-harness and datacurve-ai/deep-swe. Everything else in Evals & benchmarks.