deep-swe vs teaql-agent-kit

The two are close in size: 1.6k stars for deep-swe, 2.8k for teaql-agent-kit. Over the days we have tracked them deep-swe moved +130.1% and teaql-agent-kit +1.2%, so deep-swe is growing faster right now.

deep-swe leads on context depth and model freedom. teaql-agent-kit does not take any axis by a clear margin.

Stars and commit dates come from our own daily tracking. The six axes are read off each project's documentation by our review pipeline, so they describe what a project says about itself, not what we measured in its code.

Where they stand today

Provides long-horizon, behaviorally-graded software-engineering tasks with isolated sandbox execution and separate verifier environments (via Pier).

Stars
1.6k
Tracked growth
+130.1%
Maturity
Last commit
9d ago
Language
Python
License
Apache-2.0
Cost to run
Your API key (OpenAI/Anthropic), pay-per-run model costs

Provides a TEAQL-focused, auditable evaluation harness that measures software-engineering discipline and token-efficiency for coding agents rather than offering general-purpose agent automation.

Stars
2.8k
Tracked growth
+1.2%
Maturity
Last commit
2d ago
Language
Python
License
MIT
Cost to run
Your API key or model costs may apply
0%+130%90 tracked days
datacurve-ai/deep-sweteaql/teaql-agent-kit

Six axes, head to head

Each axis runs 0 to 5. The label under a score is what that project's own docs claim, not a category average.

Axisdeep-sweteaql-agent-kit
Context depth
How much of your codebase it sees before it answers: the open diff, the diff plus related files, or the whole repository.
Whole-repo analysis
Diff + related files
Noise control
How it keeps output volume down — severity thresholds, deduplication, incremental runs over new commits only.
Behavioral verification
Guides & checkpoints
Customization
How far it bends to your team: custom rules, prompts, style guides, per-path config.
Config + prompts
Config & prompts
Privacy
Whether your code stays on your own infrastructure: fully local, self-hostable, or cloud API only.
Cloud via API key
Cloud via API key
Model freedom
Whether you can point it at any provider, or it is wired to one.
Multiple providers
Single fixed provider
Setup ease
What it takes to get a first useful run out of it.
CLI + API key
Manual setup

Which one to pick

Pick deep-swe if…

Long-horizon evaluation — choose DeepSWE when you need realistic, multi-step engineering tasks with programmatic verifiers and sandboxed grading to measure end-to-end agent behavior.

  • Context depth: Whole-repo analysis (5/5 against 3/5)
  • Model freedom: Multiple providers (3/5 against 1/5)
Runs in cli, ci. Works with byok, openai, anthropic, gemini.

Pick teaql-agent-kit if…

Evaluation-first — pick this when you need reproducible, auditable benchmarks of coding agents working with TEAQL contracts, explicit guardrails, and token-efficiency measurement.

Runs in cli, ci. Works with other-fixed.

What people want from each one

Questions people ask

Is deep-swe better than teaql-agent-kit?

deep-swe leads on context depth and model freedom. teaql-agent-kit does not take any axis by a clear margin. deep-swe is worth picking when long-horizon evaluation — choose DeepSWE when you need realistic, multi-step engineering tasks with programmatic verifiers and sandboxed grading to measure end-to-end agent behavior.

Which of deep-swe and teaql-agent-kit keeps my code private?

deep-swe: Cloud via API key (3/5). teaql-agent-kit: Cloud via API key (3/5).

What does each one cost to run?

deep-swe: Your API key (OpenAI/Anthropic), pay-per-run model costs. teaql-agent-kit: Your API key or model costs may apply.

Full profiles: datacurve-ai/deep-swe and teaql/teaql-agent-kit. Everything else in Evals & benchmarks.