Toolbox/Running agents in production

Observability & tracing

tracing, logging and debugging agents

8 repositories, most starred first
State of the category

raga-ai-hub/RagaAI-Catalyst raga-ai-hub/RagaAI-Catalyst is the clear leader; if you want a single self‑hosted platform for multi‑agent tracing, timeline analytics, evaluation and guardrails, it’s the most mature choice. Pick other projects only when you need lighter instrumentation, OTel-native traces, production-to‑CI replay tests, per‑line provenance, or fully local developer UIs.

Lightweight cross‑framework replay
Teams that run many different agent frameworks need minimal instrumentation, session replays and LLM cost tracking without replacing their stack, so lighter, framework‑agnostic tools matter for adoption and fast debugging.
OpenTelemetry‑native tracing
Organizations standardized on OpenTelemetry want AI traces and evaluations to live in the same telemetry stack to reuse existing collectors, dashboards and automation.
Replayable production tests
Promoting real failing traces into hermetic, replayable CI tests prevents regressions and makes debugging reproducible, which is critical for shipping agents safely.
Per‑step provenance & blame
Teams that need auditable, content‑addressed provenance and per‑line ‘blame’ for generated changes require finer-grained VCS‑style tracking that most observability tools don’t provide.
Local dev UX & plan review
Some teams prioritize fully local, privacy‑first developer experiences and real‑time plan review or playful visualizations over cloud dashboards, so lightweight local UIs remain attractive.
Find your fit

Which one matches your setup?

Answer any of the questions — the shortlist updates as you go. Recommendations come from the capability passports below, nothing else.

Where should reviews happen?
Can code leave your infrastructure?
What matters most?
Model access?
Pick at least one answer to get a shortlist.
Side by side

Comparison matrix

Axes are extracted from each project's docs by our review pipeline; the maturity score is computed from stars, growth and commit activity — not an opinion. Click a column to sort.

Runs inModelsContextCost to run
2.7k +19/7d
Web appCLICoding-agent pluginBYOKOllama / localFull-stack tracesRule engine filtersSelf-hostableSelf-hosted; you pay for LLM/provider API usage.
5.8k +6/7d
CLIWeb appBYOKOpenAIAnthropicFixed providerRuntime tracingNone mentionedSelf-hostableAgentOps API key for hosted dashboard; can self-host
1.2k
GitHub ActionCICLIWeb appBYOKOpenAIAnthropicNo repo contextClustering & gatingSelf-hostable / localSelf-hosted; CI replays use recorded fixtures (no model spend).
1.2k
GitHub ActionCICLIWeb appBYOKOpenAIAnthropicPer-run onlyClustering & gatingSelf-hostableCI replay $0; live judges use your model API key.
639
CLIWeb appAnthropicOpenAIFixed providerWhole-repo analysisMinimal signalsFully localFree, local
468 +8/7d
CLIOpenAIAnthropicFixed providerSession events onlyNone mentionedFully localFree, local
16.2k +11/7d
CLIWeb appCIBYOKDiff/file onlyConfigurable thresholdsSelf-hostableRequires RagaAI account; external LLM usage billed to your provider.
789 +3/7d
CLIIDECoding-agent pluginAnthropicOpenAIFixed providerWorkspace snapshotIgnore & dedupeSelf-hostableFree local CLI; external model/API usage may incur costs
At a glance

Capability profiles

Six axes, 0–5 each. The shape tells you the strategy: a wide hexagon is a generalist, a spike is a specialist.

ContextNoiseCustomPrivacyModelsMaturity

OpenLIT provides an OpenTelemetry-native, open-source AI observability platform combining traces, built-in evaluations, a rule engine, prompt management and guardrails in one stack.

ContextNoiseCustomPrivacyModelsMaturity

Provides end-to-end agent session replays, LLM cost tracking and debugging across many agent frameworks with minimal instrumentation.

ContextNoiseCustomPrivacyModelsMaturity

Promotes real failing production traces into hermetic, replayable regression tests that run offline in CI and block PRs.

ContextNoiseCustomPrivacyModelsMaturity

Promotes real failing production traces into hermetic, replayable regression cases that run offline in CI and block PRs, instead of relying on hand-authored datasets.

ContextNoiseCustomPrivacyModelsMaturity

Provides a local, real-time observability map of coding agents' plans, tool calls, and file edits without any cloud service, account, or telemetry.

ContextNoiseCustomPrivacyModelsMaturity

Visualizes multiple coding agents as a whimsical, glanceable pixel-art terminal 'office' with animated coworkers and per-agent session telemetry.

ContextNoiseCustomPrivacyModelsMaturity

Combines agent/LLM tracing, evaluation, guardrails and red‑teaming with a self‑hosted dashboard and execution-timeline analytics for multi-agent debugging.

ContextNoiseCustomPrivacyModelsMaturity

Provides content-addressed, per-step audit trails and per-line 'blame' tied to the exact prompt and conversation that produced a change—something traditional VCS and most tools do not track.

All repositories (8)

Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm…

16.2k
+117d
Python
2 yrs

Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and a…

5.8k
+67d
Python
3 yrs

Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, E…

2.7k
+197d
TypeScript
2 yrs

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect,…

1.2k
-17d
Python
3 mo

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect,…

1.2k
-17d
Python
3 mo

Version control for AI agents — track what your agent did, blame any line to a prompt, inspect any step.

789
+37d
Go
4 mo

An infinite canvas for your AI coding agents. Every repo is a region on one zoomable live map - watch Claude C…

639
HTML
2 wk

Terminal pixel-art office for AI coding agents

468
+87d
Rust
3 mo

More in Running agents in production