Agent skill

Your coding agent already knows Tracely

One command and Claude Code, Cursor or Copilot can instrument an agent, design the evaluation columns, and wire the pull-request gate — without you keeping a docs tab open or pasting snippets it half-remembers.

install
npx skills add https://github.com/Jwuthri/Tracely-ai --skill tracely

Powered by the open-source skills CLI. Add -g to install it for every project on the machine, or --agent '*' to install it into every agent you have. It's plain Markdown either way — you can also just read it on GitHub.

Why a skill and not just docs

A model asked to “add tracing to this agent” will produce something plausible. Plausible is the problem: agent observability has a handful of conventions that fail silently when you get them wrong. Nothing errors, traces still arrive, the dashboard still fills up — and six weeks later the workspace can't answer the question it was bought for.

The skill front-loads exactly those conventions, then keeps the deep reference material out of the way until the task actually calls for it. Ask for automatic instrumentation and the manual span API never enters the conversation.

What it teaches

Automatic tracing

instrument="auto", the provider and framework extras, @observe, the non-patching drop-ins, LangGraph, LiteLLM, first-party agent SDKs, redaction and threads.

Manual spans

Every observation type, multi-agent handoffs, RAG pipelines, shared state deltas, multimodal I/O, and the record-replay seam that makes CI hermetic.

Anything not Python

The OTLP conventions Tracely reads, so a TypeScript, Go or Ruby service lands as a first-class trace with no Tracely code at all.

Evaluator design

Structural checks before judges, picking the level, @VARIABLE templates, advisory verdicts, sequential grading, targeting and sampling to control spend.

The CI gate

Scenarios against your endpoint, adversarial red-team runs, hermetic replay of promoted failures, and the GitHub Action that blocks the PR.

Troubleshooting

Symptom to cause to fix, ordered by how often it's the answer — including the failures that look exactly like success.

It defaults to the boring answer. Automatic instrumentation is one line and no span code, so that's where it starts; manual spans are presented as the escape hatch they are, for the cases the automatic path genuinely can't express — custom retrievers, guardrails, handoff edges, multimodal content.

The traps it stops you falling into

Every one of these produces a green, healthy-looking workspace that is quietly worth nothing.

  • A missing conversation id turns one support thread into twelve orphan rows, and every conversation-level evaluator has nothing to grade.
  • A swallowed tool error is invisible to failure detection, clustering and the gate at once — the run looks fine and gets promoted as a good example.
  • No flush() before exit loses the last spans of every script, test and Lambda.
  • A dropped traceparent header makes the gate blind to what your agent did, so tool expectations report SKIP instead of failing.
  • An adversarial scenario is inverted — goal achieved means the attack won. Read it the usual way round and a fully successful jailbreak passes.

Pair it with the MCP server

The skill gives your agent the know-how. The MCP server gives it your data — every backend serves one at /mcp, scoped to the workspace its key belongs to.

both, once
npx skills add https://github.com/Jwuthri/Tracely-ai --skill tracely
claude mcp add --transport http tracely https://api.tracely-ai.com/mcp \
  --header "Authorization: Bearer $TRACELY_KEY"

With both connected, the useful ask stops being a code request and becomes a product one: “look at the last 20 traces, work out what's failing, and add an evaluation column that catches it.”

Next

  • SDK documentation — the source the skill distils, with runnable examples per provider and framework.
  • LLM evaluation — the concepts behind the evaluator columns the skill helps you design.
  • Tracely on GitHub — MIT, self-hostable, and where the skill lives.