npm.io
0.1.0 • Published 3d agoCLI

efaimo

Licence
Apache-2.0
Version
0.1.0
Deps
6
Size
231 kB
Vulns
0
Weekly
0

efaimo. The audit CLI for Agent Skills and MCP servers.

ci npm license node >= 22

efaimo audits everything your agent loads: MCP servers and Agent Skills.
One CLI lints them, weighs their context cost, diffs servers against the 2026-07-28 spec,
and A/B-tests whether a skill actually helps.

Everything you plug into an agent spends two budgets before any work happens: context-window tokens, and the trigger quality that decides whether the right tool fires at all. Good tools exist for single slices of that problem; efaimo's job is the whole audit in one command, for both halves of what an agent loads. Lint a skill, weigh a server's tool definitions, get a migration diff for the 2026-07-28 stateless MCP spec, and measure whether a skill actually improves task completion (SkillsBench found 16 of 84 curated skills making agents worse; linting alone cannot catch that).

Agent Skills

npx efaimo check --skill ./skills/

Point it at one skill or a whole folder. It validates each skill against the agentskills.io spec and grades it:

check skill  claude-api
grade C (71)   1 error  2 warnings  4 info

  x S101  description is 1068 chars (spec max 1024)
  ! S104  instructions are ~18.4k tokens (spec recommends staying under 5k)
          fix: move detail into references/ files loaded on demand
  ! S104  SKILL.md is 570 lines (spec recommends under 500)
  i S104  metadata level is ~294 tokens; the spec targets ~100 and this loads at
          startup for every installed skill

(That is Anthropic's own claude-api skill.) efaimo checks frontmatter and trigger quality, collisions across an installed set, the context budget (metadata loaded every session, body loaded on trigger), reference integrity, and injection hygiene. Run it on a folder and every skill gets its own grade, not one aggregate.

We graded 36 public Agent Skills: every skill in anthropics/skills, anthropics/claude-cookbooks, and obra/superpowers, at pinned commits. 97% score an A, but even this curated set has real issues: Anthropic's own claude-api scores a C with an over-limit description and ~18k-token instructions, and the median skill's instructions run ~1,700 tokens loaded on every trigger. Full report, the corpus manifest, and the two commands that reproduce it: the Skills Quality Index.

Does the skill actually help? (experimental)

Linting tells you a skill is well-formed. It does not tell you the skill makes the agent better, and research shows some skills make it worse. efaimo test measures that directly: it runs a task with and without the skill, N trials each, and an LLM judge scores every attempt.

npx efaimo test scenario.yaml                 # dry run: validate + show the plan, no API calls
npx efaimo test scenario.yaml --live          # run it for real (Claude or GPT)
npx efaimo test scenario.yaml --live --model gpt-4o-mini

Works with Claude (ANTHROPIC_API_KEY) or GPT (OPENAI_API_KEY) models; the provider is picked from the model name. Put the key in your shell or a local .env file (copy .env.example); a real shell variable always wins.

Two real runs (claude-sonnet-5, 8 trials each) show why this matters: a generic csv-cleanup skill that a capable model does not need, and a skill that encodes a convention the model cannot guess.

efaimo test A/B results: csv-cleanup measures +0 points (no measurable effect) while contoso-crm-import measures +100 points (helps); both skills lint clean at grade A

The two runs as copyable text
test  csv-cleanup helps on a messy CSV
  with skill     8/8 pass  (100%)
  without skill  8/8 pass  (100%)
  delta          +0 points   no measurable effect
test  contoso-crm-import helps on an unknowable format
  with skill     8/8 pass  (100%)
  without skill  0/8 pass  (0%)
  delta          +100 points   helps

The first is pure context overhead, the second earns its tokens, and only the trial data tells them apart. Experimental and probabilistic: raise the trial count for confidence, and treat small deltas as noise. It is opt-in because a live run spends tokens on your key. See examples/scenario.example.yaml and examples/scenario.crm.yaml.

MCP servers

npx efaimo weigh "npx -y my-mcp-server"      # what does it cost my context window?
npx efaimo check --mcp "npx -y my-server"    # quality grade + 2026-07-28 migration diff

weigh reports the token cost of tool definitions in three real serializations, per tool, with an optional --anthropic exact Claude count.

efaimo weigh example output

check --mcp connects live (it speaks both the legacy handshake and the new stateless protocol, so a 2026-07-28 server audits fine) and reports two things, separately. First, a quality grade for what models actually experience: descriptions, schemas, annotations, tool count, token cost. Second, an ungraded migration diff for the 2026-07-28 stateless spec (which removes initialize, sessions, Sampling, Roots, and Logging, and requires server/discover, resultType, and cache fields): exactly what will break and how to fix it, each rule naming the SEP it came from. Readiness never drags the grade; not having migrated to a spec that is not final until 2026-07-28 is a to-do list, not a defect. Full list: docs/RULES.md.

efaimo check example output: quality grade A (95) plus a four-item 2026-07-28 migration diff for the reference server

(That is the official reference server: solid quality, four things to migrate before the 28th.)

The same output as copyable text
check mcp  npx -y @modelcontextprotocol/server-everything
grade A (95)   quality: 0 errors  1 warning  0 info

  ! E122  tool "echo": description misses 3/4 quality axes (length 40..600;
          says when to use it; mentions the result)

2026-07-28 readiness  4 items to migrate (a migration diff, not graded: the
spec finalizes 2026-07-28)
  ! E104  server declares the logging capability; MCP Logging is deprecated in
          2026-07-28 (SEP-2577) and logging/setLevel is removed
  ! E106  server/discover is not implemented (-32601 Method not found)
  ! E118  tools/list result omits ttlMs and/or cacheScope, which 2026-07-28
          requires on list and resource-read results (SEP-2549, CacheableResult)
  i E107  results do not carry the resultType field required in 2026-07-28
          ("complete" | "input_required")

From an agent

efaimo mcp runs efaimo as a small, read-only MCP server, so an agent can lint or weigh a skill mid-session, before it commits it to context:

npx efaimo mcp      # stdio server exposing efaimo_check_skill and efaimo_weigh_skill

It reads files only (it spawns no process and opens no socket), and the token-spending test is not exposed. Client config and recipes are in docs/INTEGRATIONS.md.

How it works

efaimo sits between your agent and everything it loads, measures the cost the way your host actually serializes it, and grades the quality, before any of it reaches your context window.

efaimo weighs, checks, and tests what your agent loads

How it compares

Focused tools already own single slices of this space, and some are very good. efaimo is the one tool that covers the whole audit surface. As of mid-2026:

efaimo skill-validator upskill mcp-spec-check conformance (official)
Agent Skills linting yes yes no no no
Skill token cost yes (3-level split) yes no no no
Skill outcome testing with/without A/B trials (experimental) static LLM scoring yes (generate + eval, code agents) no no
MCP tool-definition cost yes (3 serializations, --anthropic exact) no no no no
MCP quality rules yes no no no no
2026-07-28 readiness migration diff: what breaks and how to fix no no yes/no verdict official test suite
CI budget gate + badge yes no no no no

efaimo runs the official conformance suite for you (check --conformance on http targets) and complements security scanners such as Snyk agent-scan and the skills installer (npx skills) rather than replacing them.

Install

Nothing to install. Use npx:

npx efaimo check --skill ./skills/

Or add it to a project with npm i -D efaimo, or the GitHub Action:

- uses: efaimo-ai/efaimo@v0
  with:
    command: check --skill ./skills --strict

Gate a pull request on context-window growth, too:

npx efaimo weigh "npx -y my-server" --out base.json          # record a baseline
npx efaimo weigh "npx -y my-server" --diff base.json --allow-increase 10

--badge badge.svg writes an SVG plus a shields.io endpoint JSON for your README. These two are real efaimo output for the reference server:

context cost: 1120 tok (o200k) efaimo: A (95)

More recipes (pre-commit, GitLab, editor audit, programmatic use): docs/INTEGRATIONS.md.

Rules at a glance

family covers
S101 to S106 skills: frontmatter and trigger quality, trigger collisions, context budget, reference integrity, injection hygiene
E101 to E118 MCP 2026-07-28 readiness: deprecated primitives, statelessness, server/discover, resultType, cache fields, transport
E121 to E130 MCP quality: description quality, annotations, schema hygiene, tool-count and token-cost budgets

Quality (E12x-E13x) and skill (S) findings set the letter grade; readiness findings (E101-E118) are reported as an ungraded migration diff until the spec ratifies. Every finding carries a stable id you can suppress or link. See docs/RULES.md.

Roadmap

  • Harden efaimo test: a separately chosen judge model, confidence intervals on the delta, multi-turn tool-use trials, and judge calibration, so the experimental harness earns unqualified trust.
  • Track the 2026-07-28 spec to ratification, then start grading readiness (today it is an ungraded migration diff on purpose).
  • A public, continuously updated Agent Skills Quality Index over a broad corpus.

Stability

efaimo is 0.x. The commands are stable, but the rule set, grades, and exact output may change between minor versions until 1.0; pin a version in CI if you need reproducible thresholds. Every finding keeps a stable id.

Honest scope

efaimo is a linter and cost profiler, not a security scanner. Its injection checks are info-level heuristics that an attacker evades trivially; a clean report is not a security pass. For supply-chain safety use a dedicated scanner such as Snyk agent-scan. Token figures are estimates unless you opt into --anthropic; the method and its known bias are in docs/METHODOLOGY.md.

About

Built by efaimo ai, open tooling for the space between hosts and tools. efaimo is the flagship CLI; capabilities grow as subcommands under one name. Apache-2.0. See CONTRIBUTING.md and docs/RULES.md to add a rule.

Keywords