We grade AI skills. Every grade is signed and verifiable.

Every night we run real, published AI skills through a behavioral eval harness and publish the grades — pass, fail, or "the judge couldn't decide" — win or lose. Each grade is cryptographically signed and anchored in a public transparency log, so you never have to take our word for it.

See the results  ·  How it works  ·  Verify one yourself

Per-repo freshness — last 24 hours

One row per source repo. Each cell is one hour; its color is the decision mix of verified gate-result rows in that hour. no-data is colored as loudly as a failure — an hour we heard nothing verified is never blank, never neutral, and never back-filled with a prior value.

Per-repo hourly decision mix over the last 24 hours. Rows are source repos; columns are hours, oldest on the left.
Source ← older  ·  2026-08-14T07:00:00.000Z → 2026-08-15T06:10:29.299Z  ·  newer →
iec no-data — silent 24h no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-data
iel no-data — silent 24h no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-data
iah no-data — silent 24h no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-data
iaj no-data — silent 24h no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-data
iar no-data — silent 24h no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-data
ccp last verified 2026-08-15T03:02:12.476Z no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datapassno-datano-datano-data
jrig last verified 2026-08-15T04:27:02.030Z no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datafailno-datano-data
qmd no-data — silent 24h no-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-datano-data

Sources showing no-data across the whole window have published no verified, signed Evidence Bundle in the last 24 hours. This is the honest current state — emit-evidence is still rolling out upstream. We render the silence loudly rather than fill it.

The eval-set is the spec. Browse it, criticise it, fork the methodology. When you see a number, it will trace back through a content-addressed Evidence Bundle, signed via sigstore and anchored in the Rekor transparency log, to one of the eval-sets listed below. We refuse to publish an aggregate "PASS%" across heterogeneous predicates — that is metric laundering.

Live now — nightly grades, our own memory system, and a failing grade we gave ourselves

The nightly skill roster grades 8 real marketplace skills every night (03:30 UTC) — recent nights produced passes, honest fails, and unresolved-judge advisories, all published. The governed brain — our own production team-memory system — emits signed results for its retrieval-quality, governance-decision, and provenance evals on every change. Both are re-verified here (pinned CI identity, signature, and transparency-log inclusion, row by row) before anything renders.

The dogfood scorecard runs our own 7-layer methodology against our own published CoreWeave skills. It produced one SHIPcoreweave-gpu-node-forensics (Rekor log index 2085904207) — and one BLOCKcoreweave-gpu-cost-leak-hunter (Rekor log index 2091983416), an honest failing grade of our own skill, published only after noise-robust majority judging made the verdict reproduce 7 out of 7 runs, with the per-vote record inside the signed evidence. The Evidence Bench scorecard carries the first signed row of all: the Skill Refiner Phase A.0 baseline (PROCEED), Rekor log index 1689291334. All three Evidence Bundles are signed keyless against the public production sigstore transparency log — verify any of them with the cosign verify-blob command published on its scorecard. This is reproducibility via signature, not just transparency: the layer a memory-eval leaderboard like gbrain-evals does not have.

What's here at v0.1.0

The per-repo freshness + decision-mix strip (above) and the results browser are live. The signed dashboard-render attestation is Phase 2 work still in progress. See DR-035 for the full architectural decision record.

What this site refuses to do

From the seven adversarial seats of the ratification council, all preserved verbatim in the decision record:

RefusalRationale
No predicate URIs declared here Predicate URIs live at evals.intentsolutions.io exclusively. This site renders about them; it does not declare them.
No aggregate PASS% across mixed predicate semantics NOT_APPLICABLE is not PASS. ADVISORY is not PASS. Compositing them into a single percentage is metric laundering.
No basicauth on the public origin for operator views The operator-internal view lives on a tailnet-only hostname. Anonymity-first on this origin.
No rendering from upstream manifests without re-verification Pinned OIDC subject and workflow_ref claim per source repo, with Rekor inclusion proof checked row-by-row at ingest.
No asymmetric rendering of Phase A.0 null-hypothesis results Both arms render with identical layout, font weight, and chart axes. A null result is rendered identically to a positive result. If we cannot guarantee that symmetry, we publish in a blog post instead.

Source & references