What a useful result tells you.

The question is not whether an AI sounds confident. It is whether the work meets the rules you set.

One skill cleared the checks. One was blocked.

We tested two troubleshooting skills we had published. One produced the required work. The other failed a required check: producing a report a finance leader could use.

The decision: block that skill's release under the test's rules. A useful-looking answer did not make up for a missing requirement.

The limit: this was a defined skill test, not a test of every possible workflow. The report lists the checks that ran, the checks that were skipped and the uncertainty in judging.

Read both results and the evidence

Does an improvement loop help?

We compared a single rewrite prompt with a loop that proposes an edit, checks it and keeps it only if it improves the score. Both approaches worked on the same set of 60 skill headers.

The result: the checked loop performed better than the single-prompt approach in this experiment. Its own average improvement was modest.

The limit: only the configuration headers could change. This does not prove that the loop improves whole skills or unrelated agent workflows.

Read the comparison, method and limits

See the ongoing checks

Our skill evaluations and team-memory checks publish current technical results, including failures and runs that could not complete.

Plan a test for your workflow Browse all results