Release evidence expires quietly. A local check can be green, an evaluation can look strong, and a registry observation can still belong to a package that is no longer the candidate in front of you. SkillsBar began with a simple question: how should a maintainer see that before acting on a release?
When proof outlives the candidate
The failure mode was not a broken test. It was a plausible interface telling the wrong story. Validation, security, evaluation, registry, runtime, review, and release observations could all exist while referring to different package states.
Blending those observations into one readiness score would hide the most important fact: whether the evidence still belonged to the candidate. The interface therefore had to preserve each evidence lane instead of treating every green result as interchangeable.
Identity before readiness
SkillsBar models the release journey as nine gates. Candidate identity comes first. When the canonical package digest is missing or stale, downstream evidence remains visible as historical context, but it does not count toward the current candidate.
The interface then exposes one next corrective command. That matters more than a decorative score: the maintainer can see why the candidate is held, what evidence is current, and which action can strengthen the weakest gate.
Human ownership stays visible
I framed the stale-evidence problem, chose the nine-gate model, made the product and interaction decisions, inspected failures, and own the public claims. Codex proposed implementation options, wrote Swift and web code within that direction, and added fixtures and tests.
That boundary is part of the work. Model assistance can accelerate delivery; it cannot decide whether the evidence is sufficient, approve a release, or make an external outcome true.
The correction that made it honest
The first presentation treated the focused gate too rigidly. It made the product easier to describe, but less faithful to receipt state. The merged adaptive pipeline work derived the active section from the available evidence instead.
The same change added light and dark coverage, moved scanning off the main actor, and constrained inherited shell configuration. The correction was not cosmetic. It aligned what the interface emphasized with what the candidate could actually prove.
What the current evidence proves
A maintainer can inspect why a candidate is held, distinguish current proof from historical context, and copy the next corrective command. Merged pull request 6 records a passing Swift build and 80 executed tests for the adaptive delivery slice.
That is useful implementation evidence. It is not proof of notarized distribution, production registry integration, or external adoption. Those claims stay open until their own evidence exists.