CLI proof series
Turning small tools into developer-facing artifacts
CLI proof series
Turning small tools into developer-facing artifactsProject overview
About the organization
The audience is developers evaluating whether agent-safe CLI workflows are understandable and inspectable.
What I did
I defined the safety and operator experience, directed the TypeScript implementations, and used Codex to build, test, and document the command paths.
wSearch queries Wikidata through explicit network and user-agent gates; rSearch searches and fetches arXiv material with structured output and documented limits.
Inspectable boundary
What this work proves
- Claim
- Small TypeScript CLIs can expose useful research data with explicit safety and automation contracts.
- Evidence
- Clean-revision CLI checks cover wSearch help, denied network access, and one live Wikidata retrieval, plus two fixture-backed rSearch dry-run outcomes.
- Caveat
- One successful request is not a reliability benchmark. Installed releases, live rSearch retrieval or downloads, and adoption remain unverified here.
From problem to proof
How the delivery unfolded
Problem
A research command needs a predictable failure path.
A script consuming research data needs to distinguish a successful result from denied network access or invalid arguments. A readable error alone is not enough for another program to act on.
Discovery
The two tools expose different research sources.
wSearch queries Wikidata; rSearch provides arXiv search, fetch, and download commands. Their source contracts describe output envelopes and input limits, but those descriptions need executable checks before they become runtime claims.
Constraints
Make the network boundary visible.
- wSearch requires explicit network permission and a user agent for remote requests.
- Offline checks must not be described as successful research retrieval.
- A source checkout test does not prove the behavior of an installed release.
Decisions
Start with the operator's first two questions.
- Can I discover the command syntax without making a request?
- Can my script identify a denied request from structured output?
- Keep live wSearch retrieval separate from fixture-backed rSearch dry runs.
Delivery
Check the failure boundary, then retrieve one entity.
From a clean wSearch revision, --json entity get Q42 without network permission returns wiki.error.v1 with E_POLICY; --json --plain help exits 0 without stderr. With explicit network permission and a user agent, entity get Q42 returned wiki.entity.get.v1, status success, and ID Q42 from Wikidata. The request used isolated configuration, a ten-second timeout, and no retries.
Failure and correction
Replace implied proof with the observed boundary.
This case study originally described expected output without an executed example. Clean-revision rSearch checks now also launch the CLI with a test client: download 2101.00001 --dry-run --json exits 0 with arxiv.download.dryrun.v1; an invalid ID exits 4 with status warn and a would-fail result. This is planning-output evidence, not proof of a downloaded paper or live arXiv retrieval.
Outcome
Discovery, rejection, retrieval, and planning have distinct evidence.
Two wSearch offline checks, one live Wikidata request, and two rSearch dry-run checks passed from pinned source revisions. They show bounded operator paths, not installed-release verification, service reliability, or measured adoption. Live rSearch retrieval and download behavior remain separate work.
Human-owned delivery
Directed by Jamie. Implemented with Codex.
- Jamie
- Defined the tool direction, safety expectations, and operator experience, and owns the decision to publish the work.
- Codex
- Assisted with implementation and documentation, inspected source contracts, and ran the clean-revision CLI checks and bounded Wikidata request.
Inspect the evidence
Follow the artifact, tests, and stated limits.
All work 05
01 / 05
SkillsBar
02 / 05
Skills SDK
03 / 05
CLI Tools
04 / 05
Proof Log
05 / 05