How to judge software you did not write
16 guides on evaluating an agent in ten minutes, reading a findings report, deciding which permissions to refuse, and getting your own listing verified. No account needed for any of it.
- Guides
- 16
- Subjects
- 4
- Of reading
- 105m
Everything16
A repeatable order to look at things in, built around the fact that most of the useful signal is in three places and none of them is the description.
The access requests that are almost never justified, the ones that are, and how to tell which is which from the listing alone.
Per-run pricing looks cheap and is the easiest way to be surprised by a bill. How to work out your real monthly number before you commit.
The build case is nearly always argued on the build cost, which is the one cost that stops. Here is the arithmetic with upkeep in it, and the three questions that settle it faster than any spreadsheet.
The single most common serious weakness in this category, why any tool that accepts a URL is exposed to it, and the four requests that find it in about ten minutes.
The exact checks behind the badge on every listing, what a clean pass proves, and the four things it deliberately does not prove.
Most real software has findings. Here is how to tell the ones that should stop a purchase from the ones that are simply information.
Every badge has a date and an expiry. What changes on an update, why permission changes are treated differently, and what you get told.
The four categories our tests cannot see, written down in the same place as the ones they can, because a certification that hides its limits is worth less than none.
The checks this site runs, done by hand: fetch the tarball without installing it, check its digest, read its install scripts and tool descriptions, then keys, advisories, licence and auth.
What this site lets you try today, what it does not, and how to try the rest on your own input, with no credentials and no card.
Most of the damage an agent does happens in the first week, while nobody is quite sure what it is doing. A short routine that prevents it.
Permissions get reviewed one line at a time and damage never happens one line at a time. How to score the combination, and the four pairs that turn a scare into a disclosure.
Independent testing put average production success near 57 percent. Where the missing 43 percent goes, why the failures cluster, and the four numbers worth putting on a dashboard in week one.