Your download count is lying to you
I published three small tools and watched the download numbers roll in. Then I looked at how they were distributed across versions and realised almost none of them were people. Here is how to read your own registry stats honestly.
I published a small npm package one morning and checked the stats a few days later. Four hundred and thirty-four downloads. For a tool nobody had heard of, released with no announcement, that felt like a lot. It was, in fact, almost entirely robots, and the way I worked that out is a useful trick for anyone shipping small packages.
The number that looked good
The package went up in three versions across a single morning: the first at 07:52, a fix at 09:27, another at 10:52. Nothing unusual, just the normal shape of publishing something and immediately finding two things wrong with it. Total downloads a few days later sat at 434, with 394 of them on the publish day itself, then 29, then 11, then flat zero.
A spike on release day followed by decay is exactly what real adoption looks like, so I nearly stopped there. The thing that saved me was breaking the number down per version instead of per day.
The breakdown that gave it away
Split across versions, the 434 came out as roughly 144, 146 and 144. Three near-identical counts. That distribution is impossible for real users.
When a person installs a package they get the latest version. Nobody deliberately reaches for the release that was superseded ninety minutes later. Real traffic piles up on the newest tag and leaves a thin tail on older ones. A dead flat spread across every version in the manifest means something walked the list and fetched each entry once, which is precisely what registry mirrors, security scanners and CDN prefetchers do the moment a new package appears.
Downloads tell you a machine asked for a tarball. Only the shape of the distribution tells you whether it was a person.
The same pattern showed up on the Rust crate I published the same week: ten, ten and nine across three versions released on one day. And it ruled out the other obvious explanation, that the traffic was just me testing. My own installs would have landed on the newest version only.
PyPI hides the answer in plain sight
The Python package was the most instructive, because PyPI actually separates the two populations if you ask it to. The headline number was 1,519 downloads. The breakdown was 1,205 with mirrors and 314 without.
That gap is not a rounding error, it is eighty percent of the total. Mirror traffic is infrastructure cloning the index, not humans running pip install. Every PyPI stats query lets you exclude it, and the honest number to quote is always the one without mirrors.
GitHub release assets are the clearest signal you have
For the CLI I distribute as prebuilt binaries, the release page gave me the least ambiguous read of all, because it counts each asset separately. Eighteen downloads total, broken down like this:
- Eight Homebrew bottle pulls, all for macOS on Apple Silicon, which is the exact machine I develop on.
- Seven downloads of .sha256 checksum files with no accompanying binary.
- Two macOS ARM binaries and one Intel binary.
- Zero Linux binaries. Zero Windows binaries.
The checksum-only hits are the giveaway. No human downloads a hash without the file it verifies, and one release had a Linux checksum fetched while the Linux binary itself sat at zero. That is a scanner enumerating assets.
The zeros are what settled it. Linux and Windows downloads are the ones that could not possibly be me, and both were empty. Every single install was either my own machine or a robot.
How to read your own numbers
None of this needs tooling. It is four checks you can run in a few minutes.
- Break downloads down per version, not per day. Flat across versions means automation; concentrated on the latest means people.
- On PyPI, always exclude mirrors. The headline figure can be five times the real one.
- On GitHub releases, look for checksum files downloaded without their binaries, and look at the platforms you personally do not use.
- Ignore the publish-day spike entirely. Every new package gets a welcome burst from scanners and mirrors.
The signal worth watching is different from any of these: sustained downloads, weeks after release, concentrated on the current version. That is slow and boring and it is the only pattern that means somebody chose your tool.
Why it is worth being strict about this
The obvious reason is that you will eventually put a number in a README or a job application, and any engineer who has published a package themselves will read it correctly in about five seconds. A four-figure download count on a two-week-old project does not impress that person, it tells them you have not looked closely at your own data.
The better reason is that vanity metrics quietly steer your work. If I had believed the 434, I would have concluded the tool had found an audience and moved on to the next thing. The real number, which is approximately zero, tells me something far more useful: the distribution works, the packaging works, and nobody has discovered it yet. That is a marketing problem, and it needs an entirely different fix from a product problem.
The honest zero is more actionable than the flattering four hundred, because it points at the thing actually blocking you.
Stars, on the other hand, I do trust. Starring takes an account and a deliberate click, and one of the seven on my Rust project came from a maintainer whose work sits in the dependency tree. Seven real humans beats 434 fictional ones, and it took looking properly at the data to see which of those two numbers was worth anything.