Bot Traffic Is Splitting Your Ad Budget, and a Human Signed It Off
Most teams split budget by share of traffic. But roughly half of that traffic is automated, so the split funds machines. Compare it to authenticated data and the gap is your bot tax.
Picture a marketing team carving up a six-figure quarterly budget across three acquisition channels, and doing it the way most teams do: by share of traffic. Channel A brought half the sessions, so it takes half the money. Channels B and C split the rest. Clean, defensible, and quietly wrong, because a large slice of those sessions was never a person. This is bot traffic setting your budget, and a human signed the approval.
Run the same channels through a second meter. Not raw sessions, but authenticated ones: logged-in visits, completed checkouts, accounts that filed a support ticket. On a million recorded sessions you might count 300,000 that clear a human-only bar. Now the shares move. Channel A, heavy on cheap programmatic display, contributed half the raw sessions but only 30% of the authenticated ones. Channel C, which looked like a rounding error at a fifth of traffic, holds 40% of the humans. Weighting by raw traffic overfunded the bot magnet and starved the channel that actually converts. The numbers here are illustrative, but the mechanism is not: the moment your allocation key is 'share of sessions', you have handed part of the vote to machines.
How much of the vote? Imperva's 2024 Bad Bot Report put automated traffic at just under half of everything crossing the internet in 2023, with bad bots alone near a third. Cloudflare's Radar data, tracking the surge in AI crawlers since 2024, points the same way: a large and rising share of requests are scripts, not customers. If half of raw requests are not people, every metric distilled from raw requests inherits that contamination, and it inherits it directionally, not as random noise. That is the part that should worry a finance team.
How much of your website traffic is bot traffic?
Nobody selling you a dashboard wants to answer that in writing, because the honest answer contaminates their headline number. And the ad-budget case is just the visible one. The same fault sits under platform-support prioritisation sized off visitor operating-system mix, addressable-market estimates built from 'what browsers our sector uses', and the board slide showing audience growth that is really crawler growth with a friendlier label. Each looks like a census of humans and is quietly a census of a substrate. Wire those feeds into an automated dashboard and you have let a non-human majority vote on the roadmap.
Why do traffic trackers and purchase data disagree?
Desktop Linux is the cleanest place to watch the gap open, which is why it gets quoted to death. StatCounter's North America desktop figures through mid-2026 have shown Linux drifting up toward high single digits, in some months near a tenth. The Steam Hardware Survey for July 2026, which only counts after someone bought a game and launched it on their own machine, puts Linux in low single digits. Same platform, two meters, a gap far too wide to be rounding. Bots, monitoring probes and scrapers overwhelmingly run on open-source systems and present as Linux or Unix when they present honestly at all, while a purchase is expensive to fake and pointless to automate. Gate the metric behind money and the machines mostly fall away.
The direction of the error is the confident claim; the exact size is not. Some of the Linux gap is dual-boot machines that browse on one system and game on another, some is genuinely under-counted enthusiasts, some is the storefront proxy being biased its own way. Attributing all of it to bots would be its own bad metric. The point holds regardless: when an ungated figure and a human-gated one disagree this hard, the ungated one is inflated, and you can read roughly how much off the size of the gap.
What to do about bot-inflated metrics
Triangulate, and make it boring on purpose. For any number that drives money, find the version gated behind a human-only behaviour: an authenticated session, a completed checkout, a licensed install, a paid support ticket. Treat the raw traffic figure as a ceiling and the gated proxy as a floor, and refuse to quote either as your real audience. Where they diverge, the size of the divergence is your bot-inflation estimate for that surface, and you carry it the way an analyst carries a base rate. This is measurement risk, and it belongs in a serious technical strategy review, not filed under analytics housekeeping.
Gated proxies are not a free lunch. A storefront's hardware survey only sees people who bought on that storefront, its own narrow, self-selected slice. That is the argument for triangulation, not for swapping one flawed source for another. The aim is not the true number. It is to bound it honestly from both sides and to know which side of the bound your decision can survive being wrong on. Teams that already run human-in-the-loop checks on automated systems tend to get this in their bones: the machine's tally is an input, not a verdict.
What would change my mind? A traffic tracker publishing a validated, human-verified breakdown that held its desktop Linux share near a tenth after bot filtering would collapse the inflation reading for that source, and I would retract it. Independent bot-share estimates converging well below half would soften the general claim. Absent that, the asymmetry stands. When a raw audience figure and a human-gated one disagree by this much, the safe bet is not that your audience is bigger than you thought. It is that a good part of it was never there.
Questions people ask
How do I know if my ad-spend or channel data is counting bots?
Compare each channel's share of raw sessions to its share of authenticated ones: logged-in visits, completed checkouts, paid accounts. If a channel's share of humans sits far below its share of raw traffic, you are funding bot traffic. The gap between the two shares is a rough bot-inflation estimate for that channel.
How can I tell if a market-share stat is measuring bots?
Ask what behaviour the number was gated behind. Page requests, sessions and user-agent strings count machines and people together. A purchase, an authenticated login or a licensed install sits far closer to a human count. Ungated equals contaminated by default.
Is desktop Linux really only a few percent of users?
It depends on the meter. The Steam Hardware Survey, gated behind a game purchase, puts it in low single digits; StatCounter's traffic-based North America figures have run much higher, near a tenth in some months. The honest read is a range with a heavily bot-inflated ceiling, not one headline number.
Related
- Retention Bonuses Don't Reduce Key Person Risk. Doctrine Does.
- Your Cloud Provider Is Financing Its Own AI Revenue. Do the Financial Due Diligence
- It Wasn't the Robots: The 2017 Tax Rule Driving Tech Layoffs
- Data & Strategy
Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.