Bots in Your GA4 Data: What Google Filters and What It Misses
Open the Direct channel on a site nobody has checked for a month and find a spike: a few thousand sessions from one city, one event each, no key events, and a bounce rate near the ceiling. The first instinct is a tracking fault. The second is the claim that Google Analytics removes bot traffic, so bots cannot be the explanation. That claim is half right, and the half that is wrong has a price.
Google does filter, automatically, and you cannot switch it off. It also will not tell you how much it removed. The platform drops what it recognises, reports nothing about the volume, and leaves the rest in your reports for you to deal with.
What GA4 removes, and why you get no number
Known bot-traffic exclusion runs on every property. Google's documentation is short and unusually blunt about it: traffic from known bots and spiders is automatically excluded, you cannot disable the exclusion, and you cannot see how much known bot traffic was excluded. Three sentences that answer most questions about the setting. There is no box to tick, no report to open, and no total to reconcile against your server logs.
Identification combines Google's own research with the International Spiders and Bots List maintained by the Interactive Advertising Bureau. A user agent on that list is dropped before the traffic reaches your reports.
Treat the result as a floor rather than a measurement. Third-party guides that analyse the same filter say the same thing: it catches what the list knows, so the bot share visible in any GA4 property is an under-count. Opticks Security, writing in 2026, notes that the filter does not catch invalid traffic that mimics human behaviour, rotates IP addresses, or runs through residential proxies. A UK guide from Priority Pixels adds the detail that matters for agencies: the list covers search crawlers, SEO tools, social link previewers and major uptime monitors, while headless browsers presenting a normal Chrome user agent and scraping farms on residential proxies stay in the data.
Where the filter stops
Four categories of non-human traffic survive the built-in filter, and they behave differently enough to be worth separating.
- User agent spoofing. The filter reads a user agent string. A bot that presents an ordinary Chrome string is not on the list, so nothing happens to it.
- Headless browsers. Playwright, Puppeteer and Selenium scripts declare themselves by default and can be told not to in one line of configuration. Automated QA against production, screenshot services and price scrapers all arrive looking like people.
- Crawlers that render JavaScript. Your tags are scripts, so a crawler that only fetches HTML fires nothing and never appears in GA4. The traffic that hurts is the crawler that executes pages properly, because it fires every tag on the route it follows.
- Click fraud on paid campaigns. This is a separate budget problem that lands in the same reports. Google Ads filters invalid clicks before billing and reports the filtered volume in the Invalid clicks column. Its Help page states the position plainly: you are not charged for clicks its systems identify as invalid, and Google points advertisers to a Click Quality Form when they suspect activity its filters missed.
How to recognise the traffic that survives
No single dimension says bot. A combination does, and the patterns repeat across properties.
Look at engagement first. A session with one event, no key events and almost no engagement time is either a misconfigured tag or a machine, and a cluster of them from one location settles the question. Check the city and the country next. Data-centre regions, and single cities carrying volume no marketing activity explains, are the usual signature. Then look at landing pages. When one URL absorbs thousands of sessions while the rest of the site stays flat, the crawler or the scraper has found something it likes, usually a search results page, a feed, a sitemap entry or a product listing it can enumerate.
Timing helps as well. Human demand has a shape, roughly a working day with a slope in the evening. Bot traffic arrives in bursts, at odd hours, or as a flat line that runs for weeks. Compare GA4 against your server or CDN logs before you conclude anything, because the two count differently and the gap itself is informative. And if the traffic arrives on a hostname that should not be measuring at all, a staging domain, an old subdomain or a preview URL, that is a filtering problem with its own lever, covered in our guide to hostname include filters.
The three filters you control
Google gives you three data filter types, documented under Data filters in Admin: developer traffic, internal traffic and web hostname traffic. Two of them are the practical response to pollution in your own reports.
Internal traffic. This is a two-step job, and the first step is the one people skip. Go to Admin, then Data streams, open the web data stream, click Configure tag settings, click Show more, then Define internal traffic, then Create. The rule adds a traffic_type parameter to every incoming event, which is the only event parameter you can set a value for on that screen. The default value is internal, and you can use your own label instead, such as a location name. Match the address with one of six operators: equals, begins with, ends with, contains, in range using CIDR notation, or a regular expression. IPv4 and IPv6 both work, and there is a helper link that tells you your current public address. Multiple conditions combine with OR, not AND, so one rule can carry several offices. App users cannot be filtered this way at all.
Step two creates the filter that acts on the parameter. In Admin, open Data filters, click Create Filter, choose Internal Traffic, name it and choose Exclude. The name must be unique in the property, begin with a letter, stay inside 40 characters, and use only letters, numbers, spaces and a single symbol.
Developer traffic. This one covers debug mode, which means GTM Preview sessions and any QA pass that sets the debug parameter. It is worth having, because a preview session that fires every tag twice looks exactly like a conversion problem. The filter removes that traffic from reports while DebugView keeps working, which is the point of the split.
Filter states. Every data filter has three, and the middle one is the safety net. Testing assigns matching data to a dimension called Test data filter name, with the filter's own name as the value, so you can validate before committing. Active applies the filter to incoming data and makes permanent changes. Inactive means Google is not evaluating it. Google's own testing recipe is a free-form exploration with Test data filter name and Event name as rows and Event count as the value, filtered to your filter's name. Allow 24 to 36 hours for a new filter to take effect before you judge it.
What you cannot undo afterwards
Data filters are forward-only. Google evaluates them from the point of creation onward, they never affect historical data, and applying one is permanent in the other direction too: excluded data is never processed and never becomes available in Analytics or BigQuery. There is no recovery path, which is why the Testing state exists and why the filter cap is ten per property, with Editor access or above at property level needed to create or change one.
If the goal is to hide rows rather than delete data, use report filters instead. They are the reversible option. One more ordering rule to keep straight when you run both kinds together: all active include filters are unioned and applied first as a group, and the exclude filters run after them, in sequence.
Why conversion rate is the number that suffers
Conversion rate is key events divided by sessions, so any non-human session adds to the denominator and nothing to the numerator. The arithmetic is unforgiving. Five hundred key events across ten thousand sessions is a 5.0 per cent conversion rate. The same five hundred events with two thousand scrapers in the session count becomes 4.2 per cent. Nobody changed the site, the offer or the campaign, and the number moved by a sixth.
That single distortion spreads. A channel that attracts a scraper looks worse than a channel that does not, so budget shifts for a reason that has nothing to do with demand. A landing page test can hand the win to the variant that happened to be crawled less. Audiences and predictive metrics are built on the same polluted event stream, and any advertising system consuming GA4 conversions inherits the noise. Direct traffic is where most of it hides, which is the worst possible place, because Direct is the channel teams read as brand demand.
The consent question, briefly
Filtering bots and doing consent properly are separate jobs. The ICO's guidance on storage and access technologies, finalised in April 2026, is built around informing a person and giving them a way to object. A crawler can do neither, and it never interacts with your consent banner, so it neither respects nor violates the rules. Bot volume is a data-quality question with a commercial answer, not a compliance one. It still has to be answered, because the figures you report to a client or a board carry the same distortion either way.
A twenty-minute routine that holds up
Check Direct once a month for sessions with no engagement, and read the city before you read anything else.
Add the Invalid clicks column to your Google Ads campaign table and look at it on the same day.
Keep the internal address list in one document, and re-check it whenever the office connection changes. Remember that anyone working from home will not match it, so the rule catches a fraction of the team and that has to be acceptable.
Run every new data filter in Testing, confirm it in the exploration, then set it to Active. Never start at Active, because there is no undo and no historical data to fall back on.
Write down what you excluded and when. Filtering is permanent, so your note is the only record that will exist.
Known bot filtering gives you a floor and no counter. Teams that plan around that tend to trust their conversion rate more, and they have the evidence ready when somebody in the room insists the numbers cannot be right.
Are Non-Human Sessions Distorting Your Reports?
North Digital audits the traffic behind the dashboard, separates what is real from what is not, and sets up the filters that keep the data clean without breaking consent or your history. You get the findings, the filter configuration, and a short list of what to fix first.
Book a Measurement Review