GA4 Source Group: Why Your Traffic Sources Finally Agree
Open the Traffic acquisition report for almost any site with paid social activity and you will find the same thing: four rows for one platform. facebook, fb, m.facebook.com, Meta-facebook. Each row carries its own sessions, key events and conversion rate, and nothing in the interface tells you they are the same advertiser.
Nothing is broken. The platform is one, the labels are many, and GA4 has treated every label as a separate channel since the property was created. The number a client actually wants, the total delivered by Facebook, does not exist as a row anywhere in the report.
On 11 June 2026, Google released a fix for the reporting half of that problem. Source Group is a dimension that consolidates the variant source values of a single platform into one value. The same changelog entry updated the classifications behind the older Source Platform field and added a new type of data filter for hostnames. All three changes share one purpose: making the reports describe the business rather than the tagging.
What follows is what each change does, where the fields show up, and the parts of the problem they leave untouched. If you report on paid social for clients, this matters before the next budget conversation, not after it.
Why source values fragment in the first place
GA4 has never forced source values into a controlled list. For tagged campaigns, whatever a campaign manager types into utm_source becomes a value in the report, so fb, Facebook and facebook are three sources even when they came from one ad account. Tracking templates do the same at scale: pass %instagram% in one template and ig in another and the report keeps both.
Untagged referrals are no tidier. Analytics derives the source from the referring domain, and the referring domain varies by device, app and sharing method. Mobile web versions, link shorteners, in-app browsers and app webviews all send their own strings, so a platform that sends thousands of sessions a month can appear as a scatter of near-duplicate rows.
Default Channel Groups hide part of this by grouping paid social into a single channel, which is why the problem often goes unnoticed. It surfaces the moment somebody asks for a source-level breakdown to decide where next quarter's budget goes. That is when four rows for Facebook become an argument about which row is real.
What Source Group actually does
Source Group consolidates source values for common online platforms. Google names Facebook, Instagram and TikTok as examples, and the documentation makes the point with an example of its own: the value reads "Facebook", not "facebook", "fb" and "Meta-facebook".
- One row per platform. The consolidation happens inside GA4. There is no tagging work, no UTM cleanup project, and no data layer change to schedule.
- Populated retroactively. Historical sessions carry the consolidated values, so a year on year comparison does not break on 11 June. A fix that only applied going forward would put a step change into every chart you own.
- Available where you already report. It is a dimension you can add in standard reports, in explorations, and in the advertising section where cross-channel budget analysis happens.
- Tracking template values included. Macros that dump awkward strings into the source field, such as
%instagram%andig, roll up to Instagram.
Two boundaries are worth writing down. Source Group cleans labels, not tagging. If three teams typed three different values for the same platform, you now have one reporting row and three inconsistent processes underneath it, and the process problem is still yours to fix. Second, the consolidation covers common platforms by design. Anything outside that set keeps the value it has, so the long tail of referral strings does not disappear.
Keep the raw Source dimension in your working views. It is the field you debug tagging with. Use Source Group for anything a client reads.
Source Platform answers a different question
Source Platform is not new. It existed before this release, and Google updated its classifications to line up with the new Source Group values. The division of labour is easiest to see in Google's own example: raw values %instagram% and ig consolidate to Instagram in Source Group and Meta Ads in Source Platform.
Read it as two questions pointing at the same traffic. Source Group tells you which platform sent the visit. Source Platform tells you which advertising system delivered it. A campaign run through Meta's buying tools can reach Instagram inventory, and only the platform field distinguishes the delivery system from the destination.
The release also raised third-party platforms to the same level of granularity Google applies to its own inventory. TikTok, Pinterest and Amazon are now reported in a way that stands comparison with YouTube, Search and Maps. For an agency running mixed budgets across Google and non-Google channels, that is what makes a single cross-channel table defensible instead of a debate about which platform's numbers are worse.
Before you build anything on the new fields, run one check. Imported cost data that arrives without a paid medium, and organic traffic, can show as unlabeled values in the platform column. A rising unlabeled share is a signal about your imports rather than a platform change, and it is worth tracing before you present a cost comparison built on top of it. The campaign data import guide covers the import side, currency and all.
AI traffic is grouped now too
The Source Group release includes built-in grouping for traffic from ChatGPT and Perplexity. Until now, assistant referrals arrived as strings you had to find, filter and label yourself, and most properties never got round to it. They are grouped like any other platform from this release on.
Read that next to the AI Assistant channel GA4 added on 13 May 2026. That release assigns an ai-assistant medium, a dedicated AI Assistant channel in the Default Channel Group reports, and an (ai-assistant) campaign name when the referrer matches a recognised assistant. Channel and dimension work as a pair: the channel buckets the sessions, the dimension gives you a clean value to break the bucket down by.
Two limits worth stating plainly. A visit that arrives without a referrer, through a copied link or an app that strips it, still lands in direct or organic. And a grouped value tells you where a visitor came from, not what they read there. Treat AI referral as a trend worth watching on a monthly chart rather than a channel you can report as complete.
Hostname filters: stop collecting what you never wanted
The second item in the same release is quieter and, on audits, more useful. Hostname filters are a new type of filter in Admin, sitting alongside the developer traffic and internal traffic filters. They exclude events based on the hostname the event fired on. Google's example is a container snippet copied onto google.com, though the cases we meet are more mundane:
- Staging and development environments still firing the production measurement ID months after launch.
- QA and test builds that a developer opens every week without thinking about the property.
- Parked or holding domains left behind by a registrar or a previous agency.
- Subdomains that should never have shared the stream, such as a legacy help centre or a careers microsite.
Three rules decide whether the filter helps or hurts:
- Exclusion only. There is no hostname inclusion filter. You cannot allowlist your real domains and drop everything else, so you have to know which hostnames are polluting the property and name them.
- Not retroactive. Events already collected stay in the property. The filter shapes what arrives from the moment it is saved, which is why the right time to set it up is before a migration or a launch, not six months after one.
- Processing takes 24 to 48 hours. A filter saved this morning is not a filter you can prove this morning. Use validation mode first: it previews what would be excluded, which is the difference between a considered change and accidentally cutting off a live domain's data.
If a hostname has been feeding junk into your reports for a year, the filter will not repair the history. It only stops the bleeding. Cleaning what is already there means a segment, an exploration, or a filter in your reporting layer, and that work is worth doing once so your baseline stops moving.
The audit to run on a client property this week
- List the top 50 source values for the last 90 days with sessions and key events, and circle every near-duplicate. That list is the case for using the new field.
- Put Source next to Source Group in one exploration. The row count will drop, and the major platform totals will rise. Both are expected, and neither means the data changed.
- Check the bigger picture before you explain the smaller one. Confirm the property total is unchanged, so you can tell a client this was a labelling change rather than a traffic change.
- Look at Source Platform for unlabeled rows and trace them to their cause. Organic traffic and imports missing a paid medium are the usual answers, and both are worth fixing on their own merits.
- Decide which field belongs in each dashboard. Source Group for anything client-facing, raw Source for debugging. Mixing the two across reports is how a team ends up with two different sessions totals for one platform.
- Report the hostnames actually collecting data by session count, and mark the ones that are not the real site. Do this before every migration, not after.
- Set the hostname filter in validation mode, compare the preview against your expectations, then activate. Allow 48 hours before you promise anyone a clean number.
- Write both decisions into the measurement plan. Which field the reporting uses, which hostnames are excluded and why. The next person to open the property should not have to reverse engineer it.
What none of this fixes
Source Group is a reporting layer. It does not repair the tagging behind a messy value, it does not recover sessions that were never tagged at all, and it does not change how credit is assigned. A conversion attributed to Facebook before this release is still attributed to Facebook after it, in the same model, with the same lookback window.
Hostname filters are a collection control, not a cleanup. They do nothing about what is already stored. And neither feature addresses the other reasons a property disagrees with reality: consent limiting what can be measured, tags that fail on a template the marketing team does not control, or a data layer that sends the wrong event parameters. Those still need an audit and a fix.
What the June release does is remove two categories of noise that generate a surprising share of "the numbers look wrong" tickets. Platform names that were always the same platform, and events from hosts that should never have been measured. Both now have a supported control inside GA4 instead of a spreadsheet full of filters. That is worth taking advantage of this week, especially on the client accounts where nobody has looked at the hostname column in years.
Sources Still Do Not Add Up?
North Digital audits GA4 properties for UK agencies and brands: source fragmentation, stray hostnames, imports and consent configuration. You get a written list of what to fix and what it is worth.
Get a Free Analytics Audit