Google spent three years telling advertisers to trust the automation. This week it shipped the tools to verify it.
Read that order again, because it explains most of what has gone wrong in paid search since 2023.
The release itself is modest. Multi-campaign Search experiments arriving in September, letting you test budget and ROI target changes across several Search campaigns inside a single A/B test. AI Max experiments that keep brand and location settings enabled. A Performance Planner that shows how bidding or budget changes affect existing campaigns and applies them in one click, with undo, tracked through Bulk Actions.
The brand and location detail is the actual news
Everything else in that list is convenience. This one is a correction.
Until now, testing AI Max meant testing it without the controls you would run in production. Brand exclusions off, location settings loosened, then a comparison against a campaign that had both.
That is not an experiment. That is a demo with a control group attached.
You were measuring a configuration you had no intention of shipping, then making a shipping decision from it. Every agency running AI Max tests in the first half of this year was quietly aware of the problem and mostly shrugged, because the alternative was not testing at all.
Now the test configuration can match the deployment configuration. That is the difference between evidence and theatre, and it should reset every conclusion you drew before this month.
Portfolio testing moves the unit of decision
Campaign level A/B testing has carried the same structural flaw since it launched. Budget is fungible across an account. The test is not.
So a test campaign wins by pulling volume, impression share, and in some auctions the actual buyer, from the campaigns sitting next to it. The experiment reports a lift. The account reports nothing.
From running paid media at difrnt., this is the single most common source of a false positive I see. A test that wins locally and loses at account level, discovered six weeks later when someone finally reconciles the numbers against revenue.
Multi-campaign experiments move the unit of measurement up to where the money actually lives. That is a bigger change than it looks, because it invalidates a large amount of received wisdom that was generated by the old method.
There is a second order benefit nobody will put in a release note. Portfolio level tests reach significance faster, because they pool volume across campaigns instead of splitting one campaign's traffic in half.
For accounts in the mid five figures per month, that is the difference between a readable result in three weeks and an inconclusive one in eight. Most tests at that spend level never reached significance at all, which is why so many paid search decisions are still made on instinct wearing the costume of data.
I made the related argument about measurement targets in why ROAS is the wrong metric in AI driven auctions. The instrument and the metric have to move together, or you get precision around the wrong number.
Undo is not a rollback
One caution on the Performance Planner change, because one-click application with an undo button is going to cause damage in the next two quarters.
Undo restores your settings. It does not restore the learning phase, the auction history, or the two weeks of data the algorithm just spent adjusting to a change you reversed.
Smart Bidding treats a settings change as a new problem. Reverting the setting does not revert the model's state, and there is no button for that.
The working rule is simple. Use experiments to make decisions. Use Performance Planner to build forecasts. Do not let a forecast become a decision because the interface made it one click away.
The deselect option before applying forecasted changes, and the logging through Bulk Actions, are there for exactly this reason. Someone at Google clearly argued for them.
What this says about the direction of the platform
Automation vendors ship verification tools at a specific moment. Not when the automation starts working, but when enough sophisticated buyers refuse to expand spend without proof.
That is where paid search is right now. The mid-market accepted automation years ago. The accounts with real budget did not, and they have been asking for exactly these instruments in every quarterly review since AI Max shipped.
Expect the same sequence to repeat with every automated surface Google ships from here. Capability first, controls second, verification third, with roughly eighteen months between each stage.
Knowing that sequence is worth more than any individual feature, because it tells you when to pilot something and when to wait. Piloting at stage one means paying to be a test subject.
The practical implication for the next ninety days is straightforward. Re-run any AI Max test you ran before this release, with brand and location controls enabled this time. Treat the earlier result as unmeasured rather than negative.
Then build your budget and ROI target tests at portfolio level rather than campaign level, and accept that some of your best performing campaigns are best performing because of what they take from their neighbours.
If your team needs someone to run that reconstruction properly, difrnt. (difrnt.ro) does this kind of work for mid-market and enterprise accounts. It is unglamorous and it usually finds money.
Automation you cannot test is not automation. It is a supplier relationship with no audit rights, and you agreed to it without reading the terms.