CBO Starves Your Best Creative: Use ABO to Test, CBO to Scale
You launch five new creatives into one CBO campaign on Monday. By Wednesday morning one ad set has absorbed most of the budget, the other four have spent just enough to look bad, and the report says four creatives failed. So you kill them.
You did not test five creatives. You tested one, and the campaign picked it for you before any of them had data.
CBO and ABO, mechanically
CBO (campaign budget optimisation) means the budget sits at the campaign level and the system distributes it across ad sets in real time, pushing money toward whatever it currently predicts will deliver the cheapest result. ABO (ad set budget optimisation) means the budget is locked per ad set: each one gets what you assigned it, and the system cannot move it.
That difference sounds like an admin setting. It decides what you are able to learn.
Every major performance platform ranks roughly the same way: your bid multiplied by a predicted conversion rate. CBO's job is to keep feeding whichever ad set currently scores highest on that prediction. That is exactly what you want when the ranking is trustworthy, and exactly what you do not want when the ranking is being assembled out of two days of nearly random outcomes.
The first 48 hours are noise, and CBO reads them as signal
A new creative's first day or two produces almost no reliable information. At the spend levels most accounts actually run, each ad set gets a few thousand impressions and low single-digit conversions on day one. The gap between a 1.8% and a 2.4% click-through rate on that sample sits comfortably inside the range you would get by rerunning the identical ads on a different Monday.
CBO does not wait for that gap to become real. It acts on the ordering it can see, and the ordering it can see is mostly luck: which ad set happened to enter a cheaper auction slice, which one caught a slightly better slice of the audience in its first hour.
Then it compounds. More budget produces more data, more data produces a more confident prediction, and the more confident prediction attracts more budget. Within 48 hours you have a campaign that has already decided, on evidence that would not survive a rerun.
The creative most likely to be destroyed by this is the one you would most want to find: a cheap, high-CTR video that would have become your best asset if anything had funded it past the noise floor. It never gets the impressions to show what it does at scale. It appears in your report as a low-spend, low-result line, so you delete it. You will never know what you killed.
The rule: CBO to convert, ABO to conclude
The working discipline is one line: use CBO when you want conversions, use ABO when you want a conclusion.
During testing that means one creative per ad set, equal fixed budgets, no exceptions. Two creatives inside the same ad set do not solve the problem, because delivery concentrates inside an ad set too. Isolation is the entire point: you are buying the right to say "this creative got a fair, funded shot and still did not perform" instead of "this creative barely spent."
Two habits make the test actually readable:
- Change one variable at a time. Do not swap the character and the format in the same round. When the result comes back you will not know which one moved it.
- Write the kill line before launch. Decide up front how much a creative may spend with zero add-to-carts, sized against your break-even cost per acquisition rather than the campaign's headline efficiency ratio. Written in advance it is a rule; decided in the moment it becomes "let me give it one more day," which is how losers stay alive and winners stay unfunded. The break-even derivation is in why ROAS is lying to you.
The two-track structure experienced buyers actually run
The common misreading of "ABO to test, CBO to scale" is that it describes a one-time handoff: you test for a week, promote a winner, and go back to CBO forever. It does not. It describes two standing tracks that run at the same time, permanently.
Track one: the ABO testing group. Each new creative gets its own ad set and its own fixed budget. Its purpose is not profit, it is conclusions. You expect most of these to lose, and the losses are the price of a clean read.
Track two: the CBO scaling group. A small number of proven winners, one creative per ad set, in one campaign whose budget floats across those ad sets. The platform's aggressive reallocation is now an asset rather than a hazard, because every ad set in the race has already earned its place. Here you actively want the system chasing efficiency.
The testing track never shuts down, and that is the part people skip. Winners fatigue, on TikTok often within days and on Meta within a few weeks, so the scaling group is always slowly decaying. Your testing track exists to have a replacement qualified before the current winner dies, which is the same reason volume matters at all: see how many ad creatives you need per week for the arithmetic, and creative fatigue for the decay curve you are outrunning.
CBO only works when every candidate is already qualified
This is the failure mode that survives even after people adopt the structure. You have a promising new creative, the scaling group is performing, and it feels efficient to just drop the new one in and let it compete.
It gets starved there for exactly the same reason it would have been starved in a test CBO. CBO allocates across ad sets, and against three ad sets with weeks of accumulated performance history, the new one has the worst predicted conversion rate in the campaign by default. The system is not being unfair, it is doing its job with the evidence available.
Promotion is a one-way gate. Things enter the scaling group only after the testing track concluded something about them. A CBO campaign is a race between qualified runners, not a tryout.
The same shape shows up on TikTok Shop
Run GMV Max on TikTok Shop and you arrive at a structurally identical playbook: give new creatives a dedicated test campaign with its own budget for about a week, then promote the ones that clear your screen into the main pool. TikTok goes one further and exposes the exploration budget as a platform lever, covered in Accelerated Testing; the account-level cold-start version lives in the cold-start playbook.
The two platforms differ underneath in ways worth knowing (the weighting inside the prediction, TikTok's organic channel, how fast each audience pool depletes, all broken down in TikTok vs Meta vs organic). The structures converge anyway, because both are solving one problem: a new creative has no history, and something has to feed it data protectively before it is allowed to compete on the main battlefield. Any platform that ranks by predicted performance will underweight a creative with no record yet, not as punishment, just arithmetic. Isolating tests is how you pay that toll on purpose, with a budget you chose, instead of having it silently taken.
Two mechanics that make single-day reads worthless
Both of these quietly wreck testing discipline, and both are structural rather than fixable.
The learning phase threshold is out of reach for most accounts. Delivery systems want somewhere around fifty optimisation events per ad set in a seven-day window before performance stabilises, and the requirement multiplies with your test design: five isolated test ad sets need five times that, not fifty spread across the campaign. Work backwards from your target cost per acquisition and you will find the daily budget that implies. Most small accounts will never hit it, which means "wait for it to settle down" is not a plan available to you, and it is exactly why deliberate experiment design has to replace the system's automatic learning: isolated ad sets, one variable, pre-written kill lines.
Attribution is backdated. Orders are credited to the day the ad was seen, not the day the purchase landed, so the last day or two of any report is always under-reported and always fills in later (the same mechanism that prints zeros on new creatives). No single day is a decision basis, ever.
The testing track is only as good as its supply
Notice what the two-track structure quietly assumes: that you always have new candidates worth funding. The structure is correct and completely useless if your testing group has two creatives in it, because a test track with nothing in the queue is just a slower way to scale one winner into the ground.
That supply constraint is what Riffkit removes. Instead of writing a new concept from nothing and booking a shoot to find out it does not work, you start from a short video that already won in your category and rebuild its structure around your product: your item on screen, new footage, a fresh video in minutes. You riff the formula, not the video. Filling an ABO test group with genuine candidates stops being a production budget question and becomes an afternoon.
Keep both tracks running. Test in isolation, scale in aggregate, and never let a campaign conclude an experiment you never funded.
FAQ
Should I use CBO or ABO for creative testing?
Use ABO for testing and CBO for scaling. With campaign budget optimisation, the platform reallocates budget across ad sets within the first day or two, when the performance differences between new creatives are still statistical noise. Whichever creative gets a lucky start absorbs the budget and the rest are starved before they have enough data to prove anything. Ad set budget optimisation locks a fixed budget to each ad set, so every creative gets the same funded shot and the test produces a real conclusion.
Why is CBO spending all my budget on one ad set?
Because that is what it is designed to do. CBO continuously shifts money toward whichever ad set currently has the best predicted conversion rate, and it starts making that judgement from the first few hours of data. Early differences between fresh creatives are mostly random, so the system locks onto an early leader and compounds the advantage: more budget produces more data, which produces a more confident prediction, which attracts more budget. Nothing is broken, the campaign is just concluding a test you never actually ran.
Do I have to choose between CBO and ABO permanently?
No. Experienced buyers run both at the same time as two standing tracks. An ABO testing group funds each new creative independently, one creative per ad set, and exists to produce conclusions rather than profit. A CBO scaling group holds only creatives that already passed the test, where aggressive budget reallocation is an advantage because every candidate is qualified. The testing track keeps feeding the scaling track, because winners fatigue and you always need the next one.
How long should an ABO creative test run before I decide?
Long enough to clear two known distortions. Attribution is backdated to the day the ad was seen rather than the day the order landed, so the most recent day or two of any report is systematically under-reported and no single day is a decision basis. Small accounts also rarely accumulate the roughly fifty optimisation events per ad set per week that delivery systems want before they stabilise, and an isolated five-creative test needs that many in each of five ad sets, not fifty across the campaign. In practice that means judging on several days of accumulated data against a spend limit you wrote down before launch, not on a morning check of yesterday's numbers.
Keep reading
GMV Max Accelerated Testing: The Lever That Funds Your New Creatives
Accelerated Testing forces exploration budget onto new GMV Max creatives and accounts. How it works, how it pairs with Creative Boost, and the right order.
Your Creative Didn't Fail, Your Landing Page Did: Read the Funnel in Four Segments
Weak ROAS doesn't mean bad creative. Read CTR, add-to-cart rate, checkout conversion and repeat purchase separately to find where your funnel really breaks.
GMV Max Without 1,000 Followers: Authorize Your Own Videos as Ad Creative
Below the commonly cited 1,000-follower gate? The brand marketing account authorization path runs your videos as GMV Max ad creative with no follower minimum.