← Guides / Creative strategy

Why comparing ads in an ad set is not a controlled A/B test

Comparing ads inside a standard ad set can help you choose what to run, but it does not by itself isolate the effect of the creative. Delivery, audiences and multiple creative differences can all affect the result. Structured concept testing makes those comparisons easier to interpret.

By NewForm · Updated

Key takeaways

  • Ordinary ad-set results do not establish that ads received comparable, randomly assigned audiences.
  • Comparing entirely distinct ads introduces multiple confounding variables, making it difficult to isolate what drove performance differences.
  • A structured testing model clusters variations around a single concept or persona, testing specific combinations such as modular hooks and bodies.
  • Repeated patterns can generate better creative hypotheses; they do not by themselves prove which element caused the result.

Dynamic delivery confounds standard ad set comparisons

A controlled A/B test uses random assignment to comparable groups and defines the difference being tested. Simply placing two ads in the same ad set does not establish those conditions. The videos highlight that different creatives may receive different delivery and reach different audience segments.

An observed CPA difference is still useful for an operating decision. What it does not establish on its own is whether the creative caused the difference, or whether one ad would be better for every audience. The critique here concerns casual ad-set comparisons, not a claim that controlled advertising experiments are impossible.

The problem of unisolated variables in creative testing

Beyond delivery skew across audience pockets, casual comparisons inside an ad set often involve entirely distinct creatives with multiple unisolated variables. When two ads differ simultaneously across visual pacing, hooks, value propositions, and calls to action, it is impossible to attribute differences in performance to any single element.

An ad might register higher engagement simply because its opening frame captured immediate attention, even if its underlying value proposition was weaker. Without controlling for individual creative variables, declaring one ad the winner yields little actionable insight into what actually drove the outcome.

Structuring tests around concepts and modular components

Rather than placing unrelated ads into an ad set and treating it as a split test, our approach structures creative evaluation around distinct concepts or personas. Keeping the conceptual frame constant allows teams to methodically explore variations across modular components.

This structure organizes learning; it does not remove delivery bias or turn the workflow into a randomized experiment.

  • Dedicate each ad set to a single defined concept or customer persona to maintain messaging consistency.
  • Structure variations using a modular hook-and-body framework, such as pairing distinct hooks with consistent body segments, to systematically explore the creative space.
  • Compare hook and body combinations in context, then carry promising patterns into the next round of testing.

Detecting macro patterns across multiple test cycles

Review reports across several concepts and personas before treating an isolated result as a general creative rule. A recurring pattern is a stronger reason to investigate an idea than one unusually successful asset.

Our videos describe using language models to summarize those reports. They can help propose patterns for a strategist to check against the actual results. A model summary is not independent evidence, and repeated correlations still do not establish causation.

Adapted from NewForm’s original videos on creative strategy and paid social.

Explore NewForm Framework →
CLOSINGGet startedREPLY WITHIN 24H

Ready to build your creative intelligence layer?

End of fileNewForm · 2026