Welcome back — jump to Tutorials, Support & Docs, or What’s New. Go to PlatformNot you?Log out
Welcome back — jump to Tutorials, Support & Docs, or What’s New. Go to PlatformNot you?Log out
Welcome back — jump to Tutorials, Support & Docs, or What’s New. Go to PlatformNot you?Log out
Platform AI Agents Customers Developer Docs Services Company Careers Blog 2026 Benchmark Report Resources Comparisons Buyer's Guide Integrations Use Cases Tutorials Events Sign In Book a Demo
Campaign Experimentation

Test everything. Scale what works.

Test audiences, creatives, channels, and offers in one place. Metadata builds campaign variations, tracks their performance, and optimizes spend toward your campaign goals.

Multivariate Testing

Test audiences, creatives, and offers simultaneously

Choose audiences, ad creatives, offers, and channels. Metadata builds combinations into distinct experiment cells with their own performance reporting. For example, three audiences and three creatives produce nine combinations.

Use those results to decide what to scale and what to test next. Compare cost, lead quality, and pipeline contribution alongside click and conversion rates.

Experiment Matrix: 3 Audiences x 3 Creatives 9 Experiments AUDIENCE \ CREATIVE Creative A Creative B Creative C SaaS CMOs LinkedIn CPL $52 CTR 1.8% CPL $31 CTR 3.2% CPL $67 CTR 1.1% FinTech VPs Facebook CPL $44 CTR 2.1% CPL $48 CTR 1.9% CPL $59 CTR 0.9% Enterprise IT Google Ads CPL $71 CTR 1.3% CPL $55 CTR 1.6% CPL $83 CTR 0.7% WINNER SaaS CMOs + Creative B BEST CPL $31 REVIEW Results PIPELINE $482K

Illustrative example. Figures are not customer results or guaranteed outcomes.

Experiment Planning

Give each test a fair chance.

Define the outcome, sample size, and decision rule before you start. More variations and lower conversion rates need more data.

Plan

Choose a decision rule

For a fixed-sample test at a 5% significance level, the false-positive rate is 5% under the null hypothesis when the test assumptions hold. It does not mean there is a 5% chance the result is random. Review effect size and business value as well as statistical evidence.

Min. N

Minimum Sample Sizes

Plan the sample using your baseline conversion rate, the lift worth detecting, and your desired power. Avoid declaring a winner from a handful of conversions. The calculator below estimates a sample for a simple two-variation test.

MDE

Minimum Detectable Effect (MDE)

The minimum detectable effect is the lift you plan a test to detect. Smaller lifts generally require larger samples. Statistical significance and business value are separate questions: decide what improvement would justify more spend.

What You Can Test

Seven dimensions of experimentation

Compare the campaign variables that matter to your team. Available controls depend on the channel and campaign type.

Audience Segments

Test different firmographic, technographic, and intent-based segments against each other

Ad Creative

Images, copy, video, carousel formats, test which creative resonates with each audience

Bid Strategies

Manual CPC, target CPA, maximize conversions, find which strategy delivers the best pipeline ROI

Time-of-Day

Morning, afternoon, evening delivery, discover when your audience is most likely to convert

Channel Mix

LinkedIn vs. Google vs. Meta vs. Reddit, test which channels deliver for each audience segment

CTA Text

"Book a Demo" vs. "See It Live" vs. "Get Started", small copy changes can shift conversion rates significantly

Landing Pages

Test which page, form length, or offer converts best, connect ad experiments to downstream page performance

%

Offers

Whitepaper vs. webinar vs. free trial vs. demo request, find which offer drives the most qualified pipeline

Dynamic Budget Reallocation Continuous rebalancing based on pipeline metrics BEFORE (Equal Split) Exp A Exp B Exp C Exp D $2,500/day each Goal review complete AFTER (Optimized) Exp A Exp B Exp C Exp D Paused Winner gets 52% of budget Example allocation change 1 Equal split launch Illustrative starting allocation across four experiments 2 Review performance Assess results against your campaign goals 3 Losers paused, winners scaled Underperformers are paused and their budget flows to the top performers automatically Example: more budget allocated to stronger performance

Illustrative example. Figures are not customer results or guaranteed outcomes.

Budget Allocation

Dynamic budget shifting, winners get more, losers get paused

As experiments run, Metadata monitors performance against pipeline metrics, not just clicks. Budget automatically shifts away from underperformers and toward the experiments driving real outcomes.

Optimization uses campaign performance and your configured budget settings. Decide which outcomes matter, then review changes against those goals.

Use pipeline contribution, cost per lead, and cost per opportunity to assess results alongside engagement. Allow for your sales cycle before judging downstream performance.

Keyword Experimentation

Discover high-intent keywords at scale

For Google Ads, run keyword experiments that test hundreds of terms simultaneously. Pause low performers automatically and double down on the keywords driving qualified leads and pipeline.

Compare keyword-level results and consider traffic volume, conversion delays, and spend before changing bids or pausing terms.

Experiment Results Scale Winners ILLUSTRATIVE TOP PERFORMER COST PER MQL $38 PIPELINE $482K CTR 2.4% CONV RATE 8.1% VARIANT B (Paused) COST PER MQL $67 PIPELINE $198K CTR 1.1% CONV RATE 3.2% Pipeline Over Time Winner Variant B

Illustrative example. Figures are not customer results or guaranteed outcomes.

How It Works

Three steps to smarter experiments

1. Define Variables

Choose the audiences, creatives, offers, and channels you want to compare. A three-audience, three-creative, two-offer setup produces 18 combinations. Plan enough traffic for the comparisons you intend to make.

2. Launch and Monitor

Track experiment performance in unified reporting. Review sample sizes, conversion delays, and lead quality before drawing conclusions.

3. Scale Winners

Use performance results to guide budget allocation and the next round of creative and audience tests. Metadata can optimize campaign spend within your configured settings.

Platform-wide

Experiments Run

263K+

Experiments run across the Metadata platform.

ThoughtSpot

ThoughtSpot: campaign experimentation in practice

ThoughtSpot's customer story describes its use of Metadata to automate campaign experimentation and improve demand generation. Read the customer story.

Reported customer outcomes

LogicMonitor
155%

more pipeline with 25% lower CPC

Instruqt
88%

lower cost per lead reported by Instruqt

Forrester TEI, October 2022
$3.26M

Modeled 3-year net present value for the composite organization. Study commissioned by Metadata. Read the study.

Planning Tool

A/B Test Sample Size Calculator

Estimate the visitors needed per variation for a planned conversion-rate test. Reaching this sample does not guarantee a significant result.

Relative lift you want to detect
Replace this example with your traffic estimate
Sample Size Per Variation
,

visitors needed per arm

Estimated Duration
,

based on your daily traffic estimate

Planning assumptions: two independent, equally sized groups; one conversion outcome per visitor; 80% power; a two-sided fixed-sample test. No adjustment for repeated checks or multiple comparisons. Duration excludes conversion delays. This educational calculator does not describe Metadata's campaign optimization algorithm. Method reference.

Frequently Asked Questions

How many experiments can I run?

Run combinations of audiences, creatives, offers, and channels in parallel. Choose a test scope your traffic and budget can support; more variations divide the available sample.

What can I test?

Compare audiences, creative, copy, offers, and channel performance. Available bid and scheduling controls depend on the channel and campaign type.

How does budget optimization work?

Metadata uses campaign performance to optimize spend toward your goals within your configured budget settings. Review pipeline and lead quality alongside cost and engagement metrics.

What does statistical significance tell me?

In a fixed-sample test, a p-value measures how incompatible the data are with the specified null model. It is not the probability that the result happened by chance. Consider effect size, sample quality, and business value before acting.

How much budget and time does a reliable test need?

There is no universal spend or duration that guarantees reliable results. Start with your conversion rate, the lift worth detecting, required sample size, traffic cost, and conversion delays.

Does this calculator model Metadata optimization?

No. It is an educational planning estimate for a two-variation conversion-rate test with equal traffic, independent visitors, and 80% power. It does not account for adaptive allocation, repeated checks, or multiple comparisons.

Can I compare different channels?

Yes. Compare campaign outcomes across channels, using consistent definitions and attribution windows. Channel audiences and delivery differ, so an observed performance difference alone does not establish a causal effect.

G2 Leader
Leader
Spring 2026
4.6/5 on G2 · 34 badges · 133 reports
Experimentation from pricing scoped per team · See plans →