Geo-lift incrementality test
A geo-lift test tells you how many sales a campaign caused. The campaign runs in a few test cities, and a synthetic control built from the other cities shows what those cities would have sold without it.
Adjust inputs ↓Did the campaign cause a lift?
Counterfactual
Test cities vs. their synthetic control
Daily conversions summed over the test cities. The dashed line is a weighted blend of control cities, fitted on the pre-period only. After the start line, the gap between them is the effect.
Effect
How the lift builds day by day
The shaded band shows how big the daily gap gets in control cities that saw no campaign.
Inference
How big a gap chance alone produces
Donors
Who builds the synthetic control
| Control city | Region | Weight | Pre avg/day |
|---|
Power
Smallest lift this design can detect
Before spending, back-test the design on history. The method replays fake tests inside the pre-period, adds a known lift to the test cities, and counts how often the 90% interval excludes zero.
Geo selection
Suggest test cities
Ranks every city by how well the others can reproduce its pre-period, then picks the best-fitting ones until the test group holds 12–20% of volume. Very large cities (over 6% of volume) are skipped: they're hard to reproduce and costly to hold out.
How it works
The data
40 German cities, 132 days of daily conversions, generated in your browser. Each city is its size times a shared national trend, a weekly pattern, a regional factor (North, West, South, East), two hidden factors with city-specific loadings, its own drift and count noise. After day 90 the test cities get the lift you set. Nothing else changes.
The model
Each city is divided by its pre-period mean so big and small cities are comparable. The test cities are summed into one series. The synthetic control is a weighted average of control cities: weights are non-negative and sum to 1, chosen to minimise squared error over the pre-period. The fit is accelerated projected gradient descent onto the simplex.
After the test starts, the synthetic control is the counterfactual: what the test cities would have done without the campaign. Lift is the post-period gap divided by the counterfactual.
Inference
Placebo groups: the same number of control cities as your test group, picked at random (with a fixed seed) and preferring groups of similar volume. Each gets its own synthetic control from the remaining controls. Their post-period "lifts" show how big a gap appears by chance.
- The 90% interval is your estimate minus the 5th and 95th percentile of placebo lifts. The verdict uses this interval.
- The p-value ranks your test's post/pre RMSPE ratio among the placebo ratios (Abadie's test). It also catches effects that come and go.
- Power: fake test windows inside the pre-period, a known lift added, the same interval rule applied. MDE is the lift detected in 80% of windows.
In production
The KPI comes from the warehouse (BigQuery models in dbt), at geo-day grain. Spend comes from the ad-platform APIs. Holdouts are set in platform geo-targeting, and the test plan is fixed before launch: regions, dates, MDE. Results feed the media mix model as priors on channel effects and go into the budget plan.
Limits
- Spillover: people commute and media leaks across borders. Neighbouring control regions then understate the lift.
- Few regions give few placebos. The smallest possible p-value is 1 over (placebos + 1).
- Needs a stable pre-period. A launch, outage or tracking change before the test breaks the fit.
- One test at a time per region. Overlapping campaigns in the same regions can't be separated.
- The test group must be inside what the controls can reproduce. A city unlike any other has no good synthetic twin.