There is one question worth asking about any advertising spend: what would have happened if you hadn’t run it? Apple Ads produces a great deal of data and none of it answers that question, because none of it contains a control group. Neither does the metric the industry now reaches for instead.
This piece is about that substitution. Not because exclusive reach is a bad metric — it measures its own thing perfectly well — but because it is being read as evidence of causation on the one channel where the causal question is hardest and the platform ships no instrument for answering it.
01The number Apple publishes, and what’s missing from it
Apple’s headline performance claim for its best inventory is “more than a 60 percent average conversion rate for ads at the top of search results.” The footnote is the interesting part: it is a “tap-through install rate from search results ads across all available Apple Ads countries and regions, November 2024 – October 2025.”
Read that as a formula: installs divided by taps, conditional on someone having already tapped an ad they were shown after typing a query. It is an extraordinarily high number and it is not a causal quantity — it contains no comparison group, so it cannot distinguish an install the ad caused from one that was going to happen anyway. A 60% tap-to-install rate is exactly what you would expect from an audience that had already decided.
Apple’s companion claim, that “nearly 65 percent of downloads happen directly after a search,” carries a footnote reading “App Store data from all Apple Ads countries and regions in 2022” — four-year-old data, still on Apple’s best-practices page in 2026. It is a fact about the store, not about advertising.
02What the Exclusive Reach leaderboard counts
Singular’s 2026 ROI Index, published in April 2026, introduced multi-touch attribution leaderboards for the first time, and Apple Ads appears on the Exclusive Reach board alongside Adjoe, AppLovin, Google Ads, Meta Ads, Mintegral, Reddit, Snapchat, TikTok and Unity Ads. Singular defines the metric as measuring “installs where a network was the only engagement in the path to attribution, highlighting platforms delivering truly unique audiences.”
Take that definition apart. “The path to attribution” is a path Singular can observe. “The only engagement” means no other network in that graph claimed a touch. The metric is a description of touchpoint sparsity inside one vendor’s measurement graph — and Singular is honest about the corpus: “the findings are drawn from the company’s own client base … marketers should consider the ROI Index as part of a broader context.”
Nothing in that definition invokes a control group, a randomisation, or a counterfactual. It cannot, because multi-touch attribution is a credit-allocation model applied to observed conversions. It redistributes credit for installs that happened. Incrementality asks how many would have happened regardless — a question about a world that doesn’t exist in the data.
→ THE DISTINCTION
Attribution divides a pie. Incrementality asks how big the pie would have been.
No amount of sophistication in the first ever produces the second. A channel can top the Exclusive Reach leaderboard and be almost entirely non-incremental at the same time, with no contradiction whatsoever between the two facts.
03Why it flatters search channels specifically
Here is the part that should make an Apple Ads buyer uncomfortable. Exclusive reach measures how often no other network touched the user. On App Store search, the user arrived by typing your app’s name or your category into a search field. Of course no display network claims a touch — the user came to the store under their own power and told it what they wanted.
Exclusive reach on a bottom-of-funnel search channel is close to structural. It measures the channel’s position in the funnel, not its contribution to demand. The channels scoring badly are the ones running upper-funnel video that seeds intent someone else later harvests — a description of where they sit, not evidence they did less work.
A metric that rewards being last is not evidence of having been necessary. On a search channel it is barely evidence at all.
This does not mean Apple Ads is unproductive. Our whole thesis is that the App Store audience monetises far better per head than the alternative, and that intercepting a declared intent is the most efficient moment to spend money. But the very fact that makes the audience valuable — they already decided — is the fact that makes the incremental share hard to establish. The strongest argument for the channel and the hardest objection to it are the same observation.
04How incrementality is actually established
The methods are well developed and none are available here. A short tour, with vintages, because the canonical papers are a decade old.
Ghost ads (Johnson, Lewis and Nubbemeyer, Journal of Marketing Research, 2017). The platform “runs a second, simulated auction that includes the focal ad in the set of potential ads,” then “logs ghost ad impressions — the would-be focal ad impressions in the control group.” Only the ad platform can do this, because it happens inside the auction.
PSA placebos, the older approach, serve a public-service ad to the control group. The same authors explain why it now breaks: PSA experiments “are rendered invalid when marketers use performance-optimizing computer algorithms to deliver ads,” because “the ad platform will assign different types of users to be exposed to the PSA or treatment ad.” The optimiser destroys the randomisation it was handed.
Geo lift holds out whole regions and rebuilds the counterfactual synthetically, by finding “the combination of untreated units that most closely replicate the treated.” Its precondition is unglamorous and absolute: sub-country geographic units you can switch on and off independently.
05What Apple ships: an enumeration
The Apple Ads Platform API documentation index lists its topic groups in full: Essentials, Account Management, Search Apps, App Eligibility, Ads on Apple Maps, Campaigns, Ad Groups, Geo Targeting, Keywords, Ads, Creatives, Assets, Product Pages, Bulk Operations, Budget Orders, Reports, Insights, Recommendations, Suggestions, Change History, Changelog.
There is no experiment endpoint. No lift endpoint. No holdout construct, no control-group object, no split-test resource. Insights — the newest and most analytical of those groups — is scoped to impression share and search term popularity.
The same absence holds on the help side. Apple’s “Measuring ad performance on the App Store” page contains none of: incrementality, incremental, lift, holdout, control group, experiment, A/B test, counterfactual, causal. What it offers is attribution: “The AdServices attribution API provides a privacy-centric solution that supports campaign, placement, ad group, and keyword-level attribution.” The Best Practices index covers keywords, campaign structure, bidding, redownloads, ad variations, placements and recommendations — with no measurement-methodology section at all.
The sharpest version: Apple’s campaign-structure guidance recommends a dedicated Brand campaign as one of four campaign types, and says nothing anywhere about how you would test whether it produces anything. Brand search is precisely where the incrementality question is largest.
21
topic groups in the Apple Ads Platform API, and not one of them is an experiment. Apple documents reporting, attribution, impression share and recommendations. It documents no way to construct a control group.
06The geo holdout you probably can’t build
If the platform gives you no experiment tool, the fallback is to build one yourself out of geography. Apple’s documentation makes this genuinely uncertain.
Apple states that ad groups default to reaching all users in the chosen countries, “however, Apple Ads supports location refinement within certain App Store countries and regions” — 29 of them — then adds the constraint that matters: “Location refinements can’t be set within campaigns that are running in multiple countries and regions.” Apple never publicly names the unit. State? Metro? City? The help page does not say.
There is a live vendor disagreement here that should not be smoothed over. At least one third-party guide credits the App Store channel with “City or Metro Area,” radius and ZIP targeting. That reads like a conflation with Apple Maps ads, which genuinely do carry location and location-group objects in the Platform API. Apple’s own help page is the authority and it declines to specify.
So you cannot confirm from public documentation that a geo holdout on Apple Ads is constructible, and that uncertainty is itself the story. It is hard to name another mature ad platform where a buyer cannot determine from the docs whether a regional test is even possible.
07The same question, asked at Google
The comparison is not flattering. Google publishes a named product: “Conversion Lift is an incrementality tool that helps you measure the number of purchases, site visits, and any other conversions directly driven by people seeing your ads.” It describes the mechanism explicitly — a treatment group who see the ads, a control group who don’t, and “the difference in conversions between these 2 groups will tell you the ‘lift.’”
| Capability | Apple Ads | Google Ads |
| Named lift product | None documented | Conversion Lift |
| User-based holdout | None | Yes, aggregated user attributes |
| Geo experiment | Not confirmable from docs | Yes, Google Marketing Areas |
| App campaigns eligible | n/a | Yes, App is a listed eligible type |
| Access | n/a | Gated — via account representative |
| Documented thresholds | n/a | $5,000 budget, 1,000 conversions |
Google’s geo methodology names purpose-built units: Google Marketing Areas are “sub-country geographical regions designed to serve as experimental units in geo experiments,” with the constraint that “campaigns must target a single country to be eligible for the geo-split methodology.” Apple names no unit at all.
Two honest caveats. Google’s product is gated behind an account representative, so “documented” is not “available to you.” And neither Google nor Meta discloses an iOS-specific limitation on lift measurement in the pages we read. We covered the wider asymmetry in our Apple Ads versus Google App Campaigns comparison.
08The MMPs, and the one that says yes
If the platforms won’t give you an experiment, the measurement partners advertise one. What they document is thinner than the marketing.
AppsFlyer promises incrementality experiments “across major ad networks from one dashboard” with automated holdout groups — and names no networks on that page. Its Apple Ads integration page is otherwise admirably precise about limits, stating that “Currently Apple Ads integration doesn’t support postbacks.” Incrementality appears nowhere on it, either way.
Singular describes A/B holdout testing with “real control groups” and names no supported networks at all — which sits awkwardly beside the fact that Singular is the vendor ranking Apple Ads on the Exclusive Reach leaderboard. For Adjust we found no incrementality documentation naming the channel at all.
Kochava is the only one that names the channel explicitly — as Apple Search Ads, the platform’s pre-2025 name — and it deserves credit for describing its method plainly. Its AIM tests use “a Pulse testing methodology” that “alternates between OFF periods — during which your campaign is manually paused at the ad network level — and ON periods,” noting that “you are responsible for manually pausing and resuming the selected campaign.”
That is a switchback design, not a randomised holdout. It compares one period against another, inheriting every confound a time series has: seasonality, day-of-week effects, competitor bidding, your other channels, ranking dynamics, and — specifically here — the bidder relearning after each pause, which matters on anything using Maximize Conversions and its two-week learning guidance. A reasonable workaround for the absence of a control group. Not a control group.
→ FREE 10-POINT AUDIT
Find out how much of your Apple spend is actually incremental
We open your Apple Ads account, separate the spend that intercepts demand from the spend that follows it, and show you where a brand campaign is buying users who were already on their way. One hour. Senior strategist. No pitch.
Book my free audit →
09The branded-search evidence, and its real disagreement
The reason this matters most for brand campaigns is a famous result routinely quoted without its rebuttals. The whole file, with vintages:
Blake, Nosko and Tadelis, Econometrica, 2015, on experiments run March–July 2012. eBay switched off brand-keyword advertising on Yahoo and Microsoft while continuing to buy on Google as a control. Finding: “brand-keyword ads have no measurable short-term benefits,” and “almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search traffic.”
Simonov, Nosko and Rao, Marketing Science, 2018 — Nosko co-authored both papers. Their conclusion: results “are consistent with the findings of BNT for a company like eBay, but show that eBay’s case as a very strong brand facing no competitors is not the norm.” With the brand in position one, “competitors can steal only 1%–5% of clicks,” but “a single competitor in the top position acquires 15%–20% of searchers.”
Coviello, Gneezy and Goette ran the same shape of test on Edmunds in 2015 and got a different answer: “only about half of the traffic normally flowing through branded search ads still flowed to the site when it relied only on organic search links,” rising to about 72% lost in high-penetration markets. Their own comment is the most useful sentence in the literature: “We do not yet know why there is such a difference in results between eBay and our experiment.”
Contrary evidence from mobile points the other way entirely. A 2025 paper analysing three ad-shutoff experiments at a US mobile game developer — six apps, 500 days of 2018–2019 data, 85 ad publishers — reports that “contrary to previous findings, we found that paid advertising boosts organic installs rather than cannibalizing them.” That is mobile display, not App Store search, and the data is seven years old.
What survives: the effect is real, its size is contested, it depends heavily on competitive intensity in the auction, and no published study has ever tested branded-keyword cannibalisation inside an app store. Every transfer of the eBay result to the App Store — including ones we have made ourselves — is analogy, not evidence.
10Even with the tool, the statistics are brutal
Suppose Apple shipped a lift product tomorrow. Most advertisers still could not use it, and the reason is arithmetic rather than engineering.
Lewis and Rao, in the Quarterly Journal of Economics in 2015, examined twenty-five large field experiments with major US retailers and brokerages, most reaching millions of customers and together representing $2.8 million of spend. Informative experiments needed “more than 10 million person-weeks”; for the median campaign to distinguish a 50% return from break-even, “the median campaign would have to be nine times larger,” and to detect a 10% ROI difference from zero, “61 times larger.” For a representative campaign, R² was 0.0000054 — meaning “a very small amount of endogeneity would severely bias” any observational estimate.
That is an eleven-year-old paper about retail sales, and app installs are a higher-frequency, lower-variance outcome than revenue, which helps. But the lesson’s direction holds, and it explains why the honest answer to “what’s my incrementality?” is often “your spend is too small to know.”
11What to do without the instrument
Stop calling exclusive reach incrementality. This is free and it is the highest-value item on the list. If a slide says a channel delivers unique audiences, that is what it says. It does not say the installs were caused.
Rank your spend by how answerable the question is. Brand keywords are where the incremental share is most doubtful and the test cheapest to approximate. Generic and competitor terms are where the counterfactual is genuinely open. Put your scepticism where the doubt is, not evenly across the account.
If you run a pulse test, pre-register it. Fix the OFF and ON windows and the outcome metric in advance and in writing, match day-of-week, run several cycles, and hold your other channels flat throughout. A pre-registered switchback is far harder to fool yourself with than a retrospective one. Discard the first days after each restart while the bidder relearns.
Use organic as the read-across, carefully. Apple’s Impression Share reporting has no position dimension, so it will not do this for you. But total store downloads against paid downloads across a pulse cycle is the crudest possible substitution test, and crude beats assumed. The same caution applies to any channel whose outputs you can observe but not attribute — including Apple News, where no attribution exists at all.
Ask Apple for the tool. Apple built brand and location-group objects into the Platform API the moment Maps needed them. The absence of an experiment endpoint is a product decision, not a physical constraint.
12What nobody knows
As of 31 August 2026: no published incrementality study of Apple Ads exists — not academic, not vendor, not agency. We searched three ways and found only marketing pages describing incrementality services without a single measured figure. Apple documents no experiment, lift, holdout or control-group capability anywhere.
Apple has never publicly named the geographic unit of its location refinement, so whether a geo holdout is constructible is unresolved. No MMP documents a randomised incrementality test here; the one vendor naming the channel offers a switchback. Meta’s lift documentation is robots-disallowed, so we quote none of it rather than quote it second-hand. And no study tests branded-keyword cannibalisation inside an app store — the environment every argument above is about.
The awkward summary: the highest-converting paid placement available to an iOS marketer is also the one where the causal question is hardest to ask and the platform helps least in asking it. That is not a reason to stop buying it. It is a reason to write “unknown” in the cell where the incrementality number would go, rather than filling it with the nearest available metric that happens to look impressive.