The product page is the last surface between a user who has declared intent and an install worth several times its Android equivalent.

Why this surface carries the high-ARPU thesis

The case for buying iOS users through App Store search is that you are intercepting declared intent on the storefront where per-download consumer spending is several times Google Play’s. Apple states that nearly 65% of App Store downloads happen directly after a search, and reports more than a 60% average conversion rate for ads at the top of search results. Those two numbers are the reason Apple Ads buys the most valuable users in mobile.

But conversion rate is not a property of the ad. It is a property of what happens after the tap. A search results ad delivers a user to a product page, and that page is where an expensive, high-intent, high-ARPU visitor either becomes an install or becomes nothing. If the tap costs $1.91 at the US median and the page converts at 55% rather than 62%, the derived cost per install moves by roughly 13% — a derived illustration, using AppTweak’s published 2025 US medians, not a measured result. The page is where bid efficiency is actually decided.

Which is why the testing framework governing that page deserves more scrutiny than it gets, and why a new testable variable is a bigger event than the coverage suggested.

What WWDC 2026 actually changed

Four things, in descending order of how much they should change your plans.

The header became testable. Apple renamed the Feature Banner to Header and made it editable, uploadable and reviewable in App Store Connect. Apple’s own summary of the new Creative Assets says to use these visuals on your custom product pages and even test them using product page optimization. Previously the header was fixed: you could test screenshots, previews and the icon, and nothing else. The testable set is now four assets. All text metadata — name, subtitle, description, keywords — remains untestable.

Assets ship without a build. Apple introduced the Asset Library, described as a new centralized place in App Store Connect, and stated that developers can submit assets independently of an app submission and get them approved in advance for use in a future product page update or Apple Ads campaign. This is the change with the largest operational consequence: creative cadence decouples from release cadence.

Creative assets reach organic search results. Assets can be selected to appear in organic keyword search results, working across the main product page, custom product pages and PPO tests — so the same asset now spans organic discovery and the two-slot paid search results at once.

Custom product pages gained headers and better deep linking, on top of the limit having doubled from 35 to 70 in October 2025, with keywords assignable to individual pages. Apple documents 70 custom ads in Apple Ads, each with its own deep link.

The arithmetic of a PPO test

Now the constraint side, all of it from Apple’s own documentation.

ParameterApple’s valueConsequence
Treatments per testUp to 3, plus the originalFour-way split of allocated traffic
Maximum duration90 days, hardUnderpowered tests expire rather than resolve
Confidence threshold90%Looser than the 95% norm
Indicators appearAfter 7 daysNo early read is available even if you want one
Results displayAfter 5 first-time downloadsA display threshold, not a validity threshold
StoppingIrreversibleA stopped test must be recreated from scratch

Apple states the trade-off itself: the more treatments you add to a test, the longer the test may take to reach a conclusive result. And Apple names the failure state directly. “Likely to Be Inconclusive” means the test is not getting enough data to make a determination within the 90-day limit. That is Apple telling you, in the interface, that the test cannot finish.

The 90% threshold deserves a moment. At 90% confidence, roughly one in ten declared winners is noise. Run three treatments against one baseline and you are making three comparisons, which inflates the chance that at least one crosses the line by accident — and Apple does not state whether it corrects for multiple comparisons. Apple describes its method only as algorithms that provide confidence in the reliability of the test results. That is the complete disclosure.

The traffic ceiling nobody quotes

This is the constraint that almost never appears in the coverage, and it is the one that decides whether your test can finish.

Apple describes traffic proportion as the percentage of users randomly shown a treatment instead of your original page, split evenly across treatments: at three treatments and a 30% proportion, each treatment gets 10% of total traffic. Then Apple adds the ceiling, in its Tech Talk: no single treatment can receive more traffic than your original product page.

3 × 90 × 90%
Three treatments, ninety days, ninety percent confidence — with each treatment capped below the baseline’s share. Those four numbers, all from Apple’s documentation, define everything a PPO test can and cannot resolve.

Stack the two rules and the shape becomes clear. Your baseline always gets the largest single share. With three treatments the per-treatment allocation cannot approach half of traffic, and in practice each cell is a small minority of an app’s already-limited page views. Apple publishes no minimum impression or install threshold at all — no sample-size table, no minimum detectable effect guidance. The only instrument offered is an in-product duration estimator, which uses your app’s existing daily impressions and downloads and whose model Apple does not disclose.

The practical reading: for most apps, a four-cell test detecting a realistic single-digit conversion improvement will not resolve inside 90 days. That is not a criticism of Apple, it is the arithmetic of small samples. It does mean the correct default is two treatments, not three, and one hypothesis per test.

→ FREE 10-POINT AUDIT

Find out whether your tests can even finish

We look at your actual product page traffic, your test history and your Apple Ads volume, and tell you how many cells you can afford, which hypotheses are worth a 90-day slot, and which questions belong on a custom product page instead. One hour. Senior strategist. No pitch.

Book my free audit →

The icon trap

Apple documents one detail that quietly invalidates a lot of test planning, and it is worth quoting closely. When you apply a winning treatment, Apple states that only the app previews and screenshots from the treatment will be applied, and that to apply changes to the app icon you must set it as the default icon in your next app version.

So an icon test has a release dependency at both ends. Any icon you want to test must already be inside the shipped binary for the current App Store version, sized to 1024 × 1024. And any icon that wins does not go live when you press apply; it ships with your next build. A team that runs a 60-day icon test and then waits five weeks for a release train has spent a quarter of its annual testing capacity on a change it has not yet realised.

Two more irreversibilities sit alongside it. Applying a treatment while a test is running stops the test automatically, and the action cannot be undone. You can apply one treatment per test. And stopping a test at all is terminal — you cannot restart it, only build a new one. Combined with the 90-day ceiling, this makes each test slot a genuinely scarce, non-refundable asset.

PPO against custom product pages

These two tools are constantly conflated, and the distinction decides where a question belongs.

DimensionProduct Page OptimizationCustom product pages
What it isA test on your default pageUp to 70 additional live pages
AnswersWhich asset is betterWhich audience sees what
TestableYes, that is the pointNo — Apple excludes CPPs from PPO
Time costUp to 90 days per questionLive immediately once approved
Apple Ads roleUnclear (see below)Ad variations, one deep link each

Apple states plainly that PPO tests are not available for custom product pages. So the two are complements: PPO answers which creative wins on the page everyone lands on, and CPPs answer which creative for which query across up to 70 destinations — which, since Apple lets you assign keywords to individual custom product pages, is now an organic discovery lever as well as a paid one.

There is one question Apple genuinely does not answer, and you should not let a vendor answer it for you: whether PPO test traffic includes Apple Ads traffic. Apple describes PPO placement, not traffic source — treatments appear across the App Store to people on iOS 15 and iPadOS 15 and later — and Apple’s own Tech Talk lists paid advertising among the ways visitors arrive without excluding them from tests. Vendors contradict each other flatly: one states PPO applies exclusively to organic traffic, another that PPO takes all your traffic and distributes it. Both cannot be right. Until Apple says, treat any conversion-rate read that mixes your paid and organic assumptions as unreliable, and see the wider problem of reasoning past Apple’s documentation.

The 4+ ceiling constrains what you can test

The new creative assets arrive with a content policy attached, and it narrows the hypothesis space before you write a single test brief. Apple’s asset best practices state that assets displayed on the App Store must meet a 4+ age rating, even if your app’s rating is higher and it is intended for an older audience.

Also prohibited, in Apple’s words: specific pricing, discounts, website URLs and copyright symbols; claims that cannot be verified, such as awards the app has not received; logos or references to other platforms or marketplaces; and Apple-designated recognitions such as Editor’s Choice, App of the Day, Game of the Day or Apple Design Award. Video has its own rules — audio is muted by default, so creative must work silently first, and videos autoplay and repeat, so the loop must be seamless.

Read as a test-design constraint rather than a compliance checklist, this removes several of the highest-lift creative levers other channels rely on: price framing, social proof by award, urgency, and cross-platform credibility. It also means your header hypotheses will cluster around a narrower set of variables than your paid social ones do. The full policy read is here, and because the Asset Library is shared between storefront and campaigns, one rejection now lands on both at once.

Every published uplift number measures the old world

You will be asked what lift to expect from a header. The honest answer is that nobody knows, because headers did not exist in production when any of the published data was collected.

The numbers in circulation also disagree badly. Apple states that developers see a 2.5 percentage point increase on average when referring people to a custom product page, which it presents as a 156% increase against a 1.6% default-page conversion rate — an average with an undisclosed sample and a denominator that compares referred traffic against default-page traffic including scrolled-past impressions. AppTweak reports roughly 8% for games and 6.6% for apps. SplitMetrics data published in Singular’s 2026 index describes 65 to 80% custom product page adoption in Sports and Finance with a 20 to 50% conversion uplift band. A single SoundCloud case study reports a 58% conversion increase and 39% lower cost per install — and a case study of one is a story, not a rate, in the same way a network-reported install count is not an incrementality finding.

Those ranges are three to six times apart and not one of them discloses a sample size. Report the spread, never the midpoint — the same discipline that applies to cost benchmarks. And date every one of them: all predate the header era, and all measure screenshot, preview and promotional-text pages.

Any header uplift figure published before the format ships is not a measurement. It is a forecast wearing a percentage sign.

A test plan that fits the constraints

The creative assets format, Asset Library and Product Page Preview are all marked “coming this fall” by Apple and are expected alongside iOS 27, which had not shipped as of 26 August 2026 and is widely reported for mid-September. That gives you a short planning window. Five decisions worth making inside it.

Default to two treatments. Three is available and rarely correct. Two treatments plus baseline gives each cell a materially larger share and one fewer comparison to inflate your error rate.

Sequence by irreversibility. Icons need a build to test and another build to realise, so they go first in the calendar and last in priority. Headers and screenshots now ship without a build, so they can move at whatever cadence your review turnaround allows.

Spend the 90 days on one hypothesis. A test slot is a quarter of a year. Use it on the variable with the largest plausible effect, not the one that is easiest to produce.

Push audience questions to custom product pages. With 70 pages, keyword assignment and independent asset review, most “which message for which segment” questions belong there, live, rather than consuming a PPO slot.

Write down what you would need to see to be wrong, before the test starts. Apple shows confidence indicators from day seven and explicitly warns against acting on early results, while giving you a dashboard that updates continuously. Pre-registering your stopping rule is the only defence against your own peeking.

What nobody can tell you

Apple publishes no minimum sample size for a PPO test — no impression floor, no install floor, no minimum detectable effect table. The five-download figure is a reporting threshold, not a validity threshold, and conflating the two is the most common error in this area. Apple does not disclose its confidence algorithm, so it is unknowable from public sources whether the 90-day peeking warning reflects a statistical necessity or a design convention, and Apple does not state whether it corrects for multiple comparisons across treatments. Apple never states a maximum traffic proportion numerically; it states only that no treatment may exceed the baseline’s share. Whether PPO traffic includes Apple Ads traffic is unresolved and vendors contradict each other on it. Whether the 70-page custom product page limit and the 70 Apple Ads custom ads are the same underlying object is undocumented. Header format specifications circulate as 16:9 landscape for header and search results and 9:16 portrait for In-App Event details, but those aspect ratios come from trade coverage rather than an Apple page. And no header uplift data exists at all, because the format has not shipped.

What you can do before it does is build the test calendar, cut your default treatment count, and decide which questions were never PPO questions in the first place.