What 23 paywall tests in one month look like from the inside

Victoria Kharlan
Victoria Kharlan
11 min read
What 23 paywall tests in one month look like from the inside

TL;DR:

  • Every test at Glam AI closes on net revenue per user or LTV. Conversion is a diagnostic. One variant lifted purchase conversion by a quarter and got killed anyway, because revenue per user came in under control.
  • A win belongs to the placement and the segment it was won in. The decoy plan that took yearly conversion from 0.43% to 0.55% on the main paywall did nothing on the inner paywalls.
  • Order the backlog by distance to money: paywall and pricing first, pre-purchase funnel second, activation and retention last.

Copying a paywall used to cost an engineering sprint. Now it costs an afternoon, for you and for everyone selling against you, so shipping fast has stopped being the thing that separates you. What separates you is knowing which of the twenty things you shipped last month actually made money.

That's harder than it sounds. Paid traffic costs more every quarter, LTV/CAC goes flat every summer, and a variant that lifts conversion can still take revenue down with it.

Artem Altynov runs growth product at Glam AI, an AI photo editing app doing $50M ARR on iOS and Android with 4M monthly active users. His team completed 23 tests in July. Nine won. He walked through four of them in our webinar on Glam AI's experimentation setup: two wins and two kills, and one of the kills had raised purchase conversion by a quarter.

Why did Glam AI kill a test that lifted purchase conversion?

Because revenue per user fell. The team was testing hero content on the main paywall, static imagery against a video, with the plan structure and the $59.99 year held identical on both sides.

Variant B moved view-to-payment from 1% to 1.25% across roughly 8,300 users per group, which is a clean result on a sample big enough to trust, and the team threw it out anyway.

Two Glam AI paywall variants with identical pricing, where the higher-converting version lost on net revenue per user.Variant B lifted view-to-payment from 1% to 1.25%, and Glam AI shipped variant A anyway

The video hero in B was the stronger hook. It also pushed people toward the $8.99 weekly plan instead of the year, so yearly subscriptions dropped, and refunds went up. Net revenue per user came in under A on all traffic, then under A again when they cut the data to the US alone. They kept the control.

Glam AI splits its metrics in two. Conversion, net revenue per user, and LTV decide. Churn, refunds, reviews, and support tickets are guardrails, and a variant that wins on conversion while a guardrail slides doesn't ship, whatever the uplift looks like. Every growth team I've talked to agrees with this in principle. Then conversion comes back on day three, revenue per user needs a fortnight and a refund window before it means anything, and the variant is live before anyone checks. That gap is where paywall tests do most of their damage.

What should you test first when every screen looks testable?

Glam AI orders the backlog by distance to money: how directly a change touches revenue, and how many other things have to go right before the effect reaches the bank.

Three-tier test prioritization with monetization first, pre-purchase funnel second, activation and retention last.Glam AI ranks every test idea by how many steps sit between the change and the money

Glam AI ranks every test idea by how many steps sit between the change and the money.

Tier one is monetization. Paywall, offers, pricing, plan structure. Change a price on Monday, and you can read the result by Friday, because nothing sits between the change and the purchase.

Tier two is the pre-purchase funnel, meaning onboarding, paywall entry points, and the flow into the purchase itself. The effect is real; it just travels through more steps before it lands, so you need a bigger sample and more patience to see it.

Tier three is activation and retention. Core experience, engagement, the deeper product mechanics. Some of these are the most valuable work on the roadmap, and all of them are terrible candidates for a fast experiment, because a retention win can take a quarter to surface in revenue, and by then you've shipped forty other things on top of it.

That ordering is a scheduling decision more than a value judgement. At 23 tests a month, you need results in days, so tier one carries the volume. Tier three still gets built, on a different clock, usually by someone else, which is why onboarding testing tends to stall when nobody owns it.

Does adding an expensive plan nobody buys make the other plans sell?

On the first paywall a user sees, yes. Glam AI's control offered two options: the $59.99 year and the $8.99 week. Variant B added a third above them, Yearly Business at $119.99, which works out to $2.31 a week against the regular year's $1.15.

A two-plan paywall against a three-plan version where an expensive option sits above the plan Glam AI wants users to buy.Adding Yearly Business at $119.99 above the $59.99 year moved yearly conversion from 0.43% to 0.55%

Nobody was expected to buy Business, and the point was what it did to the option underneath it. Yearly conversion went from 0.43% to 0.55% across roughly 24,500 users per group. The weekly plan held flat, so the year gained without eating the week, and net revenue per user moved from $0.261 to $0.292. That one shipped.

Worth sitting with the size of that sample. The effect is twelve basis points on a metric already under half a percent, and it took nearly 50,000 users to read it with any confidence. Plenty of teams run this exact test on 2,000 users, see noise, and conclude the decoy doesn't work. The mechanic is old and well documented on paywall design. Whether you can measure it at your volume is a separate question, and it's the one that decides if the test is worth running.

Does the order of the plans change anything on its own?

It moved onboarding paywall conversion from 3.14% to 3.77% over about 7,800 users per group. The plans, the prices, and the copy were identical in both variants.

The same three plans at the same prices in two different orders, with the year first in the winning variant.Moving the $59.99 year to the top of the list took onboarding paywall conversion from 3.14% to 3.77%

The control led with Yearly Business at the top and put the "Most Popular" badge on the $59.99 year in the middle slot. B moved the year to the top with its badge and dropped Business underneath it. Nothing else about the screen moved.

This is the cheapest test in the deck by a distance. There's no pricing to model, no design work, and if it loses, you've lost a week on the placement that already sees more traffic than anything else in the app. Most teams go straight to redesigning the paywall without ever testing the order of what's on it.

Then the same team moved a mechanic that had just won one screen over and got nothing back.

Why does a winning mechanic stop working when you move it?

Because the win belonged to the position, not to the mechanic.

Glam AI took the decoy structure that had just lifted yearly conversion and put it on the inner paywalls, the ones a user hits after they're already in the app. Effect: zero. The three-plan layout that worked on someone's first look at pricing did nothing for someone who'd already seen the prices once, which is a fair description of what a decoy does. It gives you a reference point. If you've already got one, it has nothing to give you.

The second case is the social proof video. It won its own test, so the team put it on the new plan picker from section 4. That read 2.54% against 2.43%, which is noise. Same asset, same audience, different screen around it, no effect.

Neither of these is an argument against reusing what works. It's an argument against skipping the retest, which is a much cheaper problem to fix. A proven mechanic entering a new placement, segment or channel is a new experiment with its own control, and Glam AI's July numbers put the base rate on those at roughly four wins in ten. Personalisation hits the same wall: it multiplies a paywall that already converts and does nothing for one that doesn't.

Can you charge more to users who bought expensive phones?

Glam AI tried it and lost conversion without gaining revenue. The hypothesis is one every growth team has had: someone holding an $800+ iPhone can absorb a higher price, so raise it for that segment and keep the margin. They put the year up 30%, from $59.99 to $79.99, and the week from $8.99 to $11.99, across about 8,100 users per group.

Standard Glam AI pricing against a variant charging $79.99 for the year and $11.99 for the week.A 30% rise for owners of $800+ iPhones dropped conversion from 2.55% to 2.09% and left ARPU where it was

Conversion fell from 2.55% to 2.09%. ARPU came out flat, so the extra revenue per buyer covered the buyers they lost and stopped there.

It gets stranger. They reverted the year and kept the expensive week, and weekly sagged while yearly stayed where it was. So they tried an intro offer, the year at $59 against the standing $79, the kind of thing that usually prints money. Out of 2,500 users, not one bought the year. The team went looking for a tracking bug and didn't find one.

Two more results from the same run point the same way. Business dropped from $119.99 to $99, and nothing happened; a $21 cut nobody registered. A 50% coupon on Business took placement revenue per user down 5x.

A price cut can be invisible, and a discount can be expensive, and how the user reads it decides which. Anything loud enough to notice also tells them the original number was made up, and at that point neither number is real to them.

Device price says what someone can afford. It says nothing about what they'll pay. If you want to segment on willingness to pay, use what the user has already told you: the answers they gave in onboarding, or what they've done in the app since.

What happens when you move a web funnel that works into the app?

Glam AI's best web funnel lost four times in a row inside the app. The logic for trying was sound. The quiz-first flow converts on web; the app has more traffic, so port the thing that works and collect.

The quiz-first screen that converts on web lost onboarding completion when it opened the app.The quiz-first screen that converts on web lost onboarding completion when it opened the app

Purchase conversion didn't move. Onboarding completion dropped hard, and the worst screen was the first one, which is the tell. A user who tapped an app ad thirty seconds ago and lands on a survey hasn't agreed to anything yet. On web, that same user has already read a landing page, chosen to click through, and arrived expecting questions. Same funnel, and the person hitting it is in a different frame of mind.

So they added a welcome screen before the quiz. Completion recovered, and purchase conversion fell instead. Two months, four iterations, and the challenger never beat the control once.

This is the section 5 problem with a bigger blast radius. The mechanic was proven, the audience looked identical on paper, and the surrounding context was doing more work than anyone credited. Most broken web-to-app funnels break at exactly this seam, where a flow designed for one traffic temperature meets another.

The salvage is that two months of losing cost Glam AI two months of one person's attention and no revenue, because they never shipped any of it. That's a different outcome from finding out in production.

How do you run 23 tests a month when 14 of them lose?

You stop treating a loss as a failed test and start pricing it. Nine wins out of 23 is Glam AI's July, and Artem reads six in ten failing as the normal baseline rather than a sign anything's broken. So the useful question is how to make each loss cheap enough that a 39% hit rate still compounds.

Five questions go on every test before it launches:

  • Why should this work?
  • Which single metric makes the decision?
  • What are the guardrails: churn, refunds, reviews, support?
  • What's the minimum sample and coverage for the effect to be visible?
  • What do we do if it loses, and what's the next iteration?

Most teams skip the last one, and it's what turned the web funnel story into four iterations instead of one abandoned experiment. A test with no planned next move is a test you'll argue about for a week after the numbers land.

None of this runs at 23 a month if every variant needs a release. That's the constraint Adapty's Flow & Paywall Builder removes: onboarding and paywalls built as one flow and published from the dashboard, tests running across the whole journey rather than a single screen, and a per-screen funnel you can read to find out where users left. The web funnel test found its problem on the first screen. That's only findable if the funnel reports screen by screen.

Where should you start?

Glam AI's July looks like a lot of losing. Fourteen tests that went nowhere, a web funnel that lost four times over two months, a price experiment that ended with zero purchases from 2,500 users. It works because none of it reached production, and because every test closed on the same metric.

If your testing cadence is limited by release cycles, that's the part worth fixing first. The Flow & Paywall Builder lets you build onboarding and paywalls as one flow, ship changes without a release, and read the funnel screen by screen. Book a demo, and we'll look at yours with you.

Watch the full session with Artem for the paywall screenshots and the numbers behind each test.

FAQ

Related articles

See how Adaptycan grow your app revenue