The XL Score
The XL Score is a 0–100 rating of how well an app uses every store surface it controls. It scores seven weighted pillars — metadata and keywords (25 points), creative assets (20), localization (15), live content (15), store experimentation (10), reputation operations (10) and momentum (5). Each pillar mixes adoption (are you using the surface at all, 40%) with execution (how well you use it, scored as a percentile against a peer cohort of the same platform, category and download band, 60%). Pillars that cannot be measured are excluded and their weight redistributed rather than scored as zero, and every score is published with its coverage and version stamp.
Key takeaways
- One number out of 100, built from seven weighted pillars, comparable only within the stated score version.
- Adoption is 40% and execution is 60% of every pillar — using a surface badly scores lower than using it well.
- Execution is a percentile against peers on the same platform, category and download band, not an absolute grade.
- Unmeasurable pillars are excluded and their weight redistributed, capped so no single pillar exceeds 40% of the composite.
- Below four measured pillars or 60% of measurable weight, the headline number is withheld and a partial score is shown instead.
The seven pillars and their weights
| Pillar | Public tier | Connected tier | What it measures |
|---|---|---|---|
| Metadata & Keywords | 25 | 24 | Field utilisation, keyword breadth and description structure across every indexed text surface |
| Creative Assets | 20 | 19 | Screenshot depth, device coverage, preview video and first-impression clarity |
| Localization | 15 | 14 | Storefronts with genuinely localized metadata, not English copied everywhere |
| Live Content | 15 | 14 | In-app events and Play promotional content: how many run, how often, how widely |
| Store Experimentation | 10 | 15 | Custom product pages, product page optimization tests and store listing experiments |
| Reputation Ops | 10 | 9 | Rating level and velocity, developer reply coverage and recency |
| Momentum & Hygiene | 5 | 5 | Release cadence, release-note quality and how fresh the listing is |
Weights in v1.0 are expert priors, not fitted coefficients. They encode what moves organic performance in the great majority of categories, and they will be recalibrated per category cluster against forward organic outcomes once the scored corpus is large enough to fit them honestly. That is why the version is stamped on every report: comparing a v1.0 score with a later version is only fair if both versions are stated.
Experimentation is the one pillar that gains weight when an account is connected. Product page optimization tests are close to invisible from outside the store, so on the public tier the pillar is capped at 55 out of 100 rather than guessed at.
Adoption versus execution
Every pillar subscore is 40% adoption plus 60% execution. Adoption is a stepped measure of whether you use the surface at all: do you run in-app events, do you have screenshots on every device family, do you localize beyond your home market. Execution asks how well, and is expressed as a percentile against a peer cohort rather than an absolute standard.
The split exists because most ASO scores reward presence. Ten screenshots that all show the same settings screen score the same as ten that tell a story, and the resulting number is useless for prioritisation. Splitting adoption from execution means a listing that ticks every box but does each one poorly lands in the middle of the range, which is where it belongs.
Coverage, caps and when the score is withheld
- Missing data is never scored as zero. A pillar the engine cannot see is marked unmeasured and excluded from the composite.
- Excluded weight is redistributed over the pillars that were measurable, so the composite stays out of 100.
- No pillar may take more than 40% of the total weight after redistribution. Surplus beyond that cap is dropped, which deflates the composite rather than letting one pillar define it.
- A score built on fewer than four measured pillars, or on less than 60% of measurable weight, is not published as a headline number — the UI shows a partial score instead.
- Coverage travels with the score everywhere it appears: 'Scored on 5 of 7 pillars (82% of weight)' is part of the result, not a footnote.
Penalties sit outside the pillar maths. Keyword stuffing, a listing that has gone stale and patterns consistent with gaming the store each deduct composite points and are shown as named flags with the deduction stated, so a low score can always be traced to a specific cause.
Reading the grade
| Grade | Range | What it means |
|---|---|---|
| Elite | 80–100 | Every controllable surface is in use and executed above the cohort median |
| Strong | 60–79 | Fundamentals are solid; the gap is usually experimentation or localization depth |
| Developing | 35–59 | Core metadata and creative are in place but under-used, with obvious headroom |
| Neglected | 0–34 | Major surfaces unused — typically no localization, no live content and a default listing |
The audit also reports a connected ceiling: what the composite would be if every currently unmeasurable pillar scored at the app's own average. The gap between the composite and the ceiling is measurement headroom, not a promise of ranking movement.
Getting an XL Score for your app
- 1
Run the free audit
Paste an App Store or Google Play URL into the free ASO audit. The XL Score is computed from the public store listing, with no account required.
- 2
Read coverage first
Check how many pillars were measured before acting on the number. A partial score is a signal about visibility, not about quality.
- 3
Work the biggest wins
The audit ranks pillars by points left on the table, so the top item is always the largest available gain.
- 4
Re-score after shipping
The engine is deterministic, so any change in the score after a metadata release is attributable to that release.
Frequently asked questions
What is a good XL Score?
Anything at or above 60 is Strong: the fundamentals are in place and the remaining gap is usually experimentation or localization depth. Above 80 is Elite and rare, because it requires live content and store testing to be running continuously, not just metadata to be tidy.
How is the XL Score different from other ASO scores?
Three things: it separates adoption from execution rather than rewarding mere presence, it scores execution as a percentile against a like-for-like peer cohort rather than an absolute rubric, and it publishes coverage and a version stamp with every number so the score can be audited and reproduced.
Why is my score shown as partial?
Because fewer than four pillars were measurable, or the measurable pillars accounted for less than 60% of the weight. Rather than publish a headline number built on thin data, the engine withholds it and shows what it could measure.
Does a higher XL Score guarantee more installs?
No. The score measures how completely and how well you use the store surfaces you control, which is the part of ASO you can act on. Rankings also depend on conversion, retention and competitive pressure that no listing audit can observe.
Will the score change when the methodology is updated?
Yes, and that is why every report carries a version. Weights will be refit per category cluster against forward organic outcomes; scores are only comparable within the same stated version.
appXL Research
App Store Optimization Research Team
The appXL research team analyzes App Store and Google Play ranking data across the apps our agent manages, and publishes what it finds.