A catalogue must earn its credit
A game catalogue is a decision system: which games appear, where, to whom, and under what conditions. Game count, impressions, and gross revenue can all climb while players are steered toward weaker choices, relevant titles stay hidden, or demand piles up on a handful of providers.
A defensible framework follows eligible exposure through meaningful launch, post-launch play, return, and net value. The job is to separate catalogue effects from traffic, promotions, availability, and short-term game outcomes. Treat patterns as associations until the design supports causality: more launches after a ranking change prove nothing if that day's traffic mix also shifted.
Exposure is allocation, not demand
The framework begins where players do, with exposure. Impression Share measures storefront attention:
Impression Share = Item impressions / All eligible catalogue impressions
Eligible excludes titles unavailable for a player's geography, device, currency, or regulatory state. Count an impression only when a tile enters the viewport for a defined minimum period; anything rendered below the fold is allocated inventory, not observed attention.
High share means opportunity, not demand. A hero slot can hold high share and still convert few launches, while a low-share game may pull strongly through search, filters, or a provider page. So the metric belongs in a distribution audit across providers, categories, and game ages, not in a demand or ranking-success claim.
Launch rate tests the storefront promise
Launch Rate asks whether exposure produced an intentional start, measured at user level within a fixed attribution window:
Launch Rate = Users who launch after a valid impression / Users with a valid impression
Exclude duplicate clicks, failed launches, and launches outside the window. Match the denominator to the surface under review; comparing a homepage carousel with search only holds when eligibility, rank, and player intent are comparable.
Unlike raw launches, Launch Rate accounts for distribution, but repeating one popular game across placements will inflate it. Track unique exposed users, unique launched titles per user, and repeat-exposure frequency alongside it.
Numbers make the trap concrete. A user-level example needs unique users in both counts: 900 launching users out of 10,000 exposed users gives a 9% launch rate. After a placement change, 1,000 launches from 15,000 impressions yield 6.7%. Starts rose through wider distribution, yet each opportunity persuaded less. Both facts matter, and a single metric hides one of them.
Depth must reflect meaningful play
A launch is only the doorway to play. A launch may be accidental, fail during loading, or end after a single round. Session depth has to represent a real game experience, not time spent watching a loading spinner.
For a casino catalogue, several measures do that work together. Successful game load rate (successful game starts divided by launch attempts) separates technical failure from disinterest. Meaningful play rate counts launched sessions that reach a pre-defined threshold of settled rounds or staking activity, and rounds per meaningful session divides settled rounds by the sessions clearing that threshold. Return-to-catalogue rate captures players who come back to browse after a short or failed session. And post-launch NGR contribution reports net gaming revenue from the launch cohort after direct costs, over a stated accounting window.
Session duration is ambiguous on its own: it can reflect engagement, an inactive tab, a slow exit, or a game left open. More rounds can signal genuine interest or just different mechanics. Compare depth within game type, device, market, and lifecycle stage before you judge a ranker on it. The strongest signal is not maximum depth on one title but improved meaningful launch and return without worse load reliability, narrower player choice, or lower net value.
Provider concentration exposes hidden risk
A broad catalogue can carry narrow effective demand. Where that narrow demand pools is the next risk to size. Measure concentration in displayed impressions, launches, and value; provider share at each stage reveals whether attention is broad while demand funnels into a few providers, or whether a few control the whole journey.
Top 5 Provider Share = Launches from five largest providers / All catalogue launches
For a fuller view:
HHI = sum of squared provider shares
Hold the unit constant within a comparison: provider launch, Bets Sum, or NGR share. A higher HHI after a ranking release is not automatically bad; it can reflect real relevance for a defined audience. The risk shows when concentration climbs while discovery falls, when providers lose viable routes to exposure, or when the catalogue leans on a few titles and commercial relationships. Segment before concluding: new players often need recognisable titles, established players search differently, and an aggregate share can bury lost choice inside one cohort.
Freshness follows a decay curve
Concentration often rides on novelty. New games draw attention because they are new, heavily placed, or aimed at an existing need. Track age from the moment a game became eligible on a surface, not from global release. Age-bucketed Impression Share, Launch Rate, meaningful play, and post-launch return separate week one from the following weeks and from later life, instead of treating every title as equally mature.
A post-launch decline is normal for a novelty slot. The question is whether performance settles above comparable games once premium placement ends. Compare within genre, provider, position band, country, device, and entry channel: a fresh live game and an older slot never share demand conditions. Applied this way, the decay curve stops a merchandising model from taking credit for gains that belong to recently released titles. Hold the game-age mix constant, or report how it shifted.
RTP noise can impersonate retention
RTP = Player wins / Bets Sum
Even a clean freshness read can be fooled by luck. Observed RTP swings hard, especially over short windows or small cohorts. A run of favourable outcomes lifts near-term return; an unlucky run shortens it. Neither proves that placement caused retention or that a game holds a stable retention advantage.
Promote titles right after a brief high return or a spike in engagement and you build a noise feedback loop: the temporary winners collect more exposure, more bets, and more measured revenue. High-volume players bend the averages further through their outsized shares of bets and outcomes.
Judge retention from a fixed cohort start, with the event defined plainly: a return session with meaningful play inside seven or 28 days. Read the results with and without outcome-sensitive groups, gather enough settled activity to tame variance, and treat theoretical parameters as context rather than proof of experience. Before you touch rank logic or an offer, run through questions that expose real experimentation skill to pressure-test sample size, assignment, guardrails, and whether a result is practical or merely random.
Short-window GGR deserves the same caution. Since GGR = Bets - Wins, an RTP dip can lift GGR with no gain in catalogue quality. NGR sits closer to decision value, landing after bonuses, provider fees, payment costs, taxes, fraud, chargebacks, and the relevant commissions.
Attribute catalogue effects before scaling
With the metrics cleaned of noise, the question becomes what actually caused a move, and attribution starts with decomposition:
Launches = Eligible visitors × Valid impression reach × Launch Rate
When launches rise, find the component that moved. More eligible visitors points to traffic or availability; higher reach points to placement coverage; a higher Launch Rate points to relevance, creative, rank, or intent. Post-launch depth then tests whether those extra starts turned into meaningful play.
Randomise where you can, and run the split test in a fixed order:
- Split eligible users or sessions between treatments and log the exposure each side actually receives.
- Pick a primary outcome before anything launches, such as meaningful launches per eligible user.
- Keep acquisition traffic, offers, and eligibility identical across treatment and control.
- Guard against load failures, provider concentration, narrowed choice, and NGR as the test runs, not after it ends. Where randomisation is impossible, say so and lean on weaker designs on purpose: cohorts matched by source, market, device, and lifecycle; difference-in-differences against an unaffected surface; and pre-change trend lines. A before-and-after chart cannot pull a release apart from campaign traffic, seasonality, a provider outage, a payment failure, or a bonus change.
Every ranker is capped by data quality. Missing genres, duplicate game identities, stale availability, inconsistent provider names, and thin metadata all erode relevance, which is why personalisation starts with catalogue operations. Those controls come before any smarter ranking.
Instrument the decision path
None of that attribution is possible without the right events, so event data has to reconstruct the journey. An impression should record surface, module, rank, viewport status, game ID, provider, category, algorithm version, filter state, device, market, and the eligibility outcome. Launches need timestamps and a success-or-failure flag; gameplay needs the successful load and the meaningful-play threshold.
Keep catalogue snapshots. Historical analysis depends on the provider label, category, availability, thumbnail, and rank rule the player actually saw. Without them, analysts end up comparing yesterday's taxonomy against last month's exposure.
Audit the denominators before reading any movement. Bot traffic, duplicate device identities, missing launch callbacks, unavailable titles counted as impressions, and preloaded-carousel exposures all distort the base. These defects tend to look like conversion problems right up until someone inspects the event chain.
Build a scorecard around decisions
A scorecard exists to answer a decision, so a weekly review pairs outcomes with diagnostics and guardrails rather than listing everything measurable. The primary evidence is meaningful launches per eligible visitor: did the catalogue actually move players into real play. Launch Rate by surface and rank band separates relevance from sheer exposure volume, while successful game load rate flags technical friction masquerading as weak demand. Session depth read within comparable game types tests post-launch quality without leaning on duration. Provider HHI and top-provider share watch the concentration that merchandising creates; the freshness curve by game cohort keeps novelty apart from durable value; cohort retention after meaningful play checks whether the experience earns a return; and NGR per eligible visitor ties all of it back to margin after direct costs.
Read every one of these by acquisition channel before claiming catalogue impact. Paid, affiliate, organic, CRM, and direct traffic differ in intent, device mix, and game familiarity, so a rate that improved only because high-intent traffic gained share has earned the catalogue no credit.
The next catalogue review starts with falsification
Open with a falsifiable claim: "the new lobby order increases meaningful launches per eligible visitor for new mobile players without raising provider concentration or reducing seven-day retention." Fix the population, exposure, conversion window, and guardrails before the results land.
Then inspect the chain in order: eligibility, valid impressions, launches, successful loads, meaningful play, cohort return, NGR. When one metric moves, test traffic mix, availability, freshness, RTP variance, and tracking before you celebrate. A metric survives scrutiny when its denominator holds steady, its mechanism is visible, its segment results stay coherent, and its claimed effect is still standing after the plausible alternatives have been thrown at it.