← All posts

The Population Is the Parameter

· reSlot Research · 15 min read · Kairos

  • kairos
  • simulation
  • dynamic-rtp
  • player-modeling
  • benchmark

Abstract

Operators tune dynamic RTP control systems by editing tables, so it is natural to assume the tables determine what the system does. We test that assumption in a synthetic environment built from a production seven-module control workbook and a behavioural player generator with ten mechanisms that can be switched off one at a time: loss aversion, peak-end evaluation, big-win memory, stop rules, loss chasing, the house-money effect, multi-session days, top-ups, bonus buying and per-player heterogeneity. Across 802 runs the ranking of seven candidate configurations is stable, so the tables do decide which configuration is better. Almost nothing else is theirs. Between 96% and 98% of the variance in day-two retention and, once bet sizing is behavioural, 72% to 88% of the variance in net control cost belongs to the population. The same shipped tables produce a realized RTP of 100.5% under disciplined players and 89.8% under chasers, because the big-win suppression thresholds are written in currency and chasers bet more. A mechanism ablation shows stop rules, loss aversion, multi-session play and chasing carry the retention and cost effects; peak-end, big-win memory, top-ups and bonus buying move outcomes by under two points. Sixty populations drawn from a broad prior give day-two retention from 15% to 48% and house edge from zero to 12.5 points on one unchanged workbook. Finally, the player's next action is 82% "same bet again" and its entropy falls only from 1.01 to 0.95 bits under every immediate context we can bucket, so the predictable part of behaviour lives in latent traits and long-range state, which is the case for a sequence model with a player embedding and against scoring such a model by average likelihood. This revises the cost half of our earlier claim that tables decide cost and players decide retention. Everything here is synthetic; the environment and results are published for replication.

Contents

I. The question

A dynamic RTP control system is configured, not programmed. Somebody opens a workbook and sets compensation bands, trigger probabilities, template boundaries, suppression windows and reroll thresholds, and the operating question is always the same: what happens if this cell changes. In our earlier note we answered it with a first simulation study on a provider's real workbook and summarized the result as the tables decide cost, the players decide retention.1 That study used a synthetic player with a dozen parameters. This one asks whether the conclusion survives a player who behaves more like the ones in the gambling-behaviour literature, and what the answer implies for training a model of those players.

The short version: the retention half survives with room to spare, the cost half does not, and the reason it does not is more useful than the original claim.

II. The environment

The environment has two halves, an engine and a generator, and they compose the way the Kairos design says a player world model and an outcome engine should.2

Engine. The control state machine is the client's, loaded unchanged from the production workbook (fourteen tabs, configuration version v0602+V4): multi-template switching on a 1,000-round rolling window, a newbie template schedule, cumulative compensation over the first 200 rounds and daily compensation after, loss carry-over, four player tags, high-RTP suppression with per-band daily counters, big-win reroll on absolute payout bands, and a loss-triggered kill switch. Bonus is disabled, as in the phase-one dry run, because the workbook's bonus trigger and base RTP are mutually inconsistent and the client has not yet resolved which is right. The engine reports every round with its attribution: template, tag, control flag, pre-suppression amount.

Generator. A population is a FamilyV2: a set of behavioural means, a heterogeneity scale, and ten switches. Each player draws their own parameters as log-normal jitter around the family means. The mechanisms, with the literature each is borrowed from:

MechanismWhat it doesSource
Loss aversionReturn probability responds to the day's net outcome asymmetrically, losses weighted λ times gainsKahneman and Tversky, prospect theory3
Peak-endReturn probability keys on the session's largest win and the net of its last ten roundsFredrickson and Kahneman4
Big-win memoryA win of 50× or more lifts return probability with a three-day e-folding decayReinforcement-schedule literature, treated as a prior
Stop rulesPer-player stop-loss and stop-win limits relative to session-opening balance, honoured with a per-player adherenceVoluntary limit-setting studies5
ChasingAfter three straight losses, step the bet up once; after a bust, start the next day one denomination lowerLesieur, The Chase6
House moneyAfter a win of 10× or more, step the bet upThaler and Johnson7
Multi-sessionPoisson sessions per active day, evening-peaked start times, weekend liftOperational pattern
Top-upOn bust, replenish the balance with some probabilityOperational pattern
Buy bonusAfter a dry spell, buy the bonus at 50× bet, outside the control system as in the productProduct behaviour
HeterogeneityLog-normal jitter on every behavioural parameter per playerStandard mixed-population assumption

With every switch off the generator collapses to the earlier study's player, and the environment reproduces that study's headline numbers: day-two retention 50.4% against 45.0%, realized RTP 104.6% against 104.0%, net control cost 8.9% of turnover against 8.1%. The differences reported below are therefore attributable to the mechanisms and not to a change of harness.

Every run is 28 days at 150 registrations per day, about 4,200 players, one seed. Means and spreads are across seeds. Study script, generator and results are in the Kairos repository.8

III. Family sensitivity, revisited

Seven configurations from the phase-one audit (the shipped baseline and six recalibration candidates, including C2, which adds the compensation triggers the workbook is missing for rounds 21 to 200, and R2, which repairs the pasted daily table and upgrades the newbie template) were run under seven hand-specified populations, ten seeds each.

PopulationBaseline: cost / turnoverBaseline: RTPBaseline: D2R2: cost / turnoverR2: D2
G0 default−4.7%91.9%37.1% ± 0.6−6.7%38.1%
G1 price-insensitive−3.3%92.5%38.1% ± 0.7−5.8%39.4%
G2 loss-averse−4.0%92.0%32.4% ± 0.6−6.2%32.8%
G3 chasers−7.8%89.8%32.2% ± 0.7−8.5%33.5%
G4 disciplined+5.1%100.5%44.9% ± 0.7−1.1%45.8%
G5 loyal grinders−5.8%91.4%35.3% ± 0.9−7.0%35.9%
G6 bonus buyers−3.2%93.0%36.6% ± 0.6−5.8%37.8%

Negative cost means suppression took back more than compensation paid out. Two things are visible at once.

The rankings hold. R2 is the cheapest configuration in all seven populations. C2 has the highest D2 in six of seven (the disciplined population prefers C1 by a fraction of a point). Whoever is playing, the audit's recommendation is the same recommendation.

The levels do not hold, and now that includes cost. Under the earlier player, the shipped workbook lost money in every population we tried, at a realized RTP near 104%. Under the richer player the same workbook makes an eight-point house edge for chasers and loses half a point for disciplined players. The sign of the ledger flipped, and the tables did not change.

The variance decomposition puts numbers on it. For each configuration, the variance of each metric across the seventy runs is split into a between-population component and a within-population (seed) component:

MetricShare between populations, baselineShare between populations, R2Earlier study (baseline)
D20.970.960.98
D70.630.48—
Newbie completion0.870.880.96
Net control cost / turnover0.880.720.67
Turnover0.940.97—

Retention is unchanged at 96% to 98% population. Cost moved from two-thirds population to seven-eighths.

IV. Why the ledger belongs to the population

The mechanism ablation answers it. Starting from the default population, each switch was turned off alone, ten seeds each, under the shipped baseline:

Switched offD2Cost / turnoverRTPRoundsTurnover
Nothing (all on)37.1% ± 0.9−4.3%92.7%239k5.8M
Loss aversion43.7%−4.1%92.7%244k5.8M
Peak-end38.7%−3.0%92.9%241k5.8M
Big-win memory36.3%−3.8%92.5%238k5.6M
Stop rules26.0%−7.2%90.2%319k10.5M
Chasing43.3%+7.5%103.6%512k8.5M
House money37.7%−3.0%93.7%253k6.1M
Multi-session42.4%−1.9%95.1%196k3.9M
Top-up36.5%−3.5%92.5%236k5.6M
Buy bonus36.5%−3.4%92.6%240k5.8M
Heterogeneity35.2%−3.6%92.4%232k5.4M
Everything49.4%+8.9%104.6%709k5.7M

Four mechanisms carry the environment. Stop rules, when removed, halve day-two retention and double turnover: players without limits bust more, lose more and, through loss aversion, come back less. Loss aversion alone is worth 6.6 points of D2. Multi-session days cost five points, because more sessions on a losing product means more losing days. And chasing is the one that moves the ledger: switch it off and realized RTP jumps eleven points, from 92.7% to 103.6%, and net control cost swings from −4.3% to +7.5% of turnover.

The reason is legible in the workbook. Big-win suppression is configured on absolute payout bands in platform currency, at 5,000, 10,000 and 50,000. A player who steps the bet up after a losing streak, or after a big win, bets more, so the same multiplier lands in a higher band and is discarded or rerolled with higher probability. The control system is neutral about bet size in its compensation modules, which are written as RTP ratios, and progressive in its suppression modules, which are written in currency. Chasers therefore pay for everyone else's compensation, and a population with more chasers turns a money-losing workbook into a money-making one without a single cell changing.

The provider's own balance rule, that compensation and suppression should roughly offset (A+B against C+D in their notation), is thus not a property of the tables. It is a property of who is betting how much, and it will drift with acquisition channel, promotion calendar and the denomination mix, none of which appear in the workbook.

The remaining six mechanisms, including several with the best literature behind them, move outcomes by under two points each in this environment. That is not a finding about players; it is a finding about which mechanisms a model must get right first, and which can be learned later.

V. What the prior allows

Sixty populations were drawn from broad uniform priors over every behavioural parameter (retention base 0.20 to 0.50, loss aversion 1 to 4, chase probability 0 to 0.7, stop-rule adherence 0.1 to 0.9, and so on) and each run once under the shipped baseline.

Metricp10meanp90minmax
D221.5%30.3%39.9%15.1%48.4%
D71.1%4.5%9.3%0.6%11.9%
Realized RTP89.1%92.5%97.4%87.5%100.3%
Net control cost / turnover−8.1%−4.3%+0.7%−10.0%+3.5%
Rounds per 28 days183k253k342k167k605k

One workbook, sixty plausible populations, a house edge anywhere between zero and twelve and a half points. Any forecast of this configuration's GGR that does not come with a forecast of the population is a forecast of the prior.

VI. How predictable is the next action

The Kairos design proposes to learn the player's next action, coarse then fine, from the round history. The environment lets us ask how much of that action is predictable from immediate context alone, before any model is trained. For the default population and three others, three seeds each, every round was labelled with what the player did next: same bet, raise, lower, buy bonus, end the session, end the day.

DefaultPrice-insensitiveLoss-averseChasers
Share "same bet"82.0%81.9%82.2%80.9%
Share raise / lower7.9% / 4.4%7.8% / 4.4%7.9% / 4.4%9.0% / 4.3%
Share end session / end day2.2% / 3.5%2.2% / 3.7%2.1% / 3.4%2.0% / 3.8%
H(action)1.012 bits1.0201.0061.049
H(action | outcome bucket)0.9920.9990.9851.027
H(action | outcome, streak)0.9670.9760.9620.979
H(action | outcome, streak, position, session net)0.9490.9570.9440.957

Two readings. First, the action stream is dominated by continuation: four fifths of tokens are "same bet". A model that predicts "same" always scores an average likelihood that looks respectable, and the events an operator cares about, ending the session and ending the day, are 6% of tokens. Any evaluation of a player model has to be scored on those rare tokens, as hazard AUC or calibration on the exit events, not on mean negative log-likelihood. The Kairos evaluation protocol says this; the environment shows why.

Second, conditioning on everything we can bucket from the immediate context removes about six percent of the entropy. The generator's actions are not random, they follow deterministic rules, but the rules key on per-player latent traits (this player's stop-loss, this player's chase probability) and on state relative to session-opening balance that a few buckets do not capture. Those are exactly the two things a per-round tabular model cannot see and a sequence model with a player embedding can: the trait is inferred from the player's history, the state is carried in the context. Section VI is, in that sense, the environment agreeing with the architecture.

VII. What changes in the Kairos design

Three things, one of them a correction.

Normalization keeps the currency. The Kairos paper proposed normalizing bet by denomination and balance by session-opening balance so that every game and currency an operator runs can share one model. Section IV shows that absolute currency amounts are load-bearing: suppression bands are written in them and players respond to them. The token therefore carries both the normalized bet and a log-bucketed absolute bet in a common reference currency, and the balance token does the same. We have amended that section of the paper.2

The synthetic corpus randomizes the mechanisms that matter. Stage A pre-training draws populations from a prior. Section IV says which dimensions of that prior must be wide: stop-rule parameters and adherence, loss aversion, session count, chase and house-money probabilities, bet-size dynamics. Peak-end weights, big-win memory, top-up and buy-bonus rates can start narrow.

Rare tokens are the metric. Fidelity against the ledger and hazard AUC on session and day exits come first; mean likelihood is reported but not optimized for.

VIII. Limitations

Every player here is synthetic. The mechanisms are borrowed from the literature, their parameter values are priors, and the environment's outputs describe how the control workbook behaves under those priors, not how any real cohort behaves. The consistency check in Section II shows the harness is stable across generator versions; it does not show either generator is right. Bonus is off because the workbook's bonus configuration is internally inconsistent, so the interaction between bonus RTP and control is not exercised. And the sixty-population spread in Section V is a statement about the width of our prior as much as about the workbook. All of that is what the real ledger is for, and it is why synthetic data enters pre-training and nowhere else.

IX. Conclusion

The earlier claim was half right. Tables decide the ranking of configurations, and an audit that reads the workbook, prices each module and orders the candidates is still the cheapest useful thing a provider can do. But the tables decide neither the level of retention nor, once players size their bets the way players do, the level or sign of the control ledger. Both belong to the population, and the population is not in the workbook.

That is why the missing half of a payout engine is a model of players and not a better spreadsheet. The environment described here is the pre-training corpus for that model and the test bed for its evaluation protocol, and it comes with a built-in warning: a model that scores well against it has learned this generator, not any real player. The real players arrive with the ledger.

Footnotes

  1. What the Tables Decide, and What the Players Decide, on this blog. Its Section III is the claim revised here; its cost figure of 0.67 between-population share appears in the comparison column of Section III above. ↩

  2. Kairos: A Foundation Model for the Language of Slot Play, on this blog. Section IV there describes the control-action-outcome factorization; its normalization paragraph was amended on 2026-08-28 following this study. ↩ ↩2

  3. Daniel Kahneman and Amos Tversky, "Prospect Theory: An Analysis of Decision under Risk," Econometrica 47(2), 1979. ↩

  4. Barbara Fredrickson and Daniel Kahneman, "Duration Neglect in Retrospective Evaluations of Affective Episodes," Journal of Personality and Social Psychology 65(1), 1993. ↩

  5. Michael Auer and Mark Griffiths, "Voluntary Limit Setting and Player Choice in Most Intense Online Gamblers," Journal of Gambling Studies 29, 2013, is the reference point for adherence; the stop-loss and stop-win fractions here are priors, not their estimates. ↩

  6. Henry Lesieur, The Chase: Career of the Compulsive Gambler, 1984. The chase here is a per-streak bet step-up, a deliberately mild operationalization. ↩

  7. Richard Thaler and Eric Johnson, "Gambling with the House Money and Trying to Break Even," Management Science 36(6), 1990. ↩

  8. docs/kairos/experiments/players_v2.py, run_sim_study_v2.py and results_v2.json in the reslot-front repository. Seeds 5000–5009 per population for Section III, 6000–6009 per ablation for Section IV, 7000–7002 per population for Section VI, 8000–8059 with family seeds 900–959 for Section V. 802 runs; roughly four minutes on fourteen cores. ↩

Try the loop yourself

The same maker/checker loop runs in our simulators — set a target, run the search, and read the measured results.