I. The question
A dynamic RTP control system is configured, not programmed. Somebody opens a workbook and sets compensation bands, trigger probabilities, template boundaries, suppression windows and reroll thresholds, and the operating question is always the same: what happens if this cell changes. In our earlier note we answered it with a first simulation study on a provider's real workbook and summarized the result as the tables decide cost, the players decide retention.1 That study used a synthetic player with a dozen parameters. This one asks whether the conclusion survives a player who behaves more like the ones in the gambling-behaviour literature, and what the answer implies for training a model of those players.
The short version: the retention half survives with room to spare, the cost half does not, and the reason it does not is more useful than the original claim.
II. The environment
The environment has two halves, an engine and a generator, and they compose the way the Kairos design says a player world model and an outcome engine should.2
Engine. The control state machine is the client's, loaded unchanged from the production workbook (fourteen tabs, configuration version v0602+V4): multi-template switching on a 1,000-round rolling window, a newbie template schedule, cumulative compensation over the first 200 rounds and daily compensation after, loss carry-over, four player tags, high-RTP suppression with per-band daily counters, big-win reroll on absolute payout bands, and a loss-triggered kill switch. Bonus is disabled, as in the phase-one dry run, because the workbook's bonus trigger and base RTP are mutually inconsistent and the client has not yet resolved which is right. The engine reports every round with its attribution: template, tag, control flag, pre-suppression amount.
Generator. A population is a FamilyV2: a set of behavioural means, a heterogeneity scale, and ten switches. Each player draws their own parameters as log-normal jitter around the family means. The mechanisms, with the literature each is borrowed from:
| Mechanism | What it does | Source |
|---|---|---|
| Loss aversion | Return probability responds to the day's net outcome asymmetrically, losses weighted λ times gains | Kahneman and Tversky, prospect theory3 |
| Peak-end | Return probability keys on the session's largest win and the net of its last ten rounds | Fredrickson and Kahneman4 |
| Big-win memory | A win of 50× or more lifts return probability with a three-day e-folding decay | Reinforcement-schedule literature, treated as a prior |
| Stop rules | Per-player stop-loss and stop-win limits relative to session-opening balance, honoured with a per-player adherence | Voluntary limit-setting studies5 |
| Chasing | After three straight losses, step the bet up once; after a bust, start the next day one denomination lower | Lesieur, The Chase6 |
| House money | After a win of 10× or more, step the bet up | Thaler and Johnson7 |
| Multi-session | Poisson sessions per active day, evening-peaked start times, weekend lift | Operational pattern |
| Top-up | On bust, replenish the balance with some probability | Operational pattern |
| Buy bonus | After a dry spell, buy the bonus at 50× bet, outside the control system as in the product | Product behaviour |
| Heterogeneity | Log-normal jitter on every behavioural parameter per player | Standard mixed-population assumption |
With every switch off the generator collapses to the earlier study's player, and the environment reproduces that study's headline numbers: day-two retention 50.4% against 45.0%, realized RTP 104.6% against 104.0%, net control cost 8.9% of turnover against 8.1%. The differences reported below are therefore attributable to the mechanisms and not to a change of harness.
Every run is 28 days at 150 registrations per day, about 4,200 players, one seed. Means and spreads are across seeds. Study script, generator and results are in the Kairos repository.8
III. Family sensitivity, revisited
Seven configurations from the phase-one audit (the shipped baseline and six recalibration candidates, including C2, which adds the compensation triggers the workbook is missing for rounds 21 to 200, and R2, which repairs the pasted daily table and upgrades the newbie template) were run under seven hand-specified populations, ten seeds each.
| Population | Baseline: cost / turnover | Baseline: RTP | Baseline: D2 | R2: cost / turnover | R2: D2 |
|---|---|---|---|---|---|
| G0 default | −4.7% | 91.9% | 37.1% ± 0.6 | −6.7% | 38.1% |
| G1 price-insensitive | −3.3% | 92.5% | 38.1% ± 0.7 | −5.8% | 39.4% |
| G2 loss-averse | −4.0% | 92.0% | 32.4% ± 0.6 | −6.2% | 32.8% |
| G3 chasers | −7.8% | 89.8% | 32.2% ± 0.7 | −8.5% | 33.5% |
| G4 disciplined | +5.1% | 100.5% | 44.9% ± 0.7 | −1.1% | 45.8% |
| G5 loyal grinders | −5.8% | 91.4% | 35.3% ± 0.9 | −7.0% | 35.9% |
| G6 bonus buyers | −3.2% | 93.0% | 36.6% ± 0.6 | −5.8% | 37.8% |
Negative cost means suppression took back more than compensation paid out. Two things are visible at once.
The rankings hold. R2 is the cheapest configuration in all seven populations. C2 has the highest D2 in six of seven (the disciplined population prefers C1 by a fraction of a point). Whoever is playing, the audit's recommendation is the same recommendation.
The levels do not hold, and now that includes cost. Under the earlier player, the shipped workbook lost money in every population we tried, at a realized RTP near 104%. Under the richer player the same workbook makes an eight-point house edge for chasers and loses half a point for disciplined players. The sign of the ledger flipped, and the tables did not change.
The variance decomposition puts numbers on it. For each configuration, the variance of each metric across the seventy runs is split into a between-population component and a within-population (seed) component:
| Metric | Share between populations, baseline | Share between populations, R2 | Earlier study (baseline) |
|---|---|---|---|
| D2 | 0.97 | 0.96 | 0.98 |
| D7 | 0.63 | 0.48 | — |
| Newbie completion | 0.87 | 0.88 | 0.96 |
| Net control cost / turnover | 0.88 | 0.72 | 0.67 |
| Turnover | 0.94 | 0.97 | — |
Retention is unchanged at 96% to 98% population. Cost moved from two-thirds population to seven-eighths.
IV. Why the ledger belongs to the population
The mechanism ablation answers it. Starting from the default population, each switch was turned off alone, ten seeds each, under the shipped baseline:
| Switched off | D2 | Cost / turnover | RTP | Rounds | Turnover |
|---|---|---|---|---|---|
| Nothing (all on) | 37.1% ± 0.9 | −4.3% | 92.7% | 239k | 5.8M |
| Loss aversion | 43.7% | −4.1% | 92.7% | 244k | 5.8M |
| Peak-end | 38.7% | −3.0% | 92.9% | 241k | 5.8M |
| Big-win memory | 36.3% | −3.8% | 92.5% | 238k | 5.6M |
| Stop rules | 26.0% | −7.2% | 90.2% | 319k | 10.5M |
| Chasing | 43.3% | +7.5% | 103.6% | 512k | 8.5M |
| House money | 37.7% | −3.0% | 93.7% | 253k | 6.1M |
| Multi-session | 42.4% | −1.9% | 95.1% | 196k | 3.9M |
| Top-up | 36.5% | −3.5% | 92.5% | 236k | 5.6M |
| Buy bonus | 36.5% | −3.4% | 92.6% | 240k | 5.8M |
| Heterogeneity | 35.2% | −3.6% | 92.4% | 232k | 5.4M |
| Everything | 49.4% | +8.9% | 104.6% | 709k | 5.7M |
Four mechanisms carry the environment. Stop rules, when removed, halve day-two retention and double turnover: players without limits bust more, lose more and, through loss aversion, come back less. Loss aversion alone is worth 6.6 points of D2. Multi-session days cost five points, because more sessions on a losing product means more losing days. And chasing is the one that moves the ledger: switch it off and realized RTP jumps eleven points, from 92.7% to 103.6%, and net control cost swings from −4.3% to +7.5% of turnover.
The reason is legible in the workbook. Big-win suppression is configured on absolute payout bands in platform currency, at 5,000, 10,000 and 50,000. A player who steps the bet up after a losing streak, or after a big win, bets more, so the same multiplier lands in a higher band and is discarded or rerolled with higher probability. The control system is neutral about bet size in its compensation modules, which are written as RTP ratios, and progressive in its suppression modules, which are written in currency. Chasers therefore pay for everyone else's compensation, and a population with more chasers turns a money-losing workbook into a money-making one without a single cell changing.
The provider's own balance rule, that compensation and suppression should roughly offset (A+B against C+D in their notation), is thus not a property of the tables. It is a property of who is betting how much, and it will drift with acquisition channel, promotion calendar and the denomination mix, none of which appear in the workbook.
The remaining six mechanisms, including several with the best literature behind them, move outcomes by under two points each in this environment. That is not a finding about players; it is a finding about which mechanisms a model must get right first, and which can be learned later.
V. What the prior allows
Sixty populations were drawn from broad uniform priors over every behavioural parameter (retention base 0.20 to 0.50, loss aversion 1 to 4, chase probability 0 to 0.7, stop-rule adherence 0.1 to 0.9, and so on) and each run once under the shipped baseline.
| Metric | p10 | mean | p90 | min | max |
|---|---|---|---|---|---|
| D2 | 21.5% | 30.3% | 39.9% | 15.1% | 48.4% |
| D7 | 1.1% | 4.5% | 9.3% | 0.6% | 11.9% |
| Realized RTP | 89.1% | 92.5% | 97.4% | 87.5% | 100.3% |
| Net control cost / turnover | −8.1% | −4.3% | +0.7% | −10.0% | +3.5% |
| Rounds per 28 days | 183k | 253k | 342k | 167k | 605k |
One workbook, sixty plausible populations, a house edge anywhere between zero and twelve and a half points. Any forecast of this configuration's GGR that does not come with a forecast of the population is a forecast of the prior.
VI. How predictable is the next action
The Kairos design proposes to learn the player's next action, coarse then fine, from the round history. The environment lets us ask how much of that action is predictable from immediate context alone, before any model is trained. For the default population and three others, three seeds each, every round was labelled with what the player did next: same bet, raise, lower, buy bonus, end the session, end the day.
| Default | Price-insensitive | Loss-averse | Chasers | |
|---|---|---|---|---|
| Share "same bet" | 82.0% | 81.9% | 82.2% | 80.9% |
| Share raise / lower | 7.9% / 4.4% | 7.8% / 4.4% | 7.9% / 4.4% | 9.0% / 4.3% |
| Share end session / end day | 2.2% / 3.5% | 2.2% / 3.7% | 2.1% / 3.4% | 2.0% / 3.8% |
| H(action) | 1.012 bits | 1.020 | 1.006 | 1.049 |
| H(action | outcome bucket) | 0.992 | 0.999 | 0.985 | 1.027 |
| H(action | outcome, streak) | 0.967 | 0.976 | 0.962 | 0.979 |
| H(action | outcome, streak, position, session net) | 0.949 | 0.957 | 0.944 | 0.957 |
Two readings. First, the action stream is dominated by continuation: four fifths of tokens are "same bet". A model that predicts "same" always scores an average likelihood that looks respectable, and the events an operator cares about, ending the session and ending the day, are 6% of tokens. Any evaluation of a player model has to be scored on those rare tokens, as hazard AUC or calibration on the exit events, not on mean negative log-likelihood. The Kairos evaluation protocol says this; the environment shows why.
Second, conditioning on everything we can bucket from the immediate context removes about six percent of the entropy. The generator's actions are not random, they follow deterministic rules, but the rules key on per-player latent traits (this player's stop-loss, this player's chase probability) and on state relative to session-opening balance that a few buckets do not capture. Those are exactly the two things a per-round tabular model cannot see and a sequence model with a player embedding can: the trait is inferred from the player's history, the state is carried in the context. Section VI is, in that sense, the environment agreeing with the architecture.
VII. What changes in the Kairos design
Three things, one of them a correction.
Normalization keeps the currency. The Kairos paper proposed normalizing bet by denomination and balance by session-opening balance so that every game and currency an operator runs can share one model. Section IV shows that absolute currency amounts are load-bearing: suppression bands are written in them and players respond to them. The token therefore carries both the normalized bet and a log-bucketed absolute bet in a common reference currency, and the balance token does the same. We have amended that section of the paper.2
The synthetic corpus randomizes the mechanisms that matter. Stage A pre-training draws populations from a prior. Section IV says which dimensions of that prior must be wide: stop-rule parameters and adherence, loss aversion, session count, chase and house-money probabilities, bet-size dynamics. Peak-end weights, big-win memory, top-up and buy-bonus rates can start narrow.
Rare tokens are the metric. Fidelity against the ledger and hazard AUC on session and day exits come first; mean likelihood is reported but not optimized for.
VIII. Limitations
Every player here is synthetic. The mechanisms are borrowed from the literature, their parameter values are priors, and the environment's outputs describe how the control workbook behaves under those priors, not how any real cohort behaves. The consistency check in Section II shows the harness is stable across generator versions; it does not show either generator is right. Bonus is off because the workbook's bonus configuration is internally inconsistent, so the interaction between bonus RTP and control is not exercised. And the sixty-population spread in Section V is a statement about the width of our prior as much as about the workbook. All of that is what the real ledger is for, and it is why synthetic data enters pre-training and nowhere else.
IX. Conclusion
The earlier claim was half right. Tables decide the ranking of configurations, and an audit that reads the workbook, prices each module and orders the candidates is still the cheapest useful thing a provider can do. But the tables decide neither the level of retention nor, once players size their bets the way players do, the level or sign of the control ledger. Both belong to the population, and the population is not in the workbook.
That is why the missing half of a payout engine is a model of players and not a better spreadsheet. The environment described here is the pre-training corpus for that model and the test bed for its evaluation protocol, and it comes with a built-in warning: a model that scores well against it has learned this generator, not any real player. The real players arrive with the ledger.
Footnotes
-
What the Tables Decide, and What the Players Decide, on this blog. Its Section III is the claim revised here; its cost figure of 0.67 between-population share appears in the comparison column of Section III above. ↩
-
Kairos: A Foundation Model for the Language of Slot Play, on this blog. Section IV there describes the control-action-outcome factorization; its normalization paragraph was amended on 2026-08-28 following this study. ↩ ↩2
-
Daniel Kahneman and Amos Tversky, "Prospect Theory: An Analysis of Decision under Risk," Econometrica 47(2), 1979. ↩
-
Barbara Fredrickson and Daniel Kahneman, "Duration Neglect in Retrospective Evaluations of Affective Episodes," Journal of Personality and Social Psychology 65(1), 1993. ↩
-
Michael Auer and Mark Griffiths, "Voluntary Limit Setting and Player Choice in Most Intense Online Gamblers," Journal of Gambling Studies 29, 2013, is the reference point for adherence; the stop-loss and stop-win fractions here are priors, not their estimates. ↩
-
Henry Lesieur, The Chase: Career of the Compulsive Gambler, 1984. The chase here is a per-streak bet step-up, a deliberately mild operationalization. ↩
-
Richard Thaler and Eric Johnson, "Gambling with the House Money and Trying to Break Even," Management Science 36(6), 1990. ↩
-
docs/kairos/experiments/players_v2.py,run_sim_study_v2.pyandresults_v2.jsonin the reslot-front repository. Seeds 5000–5009 per population for Section III, 6000–6009 per ablation for Section IV, 7000–7002 per population for Section VI, 8000–8059 with family seeds 900–959 for Section V. 802 runs; roughly four minutes on fourteen cores. ↩