← All posts

What the Tables Decide, and What the Players Decide

· reSlot Research · 11 min read · Kairos

  • kairos
  • simulation
  • dynamic-rtp
  • retention
  • methodology

Abstract

We ran 769 Monte Carlo simulations of a production dynamic RTP control system, a seven-module configuration workbook from a slot provider in soft launch, driven by synthetic player populations. Three findings. First, on the shipped configuration the base game pays out more than it takes in, and one repaired table recovers most of it: net control cost falls from 8.1% to 3.5% of turnover while newbie completion and day-two retention move up, not down. Second, when the same seven configurations are run under six plausible player populations, the cost ranking never changes and the retention ranking changes every time; between 98% and 99% of the variance in day-two retention comes from the population, not from the configuration or the dice. Third, the synthetic generator that produced these runs emits half a million rounds per family per 28-day run, twenty million rounds in a few CPU-minutes, which is the pre-training corpus the Kairos player model needs and, on its own, exactly the corpus it must never be validated on. Everything here is a simulator observation on synthetic players; nothing is a production figure.

Contents

I. The setup

The control system under test is the one we have described before: multi-template switching on a rolling window, a newbie template schedule, cumulative and daily compensation with loss carry-over, four player tags, high-RTP suppression, big-win reroll and a loss-triggered kill switch.1 The configuration is the provider's real workbook, fourteen tabs, loaded unchanged. The players are synthetic, generated by a behavioral model with a dozen parameters: a base day-two return probability, an elasticity of return to session RTP, an elasticity to session hit rate, a bust penalty, a long-term decay, a session-length distribution, a streak-quit multiplier, and a resurrection probability for players who skip a day.

Every run is 28 days at 150 new registrations per day, about 4,200 players and 510,000 rounds, on one seed. Where we report a mean and a spread it is across seeds. The simulator is the phase-one dry-run tool that produced our sample deliverables; the study script and the full results are in the Kairos repository.2

Seven configurations were evaluated: the shipped baseline and six recalibration candidates from the phase-one audit. Two matter for what follows. C2 adds the compensation trigger groups the workbook is missing for rounds 21 to 200, which makes the compensation the designer intended actually fire. R2 fixes three things at once: it sets every newbie-schedule slot to the intended high-hit-rate template, trims the early compensation ceilings, and replaces the daily compensation table, which in the shipped workbook is a byte-for-byte copy of the cumulative table, with the "small top-up" the design document describes.

II. What the tables decide

Twenty seeds per configuration, default player family.

ConfigurationNet control cost / turnoverNewbie completionD2D7Realized RTP
Baseline (shipped)8.14% ± 0.6787.0%45.0% ± 0.76.6%104.0%
C1 fix newbie template 87.94% ± 0.8787.9%46.1% ± 0.76.6%104.7%
C2 add missing trigger groups9.31% ± 0.9987.8%46.3% ± 0.76.8%105.8%
C3 trim early ceilings7.69% ± 0.6487.5%45.2% ± 0.96.6%104.2%
C4 C3 + cover all tags7.85% ± 0.7987.6%45.3% ± 0.86.7%104.7%
R1 newbie template 6 + trim7.20% ± 0.8689.0%46.1% ± 0.66.7%104.2%
R2 R1 + daily table repaired3.45% ± 0.5588.9%46.2% ± 0.76.1%100.3%

The first row is the finding a spreadsheet audit cannot make on its own. The shipped configuration runs the base game at a realized RTP above 100%: the compensation modules give back more than suppression takes, and the game's gross gaming revenue over the 28 days is negative, about 4% of turnover. The provider's own reconciliation rule, that compensation and suppression should roughly balance, is violated by a factor of two, and the reason is legible in the tables. The daily compensation table was pasted from the cumulative one, so every returning player receives newbie-grade top-ups every day.

R2 repairs that and pays for the newbie template upgrade out of the savings. Cost drops 4.7 points of turnover; completion rises 1.9 points; D2 rises 1.2 points, which is well outside the seed noise; realized RTP lands at 100.3%, and the 28-day GGR moves from minus 4% of turnover to roughly break-even. The one thing that softens is D7, down half a point, because the daily top-ups that were propping up returning players on days three to seven are gone. That is a real trade, and it is the kind of trade a split experiment must adjudicate rather than a simulator.

C2 is the mirror image: it makes the compensation the designer intended actually fire, and the bill is 1.2 more points of turnover for 1.3 points of D2. Whether that is a good price depends entirely on Section III.

III. What the players decide

The same seven configurations, ten seeds each, under six player families. F0 is the default above. F1 is price-insensitive (low elasticity to both RTP and hit rate). F2 is RTP-hungry. F3 is hit-rate-driven. F4 is bust-sensitive with a strong streak-quit response. F5 is loyal, with a higher base return and slower decay.

Two rankings per family, cheapest-first and highest-D2-first:

FamilyBy costBy D2
F0 defaultR2 > R1 > C3 > C1 > C4 > base > C2C2 > R1 > R2 > C1 > C3 > base > C4
F1 price-insensitiveR2 > R1 > C3 > C4 > C1 > base > C2C2 > R1 > C1 > C4 > C3 > base > R2
F2 RTP-hungryR2 > R1 > C4 > C3 > C1 > base > C2C2 > C1 > R2 > R1 > C3 > base > C4
F3 hit-drivenR2 > R1 > C4 > C1 > C3 > base > C2R1 > R2 > C1 > C2 > C4 > C3 > base
F4 bust-sensitiveR2 > C3 > R1 > C4 > C1 > base > C2C2 > R2 > C1 > R1 > base > C4 > C3
F5 loyalR2 > R1 > C3 > base > C1 > C4 > C2C2 > R2 > R1 > C1 > base > C3 > C4

The cost column is boring, which is the point. R2 is cheapest in all six families and C2 is dearest in all six; the middle shuffles by fractions of a point. Cost is arithmetic on the tables and the population barely touches it.

The D2 column is not boring. The winner changes (C2 in five families, R1 in the hit-driven one, and R2 falls to last under price-insensitive players because its lost daily top-ups were the only thing those players responded to). The gaps between configurations within a family are one to three points. The gap between families under the same configuration is twenty: baseline D2 is 37.2% under F1 and 56.4% under F5.

A variance decomposition makes it exact. For each configuration, split the variance of each metric across the sixty runs into a between-family component and a within-family (seed) component:

ConfigurationD2: share between familiesCost: share between familiesCompletion: share between families
Baseline0.980.670.96
C20.990.680.95
R10.980.680.95
R20.990.190.96

Ninety-eight to ninety-nine percent of the variance in D2 is the population. Cost is two-thirds population under the shipped tables, because those tables spend more on players who play more, and drops to one-fifth under R2, whose repaired daily table stops spending in proportion to how long people stay.

This is the operational content of the sentence we keep repeating: the tables decide cost, the players decide retention. (Revision, 2026-08-28: a follow-up study with a richer behavioural generator, The Population Is the Parameter, keeps the retention half and the configuration rankings but overturns the cost half. Once players size bets the way players do, the population decides the level and even the sign of the control ledger, because the suppression bands are written in currency. The numbers in this section stand as reported for the simpler player.) A configuration audit, a Monte Carlo on the workbook, a cost ledger, all of that answers the first half exactly and cheaply. None of it answers the second half, because the second half is not in the workbook. It is in the population, and every one of the six families above is a plausible population for a soft-launch title in a market nobody has measured yet.

IV. The cheap mistake: screening on noise

A smaller experiment, because it is the one that bites people. Sixteen configurations, the seven above plus nine single-parameter variants of the baseline, were screened at three seeds each and ranked by D2; the top three were then re-run at thirty seeds.

VariantD2 at 3 seeds (screen)D2 at 30 seeds (verify)
S5 all newbie slots → template 647.6%47.0% ± 0.6
C2 add missing trigger groups47.1%46.4% ± 0.9
R1 newbie template 6 + trim46.7%46.4% ± 0.8
Baseline45.2%45.1% ± 0.8

Every winner regressed by about six-tenths of a point on re-verification and the baseline did not, which is the winner's curse doing exactly what it does: the screen selected the configurations whose noise happened to point up. At this cohort size the effect is smaller than the one we reported from the weight-table simulator earlier, where a coarse-grid search's leaders failed a cost cap outright at five times the player count,3 but the direction is the same and the remedy is the same. Rank on the lower confidence bound, re-verify the leaders on fresh seeds, and treat the screen as a filter rather than a result.

V. What this says about Kairos

We introduced Kairos as a player world model that learns the player instead of the RNG.4 Section III is the empirical case for why that split is the right one, made on the provider's own tables.

If the population explains 98% of D2, then a forecast of D2 from a configuration is worthless without a forecast of the population, and a forecast of the population is not a number a designer can look up. It has to be learned from the ledger. The engine already computes the cost half of every candidate exactly; the missing half is a model of how these players respond to this luck path, and Section III says that model carries almost all of the retention signal. That is the object Kairos is built to be.

Section II adds a second point. The cheap deterministic pass, load the workbook, find the pasted table, price the repair, is worth doing first and on its own. It found a configuration that loses money and a repair that stops the loss while raising completion and D2 under every population we tried. A learned player model does not replace that pass; it sits after it, on the questions the pass cannot answer.

VI. Is the generator a training corpus?

The synthetic generator behind these runs emits, per 28-day run, about 510,000 rounds across 9,150 sessions and 4,200 players, with a median session of 30 rounds and a ninetieth percentile of 132. Two percent of rounds carry a cumulative-compensation flag, half a percent a daily one, seven-tenths of a percent a discard, one-tenth a reroll. Sampling forty behavioral families from broad priors and running each once produced 20.7 million rounds in a few minutes of CPU. Corpus size is not a constraint; the generator is effectively unbounded.

Whether it is a useful corpus is a different question, and Section III answers it in both directions. Because the population carries the retention signal, a model pre-trained on one family learns one population and nothing transferable; a model pre-trained on families sampled from priors learns how behavior couples to the luck path across populations, which is a structural prior and the only thing synthetic data can honestly teach. That is the domain-randomization argument, and it is the argument for Kairos stage A.

The same fact sets the limits. The generator's behavior model has a dozen parameters and one response mechanism per parameter; real players react to bet size, time of day, bonus buys and free tickets in ways it does not encode, so a model that scores well against this generator has learned the generator. Synthetic rounds go into pre-training and nowhere else. Validation, calibration and every retention claim wait for the ledger, and fidelity against the ledger is the first number to report when it arrives.

Footnotes

  1. The modules and their interaction order are documented in the phase-one design notes and summarized in Loop Engineering for Self-Improving Slot Agents on this blog. Control acts on the base game only; bonus RTP is fixed. ↩

  2. docs/kairos/experiments/run_sim_study.py and results.json in the reslot-front repository. Configuration version v0602+V4; simulator seeds 1000–1019 for Section II, 2000–2009 per family for Section III, 3000–3002 and 3100–3129 for Section IV, 4000–4040 for Section VI. The simulator itself is docs/sample_phase1_output/tools/sim_phase1.py. Total runs: 769. ↩

  3. Loop Engineering for Self-Improving Slot Agents, Section VI: 1,000-player grid estimates biased low, top cells rejected at 5,000 players against a 2% cost cap. ↩

  4. Kairos: A Foundation Model for the Language of Slot Play, on this blog. Sections III and V there describe the control-action-outcome tokenization and the ignorability argument; this note supplies the measurement those sections lean on. ↩

Try the loop yourself

The same maker/checker loop runs in our simulators — set a target, run the search, and read the measured results.