← All posts

Kairos: A Foundation Model for the Language of Slot Play

· reSlot Research · 21 min read · Kairos

  • kairos
  • foundation-model
  • player-modeling
  • methodology
  • ai

Abstract

A slot engine emits a stream of rounds: a bet, a payout, a control state, a timestamp, millions of them per game per month. Every question an operator asks of that stream is sequential, and yet the stream has never had a model of its own. General time-series foundation models have never seen a ledger, and tabular models throw the sequence away. We present Kairos, a foundation model for slot play. Its design turns on one fact that separates a slot ledger from every other sequence people have built foundation models for: the next payout is not unknown. It is drawn from a paytable the operator wrote, filtered through a control system the operator configured, and can be computed exactly. Learning it is wasted capacity. Kairos therefore writes each round as a control token, an action token pair and an outcome token, models the action pair conditioned on everything before it, and lets a deterministic engine supply the outcome. The action is factorized coarse-then-fine, the way a player decides: first whether to stop or continue, then how much. The result is a calibrated economy simulator that answers the question a control system actually needs answered, how retention responds to where and when giveaway is spent, and a backbone whose embeddings replace hand-written player tags, exit-hazard rules and lifetime-value heuristics. We show that the operator's own dynamic RTP control makes treatment assignment ignorable given the observed history, which turns a notorious attribution problem into a coverage problem, and we give an evaluation protocol that ends in experiment replay as the only test that proves usefulness rather than fit. This is a design and protocol report; no production weights are claimed.

Contents

I. Why a slot engine needs a foundation model

Every slot operation with a dynamic RTP control system already produces the raw material for one. Each round is a record: what the player bet, what the engine paid, which template was live, which control action fired, when it happened. A soft-launch title produces tens of millions of these a year; a portfolio produces billions. And every operating question is a question about the sequence. How long will this session run. Will this player come back tomorrow. What does a compensation band buy in retention. Where should the giveaway budget flow.

The tools in use answer none of these on the sequence. Configuration audits and Monte Carlo on the workbook compute what the tables do, exactly and cheaply, and stop there.1 Tabular churn models score a player from a handful of aggregates and lose the order of events, which is where a loss-compensation band does its work. General time-series foundation models are trained on electricity, traffic and weather and have never seen a ledger; their tokenizers do not know that a multiplier of zero is a categorical event and a multiplier of 50 is a different kind of event again. The situation is the one that motivated domain-specific foundation models in other fields: a stream with its own statistics, no model trained on it, and enough data to train one.

What makes the slot case unusual, and decides the entire design, is that half of the stream is not uncertain. In a market, a sensor network or a language, the next symbol is unknown and prediction is the whole job. In a slot engine the next payout is drawn from a paytable the operator wrote, filtered through a control state machine the operator configured.2 Given the round's context, the payout distribution is a computation. A model that spends parameters learning it learns the RNG, which is noise by certification. Everything genuinely uncertain in the stream sits on the other side of the round: what the player does with the outcome. That is the language Kairos learns.

The name is Greek for the opportune moment. In a control system whose entire leverage is when giveaway lands rather than how much is spent in aggregate, the opportune moment is the thing worth modeling.3

II. Why not a general-purpose model

The obvious alternative is the model everybody already has. A studio can paste a configuration workbook into ChatGPT or Claude and ask what it will do to retention, and it will get an articulate answer. Several of the studios we talk to do exactly this. The answer is a prior drawn from text about gambling, not a posterior drawn from any ledger, and the difference is the whole product. Five reasons Kairos is a small dedicated model rather than a prompt.

The ledger is not text. A round is a dozen typed fields. Rendered as text it costs thirty or more language tokens; a single player's two-thousand-round history is sixty thousand tokens, per inference, before the model has learned anything. Kairos spends one token group per round, embeds each field in its own table, and holds several sessions of one player in a 2,048-token context. The multiplier ladder, the control flag and the template id are categorical events with known semantics, and a tokenizer built for them does not have to rediscover that a payout of zero is a different kind of symbol from a payout of fifty.

No general model has seen this distribution. Language models are trained on what people have written about slots, which is not the same as what players do in them, and no operator's ledger is on the public internet. A general model's forecast of day-two retention under a compensation band is a fluent guess. It has no calibrated probability over this domain, cannot be sampled to give a distribution of outcomes, and cannot be checked against the ledger and improved. Kairos is trained on the ledger, produces distributions by construction, and is evaluated by experiment replay.

The workload is per round, every round. A live title produces hundreds of thousands of rounds a day, and the uses of a player model (exit hazard, balance rescue, per-player pricing of giveaway) are per round. At a few million parameters Kairos runs in well under a millisecond on a CPU next to the engine. A hosted general model costs cents and seconds per call, which at this volume is neither affordable nor fast enough, and the latency alone would put it outside the payout path where a control decision has to be made.

The data cannot leave. Operators hold player records under gaming licenses and data-protection regimes across jurisdictions. A model that must be sent the ledger to be useful is a model most of them cannot use. Kairos trains and runs on the operator's own infrastructure; fine-tuning a few million parameters on a night's ledger is a single-GPU job, and the weights are theirs.

Auditability is not optional near a payout engine. A configuration decision that spends real giveaway on real players has to be reproducible: same seeds, same inputs, same gates, same answer, with a decision record a regulator or a finance team can follow. A small model with a fixed vocabulary and deterministic sampling under seed gives that. A general model's answer changes with the prompt, the day and the provider's release schedule, and nothing about it can be re-run.

None of this is an argument against general models in the operation. They are the right tool for the roles that are text: drafting a candidate table, explaining a decision in a sentence an operator can read, running the self-explaining layer we described in the loop paper.4 A studio that uses Claude to draft its math has a maker. What it does not have is players, and a maker with no players to test against ships guesses. Kairos is the player. The two compose: the general model proposes, Kairos and the engine simulate, and the checker gates on the result.

III. Design principles

Kairos follows five principles, each of which has a reason specific to slot play.

Discrete tokens, not continuous regression. A round is mostly categorical by construction: the multiplier sits on a ladder, the control flag is one of five symbols, the template and tag are ids. Discretizing the remaining continuous quantities turns the whole stream into a vocabulary, which lets us train with a standard autoregressive objective, sample from it, and suppress the noise that regression would try to fit. This is the same reason discretization has worked for other heavy-tailed, noisy sequences.5

Coarse-then-fine factorization. A player's decision has a branch and a magnitude: stop or continue, and if continue, at what stake. We factorize the action token accordingly and predict the branch first, then the magnitude conditioned on the sampled branch. A single flat head would have to learn that a stake is meaningless when the branch is "end session"; the factorized head is told. It also keeps the vocabulary small: two heads of tens of symbols each instead of one head over their product.

Normalize for pooling, keep the currency. The goal is one model for every game an operator runs, in every currency, at every denomination, so bets and balances are normalized per player and per session. But parts of the control system and of player behavior key on absolute amounts, so the token keeps those too (Section IV).

Pre-train broad, fine-tune narrow. Pre-training on synthetic populations and on an operator's whole portfolio teaches the structure of how behavior couples to the luck path; fine-tuning on the target game fits the specifics. The order matters because the target game's data is the scarce resource.

Sample paths, report intervals. Kairos is used by rolling it forward many times. The mean across sampled paths is the estimate; the spread is the confidence band a guardrail evaluates against. A point forecast would hide exactly the uncertainty that a configuration decision should be made on.

IV. The slot round as a token

The three parts of a round

A Kairos token group is one round, written in causal order:

[ control c_t ]  →  [ action a_t ]  →  [ outcome o_t ]

Control is the state the operator's dynamic RTP system consulted when it prepared the round: the live template id, the player's tag, the cumulative and daily spin-count buckets, the rolling-window RTP bucket, the daily RTP bucket, and the configuration version. Every one of these already appears in the ledger schema a Provider must deliver for phase-one work; nothing new is logged.

Action is what the player did, and it is the only part Kairos predicts:

  • the coarse action is a small categorical: continue at the same bet, raise, lower, change denomination, buy bonus, toggle autoplay or turbo, end the session, end the day, or leave for good;
  • the fine action, conditioned on the coarse one, is the magnitude: a log-bucketed bet size relative to the denomination, and a log-bucketed gap to the previous round. The gap is what defines a session boundary; sessions are never pre-cut.

Outcome is what the engine returned: the multiplier, which on this class of game is already a discrete ladder of eleven rungs times a four-level multiplier; the round type (base, bonus, buy-bonus); the control flag (none, cumulative, daily, reroll, discard); the pre-suppression amount when a win was rerolled or discarded; and the post-settlement balance-to-bet ratio, bucketed.

The factorization that makes it a world model

The joint distribution of a round factorizes into a learned part and a computed part:

p(a_t, o_t | h_<t, c_t) = p(a_t | h_<t, c_t)  ×  p(o_t | a_t, c_t, config)
       learned by Kairos            computed by the engine

During training the outcome tokens are observed, so they enter the context as inputs and carry no loss. During rollout the model samples an action, the engine draws an outcome under the candidate configuration, both are appended, and the loop continues. That single design choice is the difference between a player model and a payout model. A loss on every token would spend most of its gradient on outcome tokens, which are pure paytable noise; Kairos puts the entire loss on the action pair.

Architecture

The backbone is a decoder-only transformer with rotary position embeddings and pre-layer RMSNorm. Each round's fields are embedded separately and summed, together with calendar embeddings for hour-of-day, day-of-week and day-of-month. The coarse head reads the hidden state; the sampled coarse token is embedded and used as a query in a cross-attention layer over the hidden state; the fine head reads the result. The fine head trains on the model's own sampled coarse token rather than the ground truth, so that training matches multi-step rollout and exposure bias does not accumulate over a session.

Why no learned tokenizer in version one

Learned vector-quantized tokenizers exist for continuous multi-dimensional sequences,5 and a slot round is mostly categorical already. The continuous residue is a four-dimensional state, balance over bet, bet over denomination, inter-round gap, and daily RTP. Version one buckets it by hand; version two trains a binary spherical quantizer over that four-vector so the player's current situation becomes one coarse-plus-fine token and the context gets shorter. The reason to defer is data: a learned tokenizer is one more thing to fit on a corpus that is, for now, small.

Two time scales

Rounds are the token, but retention and lifetime value live on days. Kairos therefore has two granularities. The spin-level model has a 2048-token context, enough to hold several sessions of one player, and produces session-length distributions, bet trajectories and the within-session exit hazard. The day-level model compresses each player-day into a record that is, almost literally, a candlestick: opening and closing balance, the day's high and low, turnover as volume, round count, plus the day's session count, the day's net giveaway, and the configuration in force. The day sequence is short, so this layer is a survival-and-regression head on a sequence encoder rather than a language model, but it shares the embedding tables and the normalization of the spin model.

Normalization for pooling

Bets are divided by denomination or session-median bet, balances by session-opening balance, gaps are taken in log space, and multipliers are already dimensionless. That is what lets every game a Provider operates, in every currency, at every denomination, enter one model.

Normalization alone is not enough, and we corrected this after our second simulation study.6 Two parts of a control system are written in absolute currency rather than in ratios: big-win suppression bands and the loss-triggered kill switch. Players also respond to absolute amounts, not to multiples of their denomination. A token that carried only the normalized bet would hide the very quantity that decides whether a win is rerolled. Each round therefore carries both: the normalized bet and balance for pooling, and a log-bucketed absolute bet and balance in a common reference currency for the mechanics that depend on size. Calendar embeddings keep hour-of-day, day-of-week and day-of-month; the last is not decoration in markets where pay-day effects are visible in the turnover curve.

V. Data, and the honest gap

A single slot title in soft launch produces perhaps ten to a hundred million rounds a year, which is small by the standards of foundation models in other fields. Two things soften it. Kairos learns behavior, not RNG, and the entropy of a player's action given history is far lower than the entropy of a price or a sentence. And a model of a few million parameters is trainable on tens of millions of tokens; the hundred-million-parameter tier has no use at this corpus size.

The corpus is assembled in three layers.

LayerSourcePurpose
SyntheticThe reslot.dev engine running the real configuration tables, driven by a parameter-randomized family of behavioral playersStructural prior: how behavior couples to the luck path, without committing to any numbers
Provider-wideEvery title the operator runs, normalized as aboveCross-game transfer; the more titles, the better the prior
Target gameAt least four contiguous weeks of the game under study, with attribution fields intactFine-tuning

The synthetic layer is domain randomization in the sense the robotics literature uses it:7 the behavioral generator has a set of parameters (stop rules, chase tendency, loss aversion, session budget, a per-player heterogeneity scale), and the pre-training corpus samples those from priors instead of fixing them. The lesson that motivated the heterogeneity term is instructive: without per-player variation, a synthetic cohort's survival curve is a single geometric decay and day-seven retention collapses to a fraction of a percent, far below any real cohort. A model pre-trained only on such data would confidently reproduce that mistake. Synthetic data therefore stays in pre-training and never enters a validation set.

VI. The attribution problem is a coverage problem

Any observational analysis under dynamic RTP control is polluted by the control itself: players who received compensation received it because the system selected them for losing, so regressing retention on compensation confuses selection with effect. We have said this in earlier work and it remains true.4 But the structure of the operator's control system makes the pollution more tractable than it looks.

The control is a deterministic function of observed state (cumulative rounds, daily rounds, rolling RTP, tag), composed with configured trigger probabilities. Assignment of a control action therefore depends on nothing the ledger does not record. Given the full history, treatment is ignorable: no unobserved player trait can confound it, because the policy never consulted one.8 A sequence model that carries the complete control state in its context learns p(action | history, control) as a valid conditional distribution, and rollouts under a different configuration are valid counterfactuals wherever the training data covers the new configuration.

The remaining problem is positivity. The old configuration explored a narrow region of parameter space; a candidate configuration outside it asks the model to extrapolate. Two things help. The trigger probabilities are natural randomization and should be preserved and up-weighted in training rather than averaged away. And the split experiments that phase-one work runs anyway are, for Kairos, treatment-balanced fine-tuning data. The conclusion stands: the model replaces analysis, not experimentation; it makes each experiment's samples go further.

VII. Training recipe

StageContentData
A. Synthetic pre-trainingSpin-level decoder-only model, dual head, outcome tokens as input onlyRandomized synthetic families, unbounded
B. Provider-wide pre-trainingSame architecture, continuedAll titles, normalized
C. Target-game fine-tuningTokenizer first if learned, then predictor, at reduced learning rateThe game under study
D. Intervention fine-tuningTreatment-balanced batches; a CATE head with a doubly-robust targetSplit experiments
E. Multi-task headsSession-exit hazard, D2, lifetime value, unsupervised tag clusters on the backbone embeddingSame

Reference configurations for the first release:

ModelLayersd_modelHeadsContextParams
Kairos-mini619262048≈ 4M
Kairos-small851282048≈ 20M

Optimization is standard: AdamW, cosine decay with linear warm-up, dropout decreasing with scale. Three details are load-bearing. The loss is on action tokens only. Splits are by player first and then by time, so that late-arriving players never leak into training and configuration versions are aligned to the rounds they governed. And sessions are never pre-segmented: the previous day's net loss is precisely what a loss-compensation band is designed to act on, and a model that cannot see across the session boundary cannot learn what the band does.

VIII. Evaluation

Four families of metrics, in the order they should be trusted.

QuestionMetric
Does it predict the next action?Next-action negative log-likelihood and calibration; CRPS on session length; AUC on the session-exit and day-exit hazards
Does it predict how much people bet?Error on per-player bet volatility and on the day's turnover distribution
Is it a faithful simulator?Distance between rolled-out and real distributions of session length, turnover, realized RTP and the four giveaway-ledger accounts; a discriminative score against real sessions
Is it useful?Experiment replay: hold out a completed split experiment, estimate its lift with doubly-robust or fitted-Q off-policy evaluation from Kairos rollouts, compare to the measured lift

Two cautions built into the protocol. Rare tokens are the metric: four fifths of actions are "same bet again", so a model scored on average likelihood can look good while learning nothing about exits, and the exit hazards are what an operator pays for.6 And the order matters: fidelity comes before replay, because a model must first be a good simulator before it can be trusted as a counterfactual engine. Replay is the only entry in the table that proves usefulness rather than fit, and it is the number a Provider should ask for before letting Kairos near a configuration decision. Averaging ten sampled rollouts gives the point estimate, and the spread across them is the confidence band a guardrail evaluates against.

IX. Where Kairos sits in the operation

Kairos replaces the hand-written player in the weight-table simulator with a learned one, behind the same interface: given history and the current outcome, return the next action. The maker-checker loop we described earlier4 gains a checker whose Monte Carlo is world-model-times-engine rather than hand-model-times-engine, and whose multi-path samples yield cost and retention as intervals. Guardrails evaluate at the upper confidence bound, which is where the winner's curse in configuration search is actually caught.

For the portfolio objective, the backbone embedding replaces the operator's four hand-defined tags; the hazard head is the modeled form of balance rescue; the lifetime-value and CATE heads supply the denominator of the retention price that the budget market clears against. The market's arithmetic does not change; it acquires a source.

Three things Kairos does not do. It never sits in the payout path: the engine draws every outcome, and Kairos only forecasts what players do with it. It never modifies the math model, which stays with the operator's numerical designers. And its outputs always pass through the deterministic guardrail layer, including a long-run realized-RTP floor whose justification is commercial rather than sentimental: a model this good at luck-curve planning is a model that can over-extract, and over-extraction shows up as churn.

X. Limitations

Data volume is the binding constraint. Synthetic pre-training supplies structure; only real experiments supply elasticity.

Non-stationarity. Player mix drifts with acquisition channels, promotions and free-ticket campaigns. Fine-tuning is a schedule, not a step.

Circularity. A simulator pre-trained on data from a simulator will, if under-fine-tuned, confidently restate the generator's assumptions. The synthetic layer is fenced off from validation for exactly this reason.

Extrapolation shows up as variance. When a candidate configuration lands outside the visited region, multi-path rollouts disagree more. That is a feature: the guardrail treats disagreement as a signal to run a small split test before trusting the estimate.

Behavioral economics, weaponized. The better the model, the sharper the luck-curve planning it enables. The floor in Section IX exists because the reversal is empirical: extraction beyond the player's tolerance is paid back in retention.

XI. Conclusion

Slot play is a language, but the sentence has a part the speaker already knows. Kairos tokenizes the round as control, action and outcome, learns the action coarse-then-fine, and leaves the outcome to the engine that owns it. In exchange it gets a world model that composes with any configuration, a clean answer to the attribution objection under deterministic control, and an evaluation that ends in experiment replay rather than curve fit.

The three steps that need no Provider interface are the tokenization pipeline over the phase-one ledger schema, the parameter-randomized behavioral families for synthetic pre-training, and a Kairos-mini trained and scored against the hand-written player on the fidelity rows of Section VIII. That comparison is the skeleton of the simulator-calibration report phase one owes anyway. Everything after it depends on data only the operator can provide.

Footnotes

  1. The phase-one configuration audit we run for providers: load the control workbook, price every module by Monte Carlo, find the inconsistencies, map the cost side of the feasible region. It is exact about what the tables spend and silent about what players do with it; the simulation studies that follow this note in the series measure how large that silence is. ↩

  2. The control state machine we model is the one described in our earlier notes on dynamic RTP control: multi-template switching on a rolling window, newbie template assignment, cumulative and daily compensation with loss carry-over, player tags, high-RTP suppression and big-win reroll, and a loss-triggered kill switch. Control acts on the base game only; bonus RTP is fixed and outside its reach. ↩

  3. The word is Aristotle's and the rhetoricians'; in the sense used here it names the moment at which an action has its effect, as opposed to the measured passage of time. ↩

  4. See Loop Engineering for Self-Improving Slot Agents on this blog for the maker-checker separation, the verification hierarchy, and the attribution-drift failure mode. The winner's-curse observation in the weight-table simulator, where a coarse-grid search's leaders failed re-verification at five times the player count, is the concrete case the confidence-bound rule in Section IX is written for. ↩ ↩2 ↩3

  5. Discretizing continuous series into a vocabulary for autoregressive modeling is established practice, e.g. Ansari et al., Chronos: Learning the Language of Time Series, 2024, which uses uniform binning; binary spherical quantization, which we plan for the version-two tokenizer, is from Zhao et al., Image and Video Tokenization with Binary Spherical Quantization, 2024. ↩ ↩2

  6. The Population Is the Parameter, on this blog. Section IV there: switching off bet-size chasing in the synthetic population moved realized RTP from 92.7% to 103.6% on unchanged tables, because the suppression bands are absolute. Section VI there: the next action is 82% "same bet" and its entropy falls only from 1.01 to 0.95 bits under immediate context. The normalization paragraph above was amended on 2026-08-28 in light of the first result. ↩ ↩2

  7. Josh Tobin et al., Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World, IROS 2017. The transfer target there is a physical camera; here it is a population of players. ↩

  8. This is strong ignorability in the sense of Rosenbaum and Rubin (1983), with the observed history as the covariate set. It holds because the policy is a function of that history; it would fail the moment an operator's staff assigned compensation by hand on the basis of something unlogged. ↩

Try the loop yourself

The same maker/checker loop runs in our simulators — set a target, run the search, and read the measured results.