I. From tuning parameters by hand to designing the loop
Every slot operation with a dynamic RTP control system is already running a loop, whether or not anyone has drawn it. Somebody reads last week's ledger, notices that newbie completion looks soft or that the giveaway account is running hot, opens the configuration workbook, widens a compensation band, narrows a suppression window, moves a template boundary. The change goes live. Two weeks later somebody reads the ledger again.
That loop turns once a fortnight, and its cadence is set entirely by human attention. The math is not the bottleneck: the exact RTP of a candidate paytable is a finite computation a machine finishes in milliseconds. The bottleneck is that a person must decide what to compute, interpret the result against a mental model of the control system, and convince themselves it is safe to ship. Each turn costs a context switch, the most expensive thing a small math team owns.
Loop engineering deletes the human from the position of prompting the work. In the old arrangement you drive an agent one question at a time; it answers, stops, and waits for you. You are the clock. In the new arrangement you write the loop, define what counts as a pass, and walk away. The leverage point has moved one floor up: the agent is no longer the tool you are designing, the system that orchestrates the agent is.1
This matters more in slot math than in most domains, because a slot economy is a system of coupled constraints. Change a reel strip to buy back half a point of RTP and you have also changed hit frequency, which changes how a newbie experiences their first two hundred spins, which changes newbie completion, which changes the D2 cohort, which changes the turnover the control system spends its giveaway budget against. A human tuning one parameter at a time cannot hold that graph in working memory; a loop re-derives it every iteration and writes down what it learned.
The consequence people notice first is that headcount inside the loop falls. The interesting one is cadence. A math team cycles a paytable idea perhaps once a week; a correctly assembled loop cycles it as fast as verification allows — for exact math, milliseconds; for behavioral questions, as fast as the sample can be gathered — and the constraint relocates from how many candidates can we produce to how confidently can we reject the bad ones.
II. The six primitives, mapped to slot operations
A working loop is built from exactly six structural components. Omit one and you get a loop that fails quietly: it runs, it burns compute, it emits configuration diffs, and nothing compounds.
Automation is the heartbeat, a schedule or trigger that fires without anyone typing: a cadence loop reruns on an interval, a goal loop iterates until a verifiable stopping condition is true. Skill is a procedure manual read at the start of every session, holding the constraints nobody should have to rediscover: control acts on base game only and never touches bonus; base-game RTP under a newbie template must be lifted to cover the bonus contribution a newbie will not yet receive; hit-frequency drift across a template switch must stay imperceptible; a reroll target must be a strictly smaller positive multiplier drawn by weight. Without a skill file every iteration relearns the domain; with one, a lesson written today is a constraint tomorrow.
State is loop memory, a plain file read first on every run and written last: the agent forgets, the file does not. Verifier is the independent judge, a second agent with different instructions and ideally a different model, grading the first one's output against deterministic gates with no exposure to its reasoning. Worktree is isolation, because the moment two agents write the same configuration they collide. Connector is reach: ledger queries, the live configuration and its change history, a staged candidate write.
| Primitive | Role in the loop | In a slot operation |
|---|---|---|
| Automation | Fires the loop without a human | Nightly ledger ingest; goal loop that searches paytables until exact RTP and hit frequency clear tolerance |
| Skill | Persistent domain rules and lessons | Base-only control; newbie base RTP lifted for missing bonus; hit-frequency drift budget across template switches; reroll target must be a smaller positive multiplier |
| State | Memory across sessions | Live templates with exact RTP and hit frequency; control cost as a share of turnover; open experiments; last ten lessons |
| Verifier | Independent grading | Fresh Monte Carlo at larger player counts on unseen seeds; cost-cap and guardrail gates; rejects on any single failure |
| Worktree | Parallel isolation | Paytable search, cost simulation and guardrail monitoring in separate working copies, no shared context |
| Connector | Reach beyond local files | Ledger queries, configuration plus change history, staged config write, decision audit trail |
Two are routinely underestimated. Teams build elaborate search machinery with no durable record of what was tried, so the loop re-proposes a configuration it rejected three weeks ago. And a guardrail monitor sharing a worktree with the proposer inherits the proposer's assumptions: the first time the proposer wanders into a bad regime the monitor walks in with it, and the circuit breaker never fires, because the thing holding it agrees with the thing that needs breaking.
III. The five-stage slot loop
The loop is five sub-loops chained in sequence, each with its own skill file and its own slice of state. Each runs in its own worktree; all five share one memory layer underneath.
Observe. An automation fires nightly, because the control system's most important accounting periods are daily. The observing agent pulls the round ledger with its attribution fields intact (which template was live, which control action fired, the pre-suppression amount when a win was rerolled or discarded, cumulative and daily spin counts), reconciles the giveaway account, and computes cohort retention. It writes what it saw into state, anomalies included: a template whose realized hit frequency has drifted from theory, a compensation band pinned at its trigger ceiling, a cohort whose newbie completion moved outside its usual range.
Observe must be honest about what it cannot conclude. Under dynamic control, observational attribution is polluted by the switching mechanism itself: players who received compensation received it because the system selected them for losing, so regressing outcomes on compensation mistakes selection and mean reversion for effect. Observe labels the questions the ledger alone cannot answer, and those become experiment proposals, not conclusions.
Propose (maker). The maker reads its skill file, opens the current state, and produces a candidate change with its reasoning: a reel strip, a weight set, a template boundary, a compensation band, a suppression window. Every past lesson writes a constraint into that skill file, so the search space narrows toward proposals that have historically survived verification. A representative rule set forbids any candidate whose hit-frequency drift from the outgoing template exceeds the perceptibility budget, requires that suppression never target a compensation-issued win unless that interaction has been explicitly resolved, caps per-player giveaway, and refuses any candidate proposed without a predicted cost. That last rule turns every iteration into a calibration datapoint, which is what sets how much verification budget future proposals deserve.
Verify (checker). The candidate goes to a completely separate agent: different instructions, ideally a different model, and crucially no sight of the maker's reasoning trace. The checker runs its own Monte Carlo at larger player counts on fresh seeds, over the full round state machine including the interaction ordering between compensation, template selection, high-RTP suppression and big-win reroll, then applies fixed gates. Exact RTP within tolerance. Hit frequency and volatility inside band. Control cost as a share of turnover under the portfolio hard cap, evaluated at the upper end of its confidence interval rather than the point estimate. Per-player giveaway cap respected in the tail, not just on average. No configuration inconsistency: no overlapping intervals, no correction floor above its trigger threshold, no template id missing from the math model. Fail one gate and the candidate dies, with the reason logged.
That the checker never sees how the maker reasoned is the entire edge. A checker exposed to the maker's argument grades the argument; a checker exposed only to the artifact grades the artifact.
Apply. Only verified candidates reach application, and even then Apply is not a deploy button. A change enters through a split experiment where the question is behavioral and a staged rollout where it is arithmetic. That distinction is the verification hierarchy in operational form: a split test beats a predict-then-verify record, which beats a cohort before-and-after comparison. No amount of simulation substitutes for randomized assignment on a retention question. Every applied change carries a decision record.
Monitor. A parallel worktree runs the guardrail monitor at a much tighter cadence, watching portfolio control cost against its hard constraint, per-player giveaway against its cap, and GGR against a circuit-breaker threshold. When one trips it reverts to the last known-good configuration and writes an incident into state. Neither maker nor checker can override it, and it works only because it shares no context with what it polices.
| Stage | Role | Grades its own work? | Primary artifact |
|---|---|---|---|
| Observe | Ingest ledger, reconcile the giveaway account, compute cohorts | n/a — reports, never concludes causality | Updated state; anomalies; experiment proposals |
| Propose | Generate candidate math or control parameters, with a predicted cost | No | Candidate config plus predicted cost and retention effect |
| Verify | Independent Monte Carlo and deterministic gates | No — never sees maker reasoning | Pass/fail with gate margins and rejection reason |
| Apply | Split experiment or staged rollout, with a full decision record | No — bound by the checker's verdict | Assignment plan; audit trail entry |
| Monitor | Portfolio cost cap, per-player giveaway cap, GGR circuit breaker | No — isolated worktree, no shared context | Incident record; automatic revert |
IV. The self-improvement mechanism
After every completed iteration, whether a shipped change that ran its course or a rejected candidate, the maker and the checker each write a short retrospective: what they expected, what happened, what rule should change. The lesson is appended to the relevant skill file. A representative entry: a template pair that clears the RTP gate can still fail perceptibility if the hit-frequency step is concentrated in the low-multiplier band; constrain drift per multiplier tier, not just in aggregate. After fifty iterations the skill file is a working rulebook. After five hundred it encodes more about this game's economy than anyone could hold in working memory, and unlike a simulation report it cannot be reconstructed, because it was paid for in rejected candidates and shipped experiments.
The state file is deliberately boring. A representative slice for one game follows. The numbers below are illustrative, not measurements.
# STATE.md — SLOT LOOP MEMORY (illustrative)
## Last run: 2026-01-01 00:00 UTC | iteration 214
### Live math (exact, base only)
templates: T1 96.00 / T2 95.40 / T3 94.85 hit_freq: 25.0% / 24.6% / 24.2%
vol_index: T1 8.1 / T2 8.4 / T3 8.9 hf drift budget 2.0pp — OK
### Giveaway account (rolling 7d, share of turnover)
net_control_cost 3.9% portfolio_cap 4.5% <- 0.6pp headroom
### Open
exp-041 newbie base-RTP +1.0pt split, day 3/14, guardrail=D2
exp-039 daily band widened REJECTED: per-player cap breached in tail
### Lessons (last three)
- 09-03 coarse-grid cost estimates are biased low; re-verify winners at 5x players
- 08-29 suppression on compensation-issued wins double-counts the giveaway ledger
- 08-24 newbie base RTP must absorb the missing bonus or completion collapses
Nothing there is clever. A loop without one cannot compound: the agent forgets, and nothing carries the lesson forward.
The dominant failure mode is not in the maker but in the verifier. When gates are loose, weak candidates pass, and the loop looks like it is compounding while it accrues latent cost and latent retention risk. We call the gap between the verification a loop enforces and the verification the operation requires verification debt. It is invisible by construction: a loose gate produces the same green result as a tight one, only more often. The cure is scheduled recalibration by a third agent whose only job is to audit the verifier against the outcomes log. In slot work the debt has a specific and expensive shape: cost gates that pass on point estimates rather than confidence bounds will systematically ship configurations whose true cost sits above the estimate, because the search selected them for looking cheap.
V. Industry parallels
The maker–checker split is not an artifact of agentic systems. It is the standard organization of every serious math operation in this industry, and the loop merely automates a pattern already on the org chart. At mature studios, the mathematician who designs a paytable typically does not sign it off; a separate validation step reproduces the RTP independently, and external certification labs exist for the same reason. The team that designs a compensation mechanism does not own the risk limit that constrains it. The split exists for exactly the reason the verifier sub-agent exists: the entity that produced an output cannot be trusted to grade it.
The second parallel is the one that has changed. Running the five stages by headcount needs data engineers, mathematicians, a validation team, a release process and a risk team. That division of labor is still an advantage at the level of capital, distribution and regulatory reach. It is no longer an advantage at the level of the loop: when maker, checker, applier and monitor are all agents in isolated worktrees, the division of labor survives intact while the headcount inside the loop collapses. The studio still wins on scale. It no longer wins on cadence, which is what determines how fast anyone learns what their giveaway budget is buying.
VI. Observations from simulation
Everything here was measured in the reslot.dev simulators, on their demo games at default settings. These are simulator observations, not production figures: they describe how the loop behaves, not how any game performs in the field.2
Start with the arithmetic half, in the PAR Sheet simulator. The demo game's exact RTP is 95.77%, its hit frequency approximately 33.9%, its volatility index approximately 8.1. Those are exact-analysis numbers, not simulation estimates, and the distinction is the first thing the loop exploits: when the outcome space can be enumerated, verification needs no statistics at all. With prefix pruning, exact analysis of the demo game runs in roughly 1 ms against roughly 18 ms for brute force. A verifier that grades a candidate in a millisecond changes the economics of search, because the checker stops being the slow leg.
Running the goal loop in Tune mode against targets of 96.00% RTP and 25.0% hit frequency, it reached both targets in 15 iterations at the page's default seed, screening 80 candidate moves per iteration and confirming 3. Two things about that ratio matter. Proposing is cheap and confirming is the expense: the loop spends almost all its budget on the three candidates it takes seriously, and screening exists to make that affordable. And single moves stop working once the loop is close — as soon as an individual move overshoots one target while correcting the other, the loop must switch to paired moves that adjust two positions together, a constraint that belongs in a skill file rather than in anyone's head. Starting instead from uniform reels in Generate mode, the same loop converged in 12 to 16 iterations.
The behavioral half is where verification stops being cheap. In the Weight Table simulator, the demo configuration produces a base-only newbie D2 in the range of 50 to 55% and a control cost of approximately 4% of turnover. Searching that space with a coarse-grid optimizer at 1,000 simulated players per cell yields a ranked list of configurations that look cheaper and better than the incumbent. Re-verifying the top-ranked cells at 5,000 players rejected them against a 2% cost cap: the 1,000-player estimates were biased low, and biased low precisely because the grid search selected for cells whose noise happened to point downward. That is the winner's curse in operational form: the more candidates you screen, the more the leader of the ranking is a configuration that got lucky rather than one that is good.
That experiment is the thesis of this note stated as a measurement: the verifier, not the maker, was the binding constraint. Producing candidates was effectively free, and the cost of the loop sat entirely in the sample size required to trust a ranking. A loop that answers cheap generation with more candidates rather than a higher confirmation standard is not compounding. It is manufacturing verification debt.
VII. Four ways loops quietly fail
Loops fail more often than they succeed, and the failures are undramatic. When these loops fail, they keep running and keep emitting diffs; what stops is the compounding. The dangerous loop is not the one that crashes but the one that looks healthy while degrading.
Cognitive surrender. The operator stops reading proposals and ships everything that passes the gates. Within weeks the live configuration contains decisions nobody can defend, and diagnosis is impossible because no human holds a model of how the system behaves. The cure is a deliberate attention budget: a fixed block of time each week reading the loop's most consequential changes and explaining them to yourself, whether or not anything looks broken.
Verification rot. The gates were calibrated once, against a player mix that no longer holds. They still reject the occasional candidate, so the process feels rigorous, but nearly everything passes. Slot operations are unusually exposed, because what the gates protect, the relationship between giveaway spend and retention, is an elasticity that moves with the player mix, and gates calibrated on one cohort become theater on another. The remedy is scheduled recalibration against the outcomes log, on observed precision rather than on memory.
Attribution drift. The loop begins treating observational readouts as causal. Nothing announces this; the ledger looks the same. But the moment a proposal is justified by “players who received compensation retained better,” the loop is optimizing a selection effect, because the control system chose those players on the basis of how they were losing. The defense is structural: keep a randomized split running for the mechanisms you care about, and have the skill file forbid any retention claim whose evidence sits below split-test level.
Runaway iteration. The loop runs forever because its stopping condition was never expressed as something an outside party can check. The agent asserts convergence, the loop halts, a half-tuned configuration sits in staging. Every stopping condition must be a deterministic inequality over an observable quantity: exact RTP within tolerance, hit frequency inside band, cost under cap at the upper confidence bound, guardrails green. Never let the agent that did the work be the one that says it is done.
VIII. The economics of running the loop
The standard objection is that running several agents continuously must cost more than the optimization returns. The compute cost of a slot loop is dominated by verification, which scales with the simulated player count needed to reach a decision. That is a real bill and one worth paying, because Section VI is exactly the case where paying it separated shipping a configuration from shipping an illusion. The structural point is not the size of the bill but what it is proportional to.
Consider a hypothetical worked example. Suppose an operation runs a single game on a fixed monthly verification budget. Every artifact that budget produces, from the calibrated simulator and the skill file's constraints to the gate thresholds and the guardrail logic, is game-specific in its parameters but generic in its structure. Onboarding a second game reuses the structure and re-estimates only the parameters; the third reuses the second's onboarding lessons as well. The marginal cost of the nth game falls while the pooled evidence about how giveaway spend converts into retention rises, so the gates on game n can be tighter and cheaper than the gates on game one. The break-even is organizational, not computational: it is crossed when the loop's output stops being a report someone reads and starts being a configuration that ships on a schedule.
A second cost is not compute at all. Exploration is real money: every randomized split spends part of the giveaway budget on a configuration you already suspect is worse, to buy the right to know how much worse. That is the price of the only evidence at the top of the verification hierarchy, and it belongs in the budget rather than in a surprise on the ledger. An operation that refuses to fund exploration has simply chosen to keep paying a giveaway budget whose retention return it will never measure.
IX. Conclusion: from holding the baseline to managing the portfolio
A conventional dynamic RTP control system is, structurally, a deviation-accounting machine built around a single baseline. Template switching, cumulative and daily compensation, high-RTP suppression, big-win reroll — every module exists to pull each player back toward one line, and the system's own reconciliation says so, since what compensation paid out above baseline should roughly balance what suppression took back. Held to that job it is legitimate and worth tuning well. The ceiling is in the premise. As long as everyone is pulled toward the same line, the giveaway budget cannot flow to where it buys the most retention, because the mechanism has no concept of where it buys more. It knows only how far from the line this player is, which is a different quantity.
Loop Engineering for Self-Improving Slot Agents is how an operation crosses from one to the other without a leap of faith. Run against the existing control system, the loop first does the cheap deterministic work: it prices every module, finds the configuration inconsistencies, maps the feasible region where cost and experience trade off. Then it measures what the old framing never needed to know, the elasticity of retention to giveaway per segment and per moment, because a split experiment is a stage in the loop rather than a special project someone has to justify. Once that elasticity is measured rather than assumed, the objective can be restated honestly: spend the same giveaway budget where it buys retention, under a portfolio-level cost constraint, with per-player caps and a GGR circuit breaker as deterministic guardrails and a full decision audit trail underneath. RTP stops being the target and becomes a by-product.
Two cautions belong here plainly. The loop does not manufacture edge from nothing; it compounds whatever understanding the operation brings into it, and an operation with no thesis about its players gets a very fast machine for redistributing money at random. And the gates matter more than the generator: a loop with a brilliant maker and a loose checker learns to lose money efficiently, while a loop with an ordinary maker and a strict checker compounds slowly and survives. Where every shipped configuration spends real giveaway budget on real players, surviving is the whole game.
X. Toward a standard practice
If the framing here is right, the useful next step is convergence on shared conventions. Three are ripe.
The first is a skill-file schema for slot loops. Everyone invents their own structure today, which makes lessons non-portable between games and impossible to benchmark. A standard shape (goal, hard constraints, gate definitions, dated lessons, regime tags) would let a studio carry constraints from one title to the next.
The second is a verification suite. The gates in Section III are a starting point, not a standard. The industry has mature practice around math certification; what is missing is that practice expressed as deterministic checks any verifier agent can execute, extended to what certification does not cover: cost caps evaluated at confidence bounds, per-player giveaway caps in the tail, interaction ordering between compensation and suppression, and the requirement that a retention claim name its position in the verification hierarchy.
The third is an incident format. Every operator who has run one of these loops has watched it do something surprising, and those lessons sit locked inside private state files. A shared incident schema, in the spirit of the post-mortem format the reliability community converged on a decade ago, would let the practice learn collectively. None of this needs new technology; what is missing is the convention, and conventions are written by whoever publishes first.
The practical starting point is smaller than any of that: one game, one metric, one verification rule. Take a single title. Pick the number you are willing to be judged on, whether newbie completion, D2, or control cost as a share of turnover, and choose it before you look. Write one gate a proposal must clear, and make sure the agent grading the gate is not the agent that wrote the proposal. That is a loop, and everything above is refinement on top of it.
Both halves of that first loop are already runnable. The exact-analysis half (reel strips, paytables, RTP, hit frequency, volatility, and the goal loop that tunes them) runs in the PAR Sheet simulator. The behavioral and cost half (templates, compensation, suppression, control cost as a share of turnover, and the verification problem that comes with searching that space) runs in the Weight Table simulator. Run one iteration by hand in each, then ask what would have to be true for it to run without you. That question is the whole discipline.
Footnotes
-
The Loop Engineering framing, including the six primitives and the maker–checker separation, is adapted from a quantitative-trading practitioner note on loop engineering for self-improving hedge funds. The trading application is theirs; the slot mechanics, failure modes and measurements here are ours. ↩
-
All figures in Section VI were measured in the reslot.dev simulators on 3–4 August 2026, using the demo games and default seeds published on those pages. They are properties of the simulators and of the loop's search behavior, not production performance data for any live game. ↩