Model
Calculate expected value, tail risk, payback periods, and resource breakpoints.
Research brief · September 26, 2026
Use math to remove bad options, designed experiments to find the powerful dials, and simulation only where player decisions make the system too tangled to solve.
The strongest finding is not “stop simulating.” It is: stop spending the same number of runs on every question. Use the cheapest valid method at each stage.
Calculate expected value, tail risk, payback periods, and resource breakpoints.
Test many parameters at low/high values with a fractional factorial design.
Binary-search one strong difficulty dial, or optimize only the few that interact.
Re-test finalists on fresh seeds and report a 95% confidence interval.
Let humans decide whether the measured balance is actually fun and fair.
At a true rate near 50%, about 385 independent runs give a ±5 percentage-point margin at 95% confidence. Roughly 1,067 runs buy ±3%. Use the smaller batch for screening; reserve precision for finalists.
Two quick tools turn “how many runs?” and “which method?” into explicit choices. The estimates use the conservative worst case near a 50% win rate.
Approximate 95% margin of error for a win-rate estimate.
Formula: n ≈ 0.9604 ÷ E². A Wilson interval should be used in the actual report, especially at smaller samples.
How many interacting parameters are you tuning?
If the dial moves difficulty monotonically, each evaluation halves the remaining range. About seven evaluations narrow the dial to roughly 1% of its original range.
Start at the top. The lower methods become worthwhile only after the engine is deterministic and the basic measurements are trustworthy.
Build an outcome × probability table for actions, rewards, costs, and events. Compare expected payoff, variance, failure risk, and unacceptable tails. It exposes strictly bad options before a single simulated match runs.
Give configuration A and B the same per-game seeds. Their difference reflects the tuning change instead of two unrelated streaks of luck. Separate random streams by subsystem if a change alters how many draws occur.
Screen seven low/high parameters in eight configurations instead of all 128 combinations. The tradeoff is aliasing: a cheap Resolution III design can confuse main effects with two-way interactions.
For two or three smooth, interacting dials, fit a quadratic surface through roughly 9–15 planned configurations. It reveals tradeoff ridges without a dense grid.
A probabilistic model chooses the next promising test. Give exploratory candidates small batches and spend more runs only on finalists. Promising for noisy, interacting systems; validate the optimizer against random search.
Simplify a subsystem into states and absorbing outcomes such as victory, defeat, or stalemate. Exact probabilities become a diagnostic baseline; a wide mismatch with simulation shows where choices or feedback loops matter.
Turn a large batch of undirected test runs into a staged investigation. Simulation still does the heavy lifting—only on the questions that deserve it.
Give every simulated game a derived seed. Keep a small golden set for regressions, then use fresh seeds for the final confirmation.
Choose the rules, rewards, costs, timings, and risks most likely to affect the outcome. Test broad low/high settings before fine-tuning.
Rank parameters by elasticity. If one dial is monotonic, binary-search it toward the target band; if two interact, use a small response surface instead.
Re-run finalists on fresh seeds, inspect resource curves and failure points, then let real play decide whether losses teach rather than merely punish.
A game can hit its target and still feel miserable. These supporting signals reveal difficulty walls, false choices, and runaway economies that the final outcome hides.
A resource that collapses early marks a difficulty wall; one that grows forever reveals an economy with too many sources.
Plot the stage and cause for each ending. A sharp cluster marks a problem that an average run length can conceal.
Track what the policy chooses by state. If experts always pick the same option, the other buttons may be decoration rather than decisions.
Compare random, simple heuristic, and stronger policies. The gap shows whether knowledge and adaptation are rewarded.
Do not report only the mean. Two peaks can mean the system has split into a short hopeless game and a long comfortable one.
Measure bad extremes, not only averages. Check integer thresholds where one more heal, dollar, or spare part changes the player-visible outcome.
These survive every method change—from a spreadsheet model to a future optimizer.
If an action has structurally bad expected value, simulation can confirm the flaw but cannot explain it more cheaply than arithmetic.
Common random numbers make small before/after changes visible with fewer runs. This is the highest-return engineering change.
The machine can locate a target configuration. A human must still judge whether its tradeoffs are legible, tense, and fun.
If a parameter changes the number of random draws—such as event frequency—one shared stream can drift out of alignment. Separate streams for events, combat, rewards, and other subsystems are the safer design; the research did not find a JavaScript-specific implementation guide.
Two references were read in full on September 26, 2026. The remaining links were supported by search-result indexing and should be treated as leads until their full text is checked.