← All projects

Research brief · September 26, 2026

Smarter balance tests, fewer wasted runs.

Use math to remove bad options, designed experiments to find the powerful dials, and simulation only where player decisions make the system too tangled to solve.

A band Set a target range for the experience you want—not a single magic number.

Brute force, but with a map.

The strongest finding is not “stop simulating.” It is: stop spending the same number of runs on every question. Use the cheapest valid method at each stage.

Model

Calculate expected value, tail risk, payback periods, and resource breakpoints.

Screen

Test many parameters at low/high values with a fractional factorial design.

Tune

Binary-search one strong difficulty dial, or optimize only the few that interact.

Verify

Re-test finalists on fresh seeds and report a 95% confidence interval.

Play

Let humans decide whether the measured balance is actually fun and fair.

±
A win rate without uncertainty is hard to act on.

At a true rate near 50%, about 385 independent runs give a ±5 percentage-point margin at 95% confidence. Roughly 1,067 runs buy ±3%. Use the smaller batch for screening; reserve precision for finalists.

A tiny balance lab.

Two quick tools turn “how many runs?” and “which method?” into explicit choices. The estimates use the conservative worst case near a 50% win rate.

Run-budget calculator

Approximate 95% margin of error for a win-rate estimate.

±5%
Minimum runs385
Best useScreen

Formula: n ≈ 0.9604 ÷ E². A Wilson interval should be used in the actual report, especially at smaller samples.

Method chooser

How many interacting parameters are you tuning?

Most practical

Binary search

If the dial moves difficulty monotonically, each evaluation halves the remaining range. About seven evaluations narrow the dial to roughly 1% of its original range.

The toolkit, ranked by payoff.

Start at the top. The lower methods become worthwhile only after the engine is deterministic and the basic measurements are trustworthy.

Expected-value math

Build an outcome × probability table for actions, rewards, costs, and events. Compare expected payoff, variance, failure risk, and unacceptable tails. It exposes strictly bad options before a single simulated match runs.

Use whenRandom draws are mostly independent and policy is fixed.

Same-dice A/B tests

Give configuration A and B the same per-game seeds. Their difference reflects the tuning change instead of two unrelated streaks of luck. Separate random streams by subsystem if a change alters how many draws occur.

Use whenComparing any before/after change.

Fractional factorial

Screen seven low/high parameters in eight configurations instead of all 128 combinations. The tradeoff is aliasing: a cheap Resolution III design can confuse main effects with two-way interactions.

Use whenYou need to discover which dials matter.

Response surface

For two or three smooth, interacting dials, fit a quadratic surface through roughly 9–15 planned configurations. It reveals tradeoff ridges without a dense grid.

Use whenResponses look smooth, not breakpoint-heavy.

Bayesian optimization

A probabilistic model chooses the next promising test. Give exploratory candidates small batches and spend more runs only on finalists. Promising for noisy, interacting systems; validate the optimizer against random search.

Use whenThree to six dials interact and tests are costly.

Markov model

Simplify a subsystem into states and absorbing outcomes such as victory, defeat, or stalemate. Exact probabilities become a diagnostic baseline; a wide mismatch with simulation shows where choices or feedback loops matter.

Use whenThe state space can stay intentionally small.

A reusable plan for any game.

Turn a large batch of undirected test runs into a staged investigation. Simulation still does the heavy lifting—only on the questions that deserve it.

01 · CALIBRATE

Make the randomness repeatable

Give every simulated game a derived seed. Keep a small golden set for regressions, then use fresh seeds for the final confirmation.

same seed + new config = clean A/B
02 · SCREEN

Find the real balance levers

Choose the rules, rewards, costs, timings, and risks most likely to affect the outcome. Test broad low/high settings before fine-tuning.

screen wide first; measure effect size
03 · AIM

Search the strongest dial

Rank parameters by elasticity. If one dial is monotonic, binary-search it toward the target band; if two interact, use a small response surface instead.

set the target before testing
04 · FEEL

Verify, then hand it back

Re-run finalists on fresh seeds, inspect resource curves and failure points, then let real play decide whether losses teach rather than merely punish.

metric finds magnitude; human picks the fix

The headline metric is not enough.

A game can hit its target and still feel miserable. These supporting signals reveal difficulty walls, false choices, and runaway economies that the final outcome hides.

Resource curves over time

A resource that collapses early marks a difficulty wall; one that grows forever reveals an economy with too many sources.

Failure and quit points

Plot the stage and cause for each ending. A sharp cluster marks a problem that an average run length can conceal.

Decision diversity

Track what the policy chooses by state. If experts always pick the same option, the other buttons may be decoration rather than decisions.

Agent skill gap

Compare random, simple heuristic, and stronger policies. The gap shows whether knowledge and adaptation are rewarded.

Run-length distribution

Do not report only the mean. Two peaks can mean the system has split into a short hopeless game and a long comfortable one.

Tail risk and breakpoints

Measure bad extremes, not only averages. Check integer thresholds where one more heal, dollar, or spare part changes the player-visible outcome.

Three rules worth keeping.

These survive every method change—from a spreadsheet model to a future optimizer.

Model first, simulate second.

If an action has structurally bad expected value, simulation can confirm the flaw but cannot explain it more cheaply than arithmetic.

Compare on the same luck.

Common random numbers make small before/after changes visible with fewer runs. This is the highest-return engineering change.

Optimize outcomes; protect decisions.

The machine can locate a target configuration. A human must still judge whether its tradeoffs are legible, tense, and fun.

!
Open question: seed streams need care.

If a parameter changes the number of random draws—such as event frequency—one shared stream can drift out of alignment. Separate streams for events, combat, rewards, and other subsystems are the safer design; the research did not find a JavaScript-specific implementation guide.

Sources & confidence.

Two references were read in full on September 26, 2026. The remaining links were supported by search-result indexing and should be treated as leads until their full text is checked.

Verified in full
  • Modeling & metrics reference — expected value, curve selection, sensitivity, breakpoints, uncertainty, and RNG traps.
  • TCG testing model — seeded runner, agent ladder, Wilson intervals, same-seed comparisons, and the automated-plus-human loop.
Practical tools and implementation leads
Methods and studio practice leads