Establish the baseline
Begin with an arrangement that can be reproduced. Observe its visible state and feedback before changing it, and keep the surrounding conditions as consistent as practical. A baseline is useful because every later comparison can point back to a known configuration rather than to memory or a collection of simultaneous edits. Choose one visible variable for the next test. If several variables must change together, record the group and do not claim that one member caused the result. Restore the baseline when the comparison becomes unclear.
Observe, compare, and repeat
Run the baseline and changed version through the same available observation when possible. Record what the game actually displays before deciding what it means. Keep observation and conclusion separate: a different response is observed evidence, while an explanation of the hidden cause remains unverified unless stronger support exists. Repeat the comparison when the result matters. Consistent observations can guide the next adjustment, while conflicting results should remain inconclusive or awaiting verification. A short mental record can work, but a dated comparison is safer when updates or several sessions are involved.
Interpret a consistent result
A result is consistent when the same comparison produces the same visible pattern under sufficiently similar conditions. Treat that pattern as observed evidence tied to the tested arrangement and current version. It can justify a focused refinement or another confirmation test, but it does not establish an exact hidden value or a universal mechanic beyond the conditions recorded.
Interpret a conflicting or inconclusive result
Results conflict when comparable attempts show different visible outcomes. A test is inconclusive when several conditions changed, the baseline was lost, or the game provided no clear signal. In either case, do not choose a preferred explanation. Restore the known setup, narrow the comparison, and retest only the variable connected to the original question.
Interpret a failed test
A failed test is one that cannot perform the intended comparison—for example, the setup was not preserved or the observation could not be completed. Record why it failed so the same design is not repeated. Failure provides no evidence for or against the suspected mechanic; it calls for a corrected method, a simpler baseline, or a return to setup.
When testing changes direction
Return to Sound Arrangement when testing identifies a repeatable weak area or a change that needs refinement. Return to the baseline when no specific cause can be isolated. Move toward battle preparation only when the arrangement is stable enough that an outcome can be compared with a known starting point. Testing cannot prove an exact score, hidden component weight, universal best setup, or permanent mechanic without verified evidence. Its practical value is narrower: it shows what changed under recorded conditions and helps the player choose a controlled next step. Keep failed and inconclusive comparisons as part of the record, because they prevent the same unclear test from being repeated and show where current evidence ends. When the test question itself is too broad, rewrite it around one visible behavior before continuing.