Brytalearn.How things workFind something
INTERACTIVE EXPLANATION

Can the same pattern tell two different causal stories?

Investigate a virtual garden. Change who receives an assignment card, uncover a hidden starting difference, run fair trials and discover why a pattern alone cannot settle what caused it.

Enable JavaScript to change the conditions and run the interactive experiment.

Make a discovery

Who already received an assignment can differ from who would receive it after a fair shuffle. A causal explanation needs a defensible account of assignment and response; a chart alone does not supply that account.

  • Distinguish observing an assignment from intervening on it.
  • Identify a common cause of assignment and outcome.
  • Explain why a larger observational sample does not necessarily identify an effect.
  • Keep counts and denominators when comparing within baseline categories.
  • Separate random assignment from exact baseline balance and representative sampling.
  • Distinguish population effects, finite-cohort paired truth and sample estimates.
  • Interpret a sampling distribution without turning it into a probability a hypothesis is true.
  • Retain every trial result and state the assumptions a real application would need.

Make a prediction

Two response rulebooks have exactly the same observed W/Y population table. Will more W/Y observations alone tell you which causal effect is right?

  • Yes, any sufficiently large chart proves the cause
  • No; the missing causal information is not supplied by sample size
  • Yes, choose the rulebook with a nicer animation
Read the explanation

Precision about the same observational distribution does not choose between the two mechanisms. Additional measurements, assumptions or a suitable intervention are needed.

Understand it

Make a prediction from the pattern

In the opening recipe, 30 of 50 baseline-only plots meet a target, compared with 25 of 50 extra-water plots. Those are exact constructed proportions. Decide what the chart establishes before revealing two different response stories.

Follow the assignment card

Open the garden. W0 means baseline care; W1 means an extra-watering assignment. The card is a variable you can change. Each bed keeps its own stored baseline category and response digit when you change assignment.

Reveal a harder starting point

The default existing policy gives extra-water cards to drier plots more often. Those plots also have a harder baseline. Open the soil layer without redrawing the city. This common cause can make a beneficial rule look harmful in the pooled comparison.

Replace the assignment rule

Shuffle equal numbers of W0 and W1 cards independently of dryness. The baseline-to-assignment link is removed. Dryness still affects success, and the stipulated response effect still operates.

Compare filtering with setting

Showing only plots that already have W1 retains a selected subset. Setting every plot to W1 retains the whole original cohort and replaces assignment. Compare the row counts and baseline composition side by side.

Keep the denominators

Split the comparison into dry and less-dry categories, then average each arm’s category rates using the same 50/50 target mix. An empty sample cell has no rate; this calculation stays unavailable rather than guessing a value.

Repeat, then explain the variation

Generate fresh cities and independently shuffle their assignments. Small estimates vary even when the model effect is known. Keep all repetitions, compare them with the exact sampling distribution, and distinguish sampling more plots from running more repetitions.

Look closer at the science

Correlation and risk difference are different summaries

For binary W and Y, the workbench can calculate their correlation and the difference between success rates in the two observed arms. Neither summary alone specifies the effect of replacing an assignment mechanism. A zero-variance variable makes correlation undefined.

The original structural causal model

C marks baseline dry-soil category, W marks the assignment, and Y marks a single binary endpoint. Independent population digits RC, RW and RY are uniform on 0 through 9. In A, C=1[RC<5], W=1[RW<5+k(2C−1)], and Y=1[RY<7−5C+tW]. These are authored mathematical rules, not measured plant biology.

Bounded rules need no probability clipping

A permits t=0,1,2 and k=0,1,2,3,4. The within-category benefit is t/10; k controls how strongly the existing policy favors dry plots. Assignment probabilities remain between 0.1 and 0.9, so both assignments are possible in both categories in the population.

An observational reversal

With t=2,k=3, the extra-water population is 80% dry while the baseline-only population is 20% dry. Both within-category intervention differences are +20 percentage points, yet observed rates are 50% versus 60%. This reversal is related to Simpson’s paradox; the causal graph determines which comparison answers the intervention question.

Exactly the same W/Y law, opposite causal effects

Rulebook B retains the default assignment policy but uses Y=1[RY<6−W], with no direct C-to-Y effect. Both worlds have joint probabilities W0Y0=.20, W0Y1=.30, W1Y0=.25, W1Y1=.25. A has intervention risks .45 and .65; B has .60 and .50. These are equal population observational laws, not a claim that independent finite samples have identical counts.

Identification is not precision

More observations of only W and Y estimate their common distribution more precisely without selecting between A and B. Measuring the baseline variable or performing the stated intervention can distinguish these particular worlds in principle. This does not mean observation is useless: appropriate assumptions, measured variables and study design can support causal identification.

Conditioning versus intervention

P(Y|W=1) refers to the existing W1 group. P(Y|do(W=1)) replaces the assignment equation and keeps the full population’s background distribution. The simulator implements the latter by setting W while retaining each stored C and RY. It does not manipulate Y directly to force a preferred answer.

The declared graph justifies this adjustment

In default A, C causes both W and Y, creating a back-door path. Standardizing category-specific means to a common 50/50 city composition recovers the intervention contrast at the population level. This depends on the stipulated sufficient baseline variable and positivity. It is not permission to adjust for every variable in an arbitrary dataset.

A population possibility is not a sample observation

Positive assignment probability in each category does not guarantee all cells appear in a sample of 20. The default observational cohort has no dry baseline-only plot. Its sample cell mean and cell-mean standardized difference are unavailable. Zero rows are not a 0% success rate.

Random allocation does not guarantee a balanced realized sample

The balanced shuffle assigns exactly N/2 plots to each arm. In the reproducible default 20-plot city, randomized arms contain 6 dry plots versus 2. That chance imbalance does not prove the randomizer is broken. Outcome comparison must respect the planned allocation and sampling procedure.

Random assignment and random sampling do different jobs

Assignment governs which response is observed for each sampled unit. Sampling governs which units enter the study. Random assignment alone does not establish representativeness, transportability, absence of missing outcomes or adherence to the assigned action.

One unit, two stipulated response boxes

The oracle computes Y(0) and Y(1) using the same unit’s C and RY. This coupling defines individual response types inside the simulator. A real experiment ordinarily observes one outcome under one assignment in the same circumstances. An average effect by itself does not identify each individual’s response.

Three quantities deserve three labels

The known population effect averages over the declared probability law. The finite-cohort oracle averages paired Y(1)−Y(0) over the stored plots. The observed arm-rate difference compares different realized groups. These can differ in a finite sample without any calculation being wrong.

A zero effect can have a nonzero estimate

At t=0 in A, every unit’s paired outcomes agree, so both population and finite-cohort causal effects are zero. Different units still fall into the two randomized arms; a realized comparison can be nonzero. Under the confounded existing policy at k=3, the population observational contrast remains −30 percentage points.

An exact fresh-city distribution

Under default A, fresh randomized arm successes are independent Binomial(n,.65) and Binomial(n,.45), where n=N/2. Their difference divided by n has mean .20 and variance .475/n. The chart calculates every possible mass by convolution. The shown central sampling range uses discrete 2.5% and 97.5% quantiles; its actual included mass need not be exactly 95%.

A surprising result under a known positive effect

For 20 fresh plots, the standard deviation is about 21.79 percentage points and a zero or negative estimate occurs about 24.35% of the time. At 400 plots, standard deviation falls to about 4.87 points. Those are probabilities under a specified positive-effect model, not p-values or probabilities that the effect is zero.

A fixed paper deck has a different uncertainty law

The 20-card activity contains 9 always-success, 4 helped and 7 never-success cards under default A. Exactly 184,756 balanced assignments are possible. Their mean difference is .20 and probability of a nonpositive difference is about .23107. This differs from the fresh-city value .24353 because the response-type composition is fixed.

Intervals need a named target

The optional Wilson intervals estimate individual binomial arm rates under their sampling assumptions. They are not an interval for the difference, and overlap is not a specified difference test. Approximate 95% confidence describes repeated coverage of a method, not a 95% posterior probability attached to one fixed unknown rate.

Reproducible does not mean observed in nature

The versioned SHA-256/rejection generator uses seed, cohort, unit ID, variable stream and draw index. An independent Fisher–Yates shuffle supplies balanced assignments. Filters and reveal controls leave the data unchanged. Changing sample size rebuilds that size’s allocation; it is not appending plots with all former cards guaranteed unchanged.

Where this is used

Read a surprising headline

Ask what was measured, who received the condition, what preceded it, and which causal assumptions support the conclusion. A careful observational statement can be valuable without pretending to settle the intervention.

Plan an A/B comparison

Define the outcome and comparison before inspecting results. Record assignment, baseline information, all outcomes and all attempted trials. A randomized comparison still needs appropriate sampling, measurement and operational assumptions.

Evaluate an educational model

Check whether an arrow is a declared rule or a conclusion from evidence. The garden helps test causal reasoning; its plant-like appearance contributes no independent biological evidence.

Transfer the idea to AI

A feature that predicts a label can be a shortcut rather than a cause. A model’s attention weight, classification score and response to a specified intervention answer different questions.

Try it yourself: Shuffle a garden you can hold

Supplies

  • Twenty paper plot cards
  • Twenty paper assignment labels: ten W0 and ten W1
  • A pencil and paper for all results
  • An envelope or cup for shuffling
  • Paper flaps to cover response boxes
  1. Prepare the paired response deck

    Use the downloadable key to make ten less-dry cards and ten dry cards. Each has a response-digit label 0–9 and covered Y0/Y1 boxes. This fixes nine always-success, four helped and seven never-success cards under the stated rulebook.

  2. Assign before looking

    Mix ten baseline-only and ten extra-watering labels face down. Give one to each plot card without looking at response boxes. Keep each label in place; do not choose assignments after seeing success.

  3. Open the assigned box

    For a W0 card reveal only Y0; for a W1 card reveal only Y1. Keep the other box covered. A 1 means the synthetic target is met and a 0 means it is not.

  4. Write the counts

    Record successes, total cards and dry cards in each arm. Calculate extra-water successes/10 minus baseline successes/10. Keep this exact allocation even if the result surprises you.

  5. Audit the paired truth

    After recording the result, uncover both boxes. The whole fixed deck has 9 baseline successes and 13 extra-water successes: an average paired difference of 4/20. Explain why this audit is available for authored response cards but not ordinarily for real units.

  6. Repeat without hiding results

    Cover the boxes, reshuffle all labels and record five more complete allocations. Keep every result. These are new assignments of the same deck, not fresh city samples. Compare your results with the exact fixed-deck table, then explain the difference.

Can a result be surprising without proving that the randomizer or the known rulebook is wrong?

Use fictional paper cards only. No plant manipulation, personal data, health intervention or powered equipment is needed. The online deck is a reproducible companion; a physical shuffle will normally produce different allocations.

Check your understanding

What do the opening 60% and 50% observed rates establish by themselves?

  • Extra watering causes harm
  • The existing W1 group has a lower observed success rate
  • The two groups have equal starting conditions
Answer and explanation

The existing W1 group has a lower observed success rate The chart describes an association. The assignment and response story is additional causal information.

Which link disappears when the dryness-based policy is replaced by an independent shuffle?

  • C → W, baseline to assignment
  • C → Y, baseline to response
  • W → Y, assignment to response
Answer and explanation

C → W, baseline to assignment Randomization replaces assignment. It does not erase baseline influences on the outcome or change the response rule.

Must filtering W1 rows and assigning every plot W1 keep the same baseline composition?

  • Yes, both mention W1
  • No; filtering selects existing rows, while assignment retains the whole cohort
  • Yes, because the plot count is unimportant
Answer and explanation

No; filtering selects existing rows, while assignment retains the whole cohort In default A the population W1 subset is 80% dry, while the full population retained by intervention is 50% dry.

Ten plots per arm, but 6 dry plots versus 2: does that prove the shuffle failed?

  • Yes, equal arm sizes force equal soil counts
  • No; chance imbalance can occur under a valid balanced allocation
  • Yes, discard this result and try again
Answer and explanation

No; chance imbalance can occur under a valid balanced allocation Equal numbers of labels do not force equal baseline composition. Keep the outcome and interpret the stated design.

No dry plot received baseline-only care. What is that cell’s success rate?

  • 0%
  • The treated dry rate
  • Unavailable
Answer and explanation

Unavailable There is no denominator. The population rule does not create observations in an empty sample cell.

At t=0, must every small randomized trial show exactly equal success rates?

  • No; different units can still produce a finite estimate
  • Yes, every estimate is the population effect
  • Only if we hide the dry plots
Answer and explanation

No; different units can still produce a finite estimate Paired outcomes agree within each unit, but different units enter the two arms. The population expectation and a realized comparison differ.

A nonpositive estimate occurs about 24.35% of the time under the specified +20-point model. That means…

  • A 24.35% probability that the true effect is zero
  • A sampling probability conditional on that positive-effect model
  • A mandatory p-value for any garden
Answer and explanation

A sampling probability conditional on that positive-effect model The calculation assumes the stated model and sampling design. It is not a probability a hypothesis is true.

In A, t=1 and k=2 give which population differences?

  • Observed −10 points; intervention +10 points
  • Observed +10 points; intervention −10 points
  • Both zero because assignment is imperfect
Answer and explanation

Observed −10 points; intervention +10 points The observational contrast is t/10−k/10, while the intervention contrast is t/10. Shared runoff would violate the separate no-interference assumption rather than merely change this answer.

Sources and model limits

  • Every city, digit, coefficient and outcome here is synthetic. No actual gardening, health or policy effectiveness is inferred.
  • The endpoint is binary. Geometry and card motion are an inspection sequence, not a continuous growth model, time scale or calibrated water dose.
  • No runoff, interference between units, clustering, missingness, measurement error or nonadherence is represented.
  • The graph is stipulated, not learned from a correlation coefficient or automatically identified from arbitrary uploaded data.
  • Standardization is justified for the declared baseline variable and target mix; sample empty cells remain unavailable.
  • Fresh-city sampling and balanced reassignment of a fixed paper deck have separately calculated uncertainty laws.
  • Simulator-only paired truth is not information a real randomized trial ordinarily provides for every unit.
  • Source research and numerical verification do not certify classroom comprehension, specialist review, accessibility or device/video-export behavior.

Association, intervention, adjustment and counterfactuals

Pearl’s 2009 overview, sections 2.1, 3.2.1, 3.3.1 and 3.4. Defines the methodological distinctions used here; it does not validate the invented gardening probabilities.

Pearl · Causal inference in statistics

Completely randomized allocation and paper-label procedure

NIST/SEMATECH section 5.3.3.1. Supports randomized design, not guaranteed realized baseline equality.

NIST · Completely randomized designs

Fixed-ticket randomization and scope of inference

Stark’s Stat 240 chapter 3. Motivates keeping fixed-card allocation uncertainty separate from fresh-population sampling.

Stark · Randomization models

Fresh-city binomial probability masses and variance

NIST/SEMATECH binomial distribution formulas. All chosen probabilities and numerical data belong to our synthetic rulebook.

NIST · Binomial distribution

Repeated-use interpretation of confidence

NIST explains coverage interpretation. Its mean-specific interval formula is not substituted for the binary Wilson calculation here.

NIST · Interpreting confidence limits

Independent subject review is pending.

Read the sources and model assumptions