Parents form theories constantly. The bedtime routine works. The bedtime routine doesn't work. Screen time is fine. Screen time is the problem. Sugar makes him wild. Sugar has nothing to do with it. Within any given week, a parent will hold and abandon a dozen such theories, each generated by a single salient incident, each elevated to certainty by the brain's pattern-matching machinery, each discarded the moment a contradicting incident arrives. None of this constitutes evidence. It is noise that feels like knowledge.
Tracking what works and what doesn't is the discipline of distinguishing between actual signal and the brain's compulsive narrative-making. It requires writing things down before you have a conclusion, recording outcomes rather than feelings about outcomes, and waiting long enough for the data to accumulate before deciding what it means. Most parents will not do this because the cost-benefit calculation in the moment seems absurd: why measure something so unmeasurable as a child's response to a small change in household practice?
The answer is that the alternative is worse. Without tracking, the parent is at the mercy of recency bias, confirmation bias, and the negativity bias that makes one bad night louder than ten good ones. The bedtime routine that "isn't working" has actually worked six nights out of seven, but the parent cannot see this because the seventh night was last night. Tracking surfaces base rates. Base rates are what separate parenting wisdom from parenting superstition.
The tracking does not need to be elaborate. A simple grid — date, intervention, outcome, anomalies — kept for a few weeks at a time when a question is genuinely open is enough. The point is not to instrument parenting as a quantified self project. The point is to introduce just enough friction between observation and conclusion that the conclusion has a chance of being grounded in something more than the most recent emotional impression.
The practice has a particular structure. First, isolate one variable. Not "is my child sleeping enough" but "does the 7:30 bath produce earlier sleep onset than the 8:00 bath." Second, define what counts as evidence in advance. A single late night does not falsify a working routine; three late nights in seven might. Third, run the experiment long enough to be informative. Two days is nothing; two weeks is something; two months is enough to detect anything that matters. Fourth, accept the result even when it contradicts your intuition. This is the hardest part.
The hardest part is hard because parents have ego invested in their theories. The mother who has built her identity around being the no-screens parent will not lightly accept evidence that an hour of structured screen time correlates with calmer afternoons. The father who has insisted on early bedtimes will not lightly accept evidence that this particular child does better with a later one. The tracking confronts the parent with their own attachment to being right, which is information about the parent more than about the child.
The asymmetry matters. Tracking what works is easier to sustain than tracking what doesn't, because the former produces success stories and the latter produces evidence of error. Both are necessary. A parent who only tracks what works will accumulate a portfolio of validated interventions and remain blind to the larger landscape of things that have failed and been silently abandoned without examination. The discipline of recording failures — and naming them as failures rather than as "we tried something different" — is what makes the tracking honest.
Tracking also reveals an uncomfortable truth: a great deal of what parents do has no measurable effect at all. The elaborate behavioral charts, the carefully designed reward systems, the specific phrasing of consequences — much of it produces results indistinguishable from doing nothing in particular. This is not because parenting doesn't matter. It is because the things that matter most operate on timescales too long to capture in weekly tracking, and the things that matter least are the ones that fit conveniently into a grid. The tracking, paradoxically, is most useful when it teaches the parent to stop tracking certain things — to recognize that an intervention's measurable effect is so small that the effort of measuring it is not worth what is gained.
What remains, after the discipline matures, is a small number of genuinely tracked variables that the parent has confidence about, surrounded by a wide field of practices that the parent has decided to do for reasons other than measured effectiveness — because they reflect values, or because they make the household more bearable, or because they are inherited from a grandparent and worth preserving for that alone. The tracking does not replace these reasons. It clarifies which reasons are being given.