There’s a shelf in my office I inherited from a mentor who retired ten years ago. It holds baby books dating from the 1940s through the 1970s—none of them mine. They belonged to mothers I never met, in towns I’ve never visited. One was kept by a woman in rural Oregon who noted, in pencil, that her son “stood alone for eleven seconds before falling sideways, laughing” at nine months and three days. Another, from a mother in Baltimore in 1952, records feeding patterns in a column right alongside weather observations: “Oatmeal refused. Rain again. Cut first tooth—bottom left.” A third tracks sleep with the precision of a lab technician, but interrupts the data to note that the baby “cried differently tonight, not hungry, not wet, just lonely I think.”
These are not sentimental artifacts. They are developmental data, collected by the people best positioned to collect it—the ones present at bedtime, at meals, in the quiet hours when nothing seemed to be happening but was. And they represent a tradition we have largely forgotten: caregivers as systematic observers of their own children’s development, documenting with a rigor and specificity that our modern screening tools, for all their standardization, cannot replicate.
The Tradition That Predates the Clinic
Long before developmental pediatrics existed as a specialty, mothers kept what we would now call longitudinal observational records. The Victorian-era “baby biographies” written by mothers like Milicent Shinn in the 1890s were not casual diaries. Shinn, who studied her niece’s development from birth, produced notes detailed enough that G. Stanley Hall—then the president of the American Psychological Association—cited her work in his own publications. Shinn recorded not just when her niece reached milestones but how: the failed attempts, the regressions, the context. She noted that the baby reached for objects with both hands equally until the eleventh week, then briefly preferred the left, then returned to bilateral reaching. That kind of sequential detail—the wavering before a preference consolidates—does not appear in any standardized screening tool I have used in clinical practice.
The tradition continued through mid-century public health campaigns that distributed baby books to new mothers. In Britain, the NHS issued child health records that combined growth charts with open pages for narrative notes. In the United States, well-baby clinics often gave mothers printed booklets with spaces for both structured data—weight, immunizations—and unstructured observation. The assumption was that mothers would, and could, serve as primary documentarians. They were not given a role. They already had one. The booklets formalized it.
What strikes me about these records is not their completeness—many have gaps, sometimes months long—but their quality of attention. The Oregon mother’s note about eleven seconds of standing tells you more than a checkbox reading “stands alone: yes.” It tells you the infant was experimenting with balance, found it amusing rather than frightening, and had the motor control to stand but not yet to recover from disequilibrium. The Baltimore mother’s pairing of food refusal with weather suggests she was tracking variables, looking for patterns, doing exactly what a researcher does when building a case. She may have been wrong about the connection—weather and teething may have had nothing to do with the oatmeal refusal—but her method was sound. She was observing in context.
What Standardized Screening Gains and Loses
I am not arguing against screening tools. The Ages and Stages Questionnaire, the M-CHAT, the Denver II—these instruments exist because pediatricians needed a way to detect developmental delays reliably, quickly, and across diverse populations. Before standardized screening, detection depended on clinician experience, which varied enormously. A 1992 study published in Pediatrics found that pediatricians identified fewer than 30 percent of children with significant developmental delays before age three when relying on clinical impression alone. Structured screening improved that detection rate substantially. That is not a small thing. Children who would have been missed were found, and early intervention services, where available, changed trajectories.
But the gain in detection came with a quiet loss. Screening tools ask whether a child performs a behavior. They do not ask how, or in what context, or what happened the day before. The ASQ asks if a child stacks three blocks. It does not ask whether the child stacked them once, triumphantly, and then refused to try again for a week. It does not ask whether the child stacked them while humming, or while a sibling was watching, or while tired. The M-CHAT asks if a child responds to their name. It does not ask if the child responds differently depending on who is calling, or what they were doing when called, or whether the tone of voice was familiar.
The clinical form compresses observation into a binary. The caregiver’s notebook expanded it into a story. Both contain information. But the story contains information the binary cannot hold, and that information is sometimes the difference between a child who is delayed and a child who is bored, or frightened, or being raised in a household where the primary language differs from the screening tool’s language.
I think about this when I review screening results in clinic. A mother once told me, during a well-child visit, that her daughter had failed the ASQ’s communication section. The form said she was at risk. The mother was not dismissive—she had brought the form in, completed it honestly—but she wanted me to know something the form could not record. “She talks all day at home,” the mother said. “But only to me. If anyone else is in the room, she stops. She’s not delayed. She’s selective.” The screening tool could not distinguish between a child who does not speak and a child who does not speak to strangers. The mother’s observation could. And her observation, had it been recorded in the chart with the weight it deserved, might have prevented an unnecessary referral and the months of anxiety that followed.
The Medium Shapes the Record
There is a principle in research methodology that the instrument determines the data. Survey people with multiple-choice questions and you get multiple-choice answers. Interview them with open-ended questions and you get narratives. Neither is inherently better, but they are fundamentally different kinds of information, and pretending they are interchangeable leads to errors.
The same principle applies to developmental documentation. A baby book with blank pages invites narrative. A screening form with checkboxes invites binaries. An app that sends push notifications asking “Did your baby smile today?” invites a different kind of attention than a notebook where a mother writes, at midnight, “Smiled at the ceiling fan for the first time. Watched it for twenty minutes. I think she sees the shadows.”
The medium shapes not just what gets recorded but what gets noticed. A checkbox asks for confirmation of an expected behavior. A blank page invites observation of the unexpected. The Victorian mothers were not confirming milestones. They were discovering them, sometimes naming them for the first time. Shinn’s observation that her niece’s babbling changed in pitch when different adults entered the room is not a milestone. It is an insight. No screening tool I know of captures it.
This is where the question of documentation tools becomes practical rather than nostalgic. The caregivers I work with are not Victorians with leisure time and a scientific bent. They are working parents, often stretched, often documenting in fragments—on phone notes, in the margins of appointment cards, in texts to partners (“she said ‘mama’ today at lunch, first time, I almost cried”). The medium they use shapes what they preserve. A phone note captures the moment but not the context. A text captures the emotion but not the sequence. A checklist captures the milestone but not the meaning.
What is needed is something that supports sustained, reflective recording—not fragmented data entry, not checkbox compliance, but the kind of documentation that lets a caregiver look back over weeks and see patterns they could not see in the daily blur. The principle is the same across domains: the tool should amplify human observation, not standardize it away. A screening tool produces a generic summary of a child’s development, averaged across populations. A caregiver’s narrative produces a specific, irreplaceable record of one child, in one family, in one context. The Authors Guild’s guidance on preserving original voice in documentation makes a related point about the irreplaceable core of human authorship—original thinking and a unique perspective that automated outputs cannot replicate. The parallel to developmental documentation is direct: the caregiver’s observational voice constitutes the irreplaceable core of meaningful developmental tracking, and any system that flattens that voice into generic data entry loses something clinically valuable.
When Narrative Outperforms Data
I want to be specific about what narrative documentation captures that structured screening does not, because this is not an abstract argument. It has clinical consequences.
A colleague once described a case in which a child’s autism diagnosis was delayed by eighteen months because the screening tool kept producing borderline scores. The mother, however, had been keeping notes in a notebook—detailed, dated entries about her son’s play, his responses to sounds, his eye contact during meals, the way he lined up toy cars but only on Tuesdays after preschool. When my colleague read the notebook, the pattern was unmistakable. The screening tool had been averaging across contexts. The mother had been documenting within them. The child’s behavior varied dramatically by setting—he was more socially engaged at home, less so in unfamiliar environments—and that variability, which the screening tool smoothed into a borderline score, was itself the diagnostic signal.
This is not an argument that mothers are better diagnosticians than screening tools. It is an argument that narrative data and structured data measure different things, and that clinical practice benefits from both. The screening tool provides a population-level comparison. The narrative provides individual-level context. A clinician who has both can make a better-informed decision than one who has only the score.
The problem is that our system increasingly values only the score. Electronic health records are designed for structured data entry. Billing codes require standardized responses. Time constraints in well-child visits—fifteen minutes, often less—reward efficiency over listening. The narrative observation, when it appears at all, gets squeezed into a free-text field that no one reads and that cannot be searched or analyzed. We have built a system that is excellent at detecting gross deviations from norms and poor at understanding the child who deviates in ways the norms did not anticipate.
The Mothers Who Were Researchers
One of the most cited studies in developmental psychology is Mary Ainsworth’s work on attachment in Uganda, published in 1967. Ainsworth observed twenty-eight mother-infant pairs in their homes over nine months, producing detailed narrative records of feeding, crying, holding, and response. Her observational methodology—the Strange Situation—grew directly from those field notes. What is less often noted is that Ainsworth’s Ugandan study was itself modeled on the baby-diary tradition. She had read Shinn. She had read the British mother-infant observation literature. She understood that sustained, close observation in natural settings produced data that laboratory measures could not.
The mothers in Ainsworth’s study were not passive subjects. They observed their own children with the same attentiveness Ainsworth brought, and they told her things she would not have seen: that a baby’s cry at night meant something different during the harvest season, that a particular child had been more clingy since a neighbor’s death, that feeding patterns changed when the older sibling started school. Ainsworth valued these observations and incorporated them into her analysis. She treated the mothers as informants, not just subjects.
Ainsworth’s methodology did not stop at observation. She developed a classification system—secure, anxious-avoidant, anxious-resistant, disorganized—that was grounded in the specific behavioral sequences she had documented. Each category emerged from patterns in the narrative data, not from a predetermined checklist. She was, in effect, building a diagnostic framework from the kind of close, contextual recording that the baby-diary tradition had established. The framework proved durable enough to anchor decades of subsequent research, but its origins lie in the kind of observational detail that our current systems struggle to accommodate. We use the Strange Situation’s categories in clinical training. We rarely teach the method that produced them: sustained, narrative observation in natural settings, with caregivers treated as co-documentarians.
This is the tradition we have partially lost. Not the science of observation—that continues in research settings—but the everyday practice of caregivers as systematic documentarians of their own children’s development. The baby books on my shelf are evidence of a time when this practice was expected, supported, and valued. The blank pages in the NHS booklets were an invitation. The prompts in mid-century American baby books asked open questions: “What new sound did your baby make this month?” “What surprised you?” Those questions assumed that the mother had been paying attention and that her attention was worth recording.
What We Could Recover
I am not suggesting a return to the 1890s, or even the 1950s. The conditions that made the baby-diary tradition possible—longer postpartum stays, fewer dual-income households, different expectations of maternal labor—do not describe most families’ lives today, and I would not want to pretend they did. The caregivers I see in clinic are documenting under circumstances the Victorian mothers never faced: between shifts, during commutes, in shared custody arrangements, across languages and cultures and technologies.
But the principle underlying the tradition is portable. It is this: the person who spends the most time with a child sees things no one else sees, and if given a medium that supports rather than constrains that seeing, will produce a record more nuanced than any screening tool can generate. The question is whether we design systems that invite that record or systems that suppress it.
Right now the answer is mixed. Developmental milestone apps, which I have written about elsewhere, tend to gamify tracking in ways that produce anxiety and binary data. They ask “Is your baby crawling?” not “How does your baby move across the floor?” They reward completion, not reflection. They are designed for engagement metrics, not developmental insight.
But there are counterexamples. Some community health programs still use paper journals, distributed during home visits, that combine growth charts with open pages for narrative notes. The Purdue OWL’s creative writing guide treats observation and voice as foundational skills, not optional flourishes, and the parallel holds: a caregiver’s developmental notebook is a form of creative nonfiction, requiring the same attention to detail, the same commitment to sustained recording, the same willingness to notice what was not expected.
The Small Question Behind the Big Data
The question for those of us who work in child health is whether we are building systems that honor that kind of noticing or systems that replace it with something easier to bill for. The answer will determine not just what we know about children’s development, but whether the people who know children best are treated as informants in their own children’s care—or as data entry clerks for someone else’s screening tool. For clinicians and researchers interested in documentation tools that support rather than replace human observation, resources like the AI novel writing app that supports structured drafting without flattening voice illustrate how technology can serve sustained narrative practice—a principle that applies whether the narrative is a developmental journal or a clinical observation record.
Recent Comments