Design Frameworks & Research


What Mother-Baby Talk Teaches Us About Interaction Sound

Stop agonizing over whether a sound will work cross-culturally.

 

Perception of sound varies widely across cultures. Yet when it comes to interaction sound design, we can draw on ancient, universal patterns of human intonation, alongside the global influence of gaming and entertainment, to create communication cues that resonate across cultural boundaries.

You don't need to research what an error sounds like in Seoul versus São Paulo. The shape travels. The texture doesn't. So lock the contour and localize the texture freely.

Humans are wired to understand melody before language. From birth we respond to pitch, rhythm, and contour, which makes prosody a powerful tool for intuitive, cross-cultural communication.

The research

One of the few things in human communication that looks genuinely universal starts with a single experiment. At Stanford in the 1980s, Anne Fernald recorded mothers speaking to their twelve-month-olds across five standardized situations, then electronically filtered the recordings to destroy the words and leave only the melody. She played the results to eighty adults. They could still tell what each mother was doing, approving, warning, comforting, calling for attention, from the pitch contour alone.

Her cross-language work found the same prosodic shifts in mothers and fathers across six language communities, including Japanese and, in related work, Mandarin, a tonal language, where you might expect pitch to be spoken for already. Later researchers pushed it further: adults among the Shuar in the Ecuadorian Amazon, with no exposure to English, correctly identified the intent in English infant-directed speech. A 2022 study across 21 societies and 51,000 listeners in 187 countries found the same pattern holds regardless of how distant the listener's language is from the speaker's.

The words change everywhere. The melody doesn't.

The five contours

Five communicative contexts Fernald studied, and the contour each tends to take:

Approval — Rises, then falls. Wide excursion, smooth.
Prohibition — Short, low, flat. Abrupt onset, no tail.
Attention — High and rising. Unresolved. Often repeated.
Comfort — Long, smooth, low. Slowly descending.
Play — Rhythmic, varied, high, irregular.

These are contexts, not published waveforms. Fernald defined five situations, recorded mothers in them, and measured the prosody that resulted. The shapes above are accurate characterizations drawn from her measurements and the literature that followed, not a figure reproduced from her paper.

"Play" is my shorthand. Her label for the fifth context is Game/Telephone, the telephone part being pretend phone talk recorded as a separate playful context. Play is the plain-language version.

A research group in Finland worked most of this out in 2011. Kai Tuuri, Tuomas Eerola and Antti Pirhonen published it properly, in a real HCI journal. They recorded people vocalizing communicative functions, extracted the pitch contours, built actual sounds from them, and tested whether users understood them.

It worked. Their research project was named Grammar of Earcons,

Then it went nowhere. Fifteen years, no design system, no platform, no uptake. Some of the best thinking in our field is just sitting there.

Research cited

Fernald, A. (1989). Intonation and communicative intent in mothers' speech to infants: Is the melody the message? Child Development, 60(6), 1497–1510. The filtered-speech study. doi:10.2307/1130938

Fernald, A., Taeschner, T., Dunn, J., Papoušek, M., de Boysson-Bardies, B., & Fukui, I. (1989). A cross-language study of prosodic modifications in mothers' and fathers' speech to preverbal infants. Journal of Child Language, 16(3), 477–501. The six language communities. doi:10.1017/S0305000900010679

Grieser, D. L., & Kuhl, P. K. (1988). Maternal speech to infants in a tonal language: Support for universal prosodic features in motherese. Developmental Psychology, 24(1), 14–20. The Mandarin study. doi:10.1037/0012-1649.24.1.14

Bryant, G. A., & Barrett, H. C. (2007). Recognizing intentions in infant-directed speech: Evidence for universals. Psychological Science, 18(8), 746–751. The Shuar study. doi:10.1111/j.1467-9280.2007.01970.x

Hilton, C. B., Moser, C. J., et al. (2022). Acoustic regularities in infant-directed speech and song across cultures. Nature Human Behaviour, 6(11), 1545–1556. 21 societies, 51,065 listeners, 187 countries. doi:10.1038/s41562-022-01410-x

Tuuri, K., Eerola, T., & Pirhonen, A. (2011). Design and evaluation of prosody-based non-speech audio feedback for physical training application. International Journal of Human-Computer Studies, 69(11), 741–757. The Finnish work. doi:10.1016/j.ijhcs.2011.06.004