Methodology
All analysis is performed by the Children's Media Analysis Toolkit (CMAT), an open-source Python application. Every result in this index is reproducible from the source video files using the parameters documented here.
Sampling protocol
The following parameters are held constant across all shows to ensure cross-show comparability:
| Frame sample rate | 2 fps |
| FFC configuration | General / All Ages |
| Flashing threshold | 0.1 (luminance delta, 0–1 scale) |
| Episode sample seed | 42 |
Normalization ceilings
Each component metric is rescaled to 0–1 against a fixed ceiling before it is
weighted, using (value − min) / (max − min), clamped into [0, 1]. A
ceiling is the denominator that makes six differently-shaped numbers
addable — it is not a threshold, a limit, or a judgement about the content.
Values at or above a ceiling all score 1.0 and are not distinguished from each
other. The composite is not reproducible without these, so they are stated
here:
| Cuts per minute | 0 – 45 |
| Colour saturation (mean) | 0 – 0.85 |
| Colour contrast (mean) | 0 – 0.35 |
| Motion (mean) | 0 – 0.35 |
| Flashing events per minute | 0 – 40 |
| Audio RMS (mean) | 0 – 0.35 |
These values are a scaling convention, not a validated threshold. They were set to span the range observed across the analysed corpus rather than to each metric's theoretical maximum, and were last revised on 2026-08-14. Because ceilings differ in how much of their range real content occupies, a metric's nominal weight is not the same as its effective contribution to the composite — the component figures are published for every episode so that both can be checked. Composite scores computed before 2026-08-14 are on an earlier scale and are not comparable with these.
For shows with fewer than 15 episodes, all episodes are analyzed. For shows with 15–60 episodes, a spread sample of 10 is drawn. For shows with more than 60 episodes, a spread sample of 20 is drawn. "Spread" is CMAT's own name for a design, not a standard statistical term: the ordered run is split into n contiguous, near-equal chunks and one episode is drawn uniformly at random from each — stratified random sampling with equal-sized positional strata, one unit per stratum. Inclusion probabilities are 1/chunk-size and are not equal across the run when the episode count does not divide evenly, so any estimate assuming equal probabilities will be biased. Each draw records its seed and its full candidate frame.
For long-running shows (20+ years or a significant production format change), the show is divided into named eras and each era is sampled independently.
Metric definitions
| Metric | Method | Unit |
|---|---|---|
| Shot-boundary rate (labelled "scene pacing") | PySceneDetect ContentDetector at threshold 27 → boundaries per minute. Detected boundaries, not semantic scene changes: a cut inside one continuous scene counts. Frame-differencing misses gradual transitions by construction. | boundaries / min |
| Color saturation | Mean HSV S-channel per frame, averaged across frames sampled at 2 fps | 0–1 |
| Color contrast | Spatial standard deviation of the HSV V-channel within each frame, averaged across sampled frames. Within-frame brightness spread, not a perceptual contrast metric. | 0–1 |
| Motion | Mean absolute grayscale difference between consecutive sampled frames, rescaled to 0–1. Measures image change, not depicted movement — a cut, a camera pan and a running character all raise it — and depends on the 2 fps sampling rate. | 0–1 |
| Flashing | Consecutive frames sampled at 10 fps whose whole-frame mean luminance differs by more than 0.1. Not a photosensitivity safety assessment; see below. | events / min |
| Audio intensity | Mean per-second RMS amplitude on the track downmixed to mono and resampled to 8 kHz. Linear amplitude, not perceptual loudness — not LUFS, not EBU R128. Two files mastered to the same broadcast loudness can differ here. | linear 0–1 |
| Formal-Feature Composite (FFC) | Configurable composite of observable audio-visual production and editing features | 0–1 |
Formal-Feature Composite (FFC)
The FFC is a configurable weighted sum of normalized sub-metrics. Normalization uses fixed reference ranges — not per-corpus normalization — so scores remain comparable across separate analysis runs and future additions to the index. The "General / All Ages" configuration is used for all index entries. The FFC is a summary of the stimulus. It is not a validated measure of anything happening in a viewer — not arousal, not attention, not developmental impact. Full weight and normalization configurations are available in the CMAT repository.
Language metrics
When closed-caption or subtitle files are available for an episode, CMAT extracts three language complexity measures from the dialogue text. These numbers describe how the show uses language — not how well children will understand it. Comprehension depends on the individual child's age, vocabulary, and language experience.
| Metric | What it measures | How to read it |
|---|---|---|
| Mean WPM (Words Per Minute) |
How quickly dialogue is spoken, averaged across the episode. Calculated from the timing information in the subtitle file. | Lower numbers mean slower, more deliberate speech. A typical children's show falls in the 80–120 WPM range; adult conversation is typically 130–160 WPM. Example: 95.7 WPM is a relaxed, easy-to-follow pace. |
| Mean TTR (Type-Token Ratio) |
Vocabulary variety — the ratio of unique words (types) to total words spoken (tokens). Computed per episode then averaged. | A TTR of 1.0 would mean every word is different; a TTR near 0 means the same words repeat constantly. Children's shows typically range from 0.20 to 0.45. Example: a TTR of 0.277 indicates moderate repetition — new words appear regularly alongside familiar core vocabulary. |
| Mean Lexical Density | The share of content words — nouns, verbs, adjectives, adverbs — out of all words spoken. Function words ("the," "is," "and") are not counted. | Higher density means more information packed into each sentence. Simple conversational speech sits around 0.40–0.50; information-rich narration can reach 0.65+. Example: a lexical density of 0.585 means roughly 3 in every 5 words carry direct meaning — moderately dense for children's content. |
Language metrics require a subtitle or caption file (.srt or .vtt) for each episode. Episodes without captions show "Not available" for these fields. The metrics are derived from the caption text, which may differ slightly from the audio (simplified captioning, timing gaps, etc.).
Fantastical events (hand-coded)
Recent research has converged on fantastical content — rather than editing pace alone — as the program feature most consistently associated with short-term reductions in young children's executive function after viewing (Hinten, Scarf & Imuta, 2025, Developmental Science meta-analysis; Lillard et al., 2015). A fantastical event is a discrete impossible occurrence: a character flying unaided, teleporting, transforming, an object coming to life. Following the literature, we report fantastical events per minute.
Fantasy is a semantic judgment — whether an event violates how the world works — and cannot be computed from pixels. These metrics are therefore hand-coded using CMAT's structured event codebook: a seven-category event taxonomy (physical violations, transformations, continuity violations, impossible body events, object animacy, impossible causation, other), an explicit premise-vs-event rule (a standing impossible premise such as talking animals is not counted as events; discrete impossible occurrences are), and per-event properties (narrative relevance; repetition). The full codebook ships with the open-source toolkit.
Honest limitations: coded samples are small (per-episode windows chosen before viewing and documented alongside each show's numbers); coding is currently by a single coder (the toolkit supports two-coder inter-rater reliability, planned); the codebook is versioned and any rule change is logged. Event rates describe the stimulus, not any child's response, and are associated with — not proof of — effects on viewers.
Flashing detection note
The flashing metric measures whole-frame mean luminance change between sampled frames at 2 fps, with a detection ceiling of 1 transition per second. The medically relevant range for photosensitive epilepsy screening is 3–50 Hz. This metric is not a photosensitive epilepsy screen. A score of zero does not indicate safety; a non-zero score indicates detectable whole-frame brightness transitions useful for relative comparison across shows.
Research grounding
This index measures formal features of video — content-independent structural attributes of the presentation, such as how often the image cuts and how much it changes between frames. Huston & Wright's formal-features framework and Lang's LC4MP motivate measuring these properties; they do not establish what any measured value does to a viewer, and nothing here observes a viewer at all. All associations reported in that literature are correlational.
- Huston & Wright — formal features framework
- Lang — Limited Capacity Model of Mediated Message Processing (LC4MP)
- Lillard & Peterson (2011), Pediatrics — pacing and executive function in 4-year-olds
- Lillard et al. (2015) — fantastical content as a possible moderator
- Christakis et al. (2004), Pediatrics — early TV exposure and attention (correlational)
- Itti & Koch — bottom-up visual saliency and motion
All findings are correlational. This tool measures the stimulus, not the viewer. Age, temperament, sensory-processing profile, and viewing dose are not captured.