Observation sits at the centre of early childhood assessment. The rationale is methodological as much as pedagogical: young children show what they know in familiar settings and within their own interests, not in standardised test conditions. But saying “we observe” does not by itself mean the assessment is valid.
What the field uses
Al-Hendawi and colleagues' systematic review of 88 empirical studies groups direct observation systems in early childhood into two families: standardised systems replicated across multiple research projects, and non-standardised systems built for the needs of a single study. The review reports that standardised systems offer a reliable method for generalised behavioural assessment, while non-standardised ones remain widely used for the flexibility they give to targeted evaluation.
Xia and colleagues' scoping review of authentic social-emotional assessment for ages 0–8 identifies 43 instruments across 33 studies, 29 of which meet authentic assessment criteria. Naturalistic observation predominates, while portfolio and curriculum-embedded approaches are under-represented.
The traps
Observation's least discussed problem is the context-sensitivity of measurement. Thorpe and colleagues, analysing 11,341 observation cycles across 2,306 classrooms, report a striking pattern: classroom quality scores vary systematically across the day. Whole-group and small-group formats and science, mathematics and social science content inflate scores, while meal times, physical activity and transitions constrain them. The same classroom can yield a different picture depending on when it was observed.
The classroom implication is clear: a single observation taken in a single slot tells you about that moment, not about the child. Validity comes less from the quality of one record than from how records are distributed across time and context.
- Observe the same child at different times of day and in different activity types.
- Do not generalise from one observation; a pattern needs at least three separate records.
- Keep who observed; observer differences are a real variable.
- Regularly check children with few records — quiet children are the least observed.
Miranda and colleagues, working in community-based child care centres, examined observation tools against child outcomes and report that classroom processes were associated with children's mathematics, pre-literacy and social-emotional development. Pool and Hampshire argue that planning structured and unstructured observation together makes the process both efficient and feasible in inclusive classrooms.
Set up well, observation is early childhood's most powerful assessment tool. Set up badly, it does nothing but lend impressions a scientific appearance. What separates the two is the systematicity of the record.
References
- Al-Hendawi, M. et al. (2025). Direct observation systems for child behavior assessment in early childhood education: A systematic literature review. Discover Mental Health.
- Xia, Y. et al. (2026). Authentic assessment of social–emotional development in early childhood: A scoping review. Brain Sciences.
- Thorpe, K. et al. (2020). The when and what of measuring ECE quality: Analysis of variation in the Classroom Assessment Scoring System (CLASS) across the ECE day. Early Childhood Research Quarterly.
- Miranda, B. et al. (2023). Examining the Child Observation in Preschool and Teacher Observation in Preschool in community-based child care centers. Early Childhood Research Quarterly.
- Downer, J. T. et al. (2009). The Individualized Classroom Assessment Scoring System (inCLASS): Preliminary reliability and validity. Early Childhood Research Quarterly.
- Pool, J. & Hampshire, P. (2020). Planning for authentic assessment using unstructured and structured observation in the preschool classroom. Young Exceptional Children.
Let’s begin the journey to every child’s light.
Discover NOVA ECE at your school. Try all three roles in the live demo — nothing to install.