In July I sat in a consortium meeting and asked whether we had any data showing our training was working. We didn’t. Writing about it afterward, I landed on a harder question than the one I asked: what are we even trying to observe?

I had no answer.

The answer had been sitting in my inbox since September 2025.

The email I let go cold

After the first Quality Consortium meeting held since before COVID, ECESF put a question to the group: in what ways would you measure a child’s well-being?

A colleague on the list answered with the objection first.

I think one of the reasons we keep landing in the oh so useless “valid and reliable child outcomes” space is because it is difficult to measure well being and especially if you want to end up with numbers. BUT measuring outcomes like that ignores the fact that children develop in such complex and individual ways in the first years of life, and ends up incentivizing inappropriate teaching techniques.

Then they attached something and called it a “well being and engagement scan.” The Leuven scale.

I replied asking one question: does it need special training? The answer came back that anyone can use a tool, and that being certified in it would be a different story, one we would have to investigate.

Nobody investigated. Eleven months later, I did.

What the tool actually is

Two scales. Five points each. One for well-being, one for involvement.

Well-being asks whether a child feels at ease, spontaneous, free of emotional tension. Involvement asks whether a child is intensely engaged, which this research treats as the necessary condition for deep-level learning.

You watch a child for about two minutes and assign two numbers. Ten children takes about twenty-five minutes. That timing is the centre’s own protocol, not my compression of it.

That is the entire instrument.

Small does not mean shallow. The Research Centre for Experiential Education describes the rating this way: “The core of the rating process consists of an act of empathy in which the observer has to get into the experience of the child, in a sense has to become the child.” The same document reports inter-scorer reliability of .90 for the involvement scale, which is to say two people watching the same child tend to land on the same number.

The demand is on your attention, not on your credentials.

Less is more is the method

The instinct in quality measurement is to add. More domains, more items, more rigor, more observer training. Every addition is defensible by itself. Together they produce an instrument that only a funded researcher can carry into a room.

Look at what that costs. When CSCCE and the San Francisco Department of Early Childhood ran the 2025 SEQUAL survey here, they invited 744 family child care providers. 138 responded. That is 19%, against 30% for center-based teaching staff and 54% for directors. The demographic analysis rests on 82 providers. Those are the report’s own figures, stated plainly in it, which I respect.

606 family child care providers are not in that picture. They are disproportionately the least connected, least resourced, and most language-isolated providers in the city.

The heavier the instrument, the fewer people it reaches, and it misses the same people every time.

A two-minute scan does not have that problem. Not because it is better science, but because a provider alone with six children can actually run it. Scale is not something you add after the instrument works. Scale is a property of how small it is.

The part I did not expect

I went looking for the licensing catch. I read the primary manual, 21 pages, published by Kind & Gezin and the Research Centre for Experiential Education. It contains no mention of certification, accreditation, licence, permission, or qualification. Not one. The manual and forms are a free download from the centre’s resources page.

What I found instead was in the origin section. The Flemish agency that supervises the care sector commissioned the instrument with three requirements. It had to serve as a tool for self-assessment. It had to take the child’s own experience as the measure of quality. And it had to be “appropriate for the wide range of care provision including care for the under three’s in day care centres and family care as well as the out of school care for children up to the age of twelve.”

Family care. Written into the design brief in 2005.

Then, in the manual’s own definitions: “The term supervisor in the manual refers to everyone who is engaged in the care of children including the child care family, the child care worker, the practitioners in after-school care.”

I have spent years explaining why instruments built for centers break in a family child care home. SEQUAL measures the working conditions an employer provides to teaching staff. In family child care the provider is the employer, the employee, and the facility, so the frame collapses. That argument is correct and I am tired of making it.

This is the first quality instrument I have read that did not require me to make it.

The scan is step one of three

Here is what changes how I would use it. The self-evaluation procedure has three steps: assess the actual levels of well-being and involvement by scanning the groups, analyse the observations to explain the levels you saw, then select and implement actions to improve quality.

The scan is step one. The quality improvement lives in two and three.

That reframes the whole exercise. The number is not the product. The number is the prompt that makes a provider ask why this child, in this corner, at this hour, was at a 2. The manual claims the process itself develops practitioners, because they learn to take the perspective of the child.

A score that nobody acts on is the most expensive kind of data, since it costs the same to collect as one that changes something.

The busy work nobody can explain

There is a second reason to care how small an instrument is, and it has nothing to do with response rates.

Ask an educator what a particular piece of required documentation is for. Often they cannot tell you. They can tell you it is due, who wants it, and what happens if it is late. The purpose went missing somewhere upstream and the form kept moving. I have written before about counting hours instead of building competency and about the black rectangle in a required webinar. This is the same flaw pointed at a third subject.

I do this paperwork myself, in my own home, for my own program. What it costs is not really time. It is attention, and attention is the same resource the children are drawing on. Every form is a small transfer of a provider’s focus away from the child in the room toward a reader who will never meet that child.

Which is what makes a two-minute scan interesting well beyond its price. What it asks for is exactly what the paperwork has been taking. You have to watch one child closely enough to get inside her experience. That is the practice itself, not an administrative wrapper around it.

We have built a sector where the observation is the one thing there is no time for.

Providers are already solving this quietly

I know providers who use AI to produce the documentation their funders require. Not to think with. To get through it.

I am not going to police that, and I would ask the field to hold its reaction for a moment before it forms. Ask what the alternative is for one person who is the teacher, the cook, the bookkeeper, and the licensee of the same small business. There is no team to hand it to. She is one person, and the forms arrive anyway.

The conclusion usually drawn from this is that a workload problem got solved. I read it the other way. A requirement that only stays survivable because people quietly automate it was never sized for the people being asked to meet it.

We keep hearing the workaround and recording it as capacity.

And the automation gives away what the document was worth. If a report can be generated by something that never met the child, and nobody downstream can tell the difference, then the report was never about that child. It was about having a report.

Does it come back to the room?

Here is the test I would apply to every piece of it.

After the number is recorded, after the narrative is filed, does anything change in the room? Does one child’s Tuesday go differently because of what was written on Monday?

When the honest answer is no, that document was not serving the classroom. It was serving the file.

Paperwork that never returns to the room is a transfer of a provider’s attention to a reader who will never meet the child.

There is a cost inside that transfer which never shows up on a timesheet. Writing about a child is emotional work. You are asked to take a relationship and render it into a format, for someone with no relationship to that child and nothing at stake in how she turns out. Doing that over and over, for an audience that will not act on it, wears down the exact capacity the job runs on.

The language requirement makes the same point in a harder form. In the SEQUAL sample, 60% of San Francisco’s family child care providers speak Chinese and 55% speak English. A provider who watches a child with enormous care in Cantonese, and then has to account for it in written English, is being assessed on the language of her paperwork rather than the quality of her attention. The demand was never built around the people carrying it.

Why it keeps growing

I think the reason is plain. Providers absorb it and do not complain. Silence gets read as room to spare, and the next cycle adds another requirement on top of it.

Nobody set out to overload anyone. The signal never arrived, because the people best positioned to send it are the ones with the least margin to spend on sending it.

Compliance and consent are different things. People comply when refusing costs more than complying.

So the ask I would make of funders and system administrators costs nothing to try.

Before adding the next required document, answer one question: what will change in a classroom because of it? If you cannot name the change, do not ask for it. And when a requirement lands and no one objects, treat that as a thing to go check rather than proof the load was bearable.

Common sense gets you most of the way. Anyone who has spent an afternoon in a family child care home can see that a person caring for six children cannot also produce a research file about them. We have not been acting on something we can plainly see.

What the room is for

When time does come back, I do not want it spent producing more.

What children need from an adult is an authentic relationship, and that is always the first thing to go, because no one checks. Turn a form in late and someone calls you. Rush every pickup for a week and nobody says a word.

The educators I know are not short of commitment. They are short of margin. A sector that keeps adding requirements to people who have none left gets compliance, and then calls it quality.

Breathing room is not a pleasant side effect of a simpler tool. It is the outcome I would measure.

The catch worth naming out loud

The instrument’s full title is “Well-being and Involvement in Care Settings. A Process-oriented Self-evaluation Instrument.”

Self-evaluation. In our email thread, the enthusiasm ran the other direction. One colleague wrote that the scale “provides clear yet simple indicators that would help visitors gauge quality in well being,” and that there should be a place for it within the QRIS.

I understand the appeal and I think it is the wrong move. Hand the scan to an external rater and you keep the cheapest part of the method and throw away the two steps where quality actually improves. You also change what the tool is. A self-assessment turned into a score someone else gives you is a different object, whatever the five points say.

There is a better version available. Providers run it on themselves, as designed, and bring what they learned to the table.

What I would actually do next

  1. Download the manual and forms. They cost nothing, and the barrier everyone assumed was there is not there.
  2. Ask four or five family child care providers to run one scanning cycle in their own homes. Twenty-five minutes, once.
  3. Have two people scan the same session and compare their numbers. The published reliability is someone else’s finding. Whether our observers agree with each other is ours, and it is answerable in a morning.
  4. Do steps two and three. Pick one thing to change based on what the scan surfaced, change it, scan again.
  5. Bring that to the Quality Consortium as practice rather than as a proposal.

The question in that room was whether child well-being can be measured without distorting practice. A reliability figure and a changed routine from actual family child care homes answers it better than any argument I could write.

What I am still unsure about

Whether five points can carry the weight we would put on them if this ever became policy. Whether a provider scanning her own children sees what a stranger would see, and whether that is a weakness or the entire point. Whether what worked in Flanders survives translation into a city where the family child care workforce speaks mostly Chinese and most quality frameworks arrive in English.

I don’t know. But the tool is free, it takes twenty-five minutes, and it was written to include us. That is a cheap enough experiment that not running it is the harder position to defend.

If you have used the Leuven scales in family child care, or tried to put them inside a rating system and watched it go sideways, I would like to hear how it was built. Reach out. I am not proposing a model. I am trying to find out whether the smallest version works.

Everything I used, in one place

So you don’t have to go hunting for any of this.

  • The manual and forms from the Research Centre for Experiential Education. Look for SICS (ZiKo). Free, no registration.
  • The full manual as a PDF, 21 pages, hosted by West Sussex County Council. This is the source of everything I quoted about the design brief and about who counts as a user. Full citation: Well-being and Involvement in Care Settings. A Process-oriented Self-evaluation Instrument, Ferre Laevers (Ed.), Kind & Gezin and the Research Centre for Experiential Education, 2005, ISBN 978-90-77343-76-8.
  • The scanning method, which is where the two-minute episodes and the twenty-five-minute cycle come from.
  • An overview of the wider approach from the centre, February 2015. Source of the .90 inter-scorer reliability figure and the passage about becoming the child.
  • The 2025 SEQUAL report for San Francisco from CSCCE and the San Francisco Department of Early Childhood. Every response-rate number I used is theirs, and they state the limits of their own sample more clearly than most reports bother to.

One note on the second link. That host is fussy about direct requests, so if it fails, the same document is reachable through the centre’s own resources page.