Sponsored by

Our Sponsor

The Voice Data Behind the World's Best AI

Training data quality determines model quality. Voices delivers custom or pre-built voice datasets—directed by real professional talent, not scraped audio—across any language, accent, or emotional range. 

Choose from 800+ hours of character voice performances, 1,723 hours of QA'd conversational data across 56 emotional states, or 1,000 hours of expressive, multi-language data with full emotion tagging. Every dataset ships with clear licensing, signed consent, and usage rights baked in, so your legal team sleeps easy and your model trains on data it can actually use. 

7 of the world's top 10 AI labs already use Voices. Request free samples and hear the difference professionally directed voice data makes.

ONE WILD WORLD — NEWSLETTER

Fact-checked stories about the things you thought you knew.

Written for people who don't mind being corrected — or shocked.

ISSUE 052 · DEEP DIVE · MYTHS & MISINFORMATION

Boys Get Lower Grades From Their Own Teachers Than From Anonymous Markers. It Has Been Measured in a Dozen Countries.

The grading gap is real, it has been measured for twenty years, and almost nobody in the field disputes it. The explanation everyone reaches for — that a boy's name on the page is what triggers it — happens to be the one part that fails every direct test it has been given.

An accident of bureaucracy

Israeli sixth-formers get their matriculation papers marked twice. Not sat twice — marked twice. An external examiner scores them blind, no name attached, no idea who the child is. The pupil's own teacher scores the same material knowing exactly whose handwriting they're looking at.

Victor Lavy, an economist, spotted that this piece of administrative housekeeping had the shape of a controlled experiment. Same pupil, same subject, two markers, and only one of them knows the kid.

He went looking for bias against girls. That was the whole premise — schools as quiet little factories of stereotype, girls nudged away from physics by a thousand small signals. It's the finding everyone expected, including him.

He got the opposite. In every subject he examined, boys came out lower with the teacher than with the stranger.

That was 2008. It has since turned up in Norway, Sweden, Denmark, France, Spain, Italy, Czechia, Greece, Britain, the United States and Brazil — a dozen countries counting Israel, across three continents and two decades of data. In the Czech figures the gap is easy to state: for the same maths result, girls receive on average a third of a grade better than boys. One review counted thirteen studies on the question. Eleven found the gap.

And it matters beyond the mark itself. Follow-up work links teachers' assessments to whether pupils go on to take advanced maths at all, which feeds into degree choices, which feeds into everything after that. A third of a grade doesn't sound like much. Compound it across a school career and it stops being a rounding error.

The obvious culprit

So: why? The intuitive answer is the name at the top of the page. Teacher sees "Daniel", something unconscious fires, the mark comes out a notch lower. It's tidy. It also implies a very cheap fix — take the names off, mark by student number, done.

That's a real hypothesis, and unlike a lot of things people assert about schools, it can actually be tested. Give markers identical papers. Put a boy's name on half and a girl's on the other half, at random. See what comes back.

It's been done. Several times, in several countries.

Denmark, 2019. The same exam papers scored twice, once blind and once with names visible, plus a full-cohort natural experiment alongside it. Both designs pointed the same way — no systematic gender bias in the non-blind marking.

Roskilde University. An exam switched from open marking in 2018 to student-ID-only in 2019. No evidence of gender or ethnic bias either side of the change.

Stockholm University. Written exams blinded in stages. No substantial effect on the male–female gap.

The Netherlands. 358 teachers, identical tests, names swapped. Grading was not gender-biased on average.

Four goes at the cleanest possible test of the name hypothesis. Four times, nothing.

Except the method isn't broken

The tempting response here is that name-swap experiments are just too blunt to catch anything — artificial setting, no stakes, teachers on their best behaviour. Fair worry. There's a study that kills it.

In India, researchers randomly varied the characteristics printed on children's exam papers and handed them to local teachers to mark. On gender: nothing. On age: nothing.

On caste, the same teachers, marking the same papers, discriminated.

So the instrument works. It detects bias when bias is there. It picks up prejudice against a child it has never met, from a label on a page — which is exactly what the name hypothesis predicts should happen with gender, and exactly what doesn't.

What the two study designs actually differ on

Put the two literatures side by side and the difference isn't the country, the subject, or the decade. It's who is holding the red pen.

In the experiments, the marker is a stranger. She has a name on a page and nothing else. In the observational studies — Lavy's and everything downstream of it — the non-blind marker is the child's own teacher. Camille Terrier, who ran the French version, spells it out: the blind score comes from an external examiner and is therefore free of stereotypes about the pupils, while the teacher's grade comes from someone in permanent contact with the children they teach.

The Dutch team flagged it themselves, as a limitation of their own null result: their design had to exclude the real-world teacher–pupil interactions that might be doing the work.

It isn't the name. It's the year.

Which leaves an argument nobody has won

If the gap lives inside the relationship rather than on the paper, what is it made of? Two serious answers, both with good evidence, and they don't agree.

Lavy tested the behavioural explanation directly and couldn't make it work — the differences didn't track anything about how the pupils acted. What the size of the gap did track was the characteristics of the teacher. His conclusion was that this is coming from the teachers' side, not the pupils'.

A Brazilian study using an unusually good dataset found the reverse. Teachers there inflated the scores of better-behaved pupils and docked the worse-behaved ones, and classroom behaviour accounted for around two-thirds of the gap against boys. Not a claim about what boys are like — a claim about what teachers are measuring when they think they're measuring maths.

I don't think that's settled, and I'd be suspicious of anyone who tells you it is. The two studies use different countries, different age groups and different ways of capturing behaviour. Both could be partly right.

The fix aimed at the wrong thing

Universities across Denmark have been moving to anonymous exams for years — Aarhus, the Danish School of Education, the lot — on the reasoning that a marker who can't see who you are can't hold anything against you. It's a decent principle and there are other good reasons for it.

But if the mechanism is a marker who has known the pupil since September, stripping the name off the front page does approximately nothing. You haven't removed the information. The marker already has it.

The variable that moves the mark isn't anonymity. It's whether the person marking is the person teaching. Which is a much more awkward thing to legislate.

BY THE NUMBERS

11 of 13

studies in one review that found teachers marking boys below their blind test scores

⅓ of a grade

the girls' advantage in Czech maths for an identical result

~2/3

share of the gap that classroom behaviour explained in the Brazilian data

4

name-swap experiments that found no gender effect — Denmark, Sweden, the Netherlands, India

1

of those that did find bias, on caste, using the exact same method

Everyone assumed the prejudice was in the name. It was in the room.

SOURCES

Lavy, V. (2008). "Do gender stereotypes reduce girls' or boys' human capital outcomes?" Journal of Public Economics 92(10–11), 2083–2105.

Protivínský, T. & Münich, D. (2018). "Gender bias in teachers' grading: What is in the grade." Studies in Educational Evaluation.

Terrier, C. (2020). "Boys lag behind: How teachers' gender biases affect student achievement." Economics of Education Review 77.

Rangvid, B. S. (2019). "Gender Discrimination in Exam Grading? Double Evidence from a Natural Experiment and a Field Experiment." B.E. Journal of Economic Analysis & Policy 19(2).

Bischoff, C. S., Ejrnæs, A. & Rubin, O. (2021). "A quasi-experimental study of ethnic and gender bias in university grading." PLOS ONE 16(7).

Hanna, R. & Linden, L. (2012). "Discrimination in Grading." American Economic Journal: Economic Policy.

Ferman, B. & Fontes, L. F. (2022). "Assessing knowledge or classroom behavior? Evidence of teachers' grading bias." Journal of Public Economics 216.

Lavy, V. & Sand, E. (2018). "On the origins of gender human capital gaps." Journal of Public Economics 167, 263–279.

OneWildWorld!

Facts you never asked for. Knowledge you can't unsee.

Click to read last week's deep-dive →

Share this newsletter with someone who likes being corrected.

ONE MORE THING

You can read our MANY newsletters for free at our website:

Read our archive for free at OneWildWorld.com