Article

How to simplify Dickens without ruining him

The school programme has topics for which no living children's texts exist. Not because too little has been written, but because the grammatical construction we need is rare in real authors — and across twenty-four topics of the compulsory programme we could not find a single suitable passage.

The obvious solution is to ask a machine to compose one. We forbade ourselves that.

Composed text is plausible and dead. It has no living intonation, no unexpected word, nothing of what people read books for. A child feels that sooner than an adult.

So we took another route: take a real passage by a real author in which the construction already occurs, and lower it one step, keeping the construction intact. Downwards only: a C1 text becomes B1, never the other way round. Raising the level is composition again.

What turned out to be hard

Not what we expected.

We thought the main task would be making the model simplify carefully. That was solved almost at once: of fourteen models raced on the same passages, the winner produced eight usable texts out of ten and not one introduced grammatical error, while keeping the characters' voice. A weaker model flattened complex sentences and added faults of its own — writing put a muffin on table, losing the article.

The hard part was something else: noticing when we were fooling ourselves.

Six traps

A check does not answer the same way twice. We found a passage that got through the guard when it should not have. We re-checked, got a refusal, and nearly wrote the case off as chance. Only a series of eight repeats showed that the passage was passing three times out of four, that the hole was real, and that the cause lay in our own wording of the rule — too narrow. One check proves nothing.

We asked for the impossible and blamed the model. We set B1 as the simplification target for every text, although four rules out of five aim at B2 — that is, we were asking for the text to be lowered below the very construction it had been chosen for. After the correction the result rose from six usable out of ten to eight. The model had nothing to do with it.

The threshold was counting length instead of density. Two models preserved all four occurrences of the construction. One shortened the text by sixteen words and "passed", the other lengthened it by thirteen and "failed". The bar had to be taken from the density of the original rather than fixed absolutely.

A rule written in two places always drifts apart. We fixed the guard and forgot the candidate selection — and 266 usable passages were being silently discarded until we noticed.

A control set made of misses you have already seen tests yesterday's errors. We repaired one detector three times in a day, and each time assembled the test set from the misses found the day before. After the second round, two out of eight inspected hits were genuine. After the third — ten out of ten, and the number of places found fell from 2081 to 788. Three quarters of the earlier findings had been false.

An estimate of the remainder based on known misses is always too low. We put the leakage of one rule at 11 percent, having measured exactly the noise we knew about. A proper check showed 34 — three times more.

What reading uncovered

We read the live hits of all twenty-one rules across twenty-two thousand passages. Four turned out to be broken, and it was not the accuracy figures that gave them away but the extremes of their yield.

The rule for the possessive found it in 455 passages — impossibly few for narrative prose. The explanation: it was catching That's and What's, contractions of is, and missing the real possessive because it searched for a straight apostrophe, while books use a curly one. After the repair — fifteen thousand hits instead of four hundred. A thirty-three-fold rise.

Such an error is invisible both in the code and in the headline number: the rule "worked", it simply stayed silent almost always.

The rule for prepositions of place we quarantined deliberately. Telling in the garden from at the same time is impossible without parsing the sentence, and pretending otherwise is worse than switching the rule off honestly. The rule stays in the code and cannot fire: that is visible on reading, unlike silence.

The main finding

Having re-labelled the corpus, we saw the levels shift: B1 texts fell from 914 to 70, and almost half the corpus dropped below the lowest threshold.

The culprit was not the new parsing but the calibration. The thresholds had been tuned on exam texts — adult prose. Our corpus is children's and teaching material, different in nature, and the same scale reads differently on it.

But here is what matters: the order survived. Mean difficulty rises together with the recorded level in all six sources without exception. The instrument was working correctly; only the thresholds were wrong.

This is the same distinction we reached while measuring text level: comparing two texts with each other is far more reliable than pinning a label on one. Comparison transfers from corpus to corpus; an absolute mark does not.

Recalculating the thresholds on our own corpus also revealed the limit of the data: they do not support five levels. Under a five-class fit the B1 band is left no room at all — of 608 texts, five are recognised correctly. The cause is not the instrument: the label "B1" arrives from six sources, and each means something different by it.

We adopted three classes instead of five. There is nothing on this corpus with which to separate B1 from B2 — not because the instrument is weak, but because no available source gives labelling of that boundary we can trust. For the bagrut this matters, and it will have to be closed by labelling worth believing, not by adjusting thresholds.

What follows from this

Semi-synthetics solves a narrow problem: giving the corpus what living texts did not supply, without sliding into composition. It works only together with the guards, and the value of each guard was established by measurement rather than by reasoning.

Five of the six traps we found by reading particular cases. This is the same conclusion as in the neighbouring article on measuring level: a number that goes up proves nothing until you have read the cases where the instrument is wrong.

← All papers