The Ordnance Survey lettered its antiquities by period — the older antiquities in one blackletter hand, Norman and later in another, Roman remains in Egyptian capitals — while the living landscape took the ordinary faces. This is a real Dartmoor sheet, its ink lifted off the paper and laid over the LiDAR survey; every marked point is a real label. Hover one to read it as its engraver meant it.
The window above is the OS six-inch first edition over south-west Dartmoor — the Meavy and Burrator area — streamed from the National Library of Scotland. Each reading is fixed to OS 404 (November 1881), the circular in force when the sheet was engraved.
Readings on this page were verified by eye against the specimen plates by a single reader; independent verification is plannedEach face is a new vector outline, reconstructed glyph by glyph from the OS's 1934 alphabet plates. The three antiquity hands lead; their assignments follow the 1881 rule, not the 1934 plates the letterforms were drawn from.
The shapes are dated too. Comparing these faces against the engraved sheets shows the letterforms drifted as well as the assignments. Old English survives the half century largely intact — plate and sheet agree. The face the OS came to call Lutheran did not: the 1934 plate prints a light, open, pointed letter, while the c.1888 sheets carry a denser, more angular hand for the same class. What follows is therefore a faithful record of the 1934 plates, not a facsimile of 1888 engraving, and the German Text specimen should be read with that gap in mind.
Antiquity assignments follow the verified 1881 six-inch column; the display-face uses below are summarised from the wider specimen record. The letterform comparison above is a visual reading of a small number of specimens; it is offered as an observation, not a measured result.
Recogniser output is fragmentary — unlinked tokens, mixed with survey furniture, and misread where the blackletter forms are unfamiliar. One real Dartmoor label, carried through each stage. Counts are for this label; the crop is unaltered.
The period code was revised across editions, and the face names drifted while the periods were reassigned beneath them. Reading an 1888 sheet means fixing the rule to the instrument in force — drawn here from the primary documents held for this study.
Each figure is measured and carries its interval. A number that moved because the reference data was wrong is kept distinct from one that moved because the method was.
The raw adjudicated gold set scores 81.6%, but it was built to a quota heavy in four-word chains (15% against a true 6.5%), and precision falls with chain length. Weighting each band by its real share of the compounds gives 87.9% — the unweighted figure would misdirect the fine-tuning decision.
A stratified 297-crop pixel check on what the pipeline flagged as lost. Each rate is measured against a different denominator, so all three are shown rather than multiplied. Checked against OS Open Names, the national heritage POI set and the statutory designation records; a name may survive in sources these do not cover.
38,042 genuine lost names (95% CI 33,416–42,424); the national extrapolation falls from ~1M to ~450k. The earlier figure counted records the pipeline called lost, never checked against the pixels.
Half of all detections (1,094 of 2,094) are removed here. Audited exhaustively against the independent c.1900
transcription of the second edition: only 10 removed records contain four
or more letters at all, and exactly one is corroborated as a real name —
Tumulu#, a prehistoric burial mound the
recogniser mangled. The single false removal is an
antiquity — the one class this work exists to find. Losses are rare but not
randomly distributed.
15 over-merges — survey furniture welded into pseudo-names
(B. M. Guide Post), which then escape the abbreviation filter because the
weld makes them too long. 16 confirmed under-merges, where both halves of a name
survive apart (Smallcombe + Wood, 24 m). Both failure modes
are structural and known, not random. Population precision remains the adjudicated
87.9%.
Every kept record was re-read against the OS's own written conventions, and the conventions disagreed with the string pipeline 86,166 times. The largest single class: 31,589 bench marks arriving as two records — invisible to a string rule, recoverable structurally. 57,793 records annotated; zero removal decisions changed. The census runs in shadow: no figure here has entered a published result, and no rule graduates without a measured before/after.
Reading the survey's typography turns a picture of Britain into a queryable record of what its surveyors judged to be there.
Class search over the survey's own testimony — every workhouse, every mineral railway, every prehistoric antiquity. The engraved size of a name is the Survey's own weighting of its importance, and it is recoverable.
Surveyor-attested losses — the “Site of…” annotations — together with the genuine lost-name residue as dated leads, and the first-to-second-edition comparison that the matching layer now makes possible.
Name testimony, period-specific terrain expectation and statutory-record silence, ranked and referred privately to county archaeologists — never published as discoveries.
The conventions file and the edition-dating approach are not Dartmoor-specific: they apply to the 25-inch, to the Irish survey, and to any governed document series.
A typeface classifier is the next stage: a labelled training set drawn from the sheets themselves is in preparation, after which the readings above move from hand-verified examples to measured output. No automated classification is claimed on this page.
The aim is to read every antiquity name on the first-edition six-inch through the OS's own period code, dated to the edition in force, in collaboration with the National Library of Scotland.
A methods paper on typeface as testimony is planned, with the reconstructed OS alphabets released alongside it for reuse.