Writing systems and orthography

Which unit of language a script writes down, and what it leaves unwritten. Also called Script type, orthographic system or grammatology.

Reference/All structures

Alphabetic Abjad Abugida Syllabic Logographic Mixed

The parameter

A writing system is not a language, and the two vary independently: Turkish was written in Arabic script and is now written in Latin, Serbian uses both Cyrillic and Latin, and Japanese uses three scripts at once. The parameter measures which unit of linguistic structure the graphic signs map onto. An alphabet writes consonants and vowels as separate signs of equal status. An abjad writes consonants and leaves most vowels to the reader. An abugida writes a consonant with an inherent vowel and modifies the sign to change or cancel it. A syllabary writes syllables, or in the Japanese case morae. A logographic system writes morphemes.

The terms abjad and abugida were coined by Peter Daniels in 1990, because calling the Arabic and Devanagari systems alphabets obscured what matters most in translation: what they do not record. Most working systems are also mixed. Arabic and Hebrew write short vowels only with optional pointing, used in scripture and children's books and almost nowhere else, and Japanese runs a logographic script alongside two moraic syllabaries.

How languages vary

  • Alphabetic: separate signs for consonants and vowels. Latin for English, Dutch, Turkish and Polish; Cyrillic for Russian; Greek; Armenian; Georgian.
  • Abjad: consonantal, short vowels normally unwritten. Arabic, Hebrew.
  • Abugida: consonant signs with an inherent vowel, modified for other vowels. Devanagari for Hindi; the Ethiopic script for Amharic; Thai; Khmer.
  • Syllabic: one sign per syllable or mora. Japanese hiragana and katakana.
  • Logographic: one sign per morpheme. Chinese characters.
  • Mixed: Japanese, combining kanji with two kana syllabaries; Korean hangul, alphabetic in its letters and syllabic in its blocks.

Alphabetic systems are the most widely used, largely through the spread of the Latin script, but no reliable proportion can be given: the writing systems chapter in the World Atlas of Language Structures is illustrative rather than a survey, with only a handful of data points.

Consequences for translation

Once the scripts differ, every name has to be converted, and the conversion is not deterministic. Several published standards exist for the same script pair and they disagree by design. ISO 9 maps Cyrillic to Latin character for character using diacritics, so the original spelling can be recovered from the result. BGN/PCGN, produced jointly by the United States Board on Geographic Names and the British Permanent Committee on Geographical Names, avoids diacritics, uses digraphs aimed at an English reader and is not reversible. The ALA-LC tables, approved by the American Library Association and the Library of Congress, keep catalog records consistent across scores of scripts. DIN 31635 covers Arabic script and gives one sign per letter rather than digraphs, which is why it dominates German-language Arabic and Islamic studies. None of these is wrong. Choosing among them silently, or switching between them inside one document, is.

Above all of them sits one practical rule. ICAO Doc 9303 requires issuing states to transliterate national characters into the machine readable zone of a travel document using the permitted character set, supplies its transliteration tables as a recommendation rather than a mandate, and forbids diacritics in that zone while allowing them in the visual inspection zone. The same name can therefore appear in two forms on one passport. For certified work the spelling in the identity document governs, and the standard is a fallback for names that have no document.

Three further effects follow. An abjad leaves short vowels unwritten, so a consonant skeleton supports several defensible readings and the text does not say which is meant. Directionality complicates the file: Arabic and Hebrew run right to left while embedded numerals and Latin strings run left to right, and a target-only proofread cannot tell a display artifact from an error. And capitalization does not exist in Arabic, Hebrew, Devanagari, Chinese, Japanese or Korean, so anything English carries by an initial capital, including the difference between a defined term and a common noun, must be rebuilt by other means.

Examples

AR>EN

  • Source: كتب
  • Target: kataba
  • Comment: The same three consonants also spell kutiba, it was written, and kutub, books. Vowels are supplied by the reader from context. In a name there is no context, so only the document of record settles it.

JA>EN

  • Source: 東京都渋谷区
  • Target: Shibuya Ward, Tokyo Metropolis
  • Comment: Six characters, each a morpheme, giving an address a Japanese reader parses without spaces. English needs word division, a reordering from large unit to small, and a decision on the administrative labels.

RU>EN

  • Source: Щербаков
  • Target: Shcherbakov
  • Comment: One Cyrillic letter becomes four Latin ones under an English-facing convention and a single diacritic letter under a reversible one. The two strings sort differently, search differently and look like two people; the file should state which convention was applied.

References

  • Daniels, P. T. & Bright, W. (eds.) (1996). The World's Writing Systems. New York: Oxford University Press.
  • Daniels, P. T. (1990). "Fundamentals of Grammatology." Journal of the American Oriental Society 110(4).
  • Comrie, B. (2013). "Writing Systems." In Dryer, M. S. & Haspelmath, M. (eds.), The World Atlas of Language Structures Online, Ch. 141. Leipzig: Max Planck Institute for Evolutionary Anthropology.
  • Coulmas, F. (2003). Writing Systems: An Introduction to Their Linguistic Analysis. Cambridge: Cambridge University Press.
  • International Civil Aviation Organization (2021). Doc 9303, Machine Readable Travel Documents, Part 3: Specifications Common to all MRTDs, 8th edition. Montreal: ICAO.

Language pairs, understood structurally

Knowing which shifts a pair forces is what separates a defensible translation from a fluent one. It is the basis on which Translyta's language profiles are built.

Request a quote