Block startU+0900
ilug-cal.orgLinux in India

Code Points

Languages in the Eighth Schedule, written in a handful of scripts

Twenty-two languages, a constitutional guarantee, and a table that fits on one page — because the scripts are far fewer than the tongues.

Section 1 · Code Pointsthree pieces in this section

A bound copy of a constitutional text open on a table
FigureTwenty-two languages, barely a dozen scripts — the ratio that made one shared table thinkable.

The list and what it actually encodes

The Eighth Schedule to the Constitution of India currently recognises twenty-two languages, a roster that has grown by amendment from the original fourteen of 1950. The languages span four major families — Indo-Aryan, Dravidian, Tibeto-Burman and Austroasiatic — and represent hundreds of millions of speakers across a subcontinent of extraordinary linguistic depth. But the writing systems those twenty-two languages use number barely a dozen, and several of those scripts share enough structural ancestry to be treated, for encoding purposes, as close cousins.

The dominant system is Devanagari, the abugida used for Hindi, Sanskrit, Marathi, Nepali, Konkani, Bodo, Dogri, Maithili and Sindhi — nine of the twenty-two scheduled languages sharing a single script family. The Dravidian south contributes Tamil, Telugu, Kannada and Malayalam, each with its own distinct script but all built on the same Brahmic logic of consonant-plus-inherent-vowel that Devanagari also follows. Gujarati and Gurmukhī (for Punjabi) are closely related to Devanagari at the glyph level. Odia, Bangla and Assamese form a northeastern cluster whose letterforms share enough strokes to have been treated as variants in early encoding proposals. Urdu uses Nastaliq, a Perso-Arabic form, and Meitei uses its own script, Meitei Mayek, which was added to Unicode in 2009.

A printed code chart page of script characters
PlateA chart page shows representative glyphs. The standard’s own note says they are illustrative, not normative.

That structural convergence mattered enormously when engineers at C-DAC's Pune centre began designing ISCII — the Indian Script Code for Information Interchange — in the 1980s. ISCII put ten Brahmi-derived scripts on one code table by exploiting precisely this parallel architecture: the same code point could mean "the ka-class consonant" regardless of whether it was to be rendered in Bangla or Gujarati. The mapping worked because the underlying phonological structures aligned, even when the letterforms diverged.

What the schedule made possible — and what it required

The constitutional status of a language created an administrative obligation: government documents, court proceedings and official correspondence had to be producible in that language. That requirement is what turned a linguistic fact into a procurement and engineering problem. A single eight-bit encoding space large enough for ten Brahmic scripts and an ASCII-compatible lower half was a viable solution only because the script count was tractable. Had each of the scheduled languages used a wholly unrelated writing system, no shared table of this kind could have been designed.

Unicode absorbed and extended this logic when its Devanagari block and the adjacent South Asian blocks were assigned in the early 1990s. The Unicode Consortium allocated separate blocks for each script — Devanagari at U+0900, Bengali at U+0980, Gujarati at U+0A80 and so on — but preserved the parallel column layout that ISCII had used, so that the same relative position within each block corresponds to the same phonological value. A software shaping engine written to handle one Brahmic script can therefore be extended to another by swapping the glyph tables, because the code-point logic is isomorphic.

Chronology

  1. 1950Eighth Schedule enacted with fourteen languages
  2. 1980sC-DAC Pune designs ISCII across ten Brahmic scripts
  3. Early 1990sUnicode allocates parallel Brahmic blocks, preserving ISCII's positional logic
  4. 2001Kerala IT policy triggers first large state deployments
  5. 2009Meitei Mayek added to Unicode (Meitei, a scheduled language, gains standard encoding)

The shaping requirement itself is non-trivial and falls equally on all these scripts. Every Brahmic system uses conjuncts — ligatures formed when two or more consonants meet without an intervening vowel — and matras, the dependent vowel signs whose rendered position can differ from their stored position. A matra typed after a consonant may appear to the left of it on screen; the engine must reorder before display. This is not a quirk of one script; it is a property of the family, and the constitutional schedule, by including so many members of that family, concentrated the engineering problem in a well-defined space.

The practical consequence is visible in the state-level deployments that followed Kerala's 2001 IT policy and Tamil Nadu's subsequent school computing programmes: when engineers chose fonts and rendering libraries, the choice for one Brahmic script in the schedule had direct implications for neighbouring scripts that shared the same shaping model. The schedule was a legal document, but it also quietly described the scope of the rendering problem that the open-source Indic stack would have to solve.

Gujarati and Gurmukhī (for Punjabi) are closely related to Devanagari at the glyph level.

A keyboard with an Indic layout overlay, hands typing
InsetA layout standard is only real where it arrives installed by default.

Script-to-language mapping

  • DevanagariHindi, Sanskrit, Marathi, Nepali, Konkani, Bodo, Dogri, Maithili, Sindhi (nine scheduled languages)
  • Tamil scriptTamil
  • Telugu scriptTelugu
  • Kannada scriptKannada
  • Malayalam scriptMalayalam
  • Bangla scriptBangla, Assamese (closely related; historically treated together in early encoding)
  • GurmukhīPunjabi
  • GujaratiGujarati
  • OdiaOdia
  • Nastaliq (Perso-Arabic)Urdu
  • Meitei MayekMeitei (Manipuri)

Attributions

Read next