Block startU+0900
ilug-cal.orgLinux in India

Code Points

One code point, and then thirty years of arguing about what to draw

Assigning a character a number is the easy half; agreeing what shape it takes in every context is the part that took decades.

Section 1 · Code Pointsthree pieces in this section

A printed code chart page of script characters
FigureA chart page shows representative glyphs. The standard’s own note says they are illustrative, not normative.

The number was just the beginning

Encoding a character means giving it a number — a code point — and writing that number into a table. That part is finite. A committee meets, a proposal is tabled, a value is assigned, and the Unicode Consortium publishes the update. The Devanagari block, for instance, was present in Unicode 1.0 in 1991, its code points allocated from U+0900 to U+097F. The Kannada block, the Tamil block, the Malayalam block: all assigned in the same release. On paper, the job looked done.

It was not. The code point is an address. What gets built at that address — the glyph, its weight, its exact stroke geometry, and above all its behaviour in combination with every adjacent character — is a completely separate problem, one that the encoding standard deliberately leaves open. Unicode specifies what a character is; it does not specify what it must look like. That gap between the number and the drawn shape is where three decades of argument lived.

A bound copy of a constitutional text open on a table
PlateTwenty-two languages, barely a dozen scripts — the ratio that made one shared table thinkable.

Shaping: the problem that encoding ignores

The argument was most acute in Indic scripts because of shaping — the process by which a rendering engine transforms a sequence of stored code points into a sequence of positioned glyphs on screen. In Latin text, one code point typically maps to one glyph in one place. In Devanagari or Malayalam, the relationship is far more complex. A consonant cluster may produce a conjunct — a ligature that fuses two or more consonants into a single shape not individually present in the font. A matra (a dependent vowel sign) typed after its consonant may appear to the left of it, to the right, above, below, or split across both sides depending on the script and the specific vowel.

The Unicode Standard encodes the logical order: the consonant first, the matra after. The visual order on screen may be different. The rendering engine must know which code point sequences trigger which transformations, and the font must contain the glyphs those transformations call for. Neither Unicode nor any single government body controlled both sides of that contract simultaneously.

Chronology

  1. 1991Unicode 1.0 published; Devanagari, Tamil, Malayalam, Kannada and other Indic blocks assigned
  2. Pre-1991ISCII (IS 13194) standardised by the Bureau of Indian Standards, mapping ten scripts to one eight-bit table
  3. Early 2000sMicrosoft publishes first OpenType Indic shaping specification
  4. Mid-2000s"Indic v2" shaping model published, correcting errors in the first specification
  5. 1971Malayalam script reform reduces active conjunct inventory, creating parallel orthographic traditions sharing the same code points
  6. Mid-2010sBroad renderer agreement on Indic OpenType shaping behaviour

C-DAC, the Centre for Development of Advanced Computing based in Pune, was deep in this territory from the early 1990s. Its ISCII work — the Indian Script Code for Information Interchange, standardised as IS 13194 before Unicode existed — had already mapped ten Brahmi-derived scripts onto a single eight-bit table and had thought carefully about what logical encoding implied for rendering. When Unicode absorbed ISCII's structure, it inherited the same underlying questions about how the abstract characters would behave in a live shaping engine. The answers were not included in the transfer.

The OpenType specification became one of the main arenas where those answers were worked out. OpenType's GSUB (Glyph Substitution) and GPOS (Glyph Positioning) tables gave font designers a language for describing shaping rules: if you see this sequence of code points, substitute these glyphs; if you see this cluster, apply this kern. For Indic scripts, Microsoft published a set of shaping specifications in the early 2000s — later revised and significantly corrected in what became known informally as the "Indic v2" shaping model — that described how a conforming rendering engine should process Devanagari, Bengali, Tamil, and the other scripts.

Unicode specifies what a character is; it does not specify what it must look like.

The revision was necessary because the first specification had errors. It mishandled certain conjuncts and produced wrong output for sequences that were perfectly legal under Unicode. Fonts built against version one behaved differently from fonts built against version two, and neither set was wrong by its own rules. Type designers and software engineers in Kerala and Tamil Nadu — where Malayalam and Tamil font development was particularly active — found themselves maintaining two code paths.

The glyph underneath the encoding

The shape argument was not only technical. For many scripts, there were also deep disagreements about which historical or regional form a code point should preferentially represent. Tamil has a set of traditionally distinct letterforms that differ from the standardised modern forms; both are valid, but a font choosing one was implicitly making a cultural statement. Malayalam script underwent a systematic reform in 1971 that reduced the number of conjuncts from several hundred to a much smaller set, creating two parallel traditions — the reformed orthography and the traditional one — that shared Unicode code points but expected different glyph repertoires from their fonts.

That shared code point became a pressure point. A font designer had to choose which tradition to serve by default, because the encoding gave no instruction. Selectors — Unicode variation sequences that allow a single code point to request a specific glyph variant — were one answer, but they required font support, renderer support, and input method support to line up simultaneously. In practice, communities continued to ship separate fonts for traditional and reformed orthography rather than rely on a mechanism that could break at any layer in the stack.

A screen rendering Devanagari text at large size
InsetRendered, not stored: the head-line is drawn by the shaping engine, never encoded in the text.

Devanagari presented its own geometry disputes. The vowels, consonants and combining marks of the Devanagari Unicode block — encode a system inherited largely from ISCII's structure, but different publishing traditions within Hindi, Marathi, Sanskrit, and Nepali all had preferences about how the head-line (the horizontal bar from which letters hang), how the half-forms of consonants, and how the nukta (a combining dot that modifies certain consonants) should be rendered. A nukta applied to a consonant produces a distinct character used in loanword transcription; in some fonts it sat differently from what traditional Devanagari typesetters expected.

Thirty years and the state of the table

By the time OpenType Indic shaping had stabilised into something most major renderers agreed on — broadly by the mid-2010s — the fonts that had been shipped on government workstations and school computers through state procurement contracts in the intervening years had already established their own de facto norms. The Unicode Consortium's published charts and character notes show the code points and their representative glyphs, but the charts explicitly carry a disclaimer that their glyphs are illustrative, not normative. The normative shape lives in the interaction between the font and the rendering engine, and that interaction was the subject of argument for as long as anyone was paying attention.

What the thirty years produced was not a single answer but a layered one: a shaping engine that works correctly for the common cases, a font format expressive enough to describe complex substitutions, and a community of type designers in India who had to build the actual tables by hand, conjunct by conjunct, matra by matra, while the specifications were still being corrected underneath them. The encoding was finished in 1991. The drawing is still being revised.

The gap that caused the argument

  • code pointa number assigned to a character in the Unicode or ISCII table; says what the character is, not how it looks
  • shapingthe process a rendering engine runs to convert stored code points into positioned glyphs on screen
  • conjuncta ligature fusing two or more consonants; may have no individual-glyph equivalent in the font
  • matraa dependent vowel sign; its visual position (left, right, above, below, or split) differs from its logical storage position
  • glyphthe actual drawn shape at a given position; one code point may map to multiple glyphs depending on context
  • GSUB / GPOSOpenType table types that let a font encode its own substitution and positioning rules for shaping engines
Chart listing Devanagari vowels and consonants organized by row under "देवनागरी लिपि" heading
FigureThe varṇamālā order set out as a teaching chart: back of the throat first, sixteen to a row.Photo: Devanagari Varnamala · Wikimedia Commons

Attributions

Read next