Monday, 31 January 2022

Unicode Trivia U+09F8

Codepoint: U+09F8 "BENGALI CURRENCY NUMERATOR ONE LESS THAN THE DENOMINATOR"
Block: U+0980..09FF "Bengali"

Decimal Day. 15 February 1971. A Monday. The day the United Kingdom and the Republic of Ireland converted to decimal currency. Before that, each pound was divided into twenty shillings and each shilling into twelve pence. We'll ignore farthings.

So, if I bought something worth one penny with an old, pre-decimalisation five pound note, I'd get a dirty look and the following change:

£4 19/11 = £4. 19s. 11d = 4 pounds, 19 shillings, 11 pence

Such mixed-radix currencies were not uncommon. In British India, the rupee had been divided into sixteen annas, each anna into four pice (paisa), and each pice into three pies. The change from five rupees for a one pie item would be:

Rs. 4/15/3/2 = 4 rupees, 15 annas, 3 pice, 2 pies

In pre-decimal Bengal, the taka (rupee) had been divided into sixteen ana, and each ana into twenty ganda. The change from five taka for a one ganda item would be:

Tk. 4/15/19 = 4 taka, 15 ana, 19 ganda

Of course, that's in English using the Latin script and Western Arabic numerals. In Bengali, one could have written:

৪৲৸৶৹৻১৯

U+09EA U+09F2 U+09F8 U+09F6 U+09F9 U+09FB U+09E7 U+09EF

As Anshuman Pandey, points out, only one currency mark was actually used when multiple units were written. We'll return to this in due course, but in the meantime I've left that refinement out of the example above.

Bengali is a Brahmic script written left-to-right, so in Unicode this example is:

  1. U+09EA "BENGALI DIGIT FOUR"
  2. U+09F2 "BENGALI RUPEE MARK"
  3. U+09F8 "BENGALI CURRENCY NUMERATOR ONE LESS THAN THE DENOMINATOR"
  4. U+09F6 "BENGALI CURRENCY NUMERATOR THREE"
  5. U+09F9 "BENGALI CURRENCY DENOMINATOR SIXTEEN"
  6. U+09FB "BENGALI GANDA MARK"
  7. U+09E7 "BENGALI DIGIT ONE"
  8. U+09EF "BENGALI DIGIT NINE"

The first two glyphs ("৪৲") represent "4 taka" in decimal; the Bengali digit four just happens to look like a Western Arabic digit eight. The next three glyphs ("৸৶৹") represent "15 ana". This is complicated by the fact that, traditionally, ana were written as fractions of a taka. Finally, the last three glyphs ("৻১৯") represent "19 ganda" in decimal where, just to confuse us further, the ganda mark comes before the digits, not after them as with taka and ana.

The ana component is the most perplexing. The Unicode codepoint name for U+09F8, "BENGALI CURRENCY NUMERATOR ONE LESS THAN THE DENOMINATOR", doesn't really help. Fortunately, there's an explanation within the much later proposal to add the ganda mark in 2007.

The fifteen possible quantities of ana are:

  • ৴৹ = 1 ana (Numerator 1)
  • ৵৹ = 2 ana (Numerator 2)
  • ৶৹ = 3 ana (Numerator 3)
  • ৷৹ = 4 ana (Numerator 4)
  • ৷৴৹ = 5 ana
  • ৷৵৹ = 6 ana
  • ৷৶৹ = 7 ana
  • ৷৷৹ = 8 ana
  • ৷৷৴৹ = 9 ana
  • ৷৷৵৹ = 10 ana
  • ৷৷৶৹ = 11 ana
  • ৸৹ = 12 ana (Numerator One Less Than the Denominator)
  • ৸৴৹ = 13 ana
  • ৸৵৹ = 14 ana
  • ৸৶৹ = 15 ana

This looks like a modified base-4 tally mark system. But, thinking back to what Anshuman Pandey said about elided currency marks, I wonder if this scheme didn't originate in a finer-grained positional system.

Imagine that instead of the taka being divided directly into sixteen ana, it was divided into four virtual "beta", which were themselves divided into four virtual "alpha". Obviously:

  • ana = alpha + beta * 4

But now we have the following encoding:

  • ৴ = 1 alpha (Numerator 1)
  • ৵ = 2 alpha (Numerator 2)
  • ৶ = 3 alpha (Numerator 3)
  • ৷ = 1 beta
  • ৷৷ = 2 beta
  • ৸ = 3 beta (Numerator One Less Than the Denominator)

For beta, the "denominator" is indeed four, to the mysterious U+09F8 "BENGALI CURRENCY NUMERATOR ONE LESS THAN THE DENOMINATOR" suddenly makes sense.

We can now come up with an algorithm for writing out a currency amount according to the scheme described by Anshuman Pandey:

  • Let T, A, G be the number of taka (0 or more), ana (0 to 15), ganda (0 to 19) respectively
  • If T is not zero then
    • Write out the Bengali decimal representation of T
    • If both A and G are zero
      • Write out U+09F2 "BENGALI RUPEE MARK"
      • We're finished
  • Let α be A modulo 4 (0 to 3)
  • Let β be A divided by 4, rounded down (0 to 3)
  • If β is 1, write out U+09F7 "BENGALI CURRENCY NUMERATOR FOUR"
  • If β is 2, write out U+09F7 "BENGALI CURRENCY NUMERATOR FOUR" twice
  • If β is 3, write out U+09F8 "BENGALI CURRENCY NUMERATOR ONE LESS THAN THE DENOMINATOR"
  • If α is 1, write out U+09F4 "BENGALI CURRENCY NUMERATOR ONE"
  • If α is 2, write out U+09F5 "BENGALI CURRENCY NUMERATOR TWO"
  • If α is 3, write out U+09F6 "BENGALI CURRENCY NUMERATOR THREE"
  • If G is zero then
    • Write out U+09F9 "BENGALI CURRENCY DENOMINATOR SIXTEEN"
    • We're finished
  • Write out U+09FB "BENGALI GANDA MARK"
  • Write out the Bengali decimal representation of G
  • We're finished

For our example, T=4, A=15, G=19, α=3, β=3 and the output is:

৪৸৶৻১৯


U+09EA U+09F8 U+09F6 U+09FB U+09E7 U+09EF

This representation is surprisingly concise and totally unambiguous.

Sunday, 30 January 2022

Unicode Trivial U+0950

Codepoint: U+0950 "DEVANAGARI OM"
Block: U+0900..097F "Devanagari"

Next, we come to our first Brahmic script: Devanagari. Devanagari is the most widely-used Brahmic script. Almost 50% of the Indian population use it to write their native language. It is a left-to-right abugida used to write dozens of languages.

It is sometimes glibly called a "washing line" script because, unlike Latin/Greek/Cyrillic scripts that "sit" on top of their baselines, Devanagari also "hangs" from a head line:

Devanagari script
देवनागरी लिपि

Devanagari typography is non-trivial, even when the letterforms are isolated:

Devanagari Type Anatomy [source]

Of course, anthropomorphism in typography isn't limited to Brahmic scripts, but one can take "type anatomy" quite literally:

Design Parameters of Devanagari (Gokhale, 1983) [source]

The flipside of grammatology is phonology.

"Om" is the sound of a sacred spiritual symbol in Indic religions. In Devanagari, in its simplest form, it is written:

Om
ओम
U+0913 U+092E

However, there is also a ligatured codepoint in the Devanagari block named "DEVANAGARI OM":

Om (sign)
U+0950

There are currently many Oms encoded into Unicode. The Devanagari Om was added right at the onset in Unicode 1.0 (1991). This incorporated a wholesale import of the ISCII (Indian Script Code for Information Interchange, 1988) character sets which encode Om as a multi-byte sequence (0xA1 0xE9).  The Unicode Consortium allocated U+0950 "DEVANAGARI OM" to allow one-to-one mapping of ISCII on this basis. Of course, once you allow one Om through...

Saturday, 29 January 2022

Unicode Trivia U+08C8

Codepoint: U+08C8 "ARABIC LETTER GRAF"
Block: U+08A0..08FF "Arabic Extended-A"

Unicode blocks are not allocated sequentially. Consequently, "Arabic Extended-A" (U+08A0..08FF, originating in Unicode 6.1, January 2012) comes numerically after "Arabic Extended-B" (U+0870..089F, originating in Unicode 14.0, September 2021).

Even more confusing, codepoints within a block can be allocated at different times. For example, U+08C8 "ARABIC LETTER GRAF" was assigned in Unicode 14.0 (September 2021); but its neighbour, U+08C7 "ARABIC LETTER LAM WITH SMALL ARABIC LETTER TAH ABOVE" was assigned in Unicode 13.0 (March 2020).

U+08C8 "ARABIC LETTER GRAF" is an addition to the Arabic script for writing the Balti language:

Isolated form of U+08C8 [source]

For example, the Balti word for knife (U+08C8 U+06CC):

[source]

U+08C8 is the only specific addition required for the Arabic script to be able to write Balti, although two other codepoints (U+0F6B "TIBETAN LETTER KKA" and U+0F6C "TIBETAN LETTER RRA", added in Unicode 5.1, April 2008) are needed to write Balti using the Tibetan script.

Balti is spoken in Baltistan, or Little Tibet, a mountainous region in the Gilgit-Baltistan part of Pakistan-administered Kashmir. This may be the origins of the balti curry dish popular in the UK since the nineteen-seventies.

Friday, 28 January 2022

Unicode Trivia U+0891

Codepoint: U+0891 "ARABIC PIASTRE MARK ABOVE"
Block: U+0870..089F "Arabic Extended-B"

Consider this photograph by Tinou Bao:

Fruit Seller
"The guy asked to be photographed"

It could elicit any number of reactions:

  1. What wonderfully colourful fruit!
  2. Are those dates expensive?
  3. I hope he doesn't drop cigarette ash on that fruit
  4. His fruit are suspiciously glossy
  5. Why did he want his photo taken?
  6. That's an interesting symbol above the price

You're in the right place if your reaction was number six.

That symbol, circled in blue, is an Arabic supertending currency symbol for Egyptian piastres. The photo was used as part of the proposal for the addition of two new currency codepoints:

  • U+0890 "ARABIC POUND MARK ABOVE"
  • U+0891 "ARABIC PIASTRE MARK ABOVE"

The proposal was formally submitted in August 2020, accepted in October 2020 and released as part of Unicode 14.0 in September 2021.

[source]


Thursday, 27 January 2022

Unicode Trivia U+0861

Codepoint: U+0861 "SYRIAC LETTER MALAYALAM JA"
Block: U+0860..086F "Syriac Supplement"

The Syriac Supplement block contains letters used for writing Suriyani Malayalam, also known as Syriac Malayalam. This is an Eastern Syriac script with eleven new letters added to capture Malayalam sounds:

The Syriac and Malayalam scripts are almost entirely unrelated; the former is a right-to-left abjad from the Middle East:

Whilst the latter is a left-to-right abugida from Southern Asia:

So the "mashing together" of the two scripts is somewhat surprising and problematic.

For example, the Suriyani Malayalam letter "ja" only appears in isolated form, so the "standard" U+0D1C "MALAYALAM LETTER JA" could have been used, however, the decision was taken to encode a separate U+0861 "SYRIAC LETTER MALAYALAM JA":

Although it may be possible to use U+0D1C within a Syriac environment, a separate encoding is needed [...] so that Syriac vowel marks can be combined with the letter. Furthermore the differing directionalities of the Malayalam and Syriac scripts may cause problems for introducing a Malayalam character directly in Syriac sequences.

Anyone who has tried editing text with mixed left-to-right and right-to-left script will appreciate that last comment.

Suriyani Malayalam is used by Saint Thomas Christians of Kerala in India as a liturgical language. According to tradition, Thomas the Apostle voyaged to Muziris on the Malabar coast (Kerala) in 52 CE, bringing Christianity to the region. This may sound implausible, but Kerala had an established Jewish community at around that time, particular in Cochin. So it is possible for an Aramaic-speaking Jew, such as Saint Thomas from Galilee, to make a trip to Kerala via the maritime Silk Road routes:

[source]

Perhaps not surprisingly, after almost two thousand years, the Saint Thomas Christians have experienced schisms and (sadly fewer) reunifications:

[source]

Wednesday, 26 January 2022

Unicode Trivia U+0840

Codepoint: U+0840 "MANDAIC LETTER HALQA"
Block: U+0840..085F "Mandaic"

The Mandaic alphabet contains 22 letters (in the same order as the Aramaic alphabet) and one digraph:

The alphabet is "rounded up" to a symbolic count of 24 letters by repeating the first letter, U+0840 "MANDAIC LETTER HALQA". It is unusual for a Semitic script in being a true alphabet with letters for both consonants and vowels:

  1. U+0840 "Halqa" = a [vowel]
  2. U+0841 "Ab" = ba
  3. U+0842 "Ag" = ga
  4. U+0843 "Ad" = da
  5. U+0844 "Ah" = ha
  6. U+0845 "Ushenna" = wa [vowel]
  7. U+0846 "Az" = za
  8. U+0847 "It" = eh
  9. U+0848 "Att" = ṭa
  10. U+0849 "Aksa" = ya [vowel]
  11. U+084A "Ak" = ka
  12. U+084B "Al" = la
  13. U+084C "Am" = ma
  14. U+084D "An" = na
  15. U+084E "As" = sa
  16. U+084F "In" = e [vowel]
  17. U+0850 "Ap" = pa
  18. U+0851 "Asz" = ṣa
  19. U+0852 "Aq" = qa
  20. U+0853 "Ar" = ra
  21. U+0854 "Ash" = ša
  22. U+0855 "At" = ta
  23. U+0856 "Dushenna" = ḏ

The eighteenth letter was renamed from "Ass" to "Asz" as part of the original proposal, presumably to stop the giggling at the back of the classroom.

The Classical Mandaic language is still used by Mandaean priests in liturgical rites. It is estimated that there are about 5,500 native speakers. Neo-Mandaic is a modern evolution of Mandaic but generally unwritten. Only a few hundred Mandaeans, located mainly in Iran, speak Neo-Mandaic as a first language.

One of the unintended consequences of the 2003 invasion of Iraq was the diaspora of over 60,000 Iraqi Mandaeans. Today, Sweden has the largest community of any country.

Tuesday, 25 January 2022

Unicode Trivia U+0837

Codepoint: U+0837 "SAMARITAN PUNCTUATION MELODIC QITSA"
Block: U+0800..083F "Samaritan"

The Samaritan script was derived from the Paleo-Hebrew circa 600 BCE and was used alongside the Aramaic script in Judaism until the latter was repurposed as the Hebrew alphabet circa 100 BCE.

Samaritan is a right-to-left abjad with 22 basic consonants and diacritics to mark vowels:

Much is made of the extensive punctuation in the Samaritan script. Here are the fifteen codepoints of the "Punctuation" column (U+0830 to U+083E):