Friday, 28 January 2022

Unicode Trivia U+0891

Codepoint: U+0891 "ARABIC PIASTRE MARK ABOVE"
Block: U+0870..089F "Arabic Extended-B"

Consider this photograph by Tinou Bao:

Fruit Seller
"The guy asked to be photographed"

It could elicit any number of reactions:

  1. What wonderfully colourful fruit!
  2. Are those dates expensive?
  3. I hope he doesn't drop cigarette ash on that fruit
  4. His fruit are suspiciously glossy
  5. Why did he want his photo taken?
  6. That's an interesting symbol above the price

You're in the right place if your reaction was number six.

That symbol, circled in blue, is an Arabic supertending currency symbol for Egyptian piastres. The photo was used as part of the proposal for the addition of two new currency codepoints:

  • U+0890 "ARABIC POUND MARK ABOVE"
  • U+0891 "ARABIC PIASTRE MARK ABOVE"

The proposal was formally submitted in August 2020, accepted in October 2020 and released as part of Unicode 14.0 in September 2021.

[source]


Thursday, 27 January 2022

Unicode Trivia U+0861

Codepoint: U+0861 "SYRIAC LETTER MALAYALAM JA"
Block: U+0860..086F "Syriac Supplement"

The Syriac Supplement block contains letters used for writing Suriyani Malayalam, also known as Syriac Malayalam. This is an Eastern Syriac script with eleven new letters added to capture Malayalam sounds:

The Syriac and Malayalam scripts are almost entirely unrelated; the former is a right-to-left abjad from the Middle East:

Whilst the latter is a left-to-right abugida from Southern Asia:

So the "mashing together" of the two scripts is somewhat surprising and problematic.

For example, the Suriyani Malayalam letter "ja" only appears in isolated form, so the "standard" U+0D1C "MALAYALAM LETTER JA" could have been used, however, the decision was taken to encode a separate U+0861 "SYRIAC LETTER MALAYALAM JA":

Although it may be possible to use U+0D1C within a Syriac environment, a separate encoding is needed [...] so that Syriac vowel marks can be combined with the letter. Furthermore the differing directionalities of the Malayalam and Syriac scripts may cause problems for introducing a Malayalam character directly in Syriac sequences.

Anyone who has tried editing text with mixed left-to-right and right-to-left script will appreciate that last comment.

Suriyani Malayalam is used by Saint Thomas Christians of Kerala in India as a liturgical language. According to tradition, Thomas the Apostle voyaged to Muziris on the Malabar coast (Kerala) in 52 CE, bringing Christianity to the region. This may sound implausible, but Kerala had an established Jewish community at around that time, particular in Cochin. So it is possible for an Aramaic-speaking Jew, such as Saint Thomas from Galilee, to make a trip to Kerala via the maritime Silk Road routes:

[source]

Perhaps not surprisingly, after almost two thousand years, the Saint Thomas Christians have experienced schisms and (sadly fewer) reunifications:

[source]

Wednesday, 26 January 2022

Unicode Trivia U+0840

Codepoint: U+0840 "MANDAIC LETTER HALQA"
Block: U+0840..085F "Mandaic"

The Mandaic alphabet contains 22 letters (in the same order as the Aramaic alphabet) and one digraph:

The alphabet is "rounded up" to a symbolic count of 24 letters by repeating the first letter, U+0840 "MANDAIC LETTER HALQA". It is unusual for a Semitic script in being a true alphabet with letters for both consonants and vowels:

  1. U+0840 "Halqa" = a [vowel]
  2. U+0841 "Ab" = ba
  3. U+0842 "Ag" = ga
  4. U+0843 "Ad" = da
  5. U+0844 "Ah" = ha
  6. U+0845 "Ushenna" = wa [vowel]
  7. U+0846 "Az" = za
  8. U+0847 "It" = eh
  9. U+0848 "Att" = ṭa
  10. U+0849 "Aksa" = ya [vowel]
  11. U+084A "Ak" = ka
  12. U+084B "Al" = la
  13. U+084C "Am" = ma
  14. U+084D "An" = na
  15. U+084E "As" = sa
  16. U+084F "In" = e [vowel]
  17. U+0850 "Ap" = pa
  18. U+0851 "Asz" = ṣa
  19. U+0852 "Aq" = qa
  20. U+0853 "Ar" = ra
  21. U+0854 "Ash" = ša
  22. U+0855 "At" = ta
  23. U+0856 "Dushenna" = ḏ

The eighteenth letter was renamed from "Ass" to "Asz" as part of the original proposal, presumably to stop the giggling at the back of the classroom.

The Classical Mandaic language is still used by Mandaean priests in liturgical rites. It is estimated that there are about 5,500 native speakers. Neo-Mandaic is a modern evolution of Mandaic but generally unwritten. Only a few hundred Mandaeans, located mainly in Iran, speak Neo-Mandaic as a first language.

One of the unintended consequences of the 2003 invasion of Iraq was the diaspora of over 60,000 Iraqi Mandaeans. Today, Sweden has the largest community of any country.

Tuesday, 25 January 2022

Unicode Trivia U+0837

Codepoint: U+0837 "SAMARITAN PUNCTUATION MELODIC QITSA"
Block: U+0800..083F "Samaritan"

The Samaritan script was derived from the Paleo-Hebrew circa 600 BCE and was used alongside the Aramaic script in Judaism until the latter was repurposed as the Hebrew alphabet circa 100 BCE.

Samaritan is a right-to-left abjad with 22 basic consonants and diacritics to mark vowels:

Much is made of the extensive punctuation in the Samaritan script. Here are the fifteen codepoints of the "Punctuation" column (U+0830 to U+083E):


Monday, 24 January 2022

Unicode Trivia U+07C1

Codepoint: U+07C1 "NKO DIGIT ONE"
Block: U+07C0..07FF "NKo"

As reported by Dr Dianne White Oyler in "A Cultural Revolution in Africa: Literacy in the Republic of Guinea since Independence" (2001), the N'Ko script was developed by Souleymane Kanté in 1949, partly in response to

a 1944 challenge posed by the Lebanese journalist Kamal Marwa in an Arabic-language publication, Nahnu fi Afrikiya [We Are in Africa]. Marwa argued that Africans were inferior because they possessed no indigenous written form of communication. His statement that "African voices [languages] are like those of the birds, impossible to transcribe" reflected the prevailing views of many colonial Europeans. Although the journalist acknowledged that the Vai had created a syllabary, he discounted its cultural relevancy because he deemed it incomplete. [Page 588]

Kanté discarded both Arabic and Latin scripts as unable to transcribe all the characteristics of the Mande languages. Having developed a completely novel alphabet instead,

he called together children and illiterates and asked them to draw a line in the dirt; he noticed that seven out of ten drew the line from right to left. For that reason he chose a right-to-left orientation. In all Mande languages the pronoun n- means "I" and the verb ko represents the verb "to say". [Page 589]

So "N'Ko" means "I say" in all the target languages.

The right-to-left mantra extends not only to words, but to digits and numbers too. The ten digits zero to nine (U+07C0 "NKO DIGIT ZERO" to U+07C9 "NKO DIGIT NINE") face right:

N'Ko digits (top), Western Arabic (middle), Eastern Arabic (bottom)

This is particularly noticeable with U+07C1 "NKO DIGIT ONE": '߁'

Not only that, but the least significant digits of multi-digit N'Ko numbers are on the left, unlike almost all other writing systems. Latin, Greek, Arabic and Hebrew numbers place the least significant digit on the right, even though the latter two scripts are written right-to-left.

Consider the improbable phrase "There are 12345 eggs":

There are 12345 eggs = English

Υπάρχουν 12345 αυγά = Greek

 يوجد ١٢٣٤٥ بيضة = Arabic

יש 12345 ביצים = Hebrew

߁߂߃߄߅ ߞߟߌ߫ ߦߋ߫ ߦߋ߲߬ = N’Ko

In case of tofu:

Note that the order of the codepoints for "1", "2", "3" "4" and "5" occur in ascending memory order in all cases. For example:


At first, I wasn't sure how much "support" the Unicode standard gives for this type of anomaly. UCD's sister project CLDR (Common Locale Data Repository) has very little to say about N'Ko. There is scope for algorithmic number formatting, but I didn't find anything specific.

However, after a bit of thought I realised that, because directionality is a property of each codepoint and not of the script of the codepoints, digit ordering in N'Ko works "out of the box".

Consider these bidirectional class fields ("bc") from the UCD:

  • Latin
    • "A" (U+0041 "LATIN CAPTIAL LETTER A") =  "L" = strong left-to-right
    • "1" (U+0041 "LATIN CAPTIAL LETTER A") = "EN" = European number (left-to-right)
  • Greek
    • "α" (U+03B1 "GREEK SMALL LETTER ALPHA") =  "L" = strong left-to-right
  • Arabic
    • "ا" (U+0627 "ARABIC LETTER ALEF") = "AL" =  Arabic letter (right-to-left)
    • "١" (U+0661 "ARABIC-INDIC DIGIT ONE") = "AN" =  Arabic number (left-to-right)
  • Hebrew
    • "א" (U+05D0 "HEBREW LETTER ALEF") = "R" = strong right-to-left
  • N'Ko
    • "ߊ" (U+07CA "NKO LETTER A") = "R" = strong right-to-left
    • "߁" (U+07C1 "NKO DIGIT ONE") = "R" = strong right-to-left

Unlike the other digits, N'Ko digits are marked as strongly right-to-left. The only other examples in Unicode 14.0 I could find were Adlam digits (1989).

Another interesting codepoint from the Unicode "NKo" block is U+07F7 "NKO SYMBOL GBAKURUNEN":

It's a decorative punctuation symbol used to mark the end of a major section of text and represents the three stones holding a cooking pot over a fire:

[source]

Finally, there can't be many alphabets that have their own day: April 14.

[Many thanks to Coleman Donaldson for help with the N'Ko language]

Sunday, 23 January 2022

Unicode Trivia U+0780

Codepoint: U+0780 "THAANA LETTER HAA"
Block: U+0780..07BF "Thaana"

The Thaana script is used to write the Maldivian language. According to Wikipedia, it's an abugida with no inherent vowel. According to the ISO standard, it's a right-to-left-written alphabet (as indicated by the hundreds digit of its numeric ISO-15924 code "170").

It first appeared in about 1705 CE and seems have been developed with obfuscation in mind. The alphabet order is arbitrary and the consonant letterforms are derived from numeric figures:

On the top row, in white, are the 24 basic consonants in Thaana alphabetical order. These are the 24 consecutive Unicode codepoints U+0780 "THAANA LETTER HAA" to U+0797 "THAANA LETTER CHAVIYANI".

The second row shows the Arabic-Indic digits one to nine in blue and the Dhives Akuru digits one to six in red. Dhives Akuru was a Maldivian script used before Thaana. The main part of the alphabet looks very much like a simple replacement cipher.

An early version of the Thaana script, Gabulhi Thaana, was written scriptio continua, that is, without inter-word spacing or punctuation. This sounds like an absolute nightmare but was quite common in classical Greek and Latin. Before mechanical printing, Arabic was also written without spacing. This is, perhaps, why many writing systems have distinct letterforms for final letters in words.

According to "Scripts of Maldives", the early Thaana script, Gabulhi Thaana, got its name from the Maldivian word "gabulhi" meaning the in-between stage of a coconut, when it is neither fully ripe nor quite tender. Hence the idea of "immature" or "not fully-formed".

Saturday, 22 January 2022

Unicode Trivia U+0753

Codepoint: U+0753 "ARABIC LETTER BEH WITH THREE DOTS POINTING UPWARDS BELOW AND TWO DOTS ABOVE"
Block: U+0750..077F "Arabic Supplement"

Syriac is not the only script that makes extensive use of diacritics. The spread of the Arabic script throughout the world means it is used for diverse languages, many of which have sounds not found in Arabic. Part of the "Arabic Supplement" block contains a column "Extended Arabic letters" with the annotation:

These are primarily used in Arabic-script orthographies of African languages.

One codepoint, U+0753, has the somewhat precise name of "ARABIC LETTER BEH WITH THREE DOTS POINTING UPWARDS BELOW AND TWO DOTS ABOVE". When I render that codepoint using "Noto Sans Arabic" on my PC, I get this:

Noto Sans Arabic (2.004)

When I render it with a default local font, I get:

Arial (7.00)

Spot the difference!

There's definitely a discrepancy in the orientation of the lower dots, but which is correct? I came up with three possibilities:

  1. I have old/corrupt font files installed on my PC
  2. The name of the Unicode codepoint is incorrect
  3. The orientation of the lower dots doesn't really matter, so there is no issue
  4. One of the font glyphs is incorrect

Initially, I did indeed think it was an old version of Noto Sans Arabic installed on my machine. But I updated my local version of Noto Sans Arabic to 2.009 with the same results. Google web font specimens confirmed the issue is with Noto Sans Arabic in general:

Three of the four specimens suggest the name of the Unicode codepoint is probably correct. I checked that there are no similarly-named codepoints; there is no "ARABIC LETTER BEH WITH THREE DOTS POINTING DOWNWARDS BELOW AND TWO DOTS ABOVE"

I couldn't really imagine that, carefully named as it is, the orientation of the lower dots in U+0753 was unimportant.

I then checked Unicode Updates and Errata but found no references to this or nearby codepoints.

So the finger of suspicion fell on the glyph within the Noto Sans Arabic font being incorrect. FontForge confirmed this:


I looked through the issues reported for Noto fonts, but found nothing, so I submitted a new one.

Of course, this has only a passing connection to the Unicode standard. But one can easily imagine the amount of noise that has to be ploughed through by the committee along the lines of "My text doesn't get displayed how I expected" just to get to genuine issues with the Unicode standard itself.

According to Wiktionary, U+0753 "ARABIC LETTER BEH WITH THREE DOTS POINTING UPWARDS BELOW AND TWO DOTS ABOVE" is

The third letter of the Hausa alphabet in ajami script, equivalent to Latin script c.

I was initially a bit suspicious of this. Both Omniglot and Wikipedia suggest that the three dots go above that letter, making it more like U+062B "ARABIC LETTER THEH". However, Richard Ishida points out that there are lots of subtle local variations and the initial Unicode proposal shows an "ARABIC LETTER BEH WITH THREE DOTS POINTING UPWARDS BELOW AND TWO DOTS ABOVE" in Figure 5. The proposal cites "Using Arabic Script in Writing the Languages of the Peoples of Muslim Africa" (1992) by Mohamed Chtatou:

"Figure 5" (Chtatou, 1992)

Richard Ishida again:

Unicode policy for the Arabic script is to encode fully precomposed characters rather than to use combining characters for ijam.

It would appear that the task of supporting more obscure and/or infrequent Arabic script glyphs in Unicode (and in fonts) can only get harder.