Saturday, September 19, 2026

Test report


This following the first report at reference 1. It might be easier to follow if the reader has access to a Windows computer including the 'Character Map' tool, on my computer to be found on the task bar at the bottom of the screen.

As noted at reference one, this all stems from my having long been both slightly irritated and slightly puzzled by the absence of the set union symbol from Microsoft's 'Character Map'. Intersection present and shown in the snap above, union absent, despite diligent search of the 4,000 odd characters which are there. Are there, that is, in the default view. I did not get into geekery until later.

I make quite a lot of use of Character Map to fetch characters which are not on the keyboard, most often accented letters from the French, less often things like simple fractions. I usually paste them into Word as plain text, which avoids disturbing the Word font. 

Added to which, I keep a short list of more obscure characters in a text file, from which it is easier to retrieve them than by poking around in the depths of Character Map. I use the 'Notepad' tool for this purpose - which strips out font, which reduces the risk of disturbing font in the target document in Word.

On the present occasion, I thought to ask Gemini by set union (∪) was absent while set intersection (∩) was present. One can use the letter 'U', but it gets treated as a letter with the usual printer's decorations, which I find irritating.

So another hangover from the days of DOS. I think I have read somewhere that there is still a fair amount of the original DOS code lurking in the depths of Windows, despite all the rewrites since. The past lives on, warts and all. Or, pushing further back, from the days when Microsoft hijacked what became Windows from IBM. 

Reference 2 says something about the IBM heritage, from where the snap of code page 437 above is lifted. Set intersect is to be found in the bottom row.

Unsolicited, Gemini goes on to suggest three ways around the problem. A cunning wheeze involving a hex code and the use of 'Alt+X'. Using something called the 'Symbol Picker'. And last, poking around in the depths of Character Map. I try the first and it works, but seems a bit geeky for everyday use; the second seems terribly clumsy - so I settle for the third.

Which after a few missteps - Gemini is not word perfect on what actually happens on my laptop - turns out to involve using a special font called 'Cambria Math'.

I put the two characters into my special list. And then wondered about how they survived when pasted into the Times New Roman world of my Word document, this font not supporting one of the two characters. Gemini is still on the case. And Word is on the case, going in for some under the hood trickery.

Gemini closed with the snap at the head of reference 1 and all that remained to do was to check that all this survived into Blogger, that is to say here, which it does.

I thought Gemini had done a pretty good job. Not word perfect, but near enough to get me there without too much bother. I dare say I could have managed without, but I hadn't bothered in the all the years that this matter had been bothering me!

But it is a simple case, where the objective is to do something straightforward: it is quite clear whether one has achieved the task - in this case the insertion of the set union character - or not, and there are no politics (as it were) associated with that achievement (∪). In this, quite unlike the case when one is asking Gemini to help with some obscure problem in, say, anthropology. In this case verification is hard, if possible at all, and one is much more likely to care how Gemini arrived at his answer. Here, all you care about is whether the answer worked.

PS 1: the opening graphic was created with some more help with the text from Gemini.

PS 2: there is a digression on Unicode to come. A digression where I did not think to use Gemini and which soaked up rather a lot of time. Maybe I should have used him.

PS 3: along the way, I created the snap above, with the idea being to summarise what I had learned from Gemini about how Word and Notepad handle fonts, with named fonts being the way that Word renders characters for the display or the printer. And with both Word and Notepad documents being thought, for present purposes, as being made up of a string of characters. The key points being, first, that any particular font makes a selection from the 1,000,000+ characters in UTF-8 - with the Times New Roman that I usually use selecting under 5,000 of them. And second, that Word insists on a character being associated with a font. No floaters.

For some tutorial on elementary rendering, not involving fonts, see reference 2.

References

Reference 1: https://psmv6.blogspot.com/2026/09/test-01.html.

Reference 2: https://en.wikipedia.org/wiki/Code_page_437.

Reference 3: https://en.wikipedia.org/wiki/UTF-8. ‘… UTF-8 supports all 1,112,064 valid Unicode code points using a variable-width encoding of one to four one-byte (8-bit) code units…’. Variable length coding, with the common code points – that is to say the 128 ASCI codes – getting a short code.

Group search key: aisk.

No comments:

Post a Comment