Decadent beginnings

Happy MMX, or ˌβί, or 二千十, or if your character set will handle it, ፳፻፲. Those are the Roman, Greek, Chinese, and Ethiopic numerals for 2010. And a Happy New Year to those of you who feel like being happy. My apologies for the lengthy absence here – rest assured that I have not given up the ghost, and have a backlog of interesting posts and articles for this time of new beginning.

Looking at the numerals above, note that each only has three graphemes (characters), as opposed to the four that Western numerals require. While the Roman numerals are generally less concise for writing numbers than most other systems – one of the reasons they are stigmatized in Western thought – for round and nearly-round numbers, Western numerals generally require more symbols than other systems, because of the requirement that unused place-values have zeroes to occupy the space. Yet while each of the other four systems requires only three graphemes, they are not at all the same, but express 2010 in four distinct ways:

Roman MMX = 1000 + 1000 + 10
Greek ˌβί = (1000×2) + 10
Chinese 二千十 = (2×1000) + (1x)10
Ethiopic ፳፻፲ = (20×100) + 10

But enough of this variability – you can read all about it later this month when my book finally emerges from the depths. I want to talk about another, equally interesting form of numerical variability: How do you say the number 2010?

The great numerical question of the last decade was whether the millennium began in 2000 or 2001. (For the record, I’m very strongly in the 2000 camp, on the basis that the turning over of the calendrical odometer was by far more culturally significant.) A close second is what we ought to call the decade as a whole: the naughts, or the noughts, or the aughts, or the aught-noughts, or the naughties, or any number of more ridiculous and facetious answers. And now that the decade is done, we seem to have done quite well without an agreed-upon term – although it will probably be more important to have one when the decade is being considered retrospectively – no one knew what eighties music was until 1990 at the earliest. Similarly, whether this new decade will be the tens or the teens, as discussed at David Crystal’s blog yesterday, isn’t an issue of present concern. But what to call this very year 2010 certainly is!

If you are a speaker of English, you have several options:
a) two thousand ten
b) two thousand and ten
c) twenty hundred ten
d) twenty hundred and ten
e) twenty ten

One of the not-so-dark secrets of the English language is that virtually any number above 100 has multiple, grammatically valid readings. If I were to be extremely radical I could contend that English has two parallel lexical numeral systems – but let’s not go that far quite yet. Let’s start by reducing the five variants to three by noting that the difference between a and b, and between c and d, is only the presence/absence of the word ‘and’. Crystal asserts that the versions with and are characteristically British, while the and-less versions are American, but this runs counter to my ethnographic experience working here in Detroit, where students from an ordinary public school education are brought into a mathematics program where and is strongly stigmatized, and need to learn not to use and in contexts where it would be natural for them. I’m not denying that this national pattern may have held true at one time, or in particular contexts (e.g., radio/TV broadcasts), but it surely is not as clear-cut as Crystal makes it sound.

In particular, year-names ending in -0x tend to take and for the very sensible reason that two thousand and eight clearly delineates the end of a numeral-phrase whereas two thousand eight invites the possibility that the speaker is about to continue – e.g. two thousand eight hundred. Although I should add that the potential for real confusion is quite low, on contextual grounds, and that one does occasionally hear and-less readings of 2001-2009. But let’s admit that one can use and in these phrases, or not, as one wishes.

Let’s reduce the variability further by noting the extremely unusual nature of the numeral-phrase *twenty hundred. Many English numerals ending in 00 between 1100 and 9900 are conventionally expressed in hundreds, not thousands – eleven hundred, sixty eight hundred, etc. This is probably because they are more compact than one thousand one hundred, six thousand eight hundred, etc. That brevity is the relevant cognitive criterion is demonstrated by the exceptions, the thousands from 2000 through 9000. Two thousand is shorter (3 syllables) than *twenty hundred (4 syllables) – remember, we are talking about verbal expressions here, so syllable length is the relevant criterion. So even though we could say nineteen hundred (and) ten, *twenty hundred (and) ten sounds decidedly odd. So let’s just forget about them.

But wait! What if we omit the word hundred entirely – as it is possible to do in almost any context in English. So, for instance, I can say, “I get paid eight twenty five a week” and “I get paid eight twenty five an hour” and be understood as saying that I make $825.00 a week but $8.25 an hour, solely from contextual information. I am presumably in this case working 100 hours a week, which may be a slight exaggeration. In any event, when talking about years there is rarely even the slightest chance of lexical ambiguity, and so nineteen hundred seventy-four is almost always reduced to nineteen seventy four – in fact, the only place where the full expression is encountered is in extremely formal or prestigious contexts such as official proclamations, diplomas, etc.

Note, also, that only the and-less versions of these phrases can be so reduced: *nineteen and seventy-four or *twenty and ten are unacceptable (both in Britain and the US). The simplest explanation (though not the only one) for this phenomenon is that when one is abbreviating, it makes sense that one would abbreviate maximally, rather than adding the unnecessary and. Another potential factor is that phrases like twenty and ten are found in English sentences such as, “The first-class and regular seats cost twenty and ten dollars, respectively,” although I wouldn’t stake my professional reputation on this potential ambiguity being all that important.

So that leaves us with two thousand (and) ten or twenty ten, which of course is not a surprise. On the grounds of compactness, twenty ten clearly wins out; on the grounds of ambiguity, two thousand and ten seems preferable, and two thousand ten might represent a good compromise. But these are not the only factors to consider – also relevant is how we say (or expect to say) surrounding year-names; if we say two thousand eight and two thousand nine, then twenty ten is potentially jarring. And we decidedly do not say twenty eight and twenty nine for 2008 and 2009, for the obvious reason that they can readily be confused with twenty-eight (28) and twenty-nine (29)!

The lesson: Whatever reading you might choose will be on the basis of one or more criteria, and there is no ultimate good or ‘proper’ choice – every choice will be less than maximally preferable on at least one of those criteria. So do what you like!

And the problem gets even worse, because depending on the context in which numerals are used, they may have additional valid readings. For instance, nominal number representations – ones that serve as labels rather than as counts of things – are frequently (though not always) read digitally rather than lexically. I’m working on a paper on the reading of phone numbers, showing the ways in which digital and lexical representations of numbers interface (and interfere) with one another (see pilot data here). To illustrate the point: my paper is tentatively entitled ‘Jenny’s Revenge: Eight Hundred Sixty Seven, Five Thousand Three Hundred (and) Nine‘, which will make sense largely to those born before one thousand nine hundred (and) seventy five. But year numbers are not just labels, but rather denote a place in a series – 2010 is defined as the year after 2009 and before 2011 – and are less readily digitized. Consider the following:

two zero one zero
two oh one oh
twenty one zero
twenty one oh

None of these is even remotely acceptable as a reading of the year number 2010. More strikingly, even though nineteen oh nine is the preferred pronunciation of 1909 for most people, 2010 cannot be acceptably read as *twenty one oh, I suspect, by anyone. But if my phone number were 867-2010, at least the first two variants would be acceptable, and in fact might be preferred, because they are clear representations of each digit. When one is speaking on the phone, for instance, one tends towards maximally clear and distinct representations of each digit in the expectation that one’s listener may be writing the number down – and no one writes phone numbers out lexically. With year numbers, this expectation rarely if ever holds true. There has been no scholarship to date on the question of which numerical representations are acceptable (or not), preferred (or not) in various contexts.

To return to our beginnings, the problem of multiple representations for the same number also arises in numerical notation, although due to the structure of Western numerals, less so than in other representations. In Chinese, for instance, there is both an informal (二千十 = 2 1000 10) and formal (二千一十 = 2 1000 1 10) variant – the difference being the deletion of the morpheme/grapheme for one in the tens position. Similarly, in the classical Roman numerals one could use subtractive notation far more widely than is presently taught in schools – for instance, check out this inscription where 88 is written as XXCIIX instead of the expected (modern) LXXXVIII. And even though ‘we all know’ that the Roman numeral for 4 is IV, Roman numeral clocks even to this day read IIII (but IX for 9)!

So in summation, may you have a lexically ambiguous, but nonetheless pleasant, two-oh-one-oh through two-oh-one-nine.

Edit to add: Shortly after I posted, Mark Liberman over at Language Log offered his own mirthful take on the issue in his post, ‘2010‘, which you should all go read right now, if you haven’t already. I shall have to register a complaint, however, in that his post is (ordinally) #2012 over at the Log (see the URL) – they surely should have instructed their bloggers properly on the importance of this numerical correlation!

Thanksgiving link roundup

Today, most of my colleagues are toiling away in an attempt to cook and carve some sort of fowl. Me, well, I’m Canadian, and even though I work over in the Dark Nether Reaches and get to enjoy its three-day week, I live over here in Canada’s Deep South and get to … have a flu shot and catch up on posting some links of interest?

I don’t have much to add about the sad passing of Dell Hymes last week. I didn’t know him but I know many people who did, and no one who purports to be a linguistic anthropologist (or sociolinguist … or anthropological linguist … or …) can possibly be ignorant of his work. The NYT description of him as a “Linguist with a Wide Net” is utterly evocative and has me imagining it literally. He will be missed, but his legacy on the discipline will remain vital for decades.

While Turkey officially switched from the Arabic to the Roman alphabet in the 1920s, at the same time it prohibited the use of letters not used to represent Turkish – which includes the ‘ordinary’ Roman letters Q, W, and X. While sometimes portrayed as a ban on those letters specifically, it is a more general ban on non-Turkish characters, as far as I can tell, which would seem to prohibit all sorts of texts. Ostensibly designed to promote national unity and secular rule, the law has only been applied to Turks of Kurdish descent. As someone who until last year was a resident of a region where texts written in my native language are under severe legal constraints, this has been a matter of some interest and concern to me for a few years now. Mark Liberman tells us more over at Language Log.

Researchers at the University of Edinburgh are investigating the cultural evolution of language, arguing that language change is patterned by the biological constraints of the human brain – in other words, language changes to accomodate itself to the sorts of brains we possess. They are examining this idea experimentally using an artificial language of simple syllables used to describe alien-looking fruit … which is not as bizarre as I may have made it sound. Edinburgh is doing a lot of exciting work these days in linguistics, what with Jim Hurford, Simon Kirby, and Geoff Pullum (among others) housed there.

Relatedly, Marc Changizi claims (following up on work he has been doing for the past several years) that there are strong cognitive / evolutionary constraints on the graphemes (discrete written units) of writing systems, creating similiarites across writing systems that reflect the cultural evolution of graphemes to accomodate the needs and capacities of the human brain. I have more doubts about this one, which I may talk about in more detail – basically my concern is that the cross-cultural analysis is weak and inadequately accounts for borrowing (Galton’s problem). But it’s interesting work that deserves some attention. Hat tip to The Lousy Linguist for both this item and the previous one).

Lastly, Alun Salt has recently published a very interesting paper, ‘The Astronomical Orientation of Ancient Greek Temples‘ arguing for a more rigorous statistical approach to archaeoastronomy and establishing solar orientations. He’s not the first to use statistical analysis in archaeoastronomy but he does note with some dismay that there is generally insufficient concern with quantitative reasoning among archaeoastronomers to be able to apply statistical tests effectively. Salt highlights some of the complexities in making these determinations – leap second daters, take note! More important than the article itself, though, is its venue, the open-access PLoS ONE. Although ‘cheap’ by open-access standards, the fact that authors must pay ‘only’ $1350 to cover publication costs is, I think, problematic in humanities and social science disciplines where grants are small and getting proportionally smaller.

To my American friends, good luck with your birds, and thanks for reading!

Leap second dating

Archaeologists have long been used to being dependent on physicists for radiometric dating, but gravimetric dating? A new paper deposited last week to arXiv suggests so:

The physical origin of the leap second is discussed in terms of the new gravity model. The calculated time shift of the earth rotation around the sun for one year amounts to $\displaystyle{\Delta T \simeq 0.621 s/ year}$. According to the data, the leap second correction for one year corresponds to $\Delta T \simeq 0.63 \pm 0.03 s/ year $, which is in perfect agreement with the prediction. This shows that the leap second is not originated from the rotation of the earth in its own axis. Instead, it is the same physics as the Mercury perihelion shift. We propose a novel dating method (Leap Second Dating) which enables to determine the construction date of some archaeological objects such as Stonehenge.

So how do we get from leap seconds to Stonehenge? The authors are claiming that the predictions of general relativity allow us to estimate the time shift of the earth’s rotation around the sun at ~ 10.3 minutes / 1000 years. The same process that leads to us adding ‘leap seconds’ to the calendar allows us to measure the difference in sunrise / sunset over long time periods. Now, I’m not a physicist so I can’t follow all that other stuff, other than understanding that the shift in Mercury’s perihelion is one of the demonstrations of general relativity used by Einstein. So let’s grant it.

The authors claim that “some of the archaeological objects may well possess a special part of the building which can be pointed to the sun at the equinox.” And if you expect the alignment to occur at sunrise but you’re off by 10 minutes, well, it must be because it was built 1000 years ago, right? But with a shift of 10 minutes per millennium, you’ve got a new problem, namely that you’re going to get a whole bunch of false positive solar alignments. The authors’ assumption that we know in advance which objects are aligned to particular solar events is incorrect.

Moreover, the authors note correctly that “It should be noted that the new dating method has an important assumption that there should be no major earthquake in the region of the archaeological objects.” Indeed, one would need to ensure that there had been virtually no movement of the celestially-aligned features – post-glacial rebound, for instance, can cause massive shifts in elevation over the time scale we’re considering, not to mention garden-variety post-depositional processes. And bear in mind that an alignment requires at least two archaeological features that can be demonstrated to be associated with one another. The error bars would be HUGE.

Finally, the idea that new dating techniques allows physical scientists to ‘tell’ archaeologists the date of their stuff is incorrect. When radiocarbon dating was developed in the late 40s, it required evidentiary confirmation, confirmation which could only come from dating archaeological materials of known age – in this case, Egyptian materials dated non-radiometrically (e.g. papyri containing dates), which could confirm that the rate of C-14 formation was (more or less) constant (Trigger 2006: 382). We don’t have anything like that here.

I’m not saying that this idea is so ridiculous that no one should try it – though it might be. But my advice to astrophysicists is to take a deep breath and consult an archaeologist before claiming to have developed a new dating technique. In other words: look before you leap.

Fujita, Takehisa. and Naohira Kanda. 2009. Physics of leap second. arXiv:0911.2087v1.
Trigger, Bruce. 2006. History of archaeological thought, 2nd ed. New York: Cambridge University Press.

(Hat tip to the weird and wacky folks at Improbable Research)

Claude Lévi-Strauss, 1908-2009

Word today that the renowned anthropologist Claude Lévi-Strauss has died this past weekend at the age of 100 (NYT obituary here). I posted last year in honour of his centenary. Read often but rarely well, his influence on the discipline is enormous and it is nearly impossible to conceptualize social anthropology without his work.

Pseudo-disciplines

There is a fascinating short essay ‘Ancient History and Pseudoscholarship‘ over at Livius.org. I don’t share the author’s belief that most laypeople are able to distinguish pseudoscholarship from professional work, nor that there is an absolute decline in pseudoscience over the past few decades. I do absolutely agree that the prevalence of faulty reasoning and uncritical use of evidence by scholars in the historical and social sciences is far more problematic than the more outlandish pseudoscientific beliefs such as the ancient astronaut hypothesis. And it will come as no surprise to you that I share the author’s conviction that a robust and broad training (in my work, that would include linguistics, archaeology, history, anthropology, and cognitive science) in order to allow professionals to avoid pseudoscientific errors in their own research and teaching.