An interview with Annemarie Verkerk and Russell Gray on “Enduring constraints on grammar” (2026)

In the following interview, Annemarie Verkerk and Russell Gray answer Martin Haspelmath’s questions about their paper “Enduring constraints on grammar revealed by Bayesian spatiophylogenetic analyses” (Verkerk et al. 2026)

Martin: Congratulations on this new paper on grammatical universals based on the Grambank data! My first question is whether you were surprised by the results.

Annermarie and Russell: Yes and no. We were not really surprised that many proposed universals weakened once we properly controlled for genealogy and geography (remember, one of us was an author on the Dunn et al. 2011 paper). Languages are not independent marbles in a jar; they inherit things and borrow things. Modelling those dependencies is important. What surprised us was how many “universals” survived and which ones survived. About a third of the 191 testable universals came through quite robustly, especially hierarchical universals and narrow word order universals. So the result is neither “anything goes” nor “universal grammar everywhere”. It is the more interesting middle ground: grammar is flexible, but not unconstrained. The fact that most hierarchical universals and some but not all of the narrow word order universals survived was surprising too – there were a lot of proposals we tested aside from these, but hardly any survived scrutiny. This is interesting because the constraints we do find have certain characteristics.

Your paper does not endorse any specific type of explanation for the cross-linguistic patterns, even mentioning generative approaches such as Cinque’s and Harley & Ritter’s “feature geometry”. But is the Grambank approach compatible with generative grammar?

Yes, in principle. Grambank is not built on generative theory, and our paper does not try to choose between generative, functional, processing, or diachronic explanations for our findings. It tests whether proposed surface-level typological generalisations actually hold once we control for shared inheritance and geography. Some generative accounts, such as Cinque’s work on word order (e.g. Cinque 2013) or Harley and Ritter’s (2002) feature geometry, make predictions that could be reflected in typological patterns. But our analyses of the Grambank data will not by themselves tell you whether the mechanism is Merge, processing efficiency, grammaticalisation, or even Martian intervention. They tell us which patterns are real enough to deserve explanation in the first place.

In the “Research Briefing”, you say that “the sampling approach traditionally used by linguists should be abandoned”, but doesn’t your paper actually show that your “big data” approach converges with earlier sampling-based approaches? After all, no previous study looked at as many languages as you did, and those universals that were based on worldwide samples did best (compared to those based just on European languages).

Partly, yes. We wouldn’t want to caricature earlier typologists as people wandering around with three European languages and a butterfly net, but there has been some cherry-picking and a Eurocentric focus in the past. However, the best earlier work was often remarkably good. In fact, it is pleasing that we find strong support for certain claims, many of which have been of key interest in the field of linguistic typology for decades. But the point is that sampling is a poor substitute for diachronic modelling. Sampling throws away data, reduces statistical power, and tells us little about historical pathways. Our approach says: use all the data you can, then explicitly model the non-independence. So yes, some classic typology comes out looking good — but now we can say why, how strongly, and with what caveats. And we can also point towards why other proposals were not supported.

Your paper crucially relies on the world tree proposed by Bouckaert et al. (2022). Can you say a bit more about this tree (or these trees)? For many readers, this will be the most surprising aspect of your methodology, because it is well-known that there are hundreds of separate families (including isolates). So this is not a genealogical tree, and yet you apply your phylogenetic methods to it.

This is a very important point. The Bouckaert et al. authors do not claim that we have reconstructed a classical genealogical tree for all the world’s languages in the same sense as we reconstruct Indo-European or Austronesian family trees. That would be absurd. The tree is better thought of as a structured model of relatedness among languages: within families, it captures known genealogical information from Bayesian phylogenies and the Glottolog classification; and across families, it provides a way to model covariance in the remote past rather than pretend that all families and areas are independent. It uses archaeological and genetic information, along with a spatial diffusion model, to model possible deeper connections. This means it is more likely, in the Bouckaert et al. tree set, that the Japonic languages are more closely related to the Korean languages than to Indo-European languages, for example. Importantly, we did not rely on a single tree. We ran analyses across a large sample of trees, so the enormous uncertainty in deep linguistic relationships is built into our results. We also checked the robustness of our inferences by using the traditional categorical controls for language family. The results were very similar. So the tree is not a magical Proto-World claim. It is a statistical scaffold — a best guess given current knowledge. It is useful, imperfect, but much better than ignoring relatedness altogether or just controlling for language family or genus.

The correlated-evolution method effectively excludes inheritance as a confounding factor, because only changes on branches are counted. But how do you deal with the massive effect of contact? Few changes happen in isolation from neighbouring languages. For example, Greek, Italic and Germanic do not seem to share a common ancestor (much later than Proto-Indo-European), and yet their (hypothetical) changes from OV order to VO order can hardly be said to be independent (given that in Europe, VO order is a general areal feature).

Contact is the hard problem, and we do not pretend to have solved it completely. We explicitly model geographical proximity in the spatiophylogenetic analyses, so areal effects are not simply ignored. But, of course, geographic distance is only a proxy for a wide range of spatial processes (diffusion, trade, migration, and ecological and cultural similarities). For example, Pama-Nyungan and non-Pama-Nyungan languages of Australia are said to have converged through contact, which is exactly the kind of case where a purely genealogical model would be too optimistic about independence. That is why we used complementary analyses and why we are cautious about the mechanisms driving change. The strongest claims in the paper are therefore not “this change happened independently in every case”, but “after controlling as well as we could for genealogy and geography, these feature combinations repeatedly emerge.” New methods are being developed that will allow us to do so even more effectively. The generalised dynamic coevolutionary models, recently developed by Erik Ringen and colleagues, are an exciting development. These methods can be applied to more than two variables, handle both discrete and continuous data, and explicitly model genealogical and geographic dependencies. Given the multifaceted support for narrow word order universals, examining dependencies among more than two variables may allow us to identify ‘pivot’ word orders to which others align.

Check out Ringen et al. (2026), and for two examples of where this approach was used to test cultural evolution hypotheses, see Sheehan et al. (2023) (on coevolution of religious and political authority in Austronesian societies), and Hrnčíř et al. (2026) (on Kava consumption and the rise of sociopolitical complexity in Oceania).

Thank you very much, Annemarie and Russell!

References

Bouckaert, Remco & Redding, David & Sheehan, Oliver & Kyritsis, Thanos & Gray, Russell & Jones, Kate E & Atkinson, Quentin. 2022. Global language diversification is linked to socio-ecology and threat status. SocArXiv. (doi:10.31235/osf.io/f8tr6) (https://osf.io/preprints/socarxiv/f8tr6_v1/)

Cinque, Guglielmo. 2013. Cognition, universal grammar, and typological generalizations. Lingua 130. 50–65. (doi:10.1016/j.lingua.2012.10.007)

Dunn, Michael & Greenhill, Simon J. & Levinson, Stephen C. & Gray, Russell D. 2011. Evolved structure of language shows lineage-specific trends in word-order universals. Nature 473(7345). 79–82. (doi:10.1038/nature09923)

Harley, Heidi & Ritter, Elizabeth. 2002. Person and Number in Pronouns: A Feature-Geometric Analysis. Language 78(3). 482–526.

Hrnčíř, Václav & Sheehan, Oliver & Claessens, Scott & Gray, Russell D. 2026. Kava consumption and the rise of sociopolitical complexity in Oceania. Proceedings of the National Academy of Sciences 123(9). e2521658123. (doi:10.1073/pnas.2521658123)

Ringen, Erik J. & Claessens, Scott & Martin, Jordan S. & Jaeggi, Adrian V. 2026. Trait coevolution and causal inference using generalized dynamic phylogenetic models. Methods in Ecology and Evolution 17(6). 1818–1836. (doi:10.1111/2041-210x.70303)

Sheehan, Oliver & Watts, Joseph & Gray, Russell D. & Bulbulia, Joseph & Claessens, Scott & Ringen, Erik J. & Atkinson, Quentin D. 2023. Coevolution of religious and political authority in Austronesian societies. Nature Human Behaviour 7(1). 38–45. (doi:10.1038/s41562-022-01471-y)

Verkerk, Annemarie & Shcherbakova, Olena & Haynie, Hannah J. & Skirgård, Hedvig & Rzymski, Christoph & Atkinson, Quentin D. & Greenhill, Simon J. & Gray, Russell D. 2026. Enduring constraints on grammar revealed by Bayesian spatiophylogenetic analyses. Nature Human Behaviour. Nature Publishing Group 10(1). 126–136. (doi:10.1038/s41562-025-02325-z)

Martin Haspelmath
Martin Haspelmath

The text only may be used under licence Creative Commons Attribution 4.0 International. All other elements (illustrations, imported files) are “All rights reserved”, unless otherwise stated.


OpenEdition suggests that you cite this post as follows:
Martin Haspelmath (July 14, 2026). An interview with Annemarie Verkerk and Russell Gray on “Enduring constraints on grammar” (2026). Diversity Linguistics Comment. Retrieved September 12, 2026 from https://doi.org/10.58079/16kpe


Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.