In the study of American statutory interpretation, Nix v. Hedden (1893) exemplifies the primacy of ordinary meaning. The Court famously held that, in a tariff act, a tomato is a "vegetable," not a "fruit." However, the Court suggested that the statute was addressed to buyers and sellers of produce and that it would eschew "ordinary" meaning for a technical meaning in trade or commerce-if one existed. In this chapter, from a book on corpus linguistics and the interpretation of historical statues, we revisit Nix with corpus linguistic analysis. We compiled a specialized corpus of 19th century trade sources and find that the usage of "tomato" was decidedly mixed: a vegetable, a fruit, and sometimes both simultaneously! This new data complicates the seemingly simple holding in Nix. It also challenges the case's import. Current Supreme Court Justices invoke Nix for the principle that (1) statutes should not be interpreted "literally" (they should be interpreted contextually), by (2) giving terms their ordinary, nontechnical meanings. Our study of Nix underscores the false equivalence between contextual and ordinary interpretation. In some cases, the contextually relevant meaning is a technical one. Our study of 19th century trade meaning supports that Nix may even be one of those cases.Download the article from SSRN at the link.
Showing posts with label Corpus Linguistics. Show all posts
Showing posts with label Corpus Linguistics. Show all posts
October 19, 2025
Gales, Solan, and Tobia on Nix v. Hedden: Corpus Linguistics and the Interpretation of Statutes Over Time
Tammy Gales, Hofstra University, Lawrence M. Solan, Brooklyn Law School, and Kevin Tobia, Georgetown University Law Center, have published Nix v. Hedden: Corpus Linguistics and the Interpretation of Statutes Over Time (draft) Here is the abstract.
February 8, 2024
Phillips on A Corpus Linguistic Analysis of "Possessions" in American English, 1760-1776 @BYU
James Cleith Phillips, Brigham Young University, is publishing A Corpus Linguistic Analysis of "Possessions" in American English, 1760-1776 in the Chapman Law Review. Here is the abstract.
The U.S. Constitution’s Fourth Amendment protects against unreasonable searches and seizures of persons, houses, papers, and effects. Yet state constitutions often use different language, thus providing a different scope of protection. Specifically, starting with Pennsylvania in 1776, sixteen states have constitutional provisions that include possessions as protected from unreasonable searches and seizures. And currently there is litigation in various state courts, including the Pennsylvania State Court, over the meaning of this constitutional protection. Possessions potentially implies more than houses, papers, or effects—arguably covering anything one possesses, including private land, which would significantly expand the coverage of such constitutional protection. But traditional tools of constitutional interpretation, such as dictionaries or etymology, often fall short in uncovering the original public meaning of constitutional text. Hence, increasingly courts (including the U.S. Supreme Court) have looked to corpus linguistics to better answer the linguistic questions that judges face in interpreting the words of the law. Understandably, judges use economic tools to tackle economic questions and historical tools to answer historical questions. Should they not use linguistic tools for linguistic questions? “[W]ords are . . . the material of which laws are made. Everything depends on our understanding of them.” We can and should use the right tools for seeking this understanding. This article will proceed in four parts. Part I introduces the question at issue in the context of the first state constitution to include the term: the Pennsylvania Constitution. It does so, at least in part, because other state constitutions arguably copied the Pennsylvania Constitution, and thus the meaning of the that constitution likely sheds light on the state constitutions that followed it. Part II highlights shortcomings of the traditional tools usually employed in constitutional interpretation. Part III explains how the tools of corpus linguistics can address these shortcomings. And Part IV presents a corpus linguistic analysis of the term possessions. This approach, more rigorous than that usually undertaken, provides data on the linguistic question that undergirds the legal issue—which reading of these state constitutions is more probable than the other. After all, a “problem in [legal interpretation] can seriously bother courts only when there is a contest between probabilities of meaning.” Corpus linguistics can help with that contest. And this article finds that founding-era Americans sometimes used the word possessions to include land one owned, and sometimes not. In the context of the lemma land, a majority of the time the word possessions appeared to include land as property. More significantly, when looking more broadly at any instance of the term possessions, whether or not the lemma land was used nearby, early Americans used the term to include land approximately 86% of the time. This is evidence, then, that the Pennsylvania Constitution, and likely other state constitutions, were originally understood to protect against unreasonable searches of one’s land—thus providing broader protection than the U.S. Constitution’s Fourth Amendment.Download the article from SSRN at the link.
January 13, 2024
Buffington on Being vs. Because: New Observations on the Syntax & Semantics of the US Constitution's Second Amendment @AlbanyLaw
Joe Buffington, Albany Law School, has published Being vs. Because: New Observations on the Syntax & Semantics of the US Constitution’s Second Amendment. Here is the abstract.
The Second Amendment of the US Constitution is ambiguous due to its subordination of one clause to another without the use of an overt subordinating conjunction. Many scholars have argued that the subordination is more or less similar, if not identical, to what is seen in because-clauses. One such scholar, Karen Sullivan, has recently used corpus linguistics to conclude that the likeliest interpretation of the Second Amendment’s subordination when the Amendment was written was one of “external causation,” where the militia clause is understood as the real-world reason why the right-to-bear-arms clause is true. This essay responds to Sullivan’s significant work by presenting three synchronic differences between being-clauses and because-clauses that suggest that external causation may not be an optimal interpretation of the Amendment’s structure, after all. An alternative analysis, where the missing conjunction is modeled as a covert proform, is proposed, and consequences of the analysis are considered – in particular, I present a novel argument that the US Supreme Court’s controversial decision in District of Columbia v. Heller was, in essence, correct.Download the essay from SSRN at the link.
September 23, 2023
Baumgartner on The Meaning of "Reasonable": Evidence From a Corpus-Linguistic Study @UZH_ch @kneer @kevin_tobia @CambridgeUP
Lucien Baumgartner and Markus Kneer, both of the University of Zurich, Institute of Philosophy, are publishing The Meaning of ‘Reasonable’: Evidence From a Corpus-Linguistic Study in The Cambridge Handbook of Experimental Jurisprudence (Kevin P. Tobia, ed., Cambridge University Press, Forthcoming). Here is the abstract.
The reasonable person standard is key to both Criminal Law and Torts. What does and does not count as reasonable behavior and decision-making is frequently deter- mined by lay jurors. Hence, laypeople’s understanding of the term must be considered, especially whether they use it predominately in an evaluative fashion. In this corpus study based on supervised machine learning models, we investigate whether laypeople use the expression ‘reasonable’ mainly as a descriptive, an evaluative, or merely a value-associated term. We find that ‘reasonable’ is predicted to be an evaluative term in the majority of cases. This supports prescriptive accounts, and challenges descriptive and hybrid accounts of the term—at least given the way we operationalize the latter. Interestingly, other expressions often used interchangeably in jury instructions (e.g. ‘careful,’ ‘ordinary,’ ‘prudent,’ etc), however, are predicted to be descriptive. This indicates a discrepancy between the intended use of the term ‘reasonable’ and the understanding lay jurors might bring into the court room.Download the essay from SSRN at the link.
July 21, 2022
Choi on Computational Corpus Linguistics
Jonathan H. Choi, University of Minnesota Law School, has published Computational Corpus Linguistics. Here is the abstract.
Scholars and judges increasingly interpret legal text by studying word use in real-world documents, a method known as “corpus linguistics.” But the traditional approach to corpus linguistics encounters several problems. It focuses on word frequencies at the expense of subtler linguistic cues and presents no clear dividing line between correct and incorrect textual meanings. It also requires a variety of subjective and opaque judgment calls, allowing motivated interpreters to cherry-pick the method that supports their favored meanings. This Article proposes a new, computational approach to corpus linguistics. It uses machine learning and natural language processing to algorithmically evaluate word meaning. By measuring the semantic similarity between words, we can answer questions of legal interpretation—for example, by testing whether “judge” is similar to “representative,” and therefore whether judicial elections are governed by the Voting Rights Act. Computational approaches produce quantitative estimates of similarity that reflect the intuitive semantic relationships between words. This Article extracts qualitative implications from these quantitative estimates by benchmarking against a known scale of word similarity, based on H.L.A. Hart’s famous “vehicles in the park” hypothetical. Applying computational corpus linguistics, this Article finds that semantic questions in real-world legal cases rarely give clear answers. Borrowing Hart’s analogy, most cases are closer to asking whether a bicycle is a vehicle than whether a car is a vehicle. Moreover, estimates of similarity vary substantially between corpora, even large and reputable ones. This suggests that the choice of corpus matters more than previously recognized and that traditional corpus linguists must consult multiple corpora to decrease the risk of cherry-picking. These empirical findings have important implications for ongoing doctrinal debates outside of corpus linguistics, suggesting that text is less clear and objective than many textualists believe. The Article develops these implications with discussion on the nature of linguistic meaning in legal interpretation. Ultimately, the Article offers new insights both to theorists considering the role of legal text and to empiricists seeking to understand how text is used in the real world.Download the article from SSRN at the link.
Subscribe to:
Posts (Atom)