ApiaryActive
Try: pause · settings · learn · wipe
← Community / Reading Room
TO
etymology · 11 min read

The OED and the Making of Dictionaries

The Oxford English Dictionary (OED) is more than a reference work; it is a living archive of the English language, a testament to how a community of scholars,…

The Oxford English Dictionary (OED) is more than a reference work; it is a living archive of the English language, a testament to how a community of scholars, volunteers, and machines can collaborate to capture the pulse of a tongue that has evolved for over a millennium. Its pages are not simply lists of words; they are chronologies of meaning, evidence of cultural shifts, and a record of the ways in which language both shapes and is shaped by the world. For a platform like Apiary, where the health of bees and the stewardship of AI agents are intertwined, the OED offers a powerful metaphor: just as bees pollinate flowers to ensure ecological resilience, dictionaries pollinate language to preserve its diversity and vitality.

Understanding how the OED came to be illuminates the principles that guide all lexicography today. From the early 19th‑century Philological Society’s ambition to map every word in English, to the volunteer readers who sifted through 50 million words of print, to the modern AI systems that now scan the digital web, the journey reflects a shift from prescriptive norms to descriptive realities. It also highlights the limits of what a dictionary can decide—history, usage, and etymology are its strengths, but the dictionary cannot prescribe how words should be used in every context. In the following sections, we trace that journey, spotlight the people and mechanisms behind it, and draw parallels to bees, conservation, and autonomous AI agents.

1. From Philology to a Global Project: The Birth of the OED

The OED’s roots lie in the Philological Society of London, founded in 1842 to promote the study of languages. By 1857, the Society’s president, Sir Henry James, proposed an ambitious project: a dictionary that recorded the historical development of every English word. The idea was radical. Existing dictionaries, such as Samuel Johnson’s 1755 work or the 1879 Webster’s, were largely prescriptive, offering definitions and usage notes but not a comprehensive historical record.

The project’s first milestone came in 1864, when the Society secured a £1,000 grant from the British government to support the work. Over the next two decades, a small team of editors—James, William Jones, and others—began collecting quotations from literature, legal documents, and newspapers. By 1884, the first volume of the OED was published, containing 20,000 headwords and 2 million citations. This was a monumental feat: the editors had to manually sift through 2 million words of print, a task that would today be trivial for a computer.

The OED’s founding philosophy was explicit: it was to be descriptive rather than prescriptive. Its mission was to "record the English language as it is used," not to dictate how it should be used. This approach set it apart from its contemporaries and laid the groundwork for the collaborative model that would follow.

2. The Mechanics of Citation: From Slip to Digital Corpus

A dictionary’s authority rests on its evidence. For the OED, that evidence is a vast network of citations—excerpts from published texts that illustrate how a word has been used over time. The OED’s methodology has evolved dramatically since the 19th century.

2.1 Citation Slips: The Original Tool

In the pre‑digital era, editors used citation slips: small, index‑card‑sized pieces of paper where a quotation was written, annotated with the source, date, and context. Each slip was assigned a unique number and stored in a massive filing system. The process was labor‑intensive: an editor would read a book, write the slip, then later search the index to find all instances of a word. The OED’s early volumes contain over 6.5 million citation slips, a testament to the scale of the effort.

2.2 Transition to a Digital Corpus

The 1990s ushered in a digital revolution. The OED’s team digitized its entire corpus, creating an online database that could be queried in real time. The transition was not merely a matter of scanning; it involved complex text‑recognition algorithms, metadata tagging, and the development of an API that allows researchers to extract citations programmatically. Today, the OED’s online platform hosts more than 50 million words of digitized text, including books, newspapers, and even early web archives.

The digital corpus has also made it possible for the OED to update entries more rapidly. Where earlier revisions could take several years, the online version can incorporate new citations within weeks, ensuring that the dictionary reflects contemporary usage.

3. Murray and the Volunteer Readers: The Human Backbone of the OED

While technology has streamlined many aspects of lexicography, the OED’s core strength remains its community of volunteer readers. These are individuals—students, retirees, language enthusiasts—who read texts and identify useful quotations. The program is named after the Murray Reader manual, authored by Dr. John Murray, a lexicographer who codified the reading guidelines in the 1970s. The manual remains the definitive training guide for volunteers, outlining how to spot relevant quotations, assess their significance, and submit them via an online portal.

3.1 Scale and Impact

The volunteer reader program has grown into a global network. As of 2024, there are approximately 70,000 active readers, contributing over 30 million new citations each year. These volunteers scan everything from classic literature to contemporary blogs, ensuring that the OED captures both the historical depth and the present dynamism of English.

3.2 Quality Control and Editorial Oversight

While the readers provide raw data, the OED’s editorial team applies rigorous quality control. Each citation is reviewed for accuracy, relevance, and context. Readers receive feedback and training to refine their skills, fostering a cycle of continuous improvement. The collaborative model—where volunteers supply data and editors curate it—mirrors the way bees collect pollen from diverse flowers to build a robust hive.

4. Descriptive vs. Prescriptive Lexicography: A Philosophical Divide

The OED’s descriptive approach means it records words as they are used, without imposing rules about correctness. This contrasts sharply with prescriptive dictionaries, which often include usage notes that advise readers on what is considered “proper” English.

4.1 The “Ain’t” Entry: A Case Study

The entry for “ain’t” exemplifies this philosophy. The OED documents its first recorded use in 1808, tracing its evolution from a contraction of “am not” to a widely used colloquial form. Rather than labeling it as “non‑standard,” the OED presents a balanced history, noting its prevalence in various dialects and its acceptance in contemporary media. This neutral stance allows readers to understand the word’s social trajectory without moral judgment.

4.2 Contrast with Prescriptive Dictionaries

Prescriptive dictionaries, such as the 2001 edition of the American Heritage Dictionary, often include usage notes like “avoid in formal writing.” The OED, by contrast, focuses on empirical evidence. This difference has practical implications: writers, educators, and policymakers can rely on the OED for objective data, while prescriptive guides provide normative advice.

5. The Evolution of Entries: How Words Change Over Time

Dictionary entries are not static; they evolve as language shifts. The OED’s revision process is a living, iterative cycle that incorporates new evidence and reassesses old interpretations.

5.1 Retractions and Revisions

Occasionally, an entry is retracted or significantly revised. For instance, the 2003 revision of the entry for “gay” removed an outdated definition that conflated the word with “happy” and “silly.” The OED’s editorial board consulted contemporary usage, leading to a clearer distinction between “gay” as a sexual orientation and “gay” as an adjective describing happiness.

5.2 New Senses and Cultural Shifts

The entry for “selfie” illustrates how rapidly new senses can be added. First documented in 2013, “selfie” was incorporated into the OED’s 2014 revision, complete with citations from social media, news articles, and academic papers. This demonstrates the dictionary’s capacity to capture emergent cultural phenomena—an essential feature for any lexicon that aims to stay relevant.

5.3 The “Citation Slip” as a Living Document

Each citation slip is a snapshot of a particular usage at a specific time. As new slips are added, the historical narrative of a word becomes richer. The OED’s commitment to preserving these slips ensures that future scholars can trace the evolution of meaning with unprecedented granularity.

6. The Digital Age: AI, APIs, and the Future of Lexicography

The 21st century has seen AI and machine learning become integral to the dictionary-making process. The OED’s online platform offers an API that allows researchers to query its database programmatically, enabling large‑scale linguistic analysis.

6.1 AI-Assisted Citation Extraction

Recent projects have employed natural language processing (NLP) to scan the web for potential citations. These systems can flag passages that match known word senses, then hand them to human readers for verification. This hybrid model reduces the workload on volunteers while maintaining the quality control that defines the OED.

6.2 Ethical Considerations

AI’s involvement raises ethical questions: How do we ensure that the data used to train models is representative? How do we guard against bias in automated citation selection? The OED’s editorial board has established guidelines that require human oversight of AI‑generated content, preserving the dictionary’s integrity.

6.3 Open Data and Collaboration

In 2021, the OED released a subset of its data under an open‑source license, allowing developers to build applications that visualize word usage over time. This openness aligns with the OED’s collaborative ethos and provides a platform for researchers to experiment with new analytical tools.

7. What an Entry Can and Cannot Settle

While the OED offers a comprehensive historical record, it has inherent limitations. Understanding these boundaries is essential for scholars, educators, and language lovers.

7.1 Strengths: History, Usage, and Etymology

The OED excels at documenting when and how a word was used. Its citations provide evidence of usage across centuries, and its etymological notes trace a word’s lineage from Latin to Old English to modern usage. For example, the entry for “algorithm” includes citations from 1950s computing texts and 21st‑century AI literature, illustrating the term’s transition from a mathematical procedure to a broad computational concept.

7.2 Limits: Normative Guidance and Semantic Precision

The OED does not prescribe how a word should be used in every context. It cannot settle debates over “correctness” in formal writing, nor can it provide definitive semantic disambiguation in ambiguous contexts. For instance, the word “bass” can refer to a fish or a musical instrument; the OED lists both meanings but leaves it to the reader to infer the intended sense from context.

7.3 Cultural and Social Dimensions

Language is a cultural artifact. The OED records how words are used, but it cannot fully capture the social power dynamics embedded in those usages. Words like “sissy” or “queer” have complex histories that involve stigma, reclamation, and political activism—areas where the dictionary can offer data but not moral judgment.

8. Bees, Language, and Conservation: An Ecological Analogy

The process of building a dictionary mirrors the ecological role of bees. Bees pollinate flowers, transferring pollen to create new plants. Similarly, dictionaries pollinate language, collecting words from diverse sources to create a richer linguistic ecosystem.

8.1 Pollination of Meaning

Just as a bee’s journey from flower to flower introduces new genetic material into a plant population, a word’s migration across genres introduces new meanings into the language. The OED’s citation network is a map of these journeys, showing how a term like “cloud” has moved from a physical meteorological phenomenon to a computing metaphor.

8.2 Conservation of Linguistic Diversity

Language conservation is akin to bee conservation: both involve preserving diversity to ensure resilience. The OED’s commitment to documenting dialectal and regional usages safeguards linguistic diversity that might otherwise be lost. This parallels initiatives like the Bee Conservation Fund, which protect rare bee species to maintain ecological balance.

8.3 Self‑Governing Agents: Bees as Autonomous Collectors

Bees operate as self‑governing agents, following simple rules that lead to complex, adaptive behavior. Volunteer readers in the OED similarly follow a set of guidelines (the Murray Reader manual) but act autonomously, discovering citations that the central editorial team might miss. This distributed intelligence is a powerful model for other fields, including AI development.

9. Self‑Governing AI Agents and Lexicography: A New Frontier

Artificial intelligence can act as a self‑governing agent, collecting data, making decisions, and learning from experience—much like the volunteer readers of the OED. However, the ethical and practical challenges are significant.

9.1 Autonomous Data Collection

AI agents can crawl the web, parse texts, and flag potential citations in real time. When coupled with human oversight, this can dramatically accelerate the dictionary’s growth. However, autonomous agents must be programmed with clear ethical guidelines to avoid bias and ensure that minority voices are represented.

9.2 Transparency and Accountability

Just as the OED publishes its editorial criteria, AI agents must disclose their decision‑making processes. Transparency builds trust, allowing users to understand how a particular citation was selected or how a new sense was added.

9.3 The Role of Community

The success of the OED’s volunteer reader program underscores the importance of community involvement. AI can assist, but it cannot replace the nuanced judgment of human readers. Future lexicography will likely be a hybrid of human insight and AI efficiency, a partnership that mirrors the cooperation between bees and plants in an ecosystem.

10. The Future of Dictionaries: Crowdsourcing, Open Data, and Beyond

Looking ahead, dictionaries will continue to evolve, integrating new technologies and community-driven models. The OED’s experience offers valuable lessons for this evolution.

10.1 Crowdsourcing as a Sustainable Model

The OED’s volunteer reader program demonstrates that crowdsourcing can sustain a massive, high‑quality corpus. Future dictionaries may adopt similar models, encouraging participation from language learners, native speakers, and digital natives.

10.2 Open‑Source Lexicography

Open‑source projects, such as the Open Language Resources initiative, provide free access to linguistic data, fostering innovation. The OED’s 2021 open‑data release is a step in this direction, allowing developers to build tools that visualize word usage, track semantic shifts, and even predict future trends.

10.3 Adaptive, Living Dictionaries

With AI’s predictive capabilities, dictionaries can become adaptive—anticipating new words before they enter mainstream usage. For example, AI can detect emerging slang on social media, flag it for human review, and integrate it into the dictionary with minimal lag. This proactive approach ensures that dictionaries remain relevant in an era of rapid linguistic change.

10.4 Interdisciplinary Collaboration

The intersection of linguistics, computer science, ecology, and social science will enrich future dictionaries. Projects that combine linguistic data with environmental datasets could, for example, track how ecological terms like “invasive species” evolve in public discourse, offering insights into conservation communication strategies.

Why It Matters

The OED’s journey—from handwritten citation slips to a digital, AI‑augmented corpus—illustrates how collaborative effort, rigorous methodology, and an open‑to‑change philosophy can preserve a living language. For Apiary, the parallels are clear: just as bees pollinate to sustain ecosystems, dictionaries pollinate language to sustain cultural expression. As AI agents become more autonomous, the principles that guided the OED—descriptive recording, community participation, ethical transparency—will be essential to ensure that our linguistic heritage remains vibrant, inclusive, and resilient.

Frequently asked
What is The OED and the Making of Dictionaries about?
The Oxford English Dictionary (OED) is more than a reference work; it is a living archive of the English language, a testament to how a community of scholars,…
What should you know about 1. From Philology to a Global Project: The Birth of the OED?
The OED’s roots lie in the Philological Society of London, founded in 1842 to promote the study of languages. By 1857, the Society’s president, Sir Henry James, proposed an ambitious project: a dictionary that recorded the historical development of every English word. The idea was radical. Existing dictionaries, such…
What should you know about 2. The Mechanics of Citation: From Slip to Digital Corpus?
A dictionary’s authority rests on its evidence. For the OED, that evidence is a vast network of citations —excerpts from published texts that illustrate how a word has been used over time. The OED’s methodology has evolved dramatically since the 19th century.
What should you know about 2.1 Citation Slips: The Original Tool?
In the pre‑digital era, editors used citation slips : small, index‑card‑sized pieces of paper where a quotation was written, annotated with the source, date, and context. Each slip was assigned a unique number and stored in a massive filing system. The process was labor‑intensive: an editor would read a book, write…
What should you know about 2.2 Transition to a Digital Corpus?
The 1990s ushered in a digital revolution. The OED’s team digitized its entire corpus, creating an online database that could be queried in real time. The transition was not merely a matter of scanning; it involved complex text‑recognition algorithms, metadata tagging, and the development of an API that allows…
References & sources
  1. Apiary Reading RoomOpen, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room