Below, we explain how Lexogoth was developed, the linguistic principles that guide it, and how we compiled a rich corpus of authentic French language. This text offers language teachers, foreign-language education specialists, and curious users a deeper insight into the vision and approach behind Lexogoth.
Table of Contents
1. The Origin of Lexogoth and Its Primary Objective
2. Corpus Compilation: Which Criteria Were Used to Select Words and Structures?
2.1 The Theoretical Foundation of Lexogoth
2.2 Corpus Selection and Composition
2.2.1 An Empirically Based Approach
2.2.2 Structure of the Corpus
2.2.3 Structure of the Entries in Lexogoth
2.3 Language Registers, Geographical Variants, and Varieties
3. How Lexogoth Distinguishes Itself from Other Language Apps and Websites
3.1 Authentic Context vs. Unnatural Sentences
3.2 Relevant Words and Structures vs. Arbitrary Thematic Lists
3.3 Deep Anchoring of Lexical Knowledge vs. Gamified Distractions
3.4 Attention to Diversity in Language: Registers, Variants, and Sociocultural Context
3.5 Empirically Validated Data vs. Generative AI (LLMs)
4. How to Best Study Vocabulary, and How Lexogoth Can Help
4.1 The Lexical Toolbox in Practice
4.2 The Role of the Educator
5. An Ongoing Project
6. Bibliography
1. The Origin and Primary Goal of Lexogoth
Lexogoth arose out of a frustration familiar to every language learner who has ever had to study vocabulary: the struggle to find the right word combinations and the realization that memorizing word lists alone is not enough to participate fluently in communicative tasks.
We believe this difficult transition is largely due to poor study methods and insufficient practice with vocabulary. Most foreign language students begin their learning journey by memorizing vocabulary lists, which allows them to quickly and efficiently acquire a small base lexicon. While this approach is useful in the early stages, many textbooks continue to rely on word lists year after year. This persistent focus eventually handicaps students: when they are required to perform meaningful communicative tasks later on, they soon discover that a mix of grammar rules and isolated words is simply not enough to use the language effectively.
Learning through lists prevents students from forming connections between the words and expressions they encounter. In other words, learners end up with a collection of disconnected lexical items, without the semantic network needed to tie them together. This lack of internal structure leads to shallow memorization, which in turn makes students forget new vocabulary much faster than if it had been embedded in a network of associations. Word lists also tend to ignore the inherent complexity of words. A word is more than just a definition or a translation: to use it properly, one needs to know which complements it takes, which meanings are frequent and which are not, what register it belongs to, in which contexts it can be used, as well as its synonyms, antonyms, and derivatives.
To make matters worse, vocabulary practice in many French textbooks is still limited to crosswords, fill-in-the-blanks, or short completion exercises, far from enough to build solid lexical competence. Because of this rigid reliance on word lists, combined with a lack of rich, meaningful input that could illustrate both the meaning and use of new words, students progress in the foreign language much more slowly than they should.
In our view, vocabulary should always be taught in context, with students exposed to numerous examples that clearly demonstrate how words are used and that help them build a robust semantic network more quickly. It is from this perspective that we developed Lexogoth: our goal is to provide students of French with a representative corpus of high-frequency words, expressions, and collocations in varied contexts. This enables learners to remember and practice these lexical items more effectively and more efficiently. Below, we explain the methodology we used and how the corpus was developed.
For further reading on vocabulary acquisition, see: Nation (2001), Thornbury (2002), Laufer (2022), Bentolila (2010).
2. Composition of the corpus: based on which criteria were words and structures selected?
2.1 Theoretical Basis for Lexogoth
The development of Lexogoth is closely linked to two major trends in foreign language education that gained serious momentum in the 1990s:
- The Lexical Approach by Michael Lewis (1993)
- The use of corpora with authentic material to support the learning process (Corpus Linguistics / Data-Driven Learning)
To bridge the gap between abstract linguistic theory and classroom reality, Lexogoth’s architecture integrates these trends through three key theoretical principles: The Lexical Approach, Schmidt’s Noticing Hypothesis, and Vygotsky’s principle of Pedagogical Mediation.
Principle 1: The Lexical Approach (Language as ‘Chunks’)
One of the key principles of Michael Lewis’s (1993) approach is that, in the context of foreign language learning, language should no longer be seen as a collection of sentences built from grammatical structures into which lexical items should be inserted. Instead, language should be understood as a collection of ready-made expressions, compound words, and regular word combinations that are stored in the learner’s memory.
“Language consists of lexicalized grammar, not grammaticalized lexis.” (Lewis, 1993)
For advanced language learners (CEFR levels B and C), this is crucial. A student might know the literal meaning of confiance (trust), but idiomatic production requires the quick and effortless retrieval of pre-fabricated, ready-made formulaic sequences (or chunks) such as faire confiance à, avoir confiance en, or un manque de confiance.
There is substantial empirical evidence supporting the usefulness of learning via chunks:
- Enhanced Writing Skills: Using chunks significantly improves foreign language learners’ writing competence (Albaqami, 2022).
- Nativelike Fluency: Native speakers rely heavily on pre-formed chunks for efficient, real-time communication (Pawley & Syder, 1983).
- Oral Proficiency: Students who learn chunks demonstrate greater fluency and better speaking skills, whereas those who focus on individual, isolated words struggle to produce sentences spontaneously in communicative situations (Boers et al., 2006).
Of course, Lewis’s proposal has also drawn valid criticism:
- Curriculum Design: On what basis should word combinations be selected, and how can an entire curriculum be organized around them? (Wray, 2002).
- Cognitive Overload: How can students practically study and memorize dozens of common collocations for a single word without becoming overwhelmed? (Wray, 2002).
- The Role of Grammar: Why should syntax be sidelined when it provides the vital generative mechanism for creating new sentences? (Ellis, 2003).
Principle 2: The Noticing Hypothesis (From Input to Intake)
Simply exposing learners to authentic texts or large corpora is not enough to guarantee language acquisition. According to Richard Schmidt’s (1990) Noticing Hypothesis, linguistic input only becomes intake (processed and stored language) when the learner consciously notices specific features and patterns in the input.
During traditional reading, students often gloss over collocations and syntactical boundaries because they focus entirely on decoding the global meaning of a text. To facilitate cognitive retention, a digital learning environment must provide input enhancement (Sharwood Smith, 1993), that is visual cues and typographical formatting that actively draw the learner’s attention to the formal characteristics, registers, and collocations of a word.
Lexogoth operationalizes this by explicitly highlighting chunks and structures within its interfaces (e.g., in the Workbench), ensuring that key patterns practically “light up” for the learner.
Principle 3: Pedagogical Mediation (Scaffolding within the ZPD)
This is where Lev Vygotsky’s (1978) concept of Pedagogical Mediation and the Zone of Proximal Development (ZPD) becomes vital. Effective learning takes place in the sweet spot between what a learner can do independently and what they can achieve with guidance. This guidance is known as scaffolding: temporary, structured support that is gradually dismantled as the learner gains autonomy (Wood, Bruner & Ross, 1976).
Traditional Data-Driven Learning (DDL) often fails because throwing raw, unedited corpus data at students places a massive cognitive burden on them, pushing them far outside their ZPD. For the inductive “learner-as-detective” model to succeed, the raw data must be filtered, structured, and linguistically annotated by an expert first.
Lexogoth acts as this mediator. By providing a pre-curated, highly geannoteerde database with clear pedagogical scaffolding, it reduces cognitive friction and enables students to safely and effectively discover linguistic rules.
Corpus Linguistics & The Lexogoth Solution
Corpus linguistics, emerging in the 1960s, revolutionized our understanding of how language is actually used by systematically analyzing large-scale digital databases (like COCA or the BNC). However, when applied directly to the classroom, traditional corpustools (concordancers) present severe hurdles for learners:
- The KWIC-nightmare: Traditional tools present data in fragmented, decontextualized Key Word in Context (KWIC) lists. These one-line snippets destroy text cohesion and pragmatics, leading to cognitive fatigue. Lexogoth’s Solution: We replace KWIC lists with rich, meaningful context blocks.
- Data Deluge: Raw corpora return thousands of random sentences, mixing obscure, highly specialized, or outdated meanings without context. Lexogoth’s Solution: A hand-curated corpus with a remarkable pedagogical density of 0.76 (370,000 annotated tokens vs 486,000 French tokens), focusing on actual, contemporary usage.
- Lack of Guidance: Standard corpora do not tell you if a word is slang, formal, or a geographic variant (e.g., GSM vs. portable vs cellulaire). Lexogoth’s Solution: Systematically integrated, jargon-free annotations covering syntax (verb choices, prepositions), semantics (subtle nuances), register (formal, informal, verlan), and cultural realia.
Ultimately, Lexogoth is not a rigid language course as Lewis originally envisioned, but an interactive, mediated CALL-ecosystem. It empowers students to independently explore, notice, and internalize authentic lexical networks, turning raw data into active, productively usable language.s in context more quickly, providing authentic material with which they can systematically and independently practice and reinforce their lexical knowledge.
2.2 Selection and Construction of the Corpus
2.2.1 An Empirically Based Approach
As we mentioned above, explanatory (or translating) dictionaries are not ideal tools for foreign language learners because, in our view, too many essential meaning components of lemmas—necessary for proper usage—are not made explicit enough. For example, this concerns the selection of complements by a verb or the actual usage in everyday language, rather than the “virtual” usage found in dictionaries or textbooks. For instance, Robert (le Robert en ligne) gives the following definitions for the verbs frôler and friser:
- [frôler] 1. Toucher légèrement en glissant, en passant. 2. Passer très près de, en touchant presque. ➙ raser. La voiture a frôlé le trottoir. au figuré: Frôler le ridicule.
- [friser] … 2. Passer au ras de, effleurer. (voir: frôler, raser) 3. Approcher de très près. Elle frise la soixantaine. Cela frise le ridicule.
Based on the definitions above, someone wanting to use frôler and friser might conclude that the two words can largely be considered synonyms. However, after checking the corpora we use, we found that friser and frôler can indeed be considered synonyms in a figurative sense (with abstract complements), but not in their literal meaning (i.e., followed by concrete complements), as we could hardly find usage examples where friser is used in that sense. Although the definitions are largely identical, it would be incorrect to assume their actual usage is the same. For example, you will rarely, if ever, come across sentences like un motard alcoolisé qui frisait dangereusement les passants dans le centre-ville de….
To prevent language learners from encountering such problems and to help them learn word meanings more smoothly, we have therefore chosen to always verify the words, structures, and their different meanings included in our corpus based on frequency. We did this using two tools:
- Consultation of the French Web 2023 corpus from SketchEngine (covering all language registers)
- Consultation of the RTBF corpus (French-speaking Belgian television and news channel; standard language), and the Le Monde corpus (French newspaper; standard and polished language)
Afterwards, we consulted the following reference works to verify whether the lemma or word combination was included and to map out the main meanings of the lemma:
– Le Robert (dictionnaire en ligne)
– Larousse (dictionnaire en ligne)
– Dictionnaire de l’Académie française (en ligne)
After this step, we searched for useful examples illustrating the meanings currently used in French (regardless of register). We aimed to gather as broad a range of relevant examples as possible and consulted the following resources:
- The French Web 2023 corpus from SketchEngine
- The website www.languefrançaise.net, specifically the general dictionary and Bob, l’autre trésor de la langue for colloquial or informal language
- French-speaking Belgian media: RTBF (public broadcaster), Belgian newspapers: Le Soir, La Libre Belgique, La Dernière Heure; French newspapers: Le Monde, Charlie Hebdo, Le Figaro, Le Parisien, Libération; Swiss newspapers: 24 heures, Le Matin, La Tribune de Genève; French-language media from Madagascar, Canada, Congo (DRC), Haiti, Morocco, and Algeria
- Official government websites from Belgium and France, and to a lesser extent Canada and Switzerland; websites of companies from the Francophone world; websites of international organizations such as the European Union, United Nations, NATO, IMF, etc.
- Numerous internet forums (especially for familiar and spoken language); some of the main sources include JeuxVideo.com (mainly youth language), Forum Santé – Doctissimo, French-speaking Reddit channels, AlloCiné (film reviews), SensCritique (film reviews). This also includes numerous blogs ranging from hair styling and technology to international politics (covering all language registers)
- French subtitles of recent films and TV series; we ensured that subtitles were created by native French speakers rather than AI-generated to guarantee authenticity
- Websites offering language advice: Banque de dépannage linguistique (Canada), the Dire et ne pas dire section of the Académie française, notes from Larousse (online dictionary), Projet Voltaire, the WordReference.com forum, and the contrastive Dutch-French grammar resource Grammar Warrior
Next, we selected examples that appeared frequently in various sources and included them partially (as chunks) or fully (as sentences) in Lexogoth. We made sure the examples clearly illustrated the existing (abstract) meanings and, in the case of verbs, that a clear selection of the expected complements was used. To illustrate this, let’s return to the example of frôler:
... sombrer dans la paranoïa communiste / Cela expliquera peut-être mieux sa méfiance qui frôle la paranoïa. / frôler le ridicule; l'illégalité, le burn-out, l'overdose / frôler la mort, la faillite / une balle lui a frôlé le coeur / Une grosse pierre, lancée par un inconnu, lui a frôlé le visage et a percuté la voiture derrière lui. / Le taux de participation aux élections frôle les 60 pour cent. / La police a interpellé un motard alcoolisé qui frôlait dangereusement les passants dans le centre-ville de Lille. …
Because not everything can be conveyed through examples alone, we regularly add extensive notes concerning register, compatibility with specific complements, synonyms, etc., to better guide learners toward correct usage. For “frôler” and “friser,” for instance, we added the following information:
1. In the figurative sense of "narrowly escaping" or "bordering on," you can also use "friser" as a synonym: "friser la mort, la faillite, le burn-out, les 50 pour cent,…" 2. In the literal meaning (grazing past something, lightly touching), where the object of "frôler" (as in "la balle lui a frôlé le nez") is something concrete, the equivalence between "frôler" and "friser" seems to no longer hold, even though Larousse and the Académie still consider both words to be synonyms in this sense. In this meaning, the verb "friser" is barely used, if at all; therefore, you will not encounter phrases like: "...frisait les passants..." or "...frisé le visage…"
If a meaning did not appear in dictionaries (e.g., recent Anglicisms, vulgar language, neologisms) but had a decent frequency, we still chose to include the word in our corpus. Conversely, if a word or meaning was recognized by one or more dictionaries but appeared infrequently or not at all in our consulted corpora, we chose not to include it.
Sometimes, we also found it necessary to compress some dictionary definitions because the differences between them were too subtle to be clearly distinguished.
2.2.2 Structure of the Corpus
The Lexogoth corpus (July 2025) consists of two parts: first, we included the vocabulary from the main lexical fields for levels A1 and A2 of the Common European Framework of Reference for Languages (CEFR) as short, useful chunks (phrases or word combinations) specifically aimed at functional language skills for beginners. This part accounts for roughly 15 percent of the total corpus. We then expanded this vocabulary by adding the main and most common meanings and uses. The second and larger part of the corpus consists of an in-depth exploration of existing lemmas and their usage in familiar and everyday spoken language, covering roughly the B level of the CEFR. An example for the verb descendre:
(a) monter l'escalier / descendre l'escalier / croiser qn dans l'escalier / descendre, monter l'escalier quatre à quatre (b) L'incendie était terrifiant, nous confie une locataire. C'est la fumée qui m'a réveillée, j'ai réussi à descendre par l'escalier avec mon fils. / L'ascenseur ne fonctionne pas depuis ce matin; j'ai dû descendre par l'escalier. / descendre par l'escalier, par l'ascenseur / descendre de l'escalier, de l'échelle, d'un arbre ... / descendre du métro, du train, de la voiture, de son vélo.... / Où est ton frère? Il est dans sa chambre. Dis-lui de descendre immédiatement pour dire bonjour à mamie / monter dans sa chambre (c) prendre l'escalier pour monter à son appartement au 4e étage / Je suis monté en prenant l'escalier. / Pour monter sur une échelle, il faut monter une marche à la fois.
(REMARKS) 1. "Descendre l'escalier" or simply "descendre" without a complement means "to come down the stairs." This is the most frequent construction. Note that this is a transitive use of "descendre" and requires the auxiliary verb "avoir": "J'ai descendu l'escalier." If "descendre" does not have a direct object (COD), you should use "être": "Je suis descendu. / Elle est descendue par l'escalier. /..." The same rule applies to "monter." 2. "Descendre de..." is primarily used to mean "to come down from (an object) that you are sitting, standing, or climbing on," as in "descendre d'un arbre, d'une échelle," or "to get out of/off a vehicle," as in "descendre de son vélo, du train,..." Here, the auxiliary verb is also "être". 3. "Descendre par l'escalier,...": When specifying how (by which way) someone came down, it is often to contrast it with another method: "on descend par l'escalier ou par l'ascenseur?" 4. "Une mamie" or "mamy" mean "grandmother," while "papy" or "papi" refer to a "grandfather." For the plural, you can add an "s" to the words: "des mamies, mamys, papis, papys." 5. "Quatre à quatre" means "head over heels, very quickly," and is often used with verbs of movement like "grimper, monter, descendre, franchir,…"
For this second part, we also systematically add meanings and uses that we consider to be at an advanced level (CEFR level C). For these entries, we regularly include notes to help learners master the correct usage and study new structures. Below is another example for the verb “descendre”:
...Je risque de me faire descendre en flèche sur ce forum, mais est-ce que je suis le seul à trouver ce film débile ? / Cherche pas, à ce prix-là, c'est du vol et ton appli va se faire descendre en flèche. / La presse argentine n'a pas attendu longtemps pour descendre en flammes la formation dirigée par Jorge Sampaoli.
(REMARKS):... 2. "En flèche" can be used as an adjective phrase ("une montée en flèche") or as an adverbial phrase ("monter en flèche"). 3. "En flèche" is most often used with verbs of "rising" or "increasing," such as "monter, augmenter, grimper," and means "at rocket speed." It is also used quite frequently with verbs in the semantic cluster of "starting," such as "partir, démarrer," etc. Finally, the combination is also used with verbs that express a decrease, such as "baisser, descendre," although to a much lesser extent. See the next note, however. 4. "Descendre en flèche" is sometimes used incorrectly in place of the informal expression "descendre en flammes," which means "to heavily criticize, tear someone down, or rip them apart." For example, the text above would have been better written as: "…de me faire descendre en flammes…" and "…va se faire descendre en flammes."
In the second part, we also systematically include words and structures that frequently appear in the media, internet forums, and government websites. By being exposed to authentic and current material, language learners can understand media and literature on important social issues and also gain the tools to express themselves about these topics.
2.2.3 Structure of the Entries in Lexogoth
The way the entries in Lexogoth were designed lies at the heart of our innovative approach: we combine the functionality of a dictionary, a concordancer, and a textbook. It was a deliberate choice to bring together lexical precision, authentic usage, and a didactic dimension, enabling students of French to acquire frequent lexical combinations (chunks) more efficiently and rapidly, through intelligent repetition or targeted consultation. For this reason, users will encounter different types of entries in the corpus:
(a) Entries with an exhaustive treatment of a single lemma: these entries are closest to the dictionary function and illustrate the main uses and meanings of a word, often organized in a clear and structured way.
(b) Entries organized around a specific theme, such as fraud, climate change, or electric mobility. These entries provide a solid basis for building and expanding a lexical network.
(c) Entries focusing on a language problem at the grammatical, lexical, or pragmatic level, which students of French often encounter when using the language.
(d) Simple entries that list words or focus on a single expression or combination. These entries are primarily designed to support the development of a basic vocabulary.
Thanks to this varied structure of the entries and of the corpus as a whole, users can acquire new vocabulary and frequent lexical patterns in an effective and flexible way.
2.3 Language Registers, Geographic Variants, and Language Varieties
In Lexogoth, we strive to highlight the diversity of the French language as broadly as possible. For instance, we pay a lot of attention to the geographical variants of French, specifically Belgicisms. Belgicisms are French words, structures, etc., that are specific to the French used in Belgium but have little spread elsewhere. We dedicate less attention to other varieties of French, such as those from Canada, Switzerland, Luxembourg, Congo, and Algeria.
Anglicisms —that is, English (especially American) words and structures used in French— also receive extensive coverage in our work, and wherever possible, we provide synonyms from the standard language.
Although the focus in Lexogoth is primarily on the standard language, we devote a lot of space to informal language—the language that French speakers use in their everyday interactions. We always mention in our notes whether an expression belongs to informal language, and here too we try to provide synonyms or alternative formulations from the standard language. We also address vulgar expressions and the uniquely French phenomenon of verlan, always providing standard-language equivalents where possible.
3. How Lexogoth Stands Out from Other Language Apps and Websites
Many language apps and websites focus on the fast and easy acquisition of elementary vocabulary through a classic approach based on isolated words, often accompanied by an image, a photo, or a translation. While this method can be useful for beginners or when learning specialized jargon, these apps often fall short in developing the active, productive vocabulary needed for correct and fluent language use. Lexogoth was designed to bridge that gap, distinguishing itself on the following key points.
3.1 Authentic Context vs. Unnatural Sentences
Many popular language apps (such as Duolingo, Drops, or Babbel) base their vocabulary instruction on repeating isolated words or very simple sentences that often feel highly artificial and sterile (e.g., “I can fly”).
- The Problem: These apps fail to show learners how a word is used in its natural environment—that is, in real written and/or spoken language. They do not teach which complements often appear with a word, what its register is, which socio-cultural elements might limit its use, or which of its meanings are frequent and which are not.
- The Lexogoth Difference: Lexogoth operates on the principle that a word is only truly known when a user understands how it can be used in context. Rather than using fragmented, one-line “Key Word in Context” (KWIC) snippets common in traditional linguistic corpora—which strip away pragmatics and cause cognitive fatigue—Lexogoth presents users with exclusively authentic context blocks. We place a special focus on frequent collocations and idiomatic expressions (chunks), allowing the student to see how words naturally behave.
3.2 Relevant Words and Structures vs. Arbitrary Thematic Lists
Traditional language platforms often rely on generic, thematic lists (such as “living,” “food and drink,” “getting around,” or “daily life”) for students to memorize.
- The Problem: These static lists often contain vocabulary that is either irrelevant to advanced communication or completely detached from actual, contemporary frequency.
- The Lexogoth Difference: Lexogoth consists of a carefully, manually compiled corpus. It contains words and word combinations that are current, highly relevant, and used with a relatively high frequency in the modern Francophonie. To guarantee this, we did not rely on automated web-scraping; instead, we consistently and quantitatively verified every candidate lemma and chunk against grootschalige, recent reference corpora (such as French Web 2023 in SketchEngine, RTBF, and Le Monde).
3.3 Deep Anchoring of Lexical Knowledge vs. Gamified Distractions
To keep users hooked, traditional apps rely heavily on gamification: offering short, fast-paced sessions (such as Drops) and presenting learning as a game with streaks, badges, and achievements, combined with simple fill-in-the-blank or matching exercises.
- The Problem: While highly engaging, these gamified, reproductive exercises do not lead to deep mental representation. They encourage passive consumption rather than active linguistic processing.
- The Lexogoth Difference: Lexogoth contains no exercises in the classic sense. Grounded in Richard Schmidt’s Noticing Hypothesis, the goal is for users to deeply anchor patterns and correct word combinations through repeated and attentive reading of the examples within a structured, visually enhanced interface. This process closely resembles how native speakers naturally acquire and retrieve their language. By deeply internalizing these pre-fabricated chunks, students less often need to rely on the time-consuming and cognitively demanding word-by-word sentence construction. They can react more quickly, correctly, and fluently in real communicative situations.
3.4 Attention to Diversity in Language: Registers, Variants, and Sociocultural Context
Most mainstream language tools treat a foreign language as a homogenous, flat system, paying little to no attention to the sociolinguistic nuances of real-world communication.
Sociocultural Realia: Clarifying culture-specific references (such as historical education certificates like the certificat d’études or pop-culture allusions “Paf le chien!”) to build true cultural literacy.
The Problem: A lack of attention to these nuances leads to inappropriate, unnatural, or awkward language use (e.g., using a formal literary term in a casual conversation, or vice-versa).
The Lexogoth Difference: Lexogoth provides meticulous, systematic attention to this diversity. Our manually curated annotations explicitly guide the user through:
Language Registers: Clearly distinguishing formal, informal, colloquial, and youth language (including verlan).
Geographical Variants: Exploring lexical distribution across the Francophonie (e.g., the regional usage of GSM, portable, smartphone, or cellulaire).
3.5 Empirically Validated Data vs. Generative AI (LLMs)
With the massive rise of Large Language Models (LLMs) like ChatGPT, Gemini, and Copilot, many modern language learning apps and methods have rushed to use generative AI to quickly, cheaply, and endlessly generate learning material (DuoLingo). While this is a smart shortcut for developers, the quality of AI-generated sentences is highly unreliable for learners.
Drawing from the scientific analysis in our paper, we argue that Generative AI cannot replace a curated corpus like Lexogoth. In fact, LLMs fall short on four critical, empirically backed fronts when compared to a pedagogically mediated Data-Driven Learning (DDL) environment:
- 1. Linguistic Authenticity vs. Standardized Output AI-generated language is probabilistic; it predicts the most likely next word based on patterns in its training data. This process often produces overly stylized, uniform, and sanitized sentences that lack natural register shifts, regionalisms, or the genuine, unadulterated “rawness” of real-world language. Lexogoth, by contrast, presents 100% authentic text fragments extracted directly from actual native-speaker usage, exposing students to the true, living diversity of the French language.
- 2. Absolute Reliability vs. Hallucinations LLMs are notorious for “hallucinations”, i.e. they confidently generate factually incorrect or linguistically faulty output. In language learning, this means AI can easily fabricate non-existent idioms, propose incorrect semantic nuances, or use rare syntactic combinations that no native speaker would ever use. Every single entry, collocation, and example in Lexogoth is manually verified and cross-referenced with major linguistic reference corpora to guarantee absolute accuracy.
- 3. Active Cognitive Processing (Noticing) vs. Passive Consumption When a student asks an AI chatbot for an explanation, the AI instantly delivers a flat, pre-digested answer. While convenient, this bypasses the essential cognitive processes of noticing and active hypothesis-building. According to Schmidt’s Noticing Hypothesis, deep retention only occurs when learners actively explore data and discover rules for themselves. Lexogoth acts as a stable, guided playground where students act as “language detectives,” ensuring long-term cognitive anchoring.
- 4. Stability and Reproducibility vs. Volatility Because LLMs are probabilistic, entering the exact same prompt twice will often yield entirely different explanations or examples. For structured classroom instruction and independent study, this volatility is highly problematic. Students and teachers need a stable, consistent, and reproducible navigation system where they can repeatedly observe and test the exact same linguistic patterns, a reliability that only a fixed, curated corpus can provide.
Our Stance on AI in Lexogoth’s Workflow
To guarantee these standards, none of the French example sentences in Lexogoth are AI-generated. They are carefully and manually selected from relevant corpora of recent native speaker usage.
However, we do believe in the responsible and efficient use of technology. We utilized advanced AI tools (such as DeepL, ChatGPT, and Gemini) to accelerate the initial translation process of our French entries into Dutch / English. Crucially, every single AI-generated translation was manually checked, refined, and updated by a human expert to ensure a high level of accuracy and pedagogical nuance for the user.
4. What is the Best Way to Study Vocabulary, and How Can Lexogoth Help?
Vocabulary acquisition is a difficult and sometimes painful process for language learners. It is never enough to simply memorize the literal translation of a word; true language proficiency requires mastering a host of interconnected linguistic features, such as:
- Syntactic complements: Which type of complement follows a verb? (e.g., s’attendre à ce que [+ subjunctive] vs. s’attendre à [+ noun]).
- Collocations and prepositions: Which exact preposition must be used after a specific noun or adjective?.
- Pragmatics and register: To which social register or geographical variant does a specific meaning belong?.
We believe Lexogoth can ease and accelerate this demanding learning process. For many common French words and combinations, our corpus provides multiple, clear usage contexts that save the user immense amounts of time (such as endless searching in traditional dictionaries or deciphering abstract, academic definitions) and give them the exact means to reliably and systematically study and review vocabulary.
To achieve this, Lexogoth translates our theoretical principles directly into three practical, user-friendly modules designed to lower cognitive load and maximize retention:
4.1 The Lexical Toolbox in Practice
Database Functions (Smart Navigation & Filtering): To make our extensive corpus highly accessible, Lexogoth features intuitive database and search functions. Instead of requiring complex query codes like traditional linguistic software, users can easily search by lemma, filter results by Part-of-Speech (POS) tags, or query specific grammatical categories. This drastically reduces search friction, allowing students and teachers to instantly isolate, organize, and study exactly the linguistic material they need.
Syntactic Patterns & Collocations (Deconstructing Word Behavior): To prevent learners from making literal, unnatural translations, Lexogoth explicitly highlights recurring syntactic patterns and collocations. Rather than presenting verbs or nouns in isolation, this feature maps out exactly how words interact with one another (e.g., which preposition follows an adjective, or how a verb’s meaning shifts depending on its complement). By visualizing these structural dependencies, students learn to perceive and replicate native-like word combinations effortlessly.
The Workbench (Guided Reading & Active Noticing): Equipped with an integrated Part-of-Speech (POS) tagger, the Workbench serves as an interactive, guided reading environment. Instead of passively staring at static vocabulary lists, students can import French texts, analyze sentence structures, and actively “notice” collocations, grammar patterns, and word boundaries. This visual and typographical input enhancement helps bridge the gap between merely reading a word and actually understanding its syntactic behavior.
Context Modules (Reducing Cognitive Load): Instead of exposing learners to the unedited, chaotic data flood of traditional corpora, Lexogoth’s database is structured into digestible, pedagogically scaffolded context blocks. By presenting clean, hand-selected sentences with targeted annotations (explaining registers, verlan, or cultural nuances), we prevent cognitive overload and allow students to safely investigate and internalize lexical networks.
The Import/Export Function (The Bridge to Production): True vocabulary anchoring requires moving from receptive reading to active production. Lexogoth features a flexible import and export system, allowing students to seamlessly transfer selected collocations, chunks, and examples into their own personal study lists, digital flashcards (such as Anki), or writing portfolios.
4.2 The Role of the Educator
However, users must not fall into the trap of complacency. While Lexogoth is a powerful, pedagogically mediated tool, effective vocabulary acquisition cannot happen in a complete vacuum. To achieve genuine fluency, autonomous corpus study must always be paired with:
Highly targeted productive exercises (such as speaking, real-time interaction, and creative writing) where students can actively deploy their newly acquired chunks in communicative situations.
Quality pedagogical instruction to help contextualize complex patterns.
5. An Ongoing Project
Currently (July 2025), Lexogoth contains a little over 40,000 clickable items. We are actively working to fill the gaps that are (most certainly) still present in the work. This process will undoubtedly take several years, and we plan to release a major update every three months.
6. Bibliography
Albaqami, M. (2022). The Role of Lexical Chunks in Promoting English Writing Competence among Foreign Language Learners in Saudi Arabia. Arab World English Journal, 13(2), 1-15.
Bentolila, A. (2010). Le vocabulaire: pour dire et lire. In A. Bentolila, Parle à ceux que tu n’aimes pas: Le défi de Babel (pp. 157-174). Odile Jacob.
Boers, F., Eyckmans, J., Kappel, J., & Stengers, H. (2006). Formulaic sequences and perceived oral proficiency: putting a Lexical Approach to the test. Journal of Second Language Writing, 15(4), 227-246.
Ellis, R. (2003). Task-Based Language Learning and Teaching. Oxford University Press.
Laufer, B. (2022). Formulaic Sequences and Second Language Learning. In P. Szudarski & S. Barclay (Eds.), Vocabulary Theory, Patterning and Teaching: Studies in Honour of Norbert Schmitt (pp. 95-108). Multilingual Matters & Channel View Publications.
Lewis, M. (1993). The lexical approach: The state of ELT and a way forward. LTP.
Nation, P. (2001). Learning Vocabulary in Another Language. Cambridge University Press.
Pawley, A., & Syder, F. H. (1983). Two puzzles for linguistic theory: Nativelike selection and nativelike fluency. In J. C. Richards & R. W. Schmidt (Eds.), Language and communication (pp. 191-226). Longman.
Schmidt, R. W. (1990). The role of consciousness in second language learning. Applied Linguistics, 11(2), 129-158.
Sharwood Smith, M. (1993). Input enhancement in instructed SLA: Theoretical bases. Studies in Second Language Acquisition, 15(2), 165-179.
Thornbury, S. (2002). How to Teach Vocabulary. Pearson Education.
Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes. Harvard University Press.
Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in joint problem solving. Journal of Child Psychology and Psychiatry, 17(2), 89-100.
Wray, A. (2002). Formulaic language and the lexicon. Cambridge University Press.