Creating and Consulting Local SelfCompiled Corpora for Teaching Foreign Languages with Tools Available in CLARIN
Abstract
Poster presentation at PhD Student Session at CLARIN2025.
Full text
Creating and Consulting Local SelfCompiled Corpora for Teaching Foreign Languages with Tools Available in CLARIN Filip Zeman Autonomous University of Madrid [email protected] Thesis supervisor: Olga Batiukova CLARIAH–ES Acknowledgements This poster is part of the project Corpus of compositionality and lexical informativeness: annotation, analysis and applications (INFOLEXIS), (PI Olga Batiukova, reference PID2022-138135NB-100), financed by MCIN/AEI/10.13039/501100011033 and by the ESF+. The poster is part of the grant PREP2022-000882, financed by MCIN/AEI/10.13039/501100011033 and ESF+. This poster was funded by Comunidad de Madrid under the agreement “The node CLARIAH-CM: Digital Humanities and Language Technologies” (4180134). Project overview Self-compiled corpora: guidelines and illustrations •PhD thesis currently underway: The Use of Self-Compiled Corpora for Teaching Spanish as a Foreign Language: Selection Requirements and Lexical Selection. •Academic use of corpora: surge of proposals in 1990s, low implementation due to technical constraints. Revived attention since the 2010s. •Aim: Develop a technique for teaching Spanish vocabulary and lexical selection requirements that builds on the original data driven learning proposals and makes use of self-compiled corpora as didactic tools. Guidelines for self-compiled corpora use Minimum size: 50,000 words (100,000 words preferred for more accurate Word Sketch results; compilation time ± 1 hour). Method of Compilation: Direct search (more precise, ability to cleanse data noise) vs. crawling (faster, more noise in data). Compilation methodology: Student participation encouraged, teacher oversight necessary. Minimum CFER level: B1-B2 (some linguistic knowledge required). Search (Sk. En. specific): “BASIC” preferred. Sketch Engine + modern, large variety of tools + uses LogDice to evaluate lexical cohesion (unlike MI score does not over valuate rare combinations, is corpus-size-independent) – Paid software, free version NoSketchEngine (also available through CLARIN) offers limited functionality Lancs Box + free, relatively easy to use, automatic, variety of tools (concordances, N-grams, collocation search) – Offers limited functionality compared to Sketch Engine, uses MI for collocation stability indication Teitok + Free – Complex to use, XML/TEI knowledge needed to run CLARIN tools for corpus compilation and use Exemplary activities with Sketch Engine tools Concordance Tool: Shows keyword in context for searches in a corpus. Activities: lookup, verify existence of words, multiword units. Word Sketch Tool: Summarizes word’s syntactic and collocational behaviour. Activities: create syntactic mind maps, differentiate fixed vs. nonfixed combinations, grammatical uses Word Sketch Difference Tool: Compares Word Sketch results of two words. Activities: discern cuasisynonyms based on collocational patterns N-Grams Tool:Lists frequent word sequences of chosen length. Activities: 2+ word formulaic sequence extraction (combinations not detectable with Word Sketch), semantic mind map creation. Fig. 1. Word Sketch Difference results for problema (problem) and trastorno (disorder), Fig. 2. Word Sketch results for salud (health). From CLARIN •Corpus query tools available in CLARIN were used to design the teaching technique and will be used as learning tools for the experimental part of the thesis. For CLARIN •Guidelines on corpus compilation and didactic use with CLARIN-available software, incl. exemplary activities (LRC) •Data on the method’s effectiveness •Upload of ready-made corpora to CLARIN (future)