Full text
Straipsniai / Articles 273 CHARITONCHARITONIDIS Independentresearcher ORCIDid:orcid.org/0000-0003-0298-2629 Fieldsofresearch:wordformation,multiwordexpressions, lexicalsemantics,wordrecognition,emotion. DOI:doi.org/10.35321/all90-11 EXPLORINGSCALEDAIC WITHINENGLISHCLOSED COMPOUNDS Anglųkalbosuždarųjųjunginių (sudurtiniųžodžių)skalėsAICtyrimas ANNOTATION TheAkaikeInformationCriterion(AIC)isanestablishedgoodness-of-fitmeasurefor selectingmodelsintheanalysisofempiricaldata.However,AICissensitivetosample size.Author’spreviousresearchhasshownthatScaledAIC,i.e.AICdividedbysample size,isaneffectivetoolforassessingmodelfitandhierarchizingregressionmodels.The presentstudyexploresfurtherpropertiesofthisvariable.Theobjectofinvestigationare 66multipleregressionmodelsreferringtotheprocessingofclosed(concatenated)English compounds taken from Gagné et al.’s (2019) Large Database of English Compounds (LADEC).Inparticular,ScaledAICisjuxtaposedtotheEnglishLexiconProject(ELP) andBritishLexiconProject(BLP)assourcesofresponsetimes,thelexicaldecisionand namingtasks,compoundlength,andtransparencynorms.One-wayANOVA,maineffects analysis,andnon-parametrictestsareusedasmethods.ThefindingssuggestthatScaled AICisresponsivetoexperimentaldesign,thesourceofresponsetimes,andthelexical decision and naming tasks. At the same time, the results of this study offer empirical supportforthevalidationofmethodsemployedbyGagnéetal.(2019). KEYWORDS:Englishcompounds,ScaledAIC,lexicaldecision,naming. ANOTACIJA Akaikės informacijos kriterijus (angl. AIC) yra pastovus modelių tinkamumo matas, taikomasempiriniųduomenųanalizei.TačiauAICyrajautrusimtiesdydžiui.Ankstesni
CHARITONCHARITONIDIS 274 ActaLinguisticaLithuanicaXC autoriaus tyrimai parodė, kad skalės AIC, padalytas iš imties dydžio, yra veiksminga priemonė modelio tinkamumui įvertinti ir regresijos modeliams hierarchizuoti. Šiame tyrimenagrinėjamostolimesnėsšiokintamojoypatybės.Tyrimoobjektas–66daugialypės regresijos modeliai, susiję su uždarųjų (sudurtinių) anglų kalbos junginių, paimtų iš Gagné’ėsirkitų (2019)Anglųkalbossudurtiniųžodžių(junginių) didžiosiosduomenų bazės(angl.LADEC),apdorojimu.PirmiausiaAICsugretinamassuAnglųkalbosžodyno projektu(angl.ELP)irBritųkalbosžodynoprojektu(angl.BLP),kaipatsakolaiko,leksinių sprendimųirįvardijimoužduočių,junginiųilgioirskaidrumonormųšaltiniai.Naudojami metodai–vienpusėANOVA(angl.Analysisofvariance),pagrindiniųrezultatųanalizėir neparametriniaitestai.Išvadosrodo,kad skalės AICreaguojaįeksperimentinį projektą, atsakolaikošaltinįirleksiniųsprendimųbeiįvardijimoužduotis.Tuopačiušiotyrimo rezultataisuteikiaempirinįpagrindąGagné’ėsirkitų(2019)taikomųmetodųpatvirtinimui. ESMINIAIŽODŽIAI:anglųkalbosjunginiai(sudurtiniaižodžiai),skalėsAIC, leksinis sprendimas,įvardijimas. 1. THELARGEDATABASE OFENGLISHCOMPOUNDS (LADEC:GAGNÉETAL.2019)1 The Large Database of English Compounds (LADEC: Gagné et al. 2019) is the largest existing database of compound words. It contains over 8000nonspaced(“closed”or“concatenated”)compounds(=nouns)selected fromvarioussourcesincluding,amongothers,theCELEXdatabase(Baayen etal.1995),theEnglishLexiconProject(ELP;Balotaetal.2007),theBritish Lexicon Project (BLP; Keuleers et al. 2012), the British National Corpus (BNC),andWordnet.FromthefullsetofLADECentries,7,804compounds canbeuniquelyparsedintotwofreemorphemesconstituents.2Avastvariety ofcompoundsisconsidered,forinstancenoun-nouncompounds,e.g.buttercup, shipyard,compoundswithasecondconstituentderivedfromaverbalstem, e.g.pacemaker, painkiller,etc.(fordefinitionsofcompoundclassesseeLieber 2004:46).Thefirstnon-headconstituentreferstoawiderangeofgrammatical categories.Figure1containsabriefsampleofLADECentries. Gagnéetal.’s(2019)multiple-regressionmodelsincludeawiderangeof predictor(=independent)variables,suchascompoundlength,bigramfrequency 1 ThissectionwasadoptedfromCharitonCharitonidis(2022)withslightalterations. 2 LADECincludespluralsofalreadylistedcompoundsasseparateentries.
Straipsniai / Articles 275 Exploring Scaled AIC within English Closed Compounds at the morpheme boundary, family size, word frequency, probability and association (vector-based) measures, emotional/sentiment norms computed fromparticipantratings,etc.Thelogresponsetimesforthecompoundsfrom ELP(lexicaldecision,naming)andBLP(lexicaldecision)areusedasdependent variables.Forthemostpart,compoundlength(numberofcharacters)andlog compound(=word)frequencyfromtheSUBTLEX-UScorpus(Brysbaert,New 2009)3andBNC(BLP)areusedascontrolvariables.InGagnéetal.’s(2019) models, the predictor variables mentioned above had significant effects on lexicaldecisionandnamingtimes. FIGURE 1. LADECentries:sample afterlife aircraft ashtray daydreaming dimwit drawback pacemaker padlock painkiller backboard ballplayer buttercup earthquake egghead eyebrow shipyard shoelace shotgun caretaker castaway crossfire offspring outcasts overdrive textbook throwback turnaround The primary focus in Gagné et al.’s (2019) study was placed on various measuresofsemantic transparency.Gagnéetal.(2019)askedparticipantstorate compoundsconsideringhowpredictablethemeaningofthecompoundisfrom its parts (meaning predictability ratings, compound-based) and how much of themeaningofeachoftheconstituentsisretainedinthecompound(meaning retentionratings,constituent-based).Theauthorsfoundthatthedistribution oftransparenciesforthesecondconstituentwasmuchmorepeakedandhigher thanthedistributionoftransparenciesforthefirstconstituent(MC1:64.80[SD: 19.59]vs.MC2:71.00[SD:16.46].N = 8115). However, the ratingfor the 3 TheSUBTLEX-UScorpusisa51-million-tokencorpusbasedonsubtitlesfromUSfilmsand television programs. Several recent studies have provided evidence indicating that frequency normsobtainedfromsubtitlesofmoviesandtelevisionprogramstendtobemoreeffectivethan thosederivedfromprintedtextswhenitcomestoexplainingthedifferencesinlexicalprocessing timeand,insomecases,accuracyamongnativespeakersofvariouslanguages(seeChenetal. 2018:2andthereferencestherein).
CHARITONCHARITONIDIS 276 ActaLinguisticaLithuanicaXC first constituent was more strongly correlated with the rating for the entire compoundthanwastheratingforthesecondconstituent(c1~cmp:r=0.75, p<.001vs.c2~cmp:r=0.66,p<.001.N=429).4Mostnotably,themeaning retentionratingforthefirstconstituentandthemeaningpredictabilityrating forthecompoundpredictedallthreetypesofresponsetimes,i.e.ELPlexical decision,BLPlexicaldecision,andELPnamingtimes. Toconclude,thepeakedandhigherdistributionoftransparenciesforthe second constituent and the first constituent’s better association with the compound’s meaning predictability appear to be immediately mapped onto theheadoperationsinEnglishcompounds.Thesecondconstituent,i.e.the head,isaunitwhosetransparencyisenhancedcategoriallyandsemantically (as forthesemantic aspect, see the relationsof entailmentandhyponymy). Thefirstconstituent,i.e.themodifier,isthemostcriticalfactorinestablishing compoundreference.Asaresult,itstransparencycovarieswiththetransparency ofthecompoundmoststrongly.5 2. AKAIKEINFORMATIONCRITERION(AIC) In 1973, Hirotugu Akaike developed a method to estimate the relative expectationofKullback-Leibler distance(Kullback1959)usingFisher’smaximized log-likelihood(Fisher1922;seealsoAldrich1997).Thismeasure,commonly referredtoastheAkaike Information Criterion(AIC;Akaike1973),introduceda novelframeworkforselectingmodelsintheanalysisofempiricaldata,marking asignificantparadigmshift(Burnham,Anderson2002). AICistypicallycalculatedasfollows:–2lnL+2k,inwhich‘lnL’refersto themaximized/fulllog-likelihoodofthemodeland‘k’referstothenumber ofparametersincludingtheconstant.Asmallersetofpredictorsistypically associatedwithmoreefficientmodels(modelswithalowerinformationloss). Thelower(=morenegative)theAICvalue,thebetterthefitofthemodel.In thiscontext,AICpenalizes,asagoodness-of-fitmeasure,theuseofalarge numberofpredictorsthat,potentially,resultinhigherAICvalues(seethe‘+2k’ partoftheAICequation). 4 Steiger’s(1980)ztestshowedthatthisdifferencewassignificant,z=27.71,p<.0001(Gagné etal.2019). 5 Byreferringtopreviousresearch,Gagnéetal.(2019)reportthat“themodifier(thefirstconstituent inEnglish)tendstoplayalargerroleintheease-of-relationselectionduringtheprocessingof compoundsandnounphrases.”
Straipsniai / Articles 277 Exploring Scaled AIC within English Closed Compounds AICissensitivetosamplesize.AICc,acorrectedversionofAIC,incorporates samplesizethroughtheformula2k(k+1)/(n–k–1).However,itspecifically addressessmall samplesizesand isnotrecommendedfor models basedon largesamplesizessuchasthatinGagnéetal.(2019).6Itshouldbenotedthat researcherssuchasKennethP.Burnham&DavidR.Anderson(2002)donot offeradefinitivesolutionforcomparingAICvaluesofmodelsfittedonboth differentandlargesamplesizes.7 In particular, Burnham & Anderson (2002: 80–85, 334–335) provide a comprehensivediscussionoftheimplicationsofunequalsamplesizesformodel comparison.Theyarguethatemployinginformationcriteriatocomparemodels withdifferentsamplesizescanleadtomisleadingresults.Similarly,asnotedin anonlinediscussionbySvetunkovin2016(seereferenceafterthebibliography), all information criteria are based on the likelihood function that, in turn, dependsonsamplesize.Specifically,asthesamplesizeincreases,thelikelihood decreases.Consequently,informationcriteriawillalsoincreaseinsuchcases.8 3. PREVIOUSRESEARCH InCharitonidis(2022)theAICvaluesfor44multipleregressionmodelswith differentcombinationsofemotionvariables(valence,arousal,andconcreteness for(a)wordsand(b)wordcontexts)weredividedbysamplesize(N)toyield ScaledAIC(AIC/N)values.9Subsequently,thesevalueswereutilizedtoassess 6 ForfurtherinformationonAICc,thereaderisreferredtoBurnham&Anderson(2002:374–380). 7 OneofthesolutionsthatBurnham&Anderson(2002)proposereferstothetransformationofthe AICvaluesto“Akaikeweights”thataredefinedas“therelativelikelihoodofthemodel,giventhe data”(Burnham,Anderson2002:xiii;seealsoWagenmakers,Farrell2004). 8 Availableat:https://stats.stackexchange.com/questions/94718/model-comparison-with-aic-basedon-different-sample-size [accessed 16.06.2023]. The reader can comprehend Svetunkov’s statementbysubstitutingdifferentvaluesforthe‘lnL’componentintheAICequation,while maintainingthe‘2k’componentconstant.AdecreaseinthelnLvaluewillresultinahigher,i.e. inferior,AICvalue. 9 Intheliterature,ScaledAICisalsoreferredtoas“meanAIC”.AccordingtoSvetunkov(personal communication),thepracticeofdividingtheAkaikeInformationCriterionbythesamplesize isnotnovel.Forinstance,Hastieetal.(2009:230–231)defineAICinanon-canonicalmanner, employingNasthedenominatorintheformula.Whilethisdeviationfromtheconventional AICformulaisnotwithoutitscritics,itremainsa prevalentapproach,as exemplifiedbyits inclusioninthestatisticalsoftwarepackageStata.Forinstance,Statareports“AICdividedbyN” initsmodeloutput,asevidencedbyvariousexamplesavailableonline(Gratitudeisextendedto I.Svetunkovforprovidingthisinformation).
CHARITONCHARITONIDIS 278 ActaLinguisticaLithuanicaXC andcomparethemodels’goodness-of-fit.Theinsertionofkeypredictorsinto global, i.e.general,modelsshowedthattheBLPlexicaldecisiontimescalled forabettergoodness-of-fitthantheELPlexicaldecisiontimes.Thefitofthe ELPnamingmodelsfellwithintherangeofthoseobservedfortheELPand BLPlexicaldecisionmodels.Mostnotably,context concreteness for the second constituentemergedasasignificantpredictorinallmodelswithSUBTLEX-US frequency,acrosslexicaldecisionandnaming. InCharitonidis(2024),allsignificantcoefficientsfromtheglobalmodelswith SUBTLEX-USfrequencywerejuxtaposedtothehyponymyvariable(Gagné etal.2020).Itwasfoundthatmodelsincludingbothhyponymyandcontext concretenessforthesecondconstituentwerealwaysassociatedwiththelowest (=best)ScaledAICvalueascomparedtonested, i.e.reduced,modelsomitting eitherofthesetwovariables.ThesubsequentlyappliedWaldtestsshowedthat nestedmodels,alwaysreferredtoasignificantreduction(=deterioration)ofthe coefficientofdetermination(R2).Tables1and2displaytheScaledAICvalues andtheresultsofthecorrespondingWaldtests,respectively. TABLE1. ScaledAICvaluesfornestedmodelsomittinghyponymy(‘Model2’) orcontextconcretenessforthesecondconstituent(‘Model3’)fromfull models(‘Model1’)topredictEnglishLexiconProject(ELP)lexical decision(LD)times,BritishLexiconProject(BLP)lexicaldecision times,andELPnamingtimes Model Scaled AIC AIC N ELPLD 1 -3.36281a-3557.85 1058 2-3.30375 -4169.334 1262 3-3.34845 -3700.038 1105 BLPLD 1 -3.79552a-2903.574 765 2-3.76618 -3920.592 1041 3-3.76718 -2987.37 793 ELPnaming 1 -3.58686b-7396.108 2062 2-3.54304 -8418.27 2376 3-3.58379 -7389.784 2062
Straipsniai / Articles 279 Exploring Scaled AIC within English Closed Compounds a. Predictors: (Constant), hyponymy judgement, length of compound, SUBTLEX-USfrequency,representationvalence(cmp),contextconcreteness (c2) b. Predictors: (Constant), hyponymy judgement, length of compound, SUBTLEX-USfrequency,contextvalence(cmp),contextarousal(c1),context arousal(c2),contextconcreteness(c2) TABLE2. Waldtestsfornestedmodelsomittinghyponymy(‘Model2’)orcontext concretenessforthesecondconstituent(‘Model3’)fromfullmodels (‘Model1’)topredictEnglishLexiconProject(ELP)lexicaldecision (LD)times,BritishLexiconProject(BLP)lexicaldecisiontimes,and ELPnamingtimes Model R2 square F change df1 df2 p ELPLD 1 .184a47.506 5 1052 .000 2-.004 4.562 1 1052 .033 3-.010 13.175 1 1052 .000 BLPLD 1 .211a40.479 5 759 .000 2-.013 12.091 1 759 .001 3-.015 14.007 1 759 .000 ELPnaming 1 .288b118.711 72054 .000 2-.003 8.286 1 2054 .004 3-.003 8.309 1 2054 .004 a. Predictors:(Constant),hyponymyjudgement,lengthofcompound, SUBTLEX-USfrequency,representationvalence(cmp),contextconcreteness (c2) b. Predictors: (Constant), hyponymy judgement, length of compound, SUBTLEX-USfrequency,contextvalence(cmp),contextarousal(c1),context arousal(c2),contextconcreteness(c2) Inconclusion,twodifferenteffect-sizemeasures,namelyScaledAICand R2,hierarchizedthesameregressionmodelsidenticallywhiledemonstrating thesamepreferenceforthebestmodel.Thus,thereisstrongevidencethatthe ScaledAICmeasureisaqualitativetoolforassessingmodelfit.
CHARITONCHARITONIDIS 280 ActaLinguisticaLithuanicaXC 4. THEPRESENTSTUDY Thepresentstudybuildsupontheauthor’spreviousresearchpresentedin section3.Theresearchsubjectsare66lexicaldecisionandnamingmodelsfor theEnglishclosed(concatenated)compoundsbuiltbyGagnéetal.(2019).All modelsincludeSUBTLEX-USfrequencyascontrolvariable.Ourobjectives aretwofoldandruninparallel.First,weassessthecharacteristicsofGagnéet al.’s(ibid.)models.Second,weexploreessentialpropertiesoftheScaledAIC measure. Theresearchquestionsare: 1.IsScaledAICsensitivetothemodeldesigninGagnéetal.(2019)?Which modelgroupsarefavoured? 2.Whatistheimpactofthecontrolvariables‘compoundfrequency’and ‘compoundlength’onScaledAIC? 3.HowismorphologicaltransparencyrelatedtoScaledAIC? Ourstudyisstructuredasfollows:Section5providesanoverviewofour methods.Section6.1providesdescriptivestatisticsforScaledAICreferring tothemodelsunderconsideration.Emphasisisgiventotheparametricversus non-parametriccharacteristics ofmodelcategories.Section6.2exploresthe relationshipbetweenthesourceofresponsetimesandthelexicalprocessing tasks.Section6.3juxtaposesScaledAICtothecontrolvariables‘compound frequency’and‘compoundlength’.Insection6.4thesignificancelevelsofthe transparencycoefficientsfromGagnéetal.’s(2019)modelsaremappedonto theScaledAICvalues.Thekeyfindingsaresummarizedinsection7,followed byadiscussionoftheresultsinsection8. 5. METHODS Ourgeneralmethodwasthecomparativeanalysisofthemainparametersand characteristicsofGagnéetal.’s(2019)models,usingScaledAICasthedependent variable. Independent variables included sample characteristics (e.g. response timesourceandthelexicalprocessingtasks),studydesign(e.g.controlvariables), andthesignificanceleveloftransparencycoefficients,amongotherfactors. Thespecific statisticalmethodsemployed wereasfollows: (a)descriptive statisticspertainingtomeansandmedians,alongwiththeapplicationofthe Shapiro-Wilktesttoassessthecentraltendency,variability,anddistribution ofScaledAICacrossdifferentmodelcategoriesandgroups(sections6.1and 6.2), (b) main effects analyses conducted for the source of response times (ELP/BLP)andthelexicalprocessingtasks(lexicaldecision/naming)(section
Straipsniai / Articles 281 Exploring Scaled AIC within English Closed Compounds 6.2),(c)distinctANOVAsperformedonresponsetimesourceandthelexical processingtasks,incorporatingcompoundlengthasacovariate(section6.3), and (d) utilization of the Kruskal-Wallis and the Jonckheere-Terpstra tests to explore differences amongthe ranks of ordinally-recodedcoefficientsfor semantictransparency(section6.4).Formoreinformationonmethods,the readerisreferredtotheanalysesinsections6.1–6.4. 6. ANALYSES 6.1. ScaledAICvs.modelcategories The66AICvaluesfromGagnéetal.’s(2019)multiple-regressionmodels with SUBTLEX-US frequency as a control variable were divided by each model’ssamplesizetoyieldasetof66ScaledAICvalues. Table3belowprovidesthedescriptivestatisticsforScaledAICandFigure 2displaysthecorrespondingboxplotreferringtotheorderedsetofvalues.10 Therewerenooutliersinthesample.Theskewness(Sk)andkurtosis(Ku) valuesweretolerable.11 The mean value for Scaled AIC was -3.50030. The standard deviation was0.19052,thatistheobservationswererelativelytightlyclusteredaround themean.Theminimumandmaximumvalueswere-3.85502and-3.15754, respectively.Themedianvaluewas-3.55732,i.e.slightlylowerthanthemean value.12Themiddle50%ofthedatarangedbetween-3.66575(firstquartile, Q1) and -3.30959 (third quartile, Q3). Accordingly, the interquartile range (IQR)was0.35616. 10 Thelowerorfirst quartilelineofthebox(Q1)markstheboundarybelowwhichthebottom 25%ofthedataextends.Similarly,theupperorthird quartilelineofthebox(Q3)marksthe boundaryabovewhichtheupper25%ofthedataextends.Theshadedareashowstheboundaries ofthemiddle50%ofthedataorinterquartile range(IQR),whichcanbecomputedbysubtracting thefirstquartilefromthethirdquartile(Q3-Q1).Thehorizontallineinsidetheboxshowsthe medianormiddle quartile(Q2),i.e.thevaluethatfallsinthemiddleofthedataset. 11 With reference to the SPSS environment, the values between -1 and +1 for skewness and between -2 and +2 for kurtosis are generally considered acceptable for normal distribution assessment. It is worth noting, however, that skewness and kurtosis alone do not provide a conclusive proof of normality (see also the discussion on https://www.researchgate.net/post/ What_is_the_acceptable_range_of_skewness_and_kurtosis_for_normal_distribution_of_data). 12 Inlinewiththispattern,therewasasmallamountofpositiveskewinthedata(0.494).
CHARITONCHARITONIDIS 288 ActaLinguisticaLithuanicaXC Todetecttheinfluenceofcompoundlengthandcompoundfrequencyon ScaledAIC,the respectiveregressioncoefficientswererecodedintoordinal values according to their positivity and significance level, see Table 6. The significancelevelsweremappedontoordinalscalesbecausetheyrepresented conventionalcut-offpointsbasedontheexactsignificancevalues. TABLE6.Ordinalrecodingchartforregressioncoefficients Significance Positivity Ordinalvalues Description p<.001 negative -3 largenegativeeffect p<.01 negative -2 moderatenegativeeffect p<.05 negative -1 smallnegativeeffect p>.05 negative/positive 0 non-significanteffect p<.05 positive 1 smallpositiveeffect p<.01 positive 2moderatepositiveeffect p<.001 positive 3largepositiveeffect In Gagné et al.’s (2019) models, SUBTLEX-US frequency was always associatedwithnegative(=latency-reducing)coefficientswithalargeeffect,p< .001.Accordingly,allcoefficientswererecodedas-3,avaluethatwasperfectly collinearwiththeoutcomevariable,ScaledAIC.Forthisreason,SUBTLEX-US frequencywasexcludedfromthepresentanalysis. Asforcompound length,allsignificantregressioncoefficientsfromGagnéet al.’s(2019)modelshadalargepositive(=latency-inducing)effect,p<.001. IncontrasttotheSUBTLEX-USvariable,severalnon-significantcoefficients showedup.Giventhesepatterns,acategoricalvariablewascreatedwiththe values‘1’forpositiveeffect(=interferenceofcompoundlength)and‘0’forno effect(=nointerferenceofcompoundlength).Theresultingsamplecontained 39ScaledAICvalues.ThePearsoncorrelationtestbetweencompoundlength andScaledAICyieldedahighlysignificantcorrelationcoefficientof0.51,p= .001,indicatingamoderate-to-strongcorrelationbetweenthetwovariables. Compoundlengthwasincludedasasingleindependentvariableinalinear regression model. It was found that the predicted Scaled AIC mean for no interferenceofcompoundlengthwas3.611(=theintercept).Theinterference ofcompoundlengthresultedinahigher(=inferior)valueof-3.407(b=0.204, p=.001).Figure6belowillustratesthesepatterns.Inanutshell,ScaledAIC deteriorateswhencompoundlengthbecomesrelevantwithinmodels.
Straipsniai / Articles 289 Exploring Scaled AIC within English Closed Compounds FIGURE6.CompoundlengthandScaledAIC WeconductedseparateANOVAsforresponsetimesource(ELP/BLP)and task performance (lexical decision/naming), taking into account compound length as a covariate. The primary objective was to identify disparities in meansthatwerepreviouslyadjustedtoaccommodatethecontrollingeffectsof compoundlength. (a)ANOVAforresponsetimesource.BLPwasassignedthevalue‘0’and ELPwasassignedthevalue‘1’.ThePearsoncorrelationtestrevealedastrong collinearitybetweenresponsetimesourceandcompoundlength,r=1(N=39). Thefollowingevidencesupportsourfinding:First,theELPgroupconsistently showed significant positive correlations with compoundlength,indicatinga large effect. Second, the BLP group consistently displayed non-significant correlationswithcompoundlength.17Consequently,thepredictedScaledAIC meanforresponsetimesourcewasthesamewithorwithoutcompoundlength intheanalysis(M=3.407inbothcases).Insummary,compoundlengthdid nothaveasignificanteffectonScaledAICwhentheresponsetimesourcewas included. (b)ANOVAfortaskperformance.Namingwasassignedthevalue‘0’and lexicaldecisionwasassignedthevalue‘1’.ThePearsoncorrelationtestrevealed anegativecorrelationbetweentaskandcompoundlength,indicatingamoderate effect,r=-.5,p=.001(N=39).Thisresultsuggeststhatcompoundlengthis morerelevanttonamingthantolexicaldecision. 17 Itshouldbenotedthat11ofthe13BLPcoefficientswerenegative.
CHARITONCHARITONIDIS 290 ActaLinguisticaLithuanicaXC When considering task as the primary variable, compound length was a significantpredictorofScaledAIC,F(1,145.030),p=.000.Similarly,when consideringcompoundlengthastheprimaryvariable,taskwasasignificant predictorofScaledAIC,F(1,125.30),p=.000.ThepredictedScaledAICmean fortaskalonewassignificantlydifferentfromthepredictedScaledAICmean whencompoundlengthwastakenintoaccount(3.584vs.3.203,respectively). Likewise,thepredictedScaledAICmeanforcompoundlengthalone(3.584) wassignificantlydifferentfromthepredictedScaledAICmeanwhenthetask wastakenintoaccount(-3.23).Summarizing,intermsofcovariateadjustment, boththelexicaldecisiontaskandcompoundlengthpredictedinferiormodels. 6.4. ScaledAICvs.transparencynorms This section investigates the effect of regressioncoefficientsforsemantic transparencyinGagnéetal.(2019)modelsonScaledAIC.Theseregression coefficientswerecodedonthreeordinalscales,eachcorrespondingtooneof thethreemorphologicallevels,i.e.compound,firstconstituent,andsecond constituent.TheordinalrecodingchartcanbefoundinTable6. Table7belowdisplaysthemediansandrangesoftheordinally-transformed transparency coefficients for all three morphological levels. Τhe medians provideusefulinformationaboutthecentraltendencyanddispersionofordinal valuesandcanhelpinformanalysesbasedonordinalvariables. TABLE7.Ordinally-transformedtransparencycoefficients:Mediansandranges Median Minimum Maximum Compound -3 -3 -1 Firstconstituent 2-3 3 Secondconstituent 0 -2 3 N=18 Ascanbeseen,themedianforthecompoundwas‘-3’,themedianforthe firstconstituentwas‘2’,andthemedianforthesecondconstituentwas‘0’. ThesefindingssuggestthatinGagnéetal.’s(2019)modelswithSUBTLEX-US frequency, transparency for the compound was associated with a large negativeeffect(shorterresponsetimes),transparencyforthefirstconstituent was associated with a moderate positive effect (longer response times), and transparencyforthesecondconstituentdidnothaveasignificanteffectorhad
Straipsniai / Articles 291 Exploring Scaled AIC within English Closed Compounds anuncertainrole.Thehighertransparencyratingsforthesecondconstituent, reportedbyGagnéetal.(ibid.),suggestaninherentbiasfavouringit,leadingto theoveralltransparencyofthecompoundbeingdependentonthetransparency ofthefirstconstituent.Inthiscontext,thepositive,latency-inducing,median forthefirstconstituentindicatesitsmediating,perhapsreference-establishing, rolein this relationship(seealsosection1).Itremainsto be demonstrated whicharethesemanticfunctionsthatsufficientlyrepresent,inprocessingterms, theinherentbiasofthesecondconstituent.18 Theresearchquestiontobeaddressednowiswhetherthepositivityand significance level of transparency coefficients influence Scaled AIC. Our methodprimarilyaimsatdetectingoverfittingeffects.AsDanielJ.Navarro &JayI.Myung(2005)argue,overfittingoccurswhen“acomplexmodelwith manyparametersandhighlynonlinearformcanoftenfitdatabetterthana simple model with few parameters even if the latter generated the data” (Navarro,Myung2005:1240).Accordingly,alargenumberofparametershave thepotentialtocapturenoiseoruniquecharacteristicsoftheavailabledatabut mayhinderthemodel’sabilitytogeneralize to new data.AICmitigatestheissue ofoverfittingbyintroducingapenaltyontheinclusionofnumerousparameters inamodel,seethe‘+2k’partoftheAICequationinsection2. Regardingtheanalysistofollow,itispostulatedthatmodelsexhibitinghigher (=inferior)ScaledAICvaluesmaypossesssignificant,systematicallyderived, coefficients,i.e.coefficientsthatarerelevantaccordingtotheLADECdataset alone.Inthiscontext,ourconjecturesuggeststhatacontrastingtrendmight emergeintheconnectionbetweentransparencyandScaledAIC,ascompared totheindicationprovidedbythemediansinTable7. In particular, lower (=better) Scaled AIC values may be associated with (a) positive coefficients (longer response times) concerning the whole compound,(b)negativecoefficients(shorterresponsetimes)concerningthe first constituent, and (c) positive or negative coefficients (longer or shorter responsetimes,respectively)concerningthesecondconstituent.Itshouldbe notedthat,regardingthesecondconstituent,themedianinTable7suggestsno effect. To answer the research question, two non-parametric measures will be employed,i.e.theKruskal-WallistestandtheJonckheere-Terpstratest.The Kruskal-Wallistest,alsoknownasthe‘Htest’,isanon-parametrictestbasedon 18 InCharitonidis(2024)itisarguedthatbothhyponymyandcontextconcretenessforthesecond constituent are significant semantic predictors in lexical decision and naming. The analysis presentedthereinshowsthatincludingbothofthesepredictorsresultsinanimprovementin ScaledAICandR2ascomparedtomodelsthatomiteitherofthesevariables.
CHARITONCHARITONIDIS 292 ActaLinguisticaLithuanicaXC thechi-squaredistribution.Itrequiresthatthedependentvariablebeordinal orcontinuous.Thistestisdesignedtodeterminewhethertherearesignificant differences between the medians of two or more groups and is used as an alternativetoone-wayANOVA.Concerningtheprocedure,thevaluesofthe continuousdependentvariable,i.e.ScaledAIC,wereorderedfromlowestto highestandthescoreswereassignedranks.Theresultingrankswereentered backintothegroupsofsignificancelevel(theindependentvariable)andthe ranksforeachgroupweresummed.Theformulaforcalculating‘H’involved, amongothers,squaringthesumofranksforeachgroupandthendividingthis valuebysamplesize.19Tables8–10containtheinputdataconsideredandthe sumofranksforeachgroup.20 TABLE8.Compound Significancelevels N SumofRanks Scaled AIC 1–smallnegativeeffect 1 2 2–moderatenegativeeffect 5 46 3–largenegativeeffect 12 123 Total 18 TABLE9.Firstconstituent Significancelevels N SumofRanks Scaled AIC 1–largepositiveeffect 6 59 2–moderatepositiveeffect 4 34 3–smallpositiveeffect 1 3 4–noeffect 1 2 5–smallnegativeeffect 1 10 6–moderatenegativeeffect 1 11 7–largenegativeeffect 452 Total 18 19 FortherestofcalculationsseeField(2009:561–562). 20 Inallthreetables,thetotalsumofranksisapproximately171.Itisequaltothesumoftheintegers from1to18,seesamplesize(N).
Straipsniai / Articles 293 Exploring Scaled AIC within English Closed Compounds TABLE10.Secondconstituent Significancelevels N SumofRanks Scaled AIC 1–largepositiveeffect 6 59 2–smallpositiveeffect 1 14 3–noeffect 8 59 4–smallnegativeeffect 1 18 5–moderatenegativeeffect 221 Total 18 BeforedelvingintotheresultsoftheKruskal-Wallis(H)test,itisimportant tonotethatthistestdoesnotprovideinformationaboutthespecificdifferences betweenindividualgroups.Toaddressthisissue,theJonckheere-Terpstra(JT) testwasadditionallyemployed.Thistestprovidedinformationaboutwhether themediansofthegroupsincreased or decreased in the orderspecifiedby thecoding(=grouping)variable,specificallyfromlargepositiveeffecttolarge negativeeffect.Regardingmethods,theJTstatisticwasconvertedintoaz-score. Apositivez-valueindicatedatrendofascendingmedians,thatisthemedians increased(=higher/inferiorScaledAIC)asthevaluesofthecodingvariable increased.Anegativez-valueindicatedatrendofdescendingmedians,thatis themediansdecreased(=lower/betterScaledAIC)asthevaluesofthecoding variableincreased.Inthefollowing,theresultsoftheKruskal-Wallis(H)and Jonckheere-Terpstra(JT)testsaregivenjointly. (a)ScaledAICforthecompoundwasnotsignificantlyaffectedbysignificance level,asdeterminedbytheKruskal-Wallistest(H(2)=2.226,p>.05).Atrend ofascendingmedianswasfoundconfirmingouroverfittinghypothesis,seethe negativemedianforthecompoundinTable7.Thistrend,however,wasnot statisticallysignificantaccordingtotheJonckheere-Terpstratest(JT=50,z= 1.064,p>.05). (b) Scaled AIC for the first constituent was not significantly affected by significancelevel,asdeterminedbytheKruskal-Wallistest(H(6)=5.427, p>.05).Atrendofascendingmedianswasfoundrejectingouroverfitting hypothesis,seethepositivemedianforthefirstconstituentinTable7.This trend,however,wasnotstatisticallysignificantaccordingtotheJonckheereTerpstratest(JT=70,z=0.549,p>.05). (c)ScaledAICforthesecondconstituentwasnotsignificantlyaffectedby significancelevel,asdeterminedbytheKruskal-Wallistest(H(4)=4.607,p> .05).Atrendofascendingordescendingmedianswasnotobserved,rejecting
CHARITONCHARITONIDIS 294 ActaLinguisticaLithuanicaXC our overfitting hypothesis. In particular, the z-statistic of the JonckheereTerpstratestwasessentiallyzero,inaccordancewiththezeromedianforthe secondconstituentinTable7(JT=54,z=0.041,p>.05). Summarizing,itcanbeinferredthatthesignificancelevelofthetransparency coefficientsinGagnéetal.’s(2019)modelswithSUBTLEX-USfrequencydoes notaffectthemagnitudeofScaledAIC.Thisfindingindirectlysupportsthe qualityofGagnéetal.’s(2019)modelswithtransparencypredictors,specifically indicatingthattheoverfittinghypothesisforthesemodelsis nottenable. A limitationofthepresentstudyisthesmallsamplesizeused,withN=18.To confirmourfindings,moreresearchisneededusingawiderrangeofScaled AICvalues. Toensureclarityandcompletenessinpresentingourresearchoutcomes,we haveincorporatedadedicatedsectionfocusedonsummarizingthekeyfindings ofourstudy.Forthiscomprehensiveoverview,pleasecontinuetoSection7. 7. KEYFINDINGS Table 11 below presents a comprehensive analysis of model performance and relevant variables in the context of lexical decision and naming tasks, basedonthefindingsofGagnéetal.(2019).Eachsectionofthetabledelves intospecificsubjects,revealingwhichmodelsaremosteffective.TheANOVA andtheKruskal-Wallis/Jonckheere-Terpstratests(sections6.3and6.4)were appliedafterassigningnominal(ordinalorcategorical)valuestotheregression coefficientsfromGagnéetal’s(2019)models.Fordetailsonthespecialtests applied,pleaserefertotherespectivesections. TABLE11. ScaledAICwithinEnglishclosedcompounds:Comprehensiveanalysis ofmodelperformanceinlexicaldecisionandnamingtasks(Gagnéetal. 2019) Subjects Statistics Evaluation Section Modelcategories Descriptives Normalitytests ELPlexicaldecision BLPlexicaldecision ELPnaming NPAR/~ PAR/✓ PAR/✓ 6.1 Responsetime source Lexicalprocessing task Maineffects ELPlexicaldecision BLPlexicaldecision ELPnaming ~ ✓ ✓6.2
Straipsniai / Articles 295 Exploring Scaled AIC within English Closed Compounds Subjects Statistics Evaluation Section Controlvariables ANOVA Compound frequency Compoundlength ✓ ~ 6.3 Semantic transparency Kruskal-Wallis JonckheereTerpstra Firstconstituent Secondconstituent Compound NOF/ns NOF/ns NOF/ns 6.4 PAR:parametricdata|NPAR:non-parametricdata| ✓:bettermodels (lowerAIC)~:inferiormodels(higherAIC)|NOF/ns:nooverfitting/nonsignificanttest 8. DISCUSSION Previous research by Charitonidis (2022, 2024) has demonstrated that ScaledAICisareliablegoodness-of-fitmeasurethatcanbeemployedinmodel selection,perhapsincooperationwithothermeasuressuchastheWaldtest(see section3).WithreferencetoGagnéetal.’s(2019)multiple-regressionmodels with SUBTLEX-US frequency, the present analysis introduced additional propertiesoftheScaledAICmeasure.Whilevalidconcernshavebeenraised regarding the comparison of models fitted on different sample sizes using informationcriteria(seesection2),thefindingsofthisstudysuggestthatin certaincontexts,ScaledAICcanindeedbeavaluabletoolforassessingmodel fitandhierarchizingregressionmodels.Ourresearchhasdemonstratedthat ScaledAICisresponsivetoexperimentaldesign,responsetimesources,and specific tasks. However, it is essential to recognize that the applicability of ScaledAICmaybecontext-dependent,anditsutilityshouldbeevaluatedon acase-by-casebasis. Beforeproceedingtotheprimaryfindingsofthispaper,itisimportantto addresstheresearchquestionssetupinsection4. 1. The distributions of Scaled AIC values, along with combinations of differentsourcesofresponsetimesandprocessingtasks,suggestthatScaled AICeffectivelyidentifiesthepresenceorabsenceofwell-definedunderlying factorsinexperimentaldesignandstatisticalmodelling.Inthiscontext,BLP lexicaldecisionandELPnamingexhibitedstrongerpredictivepowerforScaled AICevenundercontrolledconditions. 2.CompoundfrequencywasunexceptionallyanegativepredictorofScaled AIC,alwaysindicatingalargeeffect.ELPlexicaldecisionconsistentlyshowed
CHARITONCHARITONIDIS 296 ActaLinguisticaLithuanicaXC significant positive correlations with compound length predicting higher (=inferior)ScaledAICvalues.BLPlexicaldecisionconsistentlyshowednonsignificant correlations with compound length. Both lexical decision and compoundlengthpredictedinferiormodelsincovariateadjustment. 3.Thepositivityandthesignificanceleveloftransparencycoefficientsin Gagnéetal.’s(2019)modelsdidnotaffectthemagnitudeofScaledAIC.This findingimpliesthatGagnéetal.’s(2019)modelswithtransparencypredictors donotintroduceoverfittingbias. Byreferencingspecificsectionsoftheanalyses,theprimaryfindingsofthis studycanbesummarizedasfollows: InSection6.1,ouranalysisfocusedonthecomparisonbetweenScaledAIC valuesacrossvariousmodelcategories.Evenafterattemptingthetransformations ‘naturallogarithm’and‘squareroot’ontheabsolutevalues,theoverallScaled AIC sample did not conform to a normal distribution. Similarly, the ELP lexicaldecisionmodelsshowcasedanon-parametricdistributionoftheirScaled AICvalues.Onthecontrary,thedatarelatedtoBLPlexicaldecisionandELP namingfollowedanormaldistributionpattern. In Section 6.2, our focus shifted to examining the relationship between ScaledAICand(a)thesourcesofresponsetimesand(b)taskperformance. Interestingly,therangesofScaledAICvaluesfortheELPandBLPlexical decision models did not overlap, signifying their distinctiveness. The test resultsrevealedsignificantmaineffectsofbothresponsetimesourceandtask performanceonScaledAIC.Notably,thepredictivecapabilityofScaledAIC wasbetterformodelsassociatedwiththeBLPlexicaldecisiontimesandthe naming task. These findingscontributetothe precisionandefficacy of the respectivemodelssignificantly. InSection6.3,ourexplorationdelvedintotherelationshipbetweenScaled AICandthecontrolvariables‘compoundfrequency’and‘compoundlength’. Compoundfrequencywasexcludedfromtheanalysisbecauseitwasperfectly collinearwithScaledAIC.Ontheotherhand,adeclineinScaledAICvalues wasobservedwhencompoundlengthbecamearelevantfactorwithinmodels. Concomitantly,compoundlengthwasmostrelevantforthenamingtask. Intermsofcovariateadjustment,boththelexicaldecisiontaskandcompound lengthwerepredictiveofinferiormodels. InSection6.4,ourfocuswasplacedontherelationshipbetweenScaledAIC andsemantictransparency.Theresearchquestionwaswhetherthepositivity andthesignificanceleveloftransparencycoefficientsinGagnéetal.’s(2019) modelshadanimpactonScaledAIC.Theprimarygoalofourmethodwas todetectpotentialoverfittingeffects.Wepostulatedthatmodelswithinferior ScaledAICvaluesmightpossesssignificantcoefficientsthatholdrelevance
Straipsniai / Articles 297 Exploring Scaled AIC within English Closed Compounds accordingtotheLADECdatasetalone.Whileweobservedatrendofincreasing (=inferior)ScaledAICvaluesforaclusterofsignificantnegativecoefficientsat thecompoundlevel–aligningwithouroverfittinghypothesis–theJonckheereTerpstratestshowedthatthistrenddidnotachievestatisticalsignificance. Inconclusion,theexplorationofdifferentparametersusingScaledAICasa dependentvariablehasilluminatedthediversewaysinwhichmodelcategories, response time source, processing tasks, control variables, and semantic transparencyimpactthegoodness-of-fitofmodels.Byrecognizingthenuanced relationshipsamongtheseelements,researchersarebetterequippedtomake informeddecisionsinmodelselection,adjustments,andinterpretation. REFERENCES Akaike Hirotugu 1973:InformationTheoryandanExtensionoftheMaximum Likelihood Principle. – Second International Symposium on Information Theory, eds. B.F.Csaki,B.N.Petrov,Budapest:AcademiaiKiado,267–281. AldrichJohnR.1997:R.A.FisherandtheMakingofMaximumLikelihood1912– 1922.–Statistical Science12(3),162–176. Baayen Harald R., Piepenbrock Richard, Gulikers Leon 1995: The CELEX Lexical Database(Dataset,Release2,CD-ROM),LinguisticDataConsortium, UniversityofPennsylvania.Availableat:https://catalog.ldc.upenn.edu. Balota David A., Yap Melvin J., Cortese Michael J., Hutchison KeithI.,KesslerBrett,LoftisBjorn,NeelyJamesH.,NelsonDouglas L.,SimpsonGregB.,TreimanRebecca2007:TheEnglishLexiconProject.– Behavior Research Methods39,445–459. Brysbaert Marc, New Boris 2009: Moving Beyond Kučera and Francis: A CriticalEvaluationofCurrentWordFrequencyNormsandtheIntroductionofaNew andImprovedWordFrequencyMeasureforAmericanEnglish.–Behavior Research Methods41,977–990.DOI:doi.org/10.3758/BRM.41.4.977. BurnhamKennethP.,AndersonDavidR.2002:Model Selection and Multimodal Inference: A Practical Information-Theoretic Approach, 2ndedition,NewYork:Springer. DOI:dx.doi.org/10.1007/b97636. Charitonidis Chariton 2022:ContextConcretenessfortheSecondConstituent SlowsDownCompound-WordProcessing.–Lexis20.DOI:doi.org/10.4000/lexis.6769.