Automated measurement of complexity, accuracy, and fluency indices for Korean language data
Haerim HwangAbstract
This study introduces K-text, a website designed to measure Korean language proficiency from second language writing and speech data through 20 complexity, accuracy, and fluency indices using natural language processing. A key advantage of K-text is its ability to measure accuracy and fluency in addition to complexity. To test the 20 indices’ preliminary validity, this study investigates whether they can predict overall criterion-based Korean proficiency from 1 to 6 (equivalent to A1 to C2) when they are used to analyze 45,855 second language learners’ essays from the Korean Learner Corpus. The findings show that all of the individual K-text indices are significant predictors of overall Korean proficiency. Together, the 20 indices accounted for 44.50% of the variation observed in proficiency scores. These results suggest that K-text is a valuable resource for analyzing Korean language production data in relation to overall proficiency.