An Introduction to the Possibility of Applying the k-means Algorithm to the Typology of Characters from Folk Tales
DOI:
https://doi.org/10.12775/LL.1.2.2026.009Keywords
folktale, charackter analysis, myth theory, clustering, computational folkloristicsAbstract
The classification of folktale characters through the application of statistical methods – particularly those associated with unsupervised machine learning – remains a largely unexplored domain within folklore studies. Several intellectual traditions may be regarded as providing the conceptual foundation for such an inquiry, including formalism, structuralism, and cognitive approaches to the study of mythical forms, all of which share an emphasis on identifying recurrent patterns in folklore texts. In recent years, clustering algorithms have undergone significant development. Their current state of advancement appears sufficient to support the basic requirements of introductory analyses in folklore studies. The study presented in this article was designed to yield unequivocal and interpretable results, and this objective was successfully achieved. While the findings suggest that the application of clustering algorithms to folklore research holds considerable promise, the issue of continuity between the resulting groups remains problematic. Addressing this challenge calls for further theoretical reflection and the development of methodological frameworks capable of accommodating the fluid and overlapping nature of folktale character classifications.
References
Atran, S. (2013). Ewolucyjny krajobraz religii. Zakład Wydawniczy Nomos.
Czeremski, M. (2011). Oswojenie bajki. Morfologia Władimira Proppa. W: W. Propp, Morfologia bajki magicznej (s. VII–XXXVII). Zakład Wydawniczy Nomos.
Czeremski, M. (2021). Mit w umyśle. Wydawnictwo UJ.
Ezugwu, A. E., Ikotun, A. M., Oyelade, O. O., Abualigah, L., Agushaka, J. O., Eke, Ch. I., Akinyelu, A. A. (2022). A comprehensive survey of clustering algorithms: State-of-the-art machine learning applications, taxonomy, challenges, and future research prospects. Engineering Applications of Artificial Intelligence, 110, 104743. https://doi.org/10.1016/j.engappai.2022.104743
Finlayson, M. A. (2017). ProppLearner: Deeply annotating a corpus of Russian folktales to enable the machine learning of a Russian formalist theory. Digital Scholarship in the Humanities, 2, 284–300.
Grundkiewicz R., Gralinski, F. (2011). How to distinguish a kidney theft from a death car? Experiments in clustering urban-legend texts. In P. Nakov, Z. Kozareva, K. Ganchev, J. Hobbs (eds), Proceedings of the RANLP 2011 Workshop on Information Extraction and Knowledge Acquisition (pp. 29–36). Association for Computational Linguistics.
Lévi-Strauss, C. (2001). Myśl nieoswojona (przeł. A. Zajączkowski). Wydawnictwo KR.
Lévi-Strauss, C. (2010). Surowe i gotowane (przeł. M. Falski). Wydawnictwo Aletheia.
Libera, Z. (2021). Etnografia jest praktykowaniem semiotyki (przykład z polskiej etnografii XIX wieku i później). Prace Etnograficzne, 3, 185–200.
Liszka, J. J. (1989). The Semiotic of Myth. A Critical Study of the Symbol. Indiana University Press.
Lu, M., Qin, Z., Cao, Y., Liu, Z., Wanga, M. (2014). Scalable news recommendation using multi-dimensional similarity and Jaccard–Kmeans clustering. Journal of Systems and Software, 95, 242–251.
Michalec, A., Niebrzegowska-Bartmińska, S. (red.) (2019). Jak chłop u diabła pieniądze pożyczał. Polska demonologia ludowa w przekazach ustnych (s. 21–85). Wydawnictwo Uniwersytetu Marii Curie-Skłodowskiej.
Rousseeuw, P. J. (1987). Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53–65.
Saxena, A., Prasad, M., Gupta, A., Bharill, N., Patel, O. P., Tiwari, A., Er, M. J., Weiping, D., Lin, Ch.-T. (2017). A review of clustering techniques and developments. Neurocomputing, 267, 664–681.
Shameem, M., Ferdous, R. (2009). An efficient k-means algorithm integrated with Jaccard distance measure for document clustering. Proceedings of the First Asian Himalayas International Conference on Internet, Kathmundu, Nepal. https://doi.org/10.1109/AHICI.2009.5340335
The R Foundation for Statistical Computing (dostęp: 2024, 28 maja). RDocumentation [dokumentacja oprogramowania]. https://www.rdocumentation.org
Wierzchoń, S., Kłopotek, M. (2017). Algorytmy analizy skupień. Wydawnictwo Naukowe PWN.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Patryk Gujda

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.
1. The authors give the publisher (Polish Ethnological Society) non-exclusive license to use the work in the following fields:a) recording of a Work / subject of a related copyright;
b) reproduction (multiplication) Work / subject of a related copyright in print and digital technique (ebook, audiobook);
c) marketing of units of reproduced Work / subject of a related copyright;
d) introduction of Work / object of related copyright to computer memory;
e) dissemination of the work in an electronic version in the formula of open access under the Creative Commons license (CC BY - ND 3.0).
2. The authors give the publisher the license free of charge.
3. The use of the work by publisher in the above mentioned aspects is not limited in time, quantitatively nor territorially.
Stats
Number of views and downloads: 6
Number of citations: 0