Text Analysis Methods

Text analysis is the process of examining and analyzing texts, whether in written or printed form, that have been digitized and made computer-readable, using various digital tools and different methods.

Hafize Tuğba Er

6/20/20266 min read

The third lecture of the Digital Humanities Workshop, organized by the Classical Divan Digital Humanities Platform to introduce the use of digital methods and tools in humanities research, was given by Ecemnur Topcu under the title "Text Analysis". The lecture, which introduced participants to the theoretical and practical applications of digital text analysis methods, consisted of three main sections: fundamental concepts of text analysis, text analysis processes, and digital tools used in text analysis.

The first part of the workshop focused on the definition of text analysis and its place within the digital humanities. It was stated that text analysis is the process of examining and analyzing texts, whether in written or printed form, that have been digitized and made computer-readable, using various digital tools and different methods. It was emphasized that this approach has gained a wider range of applications, particularly with the development of optical character recognition (OCR) and handwriting text recognition (HTR) technologies, and that their interconnected nature is crucial.

THE THIRD LESSON OF THE DIGITAL HUMANITIES WORKSHOP HAS BEEN COMPLETED: TEXT ANALYSIS METHODS

The concepts of "close reading" and "distant reading" two fundamental perspectives prominent in textual analysis studies, were discussed in detail. Close reading refers to the researcher focusing on a single text to deeply examine its origins, while distant reading involves the evaluation and analysis of numerous texts together. This allows for the revelation of cultural and historical trends on a much broader scale than the analysis of a single work.

The course covered the use of text analysis with various examples. It was noted that the concept of text is not limited solely to books and written works. It was also mentioned that social media posts, user comments, and other types of written data, which play a fundamental role in communication and information sharing today, can also be analyzed. In this context, Ecemnur Topcu stated that sentiment analysis applications – especially in the field of psychology – can identify dominant emotional tendencies in texts. Examples were given to illustrate how, particularly in literary texts like poetry, the pastoral, lyrical, or other qualities of a text can be identified through the words used and word choices employed.

One of the fundamental applications of text analysis, word frequency/corporate analysis, was also discussed. This method allows for the determination of the frequencies, percentage distributions, and preference rates of words used in a text. In connection with this, the relationship between text mining and text analysis was explained. It was stated that text mining is a technical research process involving the discovery of meaningful data within a text, while text analysis is the stage of interpreting this data and transforming it into research findings.

The workshop also included studies on discourse analysis. It was emphasized that language is a living and changing phenomenon. Based on this, it was stated that examining the changes in meaning that the same word undergoes in different periods and cultural contexts is one of the fundamental areas of study in discourse analysis. In this way, the evolution of a word over the years and its semantic context in different societies can be examined.

The contributions of period analysis, another application of textual analysis, to historical and cultural research were also discussed. For example, it was stated that social changes can be seen and their reflection in social life can be comprehensively understood through the analysis of texts from the 13th, 16th, and 19th centuries. It was emphasized that the traces of different historical periods reflected in texts, such as the 13th century, when the negative effects of the Mongol invasions were felt, the 16th century, the most brilliant period of Ottoman classical (divan) literature, and the 19th century, when social life was undergoing a major transformation, can be made visible using digital methods.

The course included examples of how text analysis can be used in conjunction with other digital humanities methods through collaborative analysis techniques. In this context, network analysis, geographic information systems, and spatial analysis applications were discussed. For example, it was stated that network analysis could be performed by mapping the relationships between characters in a mesnevi (a type of long narrative poem). Similarly, it was mentioned that spatial analysis could be carried out by displaying locations mentioned in historical or literary works on digital maps. Thus, it was explained that multiple analyses could be performed from a single dataset.

One of the workshop's notable topics was the discovery of anonymous works. It was stated that the author of anonymous works can be determined through stylistic analysis. Thanks to stylistic studies, after learning the text, the machine retains the overall meaning of a word, both before and after its use, allowing for the identification of potential authors of anonymous texts based on these characteristics. In this context, the example of Ümmî Sinan, who lived in two different periods, was examined. It was stated that it might be possible to re-separate poems that have become mixed up due to interventions made by scribes over time, using digital stylistic analysis methods.

Before moving on to the practical examples section of the course, the data preparation processes in text analysis studies were discussed. It was mentioned that the first step involves making the data processable using optical character recognition (OCR) and handwriting text recognition (HTR) methods.

Following the data collection process, pre-processing/cleaning of the texts is necessary before analysis. In this stage, it is important to adapt the tools used to the Turkish language. Particular attention was drawn to the importance of removing and organizing compound words, phrases, and letter distinctions (such as 'ı, u, ü') to prevent the computer from interpreting them as separate words, in a manner consistent with the research objective.

Word clouds created on cleaned texts are a method that visually reveals the frequency of word usage within the text. In this method, words with a high usage percentage are shown in larger and bolder fonts, while less frequently used words are shown in smaller sizes. This method is said to make it easier to recognize the text visually.

The tagging method was also introduced. It was explained that concepts such as names of people, institutions, and place names are tagged by the computer, their meanings are taught, and with this tagging, the computer can recognize elements within the text and produce structured data outputs. It was stated that this method provides significant advantages, especially in network analysis and database creation studies.

The course also focused on the ngram method, and an example project was examined. It was explained that this method, based on word filtering, allows for the analysis of single words or concepts consisting of multiple words. It was mentioned that ngram analyses offer the possibility of classifying and examining data according to centuries, authors, and types of works. One of the important advantages of the method is the ability to determine when the relevant concept was first used. It was stated that, thanks to these features, the ngram method will make significant contributions to conceptual studies and the revision of known historical information.

The final section of the workshop focused on providing information about text analysis tools. Ecemnur Tocu introduced the software she uses in her digital and humanities studies. She discussed the possibilities offered by Kate for tagging, and R and Python for graphic and data visualization. The basic features of Voyant Tools, AntConc, and Recogito were explained for text mining, word frequency analysis, and analysis. The contributions of these tools to data analysis, visualization, and interpretation were evaluated with examples.

The course also touched upon current digital humanities projects being conducted in Turkey. In particular, the Digital Ottoman Projects, led by Ayşe Tarhan, were mentioned. It was noted that these studies largely focus on compiling collections of divans (collections of poems) and conducting spatial analysis.

In addition, the internationally conducted KİTAB-OpenITI (Open Islamicate Texts Initiative) project was examined as an example. It was stated that this database, created by digitizing Islamic Arabic texts (fiqh, tafsir, kalam, hadith), offers the possibility of conducting comparative analyses. Furthermore, it was mentioned that the KİTAB project should be developed and adapted based on Ottoman literary texts.

The workshop, which concluded with participants' questions being answered and their contributions being incorporated, revealed that digital text analysis is not merely a technical process but also a powerful research method that offers new perspectives to studies in literature, history, culture, and language. The theoretical information and application examples covered in the workshop provide a comprehensive introduction for those wishing to work in the field of digital humanities – text analysis.