Network Analysis and Gephi

Network analysis is a digital humanities method used to systematically visualize scattered information about people, places, works, concepts, and relationships found in literary and historical texts.

Hasan Ergüleç

6 min read

Fourth Session of the Digital Humanities Workshop Completed

In the fourth online workshop organized by the Classical Divan Digital Humanities Platform, methodological transformations and innovative approaches in the field of Digital Humanities (DH) were discussed. In her presentation, Selin Yavuz shared the methodological dynamics of moving traditional literary text readings into a concrete, data-driven, and visual framework, focusing on network analysis and Gephi.

1. The Boundaries of Tradition in Literary Criticism and the Transition to a Concrete Data Ground

Traditional qualitative methods, long established in literary studies, bring about certain methodological limitations when faced with large datasets. In the opening section of the workshop, the direct impacts of these limitations on research outputs were emphasized. Supporting interpretation-based traditional transmission with empirical and concrete data is considered a cross-paradigmatic necessity. The limitations of human memory compared to digital memory become particularly apparent in large data structures containing hundreds of actors and complex networks of relationships. Microscopic relationships overlooked during long-term readings can cause the big picture at the macro level to be lost, ultimately resulting in studies that may remain incomplete or deficient. The digital humanities approach does not aim to eliminate subjectivity entirely; on the contrary, it seeks to substantialize invisible connections by grounding subjective judgments on a data-driven basis of legitimacy.

2. Why Gephi? The Anatomy and Building Blocks of a Literary Network

Although there are many alternative academic-scale tools in network analysis processes, such as Cytoscape, Tulip, and Pajek, Gephi is preferred due to its user-friendly interface and structural alignment with digital humanities studies. Selin Yavuz cited Gephi's open-source nature, high data processing capacity, advanced statistical analysis modules, and flexible algorithmic design capabilities as the primary reasons for this preference. To properly construct the anatomy and relational network map of a literary network, the counterparts of the system’s core conceptual metrics in the literature must be well understood:

Literary Counterpart and Functional Definition of Concepts

Node

The core actor of the network. Depending on the text's context, this could be a person, a specific book, a library, or even an abstract concept.

Edge

Any type of relational tie or link formed between nodes.

Weight

A quantitative value that shows the frequency, intensity, or overall structural strength of that specific connection.

Degree Centrality

The total number of direct connections a single actor has within the network.

Betweenness Centrality

The ability to control the flow of information and communication by serving as a crucial bridge between different groups.

Closeness Centrality

The potential of an actor to reach all other nodes in the network through the shortest and fastest paths.

Eigenvector Centrality

The status of an actor being connected not just to a high number of links, but specifically to the most powerful and central actors in the network.

Methodological Caution: Missing Data Bias

Network statistics and mathematical metrics must always be interpreted by considering the context of the literary text. For instance, a figure with a very high betweenness centrality value may not necessarily mean they are the most primary actor of the text in every scenario. The researcher needs to possess a deep awareness regarding data quality, metric interpretation, textual verification, ethical boundaries, and limitations. One should not fall into either an overly optimistic or an overly pessimistic technological reductionism; instead, a methodological balance must be established between the power of the software and the specific context of the text.

3. Ideal Source Criteria for Gephi and Unexpected Patterns

There is no obligation to model all digital-based studies in the social sciences and philology fields using the Gephi software. Texts that are small-scale or possess linear relational structures are not functional for this methodology. Selin Yavuz emphasized that for a dataset to constitute an ideal source for Gephi, it should meet the following criteria:

• Containing a large number of actors and intricate relationships at both the micro and macro levels,

• Being documented with historical records, archives, or intertextual data,

• Exhibiting structural changes over time (in a chronological process),

• Harboring hidden patterns and communities that cannot be discerned through traditional close reading.

Macro-analyses performed on large datasets enable the generation of alternative new hypotheses and questions against the well-established assumptions frequently repeated in primary sources.

4. Global Examples and Local Case Studies

Among the pioneering and prestigious global applications of network analysis are the Mapping the Republic of Letters project conducted at Stanford University, the studies of the Stanford Literary Lab, and macro-analytical investigations into Russian literature. These studies demonstrate how powerfully compelling scientific arguments can be generated when quantitative datasets and qualitative (contextual) analysis processes are harmonized.

A Local Case Study: The Sufi Lodges in 19th-Century Istanbul Project

One of the most comprehensive local case studies reflecting this global vision is Serpil Özcan’s master’s thesis, titled "19th-Century Istanbul Sufi Lodges and Their Spatial Positioning." The Yenikapı Mevlevihane, whose socio-cultural significance has long been recognized in traditional historiography but whose relational power had never been quantitatively measured, was mapped through Gephi network analysis. It emerged as the most critical, unifying, and strategic institutional actor across the entire network, boasting a betweenness centrality score of 0.52. This discovery serves as a prime example of an unexpected finding that reshapes the intuitive assumptions of traditional history writing.

5. Application Areas in Subfields of Turkish Language and Literature

The network analysis methodology has the potential to open new horizons in different subfields of Turkology research. The interdisciplinary application proposals presented in the workshop are as follows:

• Classical Turkish Literature: Relationships between poets and patrons in poet tezkireleri (biographical memoirs), biographical networks, co-occurrence maps of concepts and words in divans, stemma codicum (lineage charts) of manuscripts, and their circulation among libraries. (Example: The network analysis performed by Abdullah Karaarslan on the Nazire Compilation numbered TY9940, and Selin Yavuz's work visualizing the most acclaimed poems through the parallel poems (nazires) written for Sheykhi's poetry).

• Modern Turkish Literature: Author and intellectual circles clustering around periodical literary journals, intergenerational interaction maps, historical correspondence and communication networks, the propagation speed of literary movements, and the geographical analysis of translation and influence networks.

• Folk Literature and Folklore: Lineages of âşıks (folk poets) and master-apprentice contexts, circulation routes of oral culture products and their variants on a geographical plane, and actor-space relationships in ritual and performance contexts.

• Turkish Language and Linguistics: Vocabulary commonalities between historical dialects and contemporary Turkic literary languages, phonetic sound change networks in dialect maps, and intertextual vocabulary correlations.

6. Digital Workflow and the Anatomy of the Gephi Interface

The scholarly process extending from a literary or historical text to a digital network map necessitates a disciplined workflow. First, the researcher records the data obtained from the text in a structured format using Excel or a similar spreadsheet program. This data is then converted into the .CSV (Comma-Separated Values) format, which serves as an intermediate format. In this process, Gephi primarily requires two data files: 'Nodes' containing the actors, and 'Edges' containing the relationships. The Gephi software offers three core workspaces (interface stations) for data processing and aesthetic presentation:

  1. Overview: The central station where the network comes to life algorithmically.

  2. Data Laboratory: The management area where the background data is viewed in an Excel-like matrix structure, and where filtering, editing, and attribute-adding operations are performed.

  3. Preview: The design workshop where complex, raw network structures are transformed into aesthetic, publication-ready, and literary visualizations through typography, color palettes, and anti-aliasing.

Conclusion and General Evaluation

Looking specifically at what these studies contribute to the field, an evaluation was made highlighting that they primarily enable us to practice evidence-based criticism, present structural changes over time, and provide the researcher with opportunities to formulate new questions. Based on the questions raised during the session, it was concluded that when utilized correctly and in a balanced manner, this methodology will yield highly positive outcomes in future studies, especially as higher volumes of data are integrated into the system.