Interdisciplinary Research Methods

Digital Humanities: Text, Data, and South Asian Sources

Digital tools and methods for literary and historical research on South Asian material, from text encoding to computational analysis, and their limits.

Digital humanities applies computational tools—text encoding, database construction, network analysis, geographic information systems, and increasingly machine learning—to humanities research questions, offering South Asianists new ways to work with large corpora, fragmentary manuscript traditions, and multilingual textual variation that manual methods handle with difficulty. Projects encoding classical Sanskrit, Persian, and Urdu manuscripts in standardised digital formats like TEI (Text Encoding Initiative) XML allow scholars to track textual variants across manuscript copies, a task central to establishing critical editions of premodern Indian texts.

Corpus-based approaches allow quantitative analysis of large bodies of text—for instance, tracking the frequency and collocation of specific terms across a poet's entire corpus, or comparing formal features (metre, vocabulary, syntax) across a genre's historical development—that would be prohibitively time-consuming through manual reading alone, though scholars must remain alert to how digitisation choices (which editions are digitised, how texts are segmented) shape what patterns become visible.

Digital mapping and GIS (Geographic Information Systems) tools have been used to visualise the geography of literary production, publishing networks, or historical events such as Partition displacement, migration patterns, or the spread of specific literary or religious movements, offering spatial perspectives that complement textual and archival analysis. Network analysis tools can similarly map relationships between historical figures, correspondence networks, or literary influence, revealing patterns of intellectual community that might not be evident from reading individual texts in isolation.

South Asian digital humanities faces distinctive infrastructural challenges: many historical and literary texts exist only in non-Roman scripts (Devanagari, Perso-Arabic, Bengali, Tamil, and others) for which OCR (optical character recognition) technology remains less mature than for Roman-script texts, complicating large-scale digitisation and computational analysis; multilingual and multiscript source material also raises technical challenges for search, tagging, and cross-referencing that Euro-American digital humanities tools were not originally designed to handle.

Scholars must also weigh access and infrastructure inequities: digital humanities projects require technical training, computing resources, and often institutional funding unevenly distributed across Indian universities, and should be assessed critically for whether they genuinely democratise access to South Asian textual heritage or primarily serve well-resourced international institutions; open-access data-sharing commitments and multilingual interface design are increasingly recognised as necessary ethical components of responsible digital humanities practice in the Indian context.

The lesson at a glance

Digital Humanities: Text,…TEI XML encodingCorpus-based analysisGIS in humanitiesOCR limitations for Sou…Digital access equity
Concept map — the lesson question at the centre, the ideas you need to hold around it.

Key concepts

TEI XML encoding
A standardised markup format (Text Encoding Initiative) used to digitally represent and compare manuscript texts and their variants.
Corpus-based analysis
Quantitative study of patterns (word frequency, collocation, stylistic features) across a large digitised body of text.
GIS in humanities
The use of Geographic Information Systems to visualise spatial dimensions of literary production, migration, or historical events.
OCR limitations for South Asian scripts
The relative immaturity of optical character recognition technology for Devanagari, Perso-Arabic, and other non-Roman South Asian scripts.
Digital access equity
The ethical concern that digital humanities infrastructure and funding are unevenly distributed, raising questions about who benefits from digitisation projects.

Thinkers to know

  • Franco MorettiLiterary scholar whose 'distant reading' concept, applying quantitative methods to large literary corpora, has influenced South Asian digital humanities debates.
  • Ganesh DevyLinguist and literary scholar whose People's Linguistic Survey of India exemplifies large-scale, technology-assisted documentation of India's language diversity.

In the Indian context

  • The People's Linguistic Survey of India, led by Ganesh Devy, documented over 780 languages using extensive field and digital methodology, a landmark large-scale humanities data project.
  • Institutions like the Indira Gandhi National Centre for the Arts and various IIT digital humanities labs have undertaken manuscript digitisation projects for Sanskrit, Persian, and regional-language texts.
  • OCR development for Indic scripts remains an active and unevenly resourced research area, directly affecting the pace of large-scale South Asian text digitisation.

Glossary

TEI

Text Encoding Initiative, a standardised XML-based markup format for digitally representing texts and manuscript variants.

OCR

Optical Character Recognition, technology for converting scanned images of text into machine-readable text.

Distant reading

Franco Moretti's term for computational, pattern-based analysis of large literary corpora as opposed to close reading of individual texts.

Corpus

A large, structured collection of texts assembled for linguistic or literary computational analysis.

Sources to read

Practice — turn this into an article

  1. Select a short passage of a classical Indian text available in more than one manuscript edition or translation and encode its basic variants using simplified TEI-style tagging.

    Deliverable: A sample TEI-tagged text excerpt with a brief explanatory note.

  2. Design a small corpus-analysis question (e.g. frequency of a specific image or word across a poet's collected works) and outline how you would test it computationally.

    Deliverable: A one-page research design note for a digital humanities mini-project.

  3. Evaluate one existing South Asian digital humanities project (manuscript archive, linguistic survey, or literary database) for its accessibility, multilingual support, and open-access commitments.

    Deliverable: A critical evaluation note for a research article's methodology or literature review section.

Self-check

  • What does TEI XML encoding allow scholars to do with manuscript variants that manual comparison makes difficult?
  • Why does OCR technology pose a particular challenge for South Asian digital humanities?
  • How can GIS and network analysis tools complement traditional textual and archival research?
  • What equity concerns should scholars consider when evaluating digital humanities projects in the Indian context?