Introduction
How do we confront the ongoing epistemological violence of colonialism upon the cultural record? Periods of European and American imperialism from the fifteenth to twentieth century transformed the cultures, geographies, landscapes, and societies of much of the world through military violence and economic exploitation. Much of the historical record and institutions of recordkeeping emerged through processes of imperial extraction by European administrators, merchants, and travelers invested in categorizing Indigenous societies through western frameworks of racialization and cultural heritage. Epistemological and structural inequality continue through the global north’s domination of knowledge production, dissemination, and evaluation, ultimately steering the development of information infrastructure and digital humanities scholarship (del Rio Riande; Liu et al.). Recent digital humanities scholarship and pedagogy call to reckon with the colonial record of humanity, the predominance of colonial forms of knowledge representations, and the absence of Indigenous knowledge and its impact upon the digital cultural record (Risam; Boulay et al; Fiormonte).
This article seeks to confront this cultural record of colonialism that perpetuates an epistemological divide between the west and non-west, modern and traditional, global north and global south. I draw from an extended digital humanities case study of a French colonial visual encyclopedia of Vietnamese crafts, cultural practices, and technologies. Produced by the French colonial administrator Henri J. Oger and unnamed Vietnamese contributors (draftspeople, annotators, informants, and woodblock printers) in 1909–1910, the book Technique du peuple annamite [Kỹ thuật của người annam, Mechanics and Crafts of the Vietnamese People] (henceforth Technique du Peuple Annamite), comprises 4,356 drawings and 7,332 captions in French and Vietnamese.1 This encyclopedia purports to be a distanced observation and totalizing representation of Vietnamese material culture and society with underlying economic aims. To counter its exoticizing hegemonic form, my approach shifts toward an intimate reading of vernacular life, plural narratives, and relational interpretation of authorship and communication (Stoler). I draw from decolonial and feminist theory, Indigenous and folklore studies to argue that digital humanities and computational humanities can make visible the layers of colonial representation and surface plural voices and multidirectional narratives (Smith; Tuck and Yang). This commitment seeks to challenge top down, linear, static, hegemonic colonial representations that often materialize in static forms of data representations, the privileging of Western frameworks of authorship, and a scientific push toward empirical findings. I seek to uncover a vernacular “something else” knowledge framework that does not just refute colonial ones through a resuscitation of ‘Indigenous’ voices through binary constructions of exogenous-Indigenous. Visual artist and scholar Trinh T. Minh Ha proposes a “speaking nearby” rather than “speaking for,” an intellectual space of critical analysis that is multivocal, described as a “non-identifiable ground where boundaries are always undone, at the same time as they are accordingly assumed (Chen 85).” In this way, I interrogate coloniality not as universal concepts that predetermine all aspects of cultural representation and meaning, from authorship and production to narrative expression and historical record.
Interpretive Digital Humanities in Three Acts
This article functions as an intellectual troubling of knowledge production and interpretation, calling for critical digital humanities that weave cultural critique and computation as processes of meaning making. In honoring Saidiya Hartman’s “Venus in Two Acts,” I present three acts of interpretive digital humanities where orality, visual communication, and cultural practices take center stage. In this text, Hartman confronts the violence of the archive of Atlantic slavery, navigating the impossibilities of what can be known and generating ways of storytelling eloquently described as “critical fabulation” (11). In engaging with a primary source that was produced within conditions of violence, inequality, and exoticization within the French colonial context in Vietnam, this article centers power relations within the production, circulation, and continued analysis of this “problematic” data set. In addition, I recognize absence and unknowability within this primary source and build on the important work of absence as “power, presence, and productive” (Sherman et al.).
Furthermore, I argue that interpretive digital humanities undermine the teleological pursuit of argument-driven results, carried from disciplinary global north structures of knowledge production. In this way, I showcase the pitfalls of the “explainability” paradigm that digital humanities and computational research encounter, and I center “unknowability” and “multiplicity” as research design principles in global south digital humanities. Explainability is the privileging of research argumentation that explains instead of explores, which is further intensified by scientific and technical paper linear structures of Introduction, Literature Review, Methodology, Results, Discussion, and Conclusion. I advance that rather than thinking about argumentation, we can embrace the power of interpretation—of narrative, aesthetics, language, computational results—and transparently communicate the scope and scale of analysis fully recognizing the limitations and possibilities. This builds upon the important scholarship on interpretative “ways of knowing,” a computational hermeneutics, or what Hannah Ringler describes as “asking questions with and about computing technologies” and which Geoffrey Rockwell and Stéfan Sinclair assert that “analytical tools are instantiations of interpretive methods” (4). For example, how might we rethink data visualization and its projections into two dimensional representations (determined by varying decisions with training models and parameters) within the realm of humanistic interpretation? How could computational results and models be understood from a semiotic perspective of meaning making through probabilistic associations and measurements of similarity (Picca)? How might the structural absences in a text and unknowability be explored through imagination and speculation (Koeser and LeBlanc)? How might our analysis shift toward multivocality and multiplicity when we shift away from western frameworks of provenance, authorship attribution, and communication?
I demonstrate how exploratory research analysis and question generation can emerge through different levels of interpretation of meaning, that is, discrete close reading and clustering of embeddings (numerical representations of words and images in multi-dimensional space). Each of the following acts brings to the foreground a specific level of interpretation: Act 1 Orality: Storytelling and Literature (Close Reading); Act 2 Visual Communication: Exploring Gesture and Style (Nonlinear Clustering through Visual Embeddings); Act 3 Cultural Practice: Social Worlds of Superstition and Folk Knowledge (Relational Clustering Through Word Embeddings). In these acts and methodological explorations, I never fully produce an empirical conclusive claim but instead emphatically embrace an alternative to a scientific, results-driven justification for research. Instead, I showcase how different clustering tools and data visualizations could function as tools of “exploratory computation,” inviting space for question generation, curiosity, and imagination, which move toward a critical understanding of history and cultural practice in all of its complexity (Nguyen and Alvarado Rojas). This is a work in progress and exists as a conversation through communicating limitations and future dreams. In communicating this messiness, I advocate for publishing and teaching with evolving digital humanities work in order to expand access to and invite critical engagement with the cultural record of the “global south.” This methodological framework shifts from a format that applies computational methods to humanities questions and instead invites a fundamental engagement with the processes of knowledge production, from the political and historical conditions of primary source data to interpretive meaning making from computational results.
A Multivocal Dataset: Multi-staged Processes of Knowledge Production
The production of the original woodblock printed book between 1908 and 1909 in Hanoi (Figures 1 and 2) occurred in several stages as outlined from the introductory essay by Oger. The editors of the 2009 reedition Olivier Tessier and Philippe le Failler, as well as scholarship by Nguyễn Mạnh Hùng, Nguyễn Quảng Minh, Nguyễn Mộng Hưng, and H. Van Puten, provide additional historical context on Oger and the material qualities of Technique du Peuple Annamite. Through my close material analysis and computational work with the visual and textual description, I add the following analyses to highlight the work’s material and epistemological production process as collective and relational. In this way I advance an alternative framework against colonial authorship and binaries of exogenous-Indigenous.
Figure 1: In 1909 the woodblock-printed book, Technique du peuple annamite—Kỹ thuật của người annam—Mechanics and Crafts of the Vietnamese People, was printed in Hanoi, and page 2 is pictured here.2
Figure 2: In 1910, an accompanying 160-page Volume of Texts was published in Paris with an introductory essay authored by French colonial civil service administrator Henri J. Oger (1885–c. 1936). The above photograph is the volume of text (top) and two volumes of drawings bound for library circulation (University of California Berkeley Library; author’s photograph).
Step 1: The Draftsperson and the Production of Drawings
Oger states that these drawings were produced via a draftsperson, who observed and conversed with artisans, particularly focused on tools and use of objects. My analysis uncovers firstly that the drawings were most likely created by multiple individuals and also represent a much more expansive cultural, social, and spiritual world including popular imagery, tales, and folk knowledge. This visual body of knowledge points to visual representation that derived not just from distanced observation, but from hearsay, relationships, and oral storytelling.
Step 2: Captioners and Production of Hán Nôm Character Captions
Oger explains that the drawings were shown to a group of local inhabitants (indigènes) as “a method of verification” and to explore new meanings.3 I interpret this stage as where the production of hán nôm captions, a character-based Sinographic script of spoken Vietnamese (chữ nôm) and Sinitic (chữ hán). I interpret that the character captions emerged as a dialog and debate among several individuals.
Step 3: Production of French Captions with Local Informants
Oger describes that drawings were described together with him and unnamed informants (possibly the captioners, the draftsperson, other contributors), with broad and technical reflections, leading to Oger’s production of French language captions. It is important to note that it is uncertain, one, if Oger is consistently present at steps one through three and two, if his French captions were based primarily on the drawing and hán nôm caption or through conversation and translation with informants.
Step 4: Two-Part Printing in Hanoi and Paris
At least thirty Vietnamese woodblock carvers, some names of which have been uncovered, contributed to producing the woodblocks.4 In a rushed and unplanned manner, the 4,356 images (some of them annotated with hán nôm captions) were assembled and then printed by woodblock onto the 700 pages comprising Technique du peuple annamite, in the summer of 1909 in Vũ Thạch pagoda in Hanoi on rue de Chanvre. In 1910, the French captions were printed in a separate 160-page volume, Volume of Texts, that included a French language introduction to the project and an essay titled “Quelques vues d’ensemble sur les industries indigènes du pays d’annam: Un nouveau programme d’enseignement pour les annamites [Some overall insights into the indigenous industries of the country of Annam: A new curriculum for the annamites]” with the author byline of Henri Oger.
My close reading analyzes the unpublished manuscript edition in Keio University in Tokyo, Japan, and the woodblock-printed edition of Technique du Peuple Annamite held at University of California, Berkeley. My computational analysis of images and captions builds from a 2009 digitized re-edition of Technique du Peuple Annamite that digitized the edition held at Thư viện Khoa học Tổng hợp [General Sciences Library] in Ho Chi Minh City, Vietnam.5 In this digitized edition, the hán nôm character captions were transliterated into quốc ngữ: a writing system popularized in the twentieth century and still used today, of vernacular Vietnamese. I rely upon the quốc ngữ transliterations of the character captions for close reading, text analysis, and word embeddings given the existing language models in contemporary Vietnamese script. The 2009 re-edition also translated the French captions into quốc ngữ and English to increase contemporary international comprehension of Technique du Peuple Annamite. In my accounting and analysis, there are 4,356 individualized drawings, 2,904 Vietnamese hán nôm captions, and 4,428 French captions, of which I analyze as separate parallel components of a relational collection of 11,688 elements. It is important to note that not all drawings have both French and hán nôm captions, approximately 66 percent of the collection, or 2,878 drawings, have both sets of captions.
In previous published work on this primary source Technique du Peuple Annamite, I used mixed methodologies of book history, content analysis, and exploratory visualizations to examine gendered labor and childbearing as a networked social world at the turn of the twentieth century in Vietnam (C. Nguyen). Questioning further its interpretation and positionality, I center a feminist crisscrossing of production, circulation, and meaning making through intimacy (de Langis et al.). And thus in this piece examine how and where digital humanities and computation can interrogate the production of a historical primary source as well as the knowledge it conveys. This work moves beyond the application of computation to a global south humanities case study and towards understanding how a critical digital humanities methodology can extend interpretation, intimacy, and imagination.
Act 1 Orality: Storytelling and Literature
Close Reading of Strange Patriotic Children
Four distinct images are presented in line drawings (Figure 3): a young boy with a mouse on a leash, a group of individuals with flags facing a soldier, a woman holding a basket crouched over a stove, and a person pulling out a tooth with string. This is page 332 of the woodblock-printed 700-page volume of Technique du Peuple Annamite. The accompanying French language essay in Volume of Texts, does not mention or name the contributors who sketched the drawings, wrote the accompanying hán nôm captions, or carved and printed the 700 pages. The Volume of Texts includes descriptive captions in French, many of which differed in narrative style and content from the vernacular Vietnamese language character captions that appear next to the drawings. The French captions describe the four images as follows: “Élevage des souris [Raising mice].” “Kì Đồng [Kỳ Đồng].” “Industrie du sucre [Sugar industry].” “Arrachage des dents [Extracting a tooth].” In the French Volume of Texts, Oger includes an introductory essay that describes the project as a totalizing encyclopedia of Vietnamese crafts, technologies, and cultural practices. Oger’s essay also provides a vague rationale and research methodology for the production of Technique du Peuple Annamite, centering his intellectual stewardship and scholarly vision at the forefront of the project as a scholarly administrator.
While most of the French and hán nôm captions were relatively short, a noticeable discrepancy appears above the second image from the left (Figure 4). The French caption just lists a name, Kỳ Đồng, yet the hán nôm caption appears visibly longer and descriptive. It states the following: “Kỳ Đồng was such a troublemaker and led some organized activity with a handful of young disciples from poor families, carrying paper flags and banging drums. They made a ruckus in the streets raising flags announcing an armed uprising and were planning to attack the Nam Định citadel. For this reason, he was arrested by the provincial official and put into prison.”
Figure 4: Page 332 features Kỳ Đồng, who leads an uprising. The Vietnamese caption reads: “Kỳ Đồng rất nghịch ngợm, bày trò cùng với năm sáu đứa đệ tử con nhà nghèo vác cờ bằng giấy nghênh ngang trong phố, đánh trống hò hét hô lên là cờ khởi nghĩa của Kỳ Đồng, định đánh hạ thành Nam Định. Vì thế bị quan tỉnh bắt về tống ngục. [Kỳ Đồng was such a troublemaker and led some organized activity with a handful of young disciples from poor families, carrying paper flags and banging drums. They made a ruckus in the streets raising flags announcing an armed uprising and were planning to attack the Nam Định citadel. For this reason, he was arrested by the provincial official and put into prison.]”
In my close reading of this text, I uncovered an extensive hagiography of Kỳ Đồng. I followed stories of his early childhood, recognized for his spectacular talents of quick-wittedness and outsmarting his Confucian scholarly father and the provincial mandarins with his ability to defeat elders in a battle of Confucian proverbial wits (Figure 5). I also trace to the young Kỳ Đồng who led an uprising and was sentenced to prison. According to drawings and captions that I found in Technique, Kỳ Đồng continued to carry out political organizing while imprisoned and then led another act of treason in Nam Định (province outside the capital of Hanoi). Officials attempted to suppress Kỳ Đồng after the uprising in Nam Định by firing squad, but a mysterious cloud aided his escape (Figure 6). The story continues that officials were forced to bury him alive (Figure 7), but he was later discovered coming back to life, walking around with one shoe while reciting wise sayings (Figure 8).
Figure 5: Page 297 features Kỳ Đồng as a child genius. The Vietnamese caption reads: “Ở xã Ngọc Đình tỉnh Thái Bình có nhà nho nghèo, đã hai ba lần đi dự thi, sinh ra được một cậu con trai bẩm tính thông minh sáng láng, nhanh mồm miệng. Khi mới lên 6 tuổi, đang ngồi nghe cha đọc sách cậu bé bỗng hỏi cha về nghĩa lý, khiến cha không thể trả lời được. Người cha đem chuyện trình lên trên. Quan huyện và quan tỉnh đến hỏi chuyện, cậu bé ứng đáp trôi chảy không sai sót, bèn gọi cậu là Kỳ Đồng. [In the commune of Ngọc Đình, Ngọc Đình Province, there lived a poor scholar who, had already sat for the imperial examinations two or three times and fathered a son who was naturally endowed with brilliant intelligence and a quick wit. Only just six years old, the boy while sitting and listening to his father read suddenly posed a question to his father regarding the meaning of the text, a question that his father could not answer. The father reported this extraordinary occurrence to his superiors. When the District and Provincial Magistrates arrived to question the boy, he responded with such fluency and precision without a single error, that they bestowed upon him the title ‘Kỳ Đồng’ (child prodigy).]”
Figure 6: Page 398 features the attempt to assassinate Kỳ Đồng by firing squad and his escape. The Vietnamese caption reads, “Kỳ Đồng bị tống ngục vẫn ăn nói bừa bãi, ngạo mạn. Quan tỉnh nghị án đưa ra góc thành xử bắn. Khi súng nổ, y có phép thuật khiến trời trở nên tối tăm. Đến lúc sáng rõ nhìn ra thì không thấy y đâu nữa, vì vậy không giết được. [Even after being thrown into prison, Kỳ Đồng continued to speak recklessly and arrogantly. The provincial officials deliberated his case and sentenced him to execution by firing squad at the corner of the citadel wall. When the shots rang out, suddenly magical powers plunged the sky into darkness. When the light finally returned and visibility was restored, he was nowhere to be found; thus, he could not be killed.]”
Figure 7: Kỳ Đồng thus had to be buried alive on page 455. The Vietnamese caption reads: “Kỳ Đồng phạm tội đại nghịch, quan tỉnh Nam Định cùng với quan Công sứ nghị xử đem chôn sống, nhưng hắn lại sống trở lại. [Kỳ Đồng committed the crime of high treason; the provincial officials of Nam Định in consultation with the Resident sentenced him to be buried alive yet he again came back to life.”
Figure 8: Page 297 features Kỳ Đồng after his resurrection from the dead, when he is found walking around with one shoe. The Vietnamese caption reads: “Kỳ Đồng phạm tội đại nghịch, đã bị chôn sống. Sau sự việc đó, một đêm trời vừa tờ mờ sáng, chợt thấy có một cậu học trò đi vào công đường. Quan tỉnh hỏi trò từ đâu tới, đáp rằng tôi vừa dưới mộ lên. Quan tỉnh thấy cậu chân đi có một chiếc giày, bèn ra vế câu đối: ‘đầu che bốn lọng.’ Cậu ứng khẩu đáp ngay: ‘chân đi một giày. [Kỳ Đồng committed an act of high treason and was buried alive. After this event, just as dawn was breaking one morning, a young scholar was seen entering the provincial magistrate’s court. The provincial magistrate asked the scholar where he had come from and he replied, ‘I have just ascended from the grave.’ Seeing that the scholar only wore a single shoe the provincial magistrate issued a couplet. “Head sheltered by four parasols.” And the young scholar responded instantly, ‘Foot clad in a single shoe].’”
Kỳ Đồng was a real historical figure born Nguyễn Văn Cẩm (1875–1929), who was known as a child prodigy skilled in poetry, letters, and arts, earning the nickname “Kỳ Đồng” (strange or exceptional child). At age 12, Kỳ Đồng famously was the symbolic leader of a rebellion against French occupation in Năm Đinh province in the late nineteenth century. Given his widespread influence in organizing social movements against colonial officials, French authorities sent Kỳ Đồng into exile twice. First, he was symbolically exiled to Algeria with a government sponsored scholarship for nine years (1887–1896) after his role in the Năm Đinh rebellion. Returning to Vietnam in 1896, he was involved in several insurgency movements, imprisoned in Saigon in 1897, and then exiled to French Polynesia in 1898 where he remained for the rest of his life. Stories of Kỳ Đồng’s intellect, wit, and organizational leadership were passed down among scholar patriots during a time of anti-colonial movements in the northern provinces in the late nineteenth century. Today he is memorialized as an important anti-French patriot in Vietnamese history (Huyền).
The reason why I conduct a slow close reading of representations of Kỳ Đồng within the corpus is to showcase the dynamic range of political, cultural, historic meanings this “colonial text” carries. I do not claim that this is overtly an anti-colonial text. I seek to disrupt the binaries of pro-colonial or anti-colonial because these hegemonic constructs shroud the complexity of this text. Each drawing and caption are an expression and function as a representative assemblage of history and culture—from Vietnamese cultural practices and craft industries to hagiographies and folklore with underlying political, medical, or moralizing lessons. This archive of meanings is constructed through multiple contributors and different stages. Trinh T. Minh Ha advances “multivocality,” the decentering and diversifying processes of voicing through an idiom of “speaking nearby” a subject rather than as prescribing it ethnographic objectification. I take a multivocal analysis of this text, by examining the range of actors in the labor of textual production and showcase a range of voices, narrative structures, and expressions. Furthermore, by focusing on vernacular expressions represented within this text, my analysis seeks to move toward what Walter Mignolo calls a “pluriverse” in order to undermine universal claims of knowledge embedded within European colonialism that emphasize the myth of single authorship and a linear narrative (2018).
The narrative elements of this encyclopedic collection, such as the hagiography of Kỳ Đồng, move between mythology and history, pointing to fluid communication circuits of communal storytelling and following literary and mythological formats with moral underpinnings of scholar heroes. Other forms of narrative style also emerge in the collection, such as visual and textual representations of proverbs and well-known spirits, deities, calendar or zodiac figures, and notable literary, historic figures worthy of emulation. For example, the collection includes several representations of the Tale of Kieu [Truyện Kiều] (Figure 9), the most famous work of Vietnamese national literature, an epic poem written by Nguyễn Du (1765–1820) in chữ nôm in a rhythmic and sonic lục bát (six-eight) meter. From the nineteenth century to present day, the cast of characters from this popular work of literature has risen to become literary, aesthetic, and national symbols of heroism, justice, beauty, honor, and love, as well as moral guidance for fortune and tragedy in everyday life.
Figure 9: Here is an example of three distinct representations of folk prints on page 102. In this example there appears to be intertextual representation, where the left folk print depicts the literary work Phan Trần (number 32) and the right is Truyện Kiều (number 29), two works that are often read in conversation. The center folk print presents four Vietnamese idioms and proverbs.
Act 2 Visual Communication: Exploring Gesture and Style
Nonlinear Clustering Through Visual Embeddings and Aesthetic Multivocality
While the previous section focused on multivocality in production and narrative, this act focuses on visual communication as an aesthetic multivocality: that is, a visual analysis opening toward different registers of interpretation, away from a text-based analysis fixed on linearity and authorship. Focusing on the clusters of images, their content, and their stylistic patterns can connect a singular image in relationship to the broader dataset of visual and textual elements. Clustering thus invites in nonlinear ways of “reading” and engaging with a collection that privileges stylistic and content specific patterns within visual information. Prevalent within art historical work, a visual analysis can help to foreground style and meaning. Existing visual and historical work on popular art by the prolific scholar of Vietnamese literature, linguistics, and hán nôm, Maurice Durand (1914–1966), examines art as a conduit toward understanding beliefs, culture, expression, and worldviews. Durand’s approach reflects historical scholarship during a transitional time within colonial and postcolonial Vietnam, focused on typologies that structure themes in Vietnamese visual expression, such as the banalities of everyday life and rhythms of nature, religions and beliefs, and literature (Durand, 1960; 2011). Alternatively, computational approaches drawing from unsupervised machine learning can move away from hierarchical typologies focused on expert driven evaluation of content and style and open up alternative non-linear visual analysis based on algorithmically detected aesthetic patterns and features.
Visual Clustering Methodology
For this project we, first, undertook a process of extracting features of visual and textual data embeddings and second, conducting exploratory visualization. Visual embeddings and word embeddings are a type of vector space model that represents images and text as dense numerical vectors in high-dimensional space, which can then be projected to examine semantic relationships as geometric distances of similarity. For the first attempt of visual clustering work, I used PixPlot, a visualization tool created by Yale Digital Humanities Lab developer Douglas Duhaime in 2019.6 PixPlot provides an out-of-the-box interactive web-based (WebGL scene) visualization of the drawings from our dataset as clusters, where each image was processed with the Inception model convolutional neural network (CNN, trained on ImageNet 2012) to embed each image into a high-dimensional space (Szegedy et al.).7 Then the visual embeddings were projected into a two-dimensional manifold with the Uniform Manifold Approximation and Projection (UMAP) algorithm for dimensionality reduction such that similar images appear proximate to one another.8 Pixplot uses TensorFlow’s Inception bindings for image analysis, and the visualization layer uses a custom WebGL viewer to provide a navigable browser based interface (Figure 10). While visual embeddings are a prevalent approach for complex data, it is important to recognize the legacies and biases of the original training data of ImageNet which classified race and gender from contemporary photographs mainly from the global north (Crawford; Crawford and Paglen). This carefully considered interrogation into algorithmic models and context of training data is of crucial importance in global south scholarship where classification is drawn from global north photographs, which are not attuned to the historical and material structures of this specific early twentieth-century dataset. A next step for continued critical computation will explore Alex Hock’s PixPlotML built on the original PixPlot that allows for training a classification machine learning model on our historical data using Pytorch-Accelerated.9 This type of machine learning trained on this Vietnamese woodblock dataset could lead to understanding visual relationships between this corpus and a larger corpus of Vietnamese, Sinitic, and East Asian woodblock-printed art could be a possible future direction for this work, taking inspiration from examinations of Ukiyo-E Japanese art form (Shinichi and Matsui) as well as ongoing work on the “Digital Florentine Codex” that examines visual and textual knowledge of sixteenth-century indigenous Mexico (Sahagún et al.). Another updated option would be to consider a transfer learning approach with a multimodal vision and language models such as Contrastive Language-Image Pre-training (CLIP), where we can explore our image dataset through training on our own carefully labelled and curated data (Radford et al.; Smits and Wevers). The multimodal vision and language model would need to consider both the French and Vietnamese language captions as part of the natural language used to reference learned visual concepts in the images, where we redefine “ground truth” to be multilingual and multivocal.
Besides tracing the design intentions and biases of algorithmic models, an intentional transparency on computational methodologies within specific research task is important. This act traces a close visual analysis of computationally produced clusters through PixPlot. This workflow explores how an interface of visual clusters might transform the process of research inquiry at an initial level, where humanistic interpretations of the generated clusters are not definitive for a classification purpose, but to explore visual similarities and raise new questions. From the corpus-level UMAP dimensionality reduction clustering and projection, I can zoom into smaller groupings where images are clustered together by proximity, examine visual characteristics in comparison, as well as zoom in on an individual image for close analysis. The following are some examples of insightful clusters (Figures 11, 12, 13).
By default, PixPlot proposes the top ten most similar clusters based on embeddings generated from the Inception model and through UMAP dimensionality reduction (Table 1). With these suggested ten clusters that are excerpted above, I can quickly explore at the corpus level a visual breakdown of subject matter such as human subjects, complex scenes such as nature and folk prints (Figure 14), and simple diagrams (Figures 15 and 16). Of interest is that clusters 1 and 2 form a large subset of the visual collection, which is focused on human subjects while the clusters 7 and 10 are simpler diagrams of objects, tools, and shapes. Through this visualization, I am able to visually explore patterns in such as the representation of a scene: if subjects are solitary or social, human or embedded within context, diagrammatic or complex. Based on visual reading of bodies and gestures informed by the clusters, we can rethink knowledge as performance and embodied; wherein these depictions reflect some type of relationship between the draftspeople and subjects themselves whose gestures and bodies are recorded as some type of observation, abstraction, or interpretation. For example, the bodies crouched over a basket (cluster 2) showcase a range of everyday labor practices around preparing and selling a variety of goods. Basket vendors, many of whom are women, reflect a mobile labor force moving across countryside and town, or possibly within the trade industry district 36 streets (36 phố phường) of urbanizing Hanoi at the turn of the twentieth century (see Figure 11) (Tessier). Cluster 1 often depicts a solitary figure standing and could reflect a stylistic approach to a demographic representation of humans through clothing, hair styles, and possible associated professions. Yet by examining these clusters more closely, not all the clusters could be solitary figures but instead a distinct vertical scene with two to three figures.
Table 1: Above is a subsection of PixPlot’s top ten clusters and my close visual analysis.
| Cluster Number Highlighted from Entire Corpus in Grid View | My Close Visual Analysis | Image Example and Accompanying French Caption (Translated into English) |
![]() |
- person figure standing - not always solitary but can be a distinctly simplified single scene with multiple figures - vertical - close to 20% of total drawings |
![]() Children’s ankle rings |
![]() |
- person figure crouching - usually bent over a rounded object such as a basket - Close to 15% of total drawings |
![]() Scraping a coconut |
![]() |
- complex, textured (leaves, nature, bamboo, horse, hair, folk prints) - examples of distinct line styles, shaded, frame around storytelling |
![]() Dragon dance (toy of the rich in paper and bamboo or painted tin wire) |
![]() |
- bowls, shoes, knives, glasses, simple diagrams, horizontal |
![]() A shoe especially used by the Chinese |
![]() |
- instruments, knives, swords, implements, vertical |
![]() Two-string guitar (large format) |
Interpretive Analysis of Results: Limitations and Possibilities
With PixPlot there are clear limitations for both exploration and explanation. Taking into consideration the limited training dataset of the model, it could be useful for first-level exploration of the prevalence of certain visual information such as body gestures, position, and style (in complex or simplified diagrams). For example, in the case of diagrams (Figures 15 and 16) these clusters can help to generate questions around how visual information might be communicated to imagined or perceived audiences for the purposes of instruction. At first, I make assumptions that image clustering could lead to “answers” regarding artistic style, in line with western concepts of authorship and attribution. Yet I quickly realize that, just because something is similar does not necessarily mean it is produced by the same hand, and just because it looks different does not mean it is produced by different hands. Woodblock printing in itself was and continues to be a multistep, multi-person process involving artists, carvers, printers, and publishers. Furthermore, the act of copying has had a long history in woodblock printing as a fundamental part of learning techniques, experimentation, and paying homage to earlier masters. Woodblock printing can also be embedded within longer manuscript practices of hand copying for education, dissemination, and artistic practice. This close focus on visual production points to how a woodblock-printed image could inherently be multivocal—where an artist’s hand mimics another’s style, or where the woodblock carver adds their own subtle mark in rendering a drawing to a woodblock and its impression onto paper. In widening the question of visual communication, a relational framework between artists and audiences must also be considered, where an image reflects a negotiation between individualized style, social aesthetic norms, as well as pragmatic information encoded within the visual depiction for an intended audience.
PixPlot invited a different level of engagement with the visual information of the text, patterns in content and perspective, and laid the groundwork for new research questions and visual analysis. While style is complicated through the collective process of woodblock printing, there are creative ways to examine style not as identification of hands and authorship, but as larger aesthetic patterns around representations of faces, perspectives, and caricatures. We can take a more piecemeal approach to style by looking at specific elements within an image drawing from work that examines brushstrokes in collective artists workshops (Ji et al.). A next step would be to explore if there are stylistic patterns within the visual images by looking at facial representations of specific representations such as race (Chinese, Western, ethnic minorities) and they way they complement textual information. In my close analysis of the folk prints subsection, I also note that some images bear stylistic similarities to a popular Vietnamese folk painting Đông Hồ folk woodcut painting (Tranh khắc gỗ dân gian Đông Hồ), originating in Đông Hồ village in Bắc Ninh Province, which dates back at least to the seventeenth century. These popular images are particularly popular during holidays and circulated widely as representative images of zodiac characters, folk stories, and good wishes, and they continue to circulate in Vietnam today (Figure 18). Other approaches can look more closely at carving quality, consistency, line thickness, presence of borders, and perspectives of scenes (Figure 17). Even without computational approaches, there appear to be distinct styles in the penmanship of the Arabic numbering, character numbering, and hán nôm character caption based on weight, size, neatness, and alignment of the character script. A more systematic analysis could shed light on the production processes of the multiple contributors.
Act 3 Cultural Practice: Social Worlds of Superstition and Folk Knowledge
Relational Clustering of French and Vietnamese Captions
On page 125, the fourth image from the left is accompanied by the French caption “superstition” while the Vietnamese character caption goes into narrative detail: “Pick up meat and blow on it. When slicing sausages or food that accidentally falls on the ground, pick it up, bring it to your mouth and blow on it to eliminate the bad air, then one can eat it. [Nhặt thịt thổi phù. Khi thái giò chả, thức ăn mà nhỡ để rơi xuống đất, bèn nhặt lên đưa lên miệng thổi phù để khử khí bẩn, rồi sau lại ăn] (Figure 19). Taking both sets of captions as layers, together with the visual depiction, we can distinctly see the ways in which the Vietnamese cultural practice is oversimplified in its French text description. While all levels of linguistic and visual representation simplify in order to communicate, how might we understand cultural practices as part of larger social worlds? Continuing my previous approach on the possibilities of exploratory visual clustering, this section considers how computational clustering of the text captions can provide context through semantic similarity and prompt a multimodal and multilingual, layered analysis. I note stylistic and content differences in the two sets of captions while recognizing the linguistic structural differences in the languages and different contributors who shaped the production of the two sets of captions. Instead I recognize the captions and images as a layered archive of meaning, produced in collaboration and conversation, with different intentions for textual explanation and communication.
Figure 19: Above is page 125, with the accompanying French captions translated to English from left to right as follows: 1) “Furnishings of blacksmith’s workshop” 2 ) “Alcohol saleswoman” 3) “Adjusting pants” and 4) “Superstition” The far right image highlighted in red clearly showcases a divergence where the Vietnamese character caption is significantly longer and more detailed, compared to the French simple characterization of the scene as “superstition.”
Act 3 begins by confronting the oversimplified and exoticized colonial representation of so-called superstition in this corpus and extends beyond colonial knowledge to “critically fabulate” and understand everyday folk life. The decision to “stay with the trouble” of an exogenous Eurocentric categorization of superstition dwells in the possibility that there might be other ways of reading, listening to, and attending to vernacular life (Haraway). While a large portion of this corpus focuses on ethnographic observation from a material cultural perspective, a substantive portion also focuses on contextualized social interactions and cultural practices. A subset analysis beginning with “superstition” might draw attention to other types of collective communication conveyed through storytelling, hearsay, beliefs, customs, and everyday practices.
Word Embeddings Methodology
While the previous act extracted image embeddings, this act focuses on the extracted word embeddings of the French and Vietnamese captions for microlevel exploration of superstition. With the ability to examine contextualized word relationships in high dimensional space, vector space models “reflect common habits of language that permeate a corpus, but are largely invisible to writers and readers, resulting in linguistic patterns that transcend the subject experience of communication” (Gavin et al. 247). In earlier work together with Kailiang Fu, Tyler Gurth, and David H. Laidlaw, we generated word embeddings to examine the entire text corpus and to evaluate data visualizations for exploring a subtopic on female labor and childbearing (Fu et al; C. Nguyen). In this current work and section, I analyze a subset of these word embeddings focused on superstition and folk representations and examine the French and Vietnamese captions. For the recalculated dataset of 4,454 captions (including blanks and multi-text descriptions of some of the images), we used paraphrase multilingual mpnet base v2 (pm mpnet base v2), a top performing multilingual sentence transformer model to convert the French language captions into numerical vectors, or word embeddings.10 We selected a sentence transformer framework compared to Word2Vec at the time given the improved ability to be more attuned to context bidirectionally and sentence representations.11 We decided on pm mpnet base v2 sentence transformers for the ability to conduct semantic textual similarity tasks (for example clustering and vector similarity) and to work with multilingual approaches to include the Vietnamese-, French-, and English-language captions (Reimers and Gurevych). By converting entire captions into embedding vectors, we are then able to plot our captions to a 768-dimensional dense vector space and calculate the distance between pairs of captions based on cosine similarity. We did the same approach on the Vietnamese captions, using a variety of monolingual and multilingual models, yet have not settled on an effective model that best captures the cultural nuance and semantic meaning of our captions. Future work will consider recent success in models trained on Vietnamese as well as the integration of newer large language models (Tran et al.). From the embeddings generated through pm mpnet base v2, we calculated their similarity by cosine similarity. We then created a Microsoft Excel–viewable 4,454 × 4,454 similarity matrix with cosine distance (subtracting cosine similarity from 1) in order to show the degree of similarity between two captions. This is to say, a value of 0 means that two vectors (i.e., captions) are most similar to itself, and values close to 0 are considered the most similar according to the vector space model.
We then created a visual layout of the most similar caption in a minimum spanning tree (MST) layout (Figure 20). An MST is constructed through iteratively linking together the most similar pair of captions into a single tree with the smallest possible edge lengths summed across the connections. An MST includes only a smaller subset of possible connections and instead offers a visual guide to explore the captions at the corpus level and to quickly reference their most similar captions. We then added the associated drawing paired with the text caption as well as the other set of captions within the visualization for quick comparison.
With the case example of the French language caption “136_superstition” (highlighted in in Figure 20), the four nearest neighbors (highlighted in green) include the following captions: “134_superstition sur le mortier et le pilon [Superstitions regarding the mortar and pestle],” “135_superstition populaire pour la maladie [Popular superstitious practices for sickness],” and “1035_divination par les sapèques [Divining using coins].” This visual layout allows me to quickly close read the French and Vietnamese sets of captions together with the images as a semantic concept cluster around “superstition.”
As a parallel step with the MST analysis, I closely examine the similarity matrix that we created from the word embeddings on the French-language caption dataset to understand how the vector space model considers similarity on a short caption such as “superstition.” I analyzed the top twenty most similar captions according to the French embeddings and noted a few structural observations (Table 2). Out of twenty, the majority, fifteen, also have character captions.
Table 2: Above are an excerpt of the similarity matrix with the top twenty most similar captions according to French word embeddings. Cosine distance scores closest to 0 are the most similar captions, with 0 being the closest in similarity to itself.
| Table Image Number | Superstition | French Caption | English Translation of French Caption | Character Caption | English Translation of Character Caption |
| 1 | 0 | Superstition. | Superstition. | Nhặt thịt thổi phù. Khi thái giò chả, thức ăn mà nhỡ để rơi xuống đất, bèn nhặt lên đưa lên miệng thổi phù để khử khí bẩn, rồi sau lại ăn. | Pick up meat and blow on it. When slicing sausages or food that accidentally falls on the ground, pick it up, bring it to your mouth and blow on it to eliminate the bad air, then one can eat it. |
| 2 | 0.163 | Superstition populaire sur la maladie. | Popular superstitious practices for sickness. | Đàn bà khó đẻ, chồng lấy củi cháy dở treo xà nhà. | When a woman is having a difficult birth, the husband should hang burnt wood from the rafters of the house. |
| 3 | 0.251 | Superstitions sur le mortier et le pilon. | Superstitions regarding the mortar and pestle. | Đón dâu ngoài cửa | Receiving the wife outside the door |
| 4 | 0.31 | Divination par les sapèques. | Divining using coins. | Xin âm dương | Asking for yin and yang |
| 5 | 0.331 | Divination par les baguettes. | Divination with sticks. | Xin quẻ thẻ | Asking through cards |
| 6 | 0.334 | Divination par la sapèque. | Divining by coin. | Thày bói gieo quẻ | Fortune-teller draws lots |
| 7 | 0.348 | Divination. | Divining. | Đánh đố thị | Method of divination by riddles |
| 8 | 0.403 | Pratiques de divination. | Divining. | N/A | N/A |
| 9 | 0.433 | Culte. | Worship. | N/A | N/A |
| 10 | 0.438 | Divination par les pattes de coq. | Divining by chicken feet. | Cái biển toán bốc [xem bói] | A sign for fortune-telling |
| 11 | 0.453 | Cérémonie de sorcière. | Shaman’s ceremony. | Người đàn bà lên đồng | Women entering a ceremonial trance |
| 12 | 0.466 | Annonce de sorcier. | Shaman’s announcement. | Cái biển xem số | A fortune reading board |
| 13 | 0.466 | Cérémonies magiques. | Magic ceremony. | Tiễn hỏa | Fire ritual (guidance for spirits) |
| 14 | 0.466 | Cérémonies magiques. | Magic ceremony. | Tẩy uế trong bếp | Cleansing the kitchen |
| 15 | 0.466 | Cérémonies magiques. | Magic ceremonies. | Đeo túi hạt mùi cho con | Hanging a bag of coriander seeds on a child |
| 16 | 0.466 | Cérémonies magiques. | Magic ceremonies. | Ném đất xuống ruộng mà thề | Throwing dirt down the field as an oath |
| 17 | 0.473 | Caisse portée par les sorciers ou devins en tournée. | Trunk carried by mediums and sorcerers on their rounds. | N/A | N/A |
| 18 | 0.474 | Thánh Mẫu. | Thánh Mẫu. | Đệ nhị Thánh Mẫu | Second Holy Mother Queen |
| 19 | 0.481 | Divination de la destinée d’un enfant. | Divining a child’s fate. | N/A | N/A |
| 20 | 0.484 | Théière en faïence. | Faience tea pot. | N/A | N/A |
Interpretive Analysis of Results: Limitations and Possibilities
As an approach to exploring similarity across the captions, word embeddings require extensive testing and evaluation to select their models—either monolingual or multilingual—sentence transformers, and measurements of similarity. It is important to explore the training data and model architecture to contextualize their capabilities for certain tasks. For example, the original BERT (2018), a novel, bidirectional, contextual neural network architecture had been trained on the entirety of the English-language Wikipedia and Google Books. Many technical advances have occurred since BERT. The model we selected for use, pm multilingual mpnet base v2 (2019), emerged from fine-tuning the MPNet base model using billions of sentence pairs that drew primarily from popular natural language processing datasets such as Reddit comments, scientific papers, WikiAnswers.12 Recognizing the language bias of model development on higher resource languages such as English, machine learning techniques such as knowledge distillation have been employed for multilingual models (Reimers and Gurevych, 2020). While natural language processing has more recently transitioned toward GPT-4 and large language model approaches, sentence transformer models can still be used effectively as lower cost option for smaller scale exploratory text tasks to examine textual structures and semantic clustering within a corpus.
Furthermore, given the complexity of this multilingual dataset, a universal multilingual embedding that considers both French and Vietnamese captions at the same time was also evaluated. I decided to instead approach word embeddings of the French and Vietnamese captions in a multistaged hybrid analysis of word embeddings with close readings because it invited in a slower, more transparent task specific workflow to make semantic comparisons between sets of language captions. I explore here the interpretive workflow of analyzing the word embeddings from French language together with close analysis of the Vietnamese language character captions in parallel to show the possibilities and limitations of a computationally assisted multilingual interpretation. From the MST and similarity matrices generated from French-language word embeddings, I can discern a few patterns. Firstly, within this corpus, French concepts of superstition are closely understood as semantically similar to divination. Secondly, the accompanying Vietnamese character captions to the same image labeled as “divination” in French provides much more cultural detail such as situational context of certain rituals such as receiving a bride outside the door and plural medical rituals.
Shifting away from a comparative, one-to-one linguistic approach, I emphasize how embeddings might be able to identify relational concepts through clusters of similarity through the similarity matrix and minimum spanning tree visualization. Looking back at the top twenty results of the French captions (see Table 2), that ultimately clustered concepts of ceremonies, divination, and superstition from orientalist and exogenous perspective of a different supernatural world. Yet, by reading instead the clustered Vietnamese character text along with the visual depiction, we can start to reimagine a different type of social world that is much more nuanced and dynamic. I highlight below a closer, slower reading of the narrative longer captions and images (Table 2) can center a social cosmological world of magic, divination, and sorcery that is part of everyday realities instead of exoticized, tokenized, and differentiated “ethno-cultural” colonial representations.
The closest caption in similarity to the original French one is the following: “When a woman is having a difficult birth, the husband should hang burnt wood from the rafters of the house [Đàn bà khó đẻ, chồng lấy củi cháy dở treo xà nhà],” where the French caption describes this as “Superstition populaire sur la maladie [Popular superstition for sickness]” (see Table 2 image number 2 and Figure 21). By examining closely the visual depiction and Vietnamese-character captions, we might be able to understand this representation as social problem-solving and various roles during challenging times such as childbearing. Certain ritual practices such as this could ward off malevolent spirits and connect to other acts of social care. This can be seen in an image where amulets and charms are used for protecting children or the many instances of fortune-telling that provided guidance in decision-making (Figure 22). These culturally detailed textual descriptions and accompanying visual depictions start to move toward a more nuanced analysis of “superstition” not as a separate spiritual or religious world, but one that is integrated in holistic social worlds of relationships and everyday practices.
Figure 21: Above is image number 2 from Table 2 which has a .163 similarity score to the caption “superstition.” The original image appears on page 124 and the Vietnamese caption reads: “When a woman is having a difficult birth, the husband should hang burnt wood from the rafters of the house [Đàn bà khó đẻ, chồng lấy củi cháy dở treo xà nhà].” The French caption reads, “Superstition populaire sur la maladie” [Popular superstition for sickness].”
Figure 22: Above is image number 15 from Table 2 which has a .466 similarity score to the caption “superstition.” The original image appears on page 196 and the Vietnamese caption reads: “Đeo túi hạt mùi cho con [Hanging a bag of coriander seeds on a child].” The French caption reads: “Cérémonies magiques [Magic ceremonies].” Note the visual depictions of facial expressions in the two subjects in the middle of the act, as one that could be interpreted as caretaking.
By exploring in relationship these narrative clusters generated through word embeddings in the French and layered with Vietnamese and visual analysis, a contextualized social world of spirituality and religiosity emerges in this historical corpus. Situated within the historical literature on magic and sorcery in Southeast Asia, these visual-textual representations show indigenous systems of beliefs as social relations and cosmologies (Watson and Ellen). Mythologies and folklore become part of everyday social relations, through ceremonies, festivals, fortunetelling, yin and yang medicinal practices, easing uncertainty and moving through life. Vietnamese notions of “belief [sự tin tưởng]” convey more nuanced concept of supernatural interwoven into the everyday. Thus, as seen in these visual-textual representations, cultural practice and folk belief are all encompassing arenas instead of separate spheres. For Vietnamese society, an embedded spiritual-temporal life was and continues to be prevalent through ancestral worship, spirit possession, and geomancy and phong thủy (feng shui), fortune-telling, and life cycle rituals (Figure 23, 24, 25). Throughout the twentieth century, cultural practice and religiosity was a debated practice by various actors and political authorities. In the same time period within which this corpus was created, Vietnamese intellectuals were actively involved in cultural debates on what comprised Vietnamese customs and culture, including rituals considered outmoded or part of a longer cultural-social identity. Later on, religious practices and superstition had been suppressed by governmental authorities throughout Vietnam from 1954 to 1990s, yet “black market” folk activities, often led by women, and cultural rituals persisted in contemporary Vietnam (Taylor). After the socialist reform period of đổi mới [renovation] beginning in 1986, folk beliefs have been redefined as part of national cultural heritage and reconstituted with new meanings within contemporary contexts (Dickhardt; Trần).
Figure 23: Above is image number 6 from Table 2 which has a .334 similarity score to the caption “superstition.” The original image appears on page 361 and the Vietnamese caption reads: “Thày bói gieo quẻ [Fortune-teller draws lots].” The French caption reads: “Divination par la sapèque [Divination by coin].”
Figure 24: Above is image number 11 from Table 2 which has a .453 similarity score to the caption “superstition.” The original image appears on page 350 and the Vietnamese caption reads: “Người đàn bà lên đồng [Women entering a ceremonial trance].” The French caption reads: “Cérémonie de sorcière [Shaman’s ceremony].”
Figure 25: Above is an image of page 571. The top Vietnamese character caption reads: “Các bà đồng rước thánh [Women mediums in procession],” and other captions depict certain detailed actions and elements of the scene. In comparison the French caption categorizes the entire scene as “Procession de femmes [Procession of women].” This and the previous image could be a representation of the popular Vietnamese spiritual culture and mother goddess religion of Đậo Mẫu, and shaman mediumship rituals of lên đồng or hậu đồng, which also involve recreating the lives of heroic figures and deities through possession, trance, and music (Durand; Cadière).
One might argue that a western gaze and exoticization of “superstition” is still present in the French caption or in the visual depiction. I argue, however, that clustering techniques begin to embed these fragmented representations, socially and situationally, into more dynamic social worlds with plural authors and contributors who relationally produced these hybrid representations of conversation, witness, and retelling. This relational form of knowledge production and communication moves us away from a one-directional understanding of external observer ethnographically recording an observed moment (such as in a controlled environment of a scientific laboratory). Through an interpretive digital humanities, a static description of superstition is a conduit for a multimodal, multivocal dynamic social world of belief, ceremonies, and other mediators that provide everyday guidance between the present and future, uncertainty and luck. Recognizing that these drawings and captions are fragmented imprints of a time, place, and set of relationships, we can imagine how these practices were debated, transmitted, and performed. This is particularly evident through the continuity of beliefs and rituals, repetition of proverbs and ruses that remain prevalent today as repeated folk knowledge or proverbial wisdom. Instead of simply a static collection of cultural information objects, this visual-textual archive is part of a wider communication circuit of everyday epistemologies, or what Juno Salazar Parreñas describes as a “historiography and ethnography of vernaculars” that capture relationships across kinships, senses, and human-non-human. Future work can embed this singular dataset within a wider landscape of community knowledge transmission, centering for example other forms of epistemological traditions such as storytelling, language, oral history, literature, and music.
Conclusion: Toward Transparent Communication and Embrace of Humanistic Unknowability
When and where should computation occur? In the important work Theater as Data, Miguel Escobar Varela differentiates between data-driven and data-assisted methodologies centered on an alignment of data structures, measurable tasks, and possible findings. I weave Escobar Varela’s framework of data-assisted methodologies into three “acts” of interpretive digital humanities in which data-assisted methodologies were able to problematize my assumptions, provide a multiplicity of views, and add alternative interpretations (Escobar Varela 17–18). Computational work can concretize and communicate uncertainty when transparently expressed and deployed for specific tasks. Rather than empirical results, my approach outlines an interpretive digital humanities for generating questions, making sense of results, and embedding humanistic analysis. To do so, I embrace a hybrid approach of close contextualized reading of image and word embeddings and consider how visual projections could align with interpretative speculation. Importantly, I also showcase moments of limitations within the dataset itself or visualization method, communicating the types of questions and relationships for exploration and putting forth a workflow for humanistically interpreting computational results. Taking the visual information of the diagrams (style, facial expressions, hand gestures), together with the French captions and Vietnamese hán nôm captions, I reimagine the historical and social ways in which knowledge circulated through orality, visual communication, and everyday cultural practice. I uncover the occurrences of narrative longer captions that exemplify extensive storytelling such as in hagiographies (Kỳ Đồng), literature (Truyện Kiều), and explanatory practices (ceremony, medical ruses, and rituals).
As a case example for critical global south scholarship, this article showcases alternative interpretive and relational frameworks of knowledge representations through nonlinear visual communication and social, multivocal expression. Moving toward global south digital humanities, all data must be situated within relationships of power and historical contexts of production, and data analysis must recognize that multivocality does not imply equality. Colonial data fragments require an intentional centering of alternative frameworks of knowledge representation that ultimately might mean departing from the current dataset, recognizing its complicity and inherent limitations.
Notes
- The term ‘Annam’ had been used throughout the early history of Vietnam, often in relationship to China as the southern territory. During the French colonial period, the term “Annamite” was used at times to refer to the Vietnamese colonial subject, and in certain contexts considered derogatory. Instead I translate this to the contemporary term Vietnamese. ⮭
- This is an image retrieved from the Library of Congress, which also cites the National Library of Vietnam https://www.loc.gov/item/2021666950/. ⮭
- During the French colonial period in Indochina (1858–1954), the use and construct of the term indigène was used to describe local, native inhabitants to the land as well as concepts of race and legal rights for colonial subjects of the population residing in the territories of Cochinchina, Tonkin, Annam (the three territories which today comprise Vietnam), Laos, and Cambodia. As a colonial era term, indigène was used as a homogenizing classification of vernacular, local culture often as a contrast to Western civilizational hierarchies. This article seeks to recenter vernacular culture and life by shifting towards relational multivocal knowledge production. I use the term ‘local’ inhabitants or informants to reflect the laborers involved in knowledge production of this specific text. I use ‘Vietnamese’ when discussing cultural practice, recognizing the ethno-racial classification of the local population, specifically of the ethnic majority Kinh (Việt) in relation to other ethnic groups was one of the core intentions of this representational encyclopedia. ⮭
- Previous scholars attempted the challenging task of uncovering some of the carvers names, who included the following: Nguyễn Văn Giai, Nguyễn Văn Đăng, and Phạm Văn Tiêu from Thanh Liễu Village, Gia Lộc District, Hải Dương Province, and Phạm (Trọng) Hải from Nhân Dục Village, Kim Động District, Hưng Yên Province. Tessier, Olivier. “Henri Oger’s Mechanics and Crafts of the Vietnamese People (1909): Sketches of Hanoians’ Vibrant Life.” Viet Nam: Tradition and Change, edited by Hữu Ngọc, Lady Borton, and Elizabeth Collins. Ohio University Press, 2016: 304–305. ⮭
- The following images and excerpts, and my subsequent analytical visualizations were drawn from the 2009 re-edition. Oger, Henri. Technique du peuple annamite – Mechanics and crafts of the Annamites – Kỹ thuật của người An Nam. First original edition 1909, reedition edited by Olivier Tessier and Philippe le Failler, Hanoi: École française d’Extrême-Orient, 2009. ⮭
- See updated versions of PixPlot, that includes TensorFlow 2.0 as well as Docker support and API: Peter Leonard. PixPlot. GitHub. https://github.com/pleonard212/pix-plot and Digital Methods Initiative. DMI PixPlot. GitHub. https://github.com/digitalmethodsinitiative/dmi_pix_plot. ⮭
- Stanford Vision Lab. “ImageNet Large Scale Visual Recognition Challenge 2012 (ILSVRC2012).” https://image-net.org/challenges/LSVRC/2012/. Accessed 9 May 2026. ⮭
- Leland McInnes, Uniform Manifold Approximation and Projection. GitHub. https://github.com/lmcinnes/umap. ⮭
- Alex Hock. PixPlot ML. GitHub. https://github.com/alexhock/pixplotml. ⮭
- Paraphrase-multilingual-mpnet-base-v2. HuggingFace. https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2. The base model is masked and permuted pre-traiing ofr language (MPNet), which combines features from bidirectional encoder representations from transformers (BERT) and XLNet. ⮭
- The framework of the model selected (pm mpnet base v2) uses the Sentence Bert (SBERT) framework to train models for sentence embeddings. See SBERT documentation at https://www.sbert.net/. ⮭
- For longer list of training data of sentence pairs training data, see All-mpnet-base-v2. Hugging Face. https://huggingface.co/sentence-transformers/all-mpnet-base-v2. ⮭
Acknowledgements
I gratefully acknowledge Kailiang Fu, Tyler Gurth, and David D. Laidlaw, who were collaborators in an earlier stage of this project that shaped the technical foundations for this paper. I also extend gratitude to the countless students from my Decolonial Data and Information Visualization courses and other co-thinkers who have been part of this project, namely Bella Hoang, Malery Nguyen, and Justin Han. I am also thankful for the two anonymous peer reviewers for their close reading and constructive comments.
Competing Interests
The author has no competing interests to declare.
- Boulay, Nadine, Ashley Caranto Morford, Arun Jacob, Kush Patel, and Kimberly O’Donnell. “Transforming DH Pedagogy.” Digital Studies/Le champ numérique, vol. 11, no. 1, art. 5, 2020, pp. 1–43. DOI: http://doi.org/10.16995/dscn.379.
- Cadière, Léopold. Croyances et pratiques religieuses des Viêtnamiens, Tome 1. 2nd ed., École française d’Extrême-Orient, 1992.
- Crawford, Kate. Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press, 2021.
- Crawford, Kate, and Trevor Paglen. “Excavating AI: The Politics of Training Sets for Machine Learning.” Stages—Liverpool Biennial, no. 9, 2021, pp. 1–22. https://excavating.ai.
- de Langis, Theresa, Nicole Yow Wei, Tara Tran, and Cindy Anh Nguyen. “A&Q: Feminist Trouble in Southeast Asia, an Invitation.” Verge: Studies in Global Asia, vol. 11, no. 1, 2025, pp. 105–30. DOI: http://doi.org/10.1353/vrg.2025.a951540.
- del Rio Riande, Gimena. “Digital Humanities and Visible and Invisible Infrastructures.” Global Debates in the Digital Humanities, edited by Domenico Fiormonte, Sukanta Chaudhuri, and Paolo Ricuarte. University of Minnesota Press, 2022. https://dhdebates.gc.cuny.edu/read/global-debates-in-the-digital-humanities/section/7383ee6b-52ce-48ff-b9e8-f63abdafc9a3.
- Dickhardt, Michael. “The Social Placing of Religion and Spirituality in Vietnam in the Context of Asian Modernity: Perspectives for Research.” Dynamics of Religion in Southeast Asia: Magic and Modernity, edited by Volker Gottowik. Amsterdam University Press, 2014, pp. 55–74.
- Durand, Maurice. Điện thần và nghi thức hầu đồng Việt Nam, edited by Oliver Tessier, translated and preface by Nguyễn Thị Hiệp, Marcus Durand, and Philippe Papin. Nhà Sách Tổng Hợp, 2019.
- Durand, Maurice. Imagerie Populaire Vietnamienne, edited by Philippe Papin. École française d’Extrême-Orient, 2011 (Original work published in 1960).
- Escobar Varela, Miguel. Theater as Data: Computational Journeys into Theater Research. University of Michigan Press, 2021.
- Fiormonte, Domenico. “Digital Humanities from a Global Perspective.” Laboratorio dell’ISPF, vol. 11, 2014. DOI: http://doi.org/10.12862/ispf14L203.
- Fu, Kailiang, Tyler Gurth, David H. Laidlaw and Cindy Anh Nguyen. “Visual Exploration of a Historical Vietnamese Corpus of Captioned Drawings: A Case Study” IEEE Computer Graphics and Applications. 2 February 2026. DOI: http://doi.org/10.1109/MCG.2026.3660122.
- Gavin, Michael, Collin Jennings, Lauren Kersey, and Brad Pasanek. “Spaces of Meaning: Conceptual History, Vector Semantics, and Close Reading.” Debates in the Digital Humanities, edited by Matthew Gold and Lauren Klein. University of Minnesota Press, 2019.
- Haraway, Donna J. Staying with the Trouble: Making Kin the Chthulucene. Duke University Press, 2016.
- Hartman, Saidiya. “Venus, in Two Acts.” Small Axe, vol. 12, no. 2, 2008, pp. 1–14.
- Huyền, Châm. “Kỳ Đồng Nguyễn Văn Cẩm: Từ cậu bé kỳ lạ đến thủ lĩnh phong trào yêu nước [Kỳ Đồng Nguyễn Văn Cẩm: From an extraordinary boy to a leader of a patriotic movement].” Công an nhân dân, 20 July 2019. https://cand.com.vn/Nhan-vat/Ky-dong-Nguyen-Van-Cam-Tu-cau-be-ky-la-den-thu-linh-phong-trao-yeu-nuoc-i529231/.
- Ji, F. M.S. McMaster, S. Schwab, G. Singh, L.N. Smith, S. Adhikari, M. O’Dwyer. “Discerning the Painter’s Hand: Machine Learning on Surface Topography.” Heritage Science, vol. 9, no. 152, 2021. DOI: http://doi.org/10.1186/s40494-021-00618-w.
- Koeser, Rebecca Sutton, and Zoe LeBlanc. “Missing Data, Speculative Reading.” Journal of Cultural Analytics, vol. 9, no. 2, 2024. DOI: http://doi.org/10.22148/001c.116926.
- Liu, Alan, Urszula Pawlicka-Deger, and James Smithies, editors. Critical Infrastructure Studies and Digital Humanities. University of Minnesota Press, 2026.
- Mignolo, Walter D. “On Pluriversality and Multipolar World Order: Decoloniality after Decolonization; Dewesternization after the Cold War.” Constructing the Pluriverse: The Geopolitics of Knowledge, edited by Bernd Reiter. Duke University Press, 2018, pp. 90–116. DOI: http://doi.org/10.1515/9781478002017-006.
- Nguyen, Cindy Anh. “Critical Readings of a French Colonial Vietnamese Text: Reimagining Social Worlds through Visual-Textual Representations of Female Labor.” DH+BH: An Interdisciplinary Collection on Digital Humanities and Book History, edited by Spencer Keralis and Cait Coker. Illinois Open Publishing Network, 2025. DOI: http://doi.org/10.21900/pww.25.
- Nguyen, Cindy Anh and Alejandro Alvarado Rojas. “Exploratory Computation in Digital Humanities: A Qualitative Evaluation Framework.” Journal of Open Humanities Data, vol. 12, no. 50, 2026, pp. 1–13. DOI: http://doi.org/10.5334/johd.500
- Nguyễn Mạnh Hùng. Ký Hoạ Việt Nam Đầu Thế Kỷ 20 [Vietnamese Sketches at the Beginning of the 20th Century]. Nhà Xuất Bản Trẻ, 1989.
- Nguyễn Mạnh Hùng. Xã hội Viẹt Nam cưới thế kỷ 19 đầu thế kỷ 20 qua bộ tư liệu kỹ thuật người An Nam của Henri Oger [Vietnamese Society at the end of the 19th century to the beginning of the 20th century through the Documents from Henri Oger’s Mechanics of the Vietnamese People]. 1996. Đại học khoa học xã hội và Nhân văn, PhD dissertation.
- Nguyễn Quảng Minh, and Nguyễn Mộng Hưng. “Những ấn phẩm chính viết về tác phẩm Kỹ thuật của dân Nam trong hơn trăm năm qua” [Major publications written about the book Technique du peuple annamite over the course of 100 years]. Tạp chí Nghiên cứu và Phát triển, vol. 96, no. 7, 2012, pp. 16–45. https://vjol.info.vn/index.php/ncpt-hue/article/view/8371.
- Nguyễn Quảng Minh, Puten H., and Nguyễn Mộng Hưng. “Vài điều mới biết về bộ sách Kỹ thuật người Nam/A New Perspective On the Book Technique du peuple Annamite.” Tạp chí Nghiên cứu và Phát triển 88, no. 5 (October 31, 2011): 128–42.
- Oger, Henri. Introduction générale à l’étude de la technique du peuple annamite. Geuthner Libraire-Éditeur and Jouve et Cie Imprimeurs-Éditeurs, 1910.
- Oger, Henri. Technique du peuple annamite – Mechanics and crafts of the Annamites – Kỹ thuật của người An Nam. First original edition 1909, reedition edited by Olivier Tessier and Philippe le Failler, Hanoi: École française d’Extrême-Orient, 2009.
- Parreñas, Juno Salazar. “From Decolonial Indigenous Knowledges to Vernacular Ideas in Southeast Asia.” History and Theory, vol. 59, no. 3, 2020, pp. 413–20. DOI: http://doi.org/10.1111/hith.12169.
- Picca, Davide. “Not Minds, but Signs: Reframing LLMs through Semiotics.” arXiv, 2025. DOI: http://doi.org/10.48550/arXiv.2505.17080.
- Radford, Alec, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, et al. “Learning Transferable Visual Models From Natural Language Supervision.” Proceedings of the 38th International Conference on Machine Learning, PMLR, 2021, pp. 8748–63.
- Reimers, Nils, and Iryna Gurevych. “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks.” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Suzhou, China. DOI: http://doi.org/10.48550/arXiv.1908.10084.
- Reimers, Nil and Iryna Gurevych. “Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation.” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2020, pp. 4512–4525.
- Ringler, Hannah. “Computation and Hermeneutics: Why We Still Need Interpretation to Be (Computational) Humanists.” Computational Humanities, edited by Jessica Marie Johnson, David Mimno, and Lauren Tilton. University of Minnesota Press, 2024, pp. 3–17.
- Risam, Roopika. New Digital Worlds: Postcolonial Digital Humanities in Theory, Praxis, and Pedagogy. Northwestern UP, 2018.
- Rockwell, Geoffrey, and Stéfan Sinclair. Hermeneutica: Computer-Assisted Interpretation in the Humanities. MIT Press, 2016.
- Sahagún, Bernardino de, Antonio Valeriano, Alonso Vegerano, Martín Jacobita, Pedro de San Buenaventura, Diego de Grado, Bonifacio Maximiliano, et al. Historia general de las cosas de Nueva España (Florentine Codex). Ms. Mediceo Palatino 218–20, Biblioteca Medicea Laurenziana, Florence, MiBACT, 1577. Digital Florentine Codex/Códice Florentino Digital, edited by Kim N. Richter, Alicia Maria Houtrouw, Kevin Terraciano, Jeanette Favrot Peterson, Diana Magaloni, and Lisa Sousa. Getty Research Institute, 2023. https://florentinecodex.getty.edu. Accessed May 23, 2025.
- Sherman, Jihan, Romi Morrison, Lauren Klein, and Daniela Rosner. “The Power of Absence: Thinking with Archival Theory in Algorithmic Design.” Designing Interactive Systems Conference, ACM, 2024, pp 214–23. DOI: http://doi.org/10.1145/3643834.3660690.
- Shinichi, Honna, and Akira Matsui. “A Deep Learning Approach to Quantitative Analysis of Japanese Artworks: Exploring the Latent Features of Ukiyo-e Artists.” SSRN, January 25, 2024, pp. 1–15. DOI: http://doi.org/10.2139/ssrn.4706101.
- Smith, Linda Tuhiwai. Decolonizing Methodologies: Research and Indigenous Peoples. 2nd ed., Zed Books, 2012.
- Smits, Thomas, and Melvin Wevers. “A Multimodal Turn in Digital Humanities: Using Contrastive Machine Learning Models to Explore, Enrich, and Analyze Digital Visual Historical Collections.” Digital Scholarship in the Humanities, vol. 38, no. 3, 2023, pp. 1267–80. DOI: http://doi.org/10.1093/llc/fqad008.
- Stoler, Ann Laura, editor. Haunted by Empire: Geographies of Intimacy in North American History. Duke University Press, 2006.
- Szegedy, Christian, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. “Going Deeper with Convolutions.” arXiv, September 17, 2014. http://doi.org/10.48550/arXiv.1409.4842.
- Taylor, Philip. Goddess on the Rise: Pilgrimage and Popular Religion in Vietnam. University of Hawaii Press, 2004.
- Tessier, Olivier. “Henri Oger’s Mechanics and Crafts of the Vietnamese People (1909): Sketches of Hanoians’ Vibrant Life.” Viet Nam: Tradition and Change, edited by Hữu Ngọc, Lady Borton, and Elizabeth Collins. Ohio University Press, 2016, pp. 301–42.
- Tran, Cong Dao, Nhut Huy Pham, Anh Tuan Nguyen, Truong Son Hy, Tu Vu. “ViDeBERTa: A Powerful Pre-trained Language Model for Vietnamese.” Findings of the Association for Computational Linguistics: EACL 2023, 2023, pp. 1071–78. DOI: http://doi.org/10.18653/v1/2023.findings-eacl.79.
- Trần, Phương. “Việt Nam sau 30/4/1975: “Me Tín dị đoan” trở thàn “bản sắc dân tộc” như thế nào – Kỳ 1 [Cultural Metamorphosis: The Journey of Superstition to National Identity in Post-War Vietnam, Part 1],” Luật Khoa Tạp Chí, April 26, 2020. https://luatkhoa.com/2020/04/viet-nam-sau-30-4-1975-me-tin-di-doan-tro-thanh-ban-sac-dan-toc-nhu-the-nao-ky-1/.
- Trinh, T. Minh-ha. “‘Speaking Nearby’: A Conversation with Trinh T. Minh-ha.” Interview by Nancy N. Chen. Visual Anthropology Review, vol. 8, no. 1, 1992, pp. 82–91.
- Tuck, Eve, and Wayne Yang. “R- Words: Refusing Research.” Humanizing Research: Decolonizing the Qualitative Inquiry with Youth and Communities, edited by Django Paris and Maisha Winn, SAGE, 2013, pp. 223–48.
- Watson, C.W. and Roy Ellen, editors. Understanding Witchcraft and Sorcery in Southeast Asia. University of Hawaii Press, 1993.


































