Staatsbibliothek zu Berlin | CrossAsia and HKUST Launch Collaboration on Manchu OCR
The Staatsbibliothek zu Berlin | CrossAsia has recently signed a Memorandum of Understanding (MoU) with the Hong Kong University of Science and Technology (HKUST) to collaborate on the OCR processing of its Manchu collection.
A Major Manchu Collection and a New Collaboration
The Staatsbibliothek zu Berlin holds one of the world’s largest Manchu collections, comprising around 550 titles. Following several years of digitization efforts, nearly the entire collection has now been digitized. To date, 406 Manchu works in 584 volumes have been made openly accessible through the Digitized Collections of the Staatsbibliothek zu Berlin. (For a brief history of the collection, see CrossAsia’s Theme Portal about the Manchu Collection.)
The collaboration with HKUST aims to advance OCR technologies for Manchu materials and thereby enable a wider range of services and research possibilities based on full-text data. The project at HKUST is led by Dr. Michael Yan Hon Chung, Assistant Professor in Digital Humanities at HKUST’s Division of Humanities. As a historian specializing in the early Qing dynasty, Dr. Chung has been developing a Manchu OCR system based on fine-tuned Vision Language Models (VLM). Remarkably, it was the Staatsbibliothek’s digital Manchu collection itself that became the starting point for his exploration into Manchu OCR.
A Research Need Behind Manchu OCR
We interviewed Dr. Chung about how the project began and the development of his Manchu OCR system:
“It is a great pleasure to see the Hong Kong University of Science and Technology (HKUST) and the Staatsbibliothek zu Berlin (SBB) sign a Memorandum of Understanding to advance the digitization and OCR processing of SBB’s substantial collection of Manchu documents. The significance of this collaboration is deeper than it may first appear, for the Manchu OCR project itself began with SBB.
In fall 2023, while conducting research for my doctoral dissertation, I came across the Digitized Collections of the Staatsbibliothek zu Berlin. I was immediately struck by the quality and scale of SBB’s publicly available scans of Manchu materials, including the Manchu version of the Baqi Tongzhi 八旗通志 (Han i araha jakūn gūsai tung jy bithe), which was crucial to my research. Yet I soon realized that I was also lost in the vastness of the collection. With roughly 250 volumes to consult, I needed a computer-assisted method to identify Manchu keywords across large quantities of scanned documents. This was the beginning of the Manchu OCR project.
The first Manchu OCR system, developed in collaboration with Mr. Martin Leong, was modest in scope. It could identify only two Manchu words: poo, meaning “artillery,” and cooha, meaning “army” or “military.” These two terms were chosen specifically for my research on the Hanjun Eight Banners (baqi hanjun). What began as a small, research-driven tool soon grew into a much larger project. With Dr. Donghyeok Choi joining the team, we turned to fine-tuning state-of-the-art machine learning models for Manchu OCR. Two and a half years later, we have developed a general Manchu OCR system that achieves over 96% word-recognition accuracy on woodblock-printed texts and well-written Qing documents and Manchu books. We are now excited to digitize SBB’s Manchu corpus so that it can support keyword search, copy-and-paste access, computational analysis, and other forms of digital humanities research. We also plan to launch a public-facing Manchu OCR web service in the near future.
The Manchu OCR project demonstrates the importance of collaboration among libraries, archives, and universities, especially in the field of digital humanities. Without SBB’s initial digitization work and its decision to make high-quality scans openly accessible, this project would not have taken off. I sincerely hope that this collaboration will become one example among many of how open cooperation between cultural institutions and universities can prepare the humanities for the emerging age of artificial intelligence.”
From Digitized Collections to Full-Text Research
We believe that this collaboration marks an important step toward making Manchu-language sources more accessible to scholars worldwide. As a next step, Dr. Chung plans to make the OCR platform (https://ocr.manchu.ai) publicly available by the end of this year. The OCR output will be integrated into our existing digital collections and made openly accessible to the public. This will enable full-text searching across our Manchu holdings. Full texts will be freely available for download and will significantly improve the usability of these materials for research and teaching.
Meanwhile, we warmly welcome scholars around the world to make use of our open-access digital resources and to explore new possibilities for research through digital methods. We also look forward to further collaborations that advance digital scholarship and foster new connections between libraries, universities, and the wider research community. Please feel free to contact us with any questions or inquiries about potential collaboration!



Diskutieren Sie hierzu im CrossAsia Forum