رکورد قبلیرکورد بعدی

" Namsel: An Optical Character Recognition System for Tibetan Text "


Document Type : AL
Record Number : 940032
Doc. No : LA6d5781k5
Language of Document : English
Main Entry : Rowinski, Zach; Keutzer, Kurt
Title & Author : Namsel: An Optical Character Recognition System for Tibetan Text [Article]\ Rowinski, Zach; Keutzer, Kurt
Title of Periodical : Himalayan Linguistics
Volume/ Issue Number : 15/1
Date : 2016
Abstract : The use of advanced computational methods for the analysis of large corpora of electronic texts is becoming increasingly popular in humanities and social science research. Unfortunately, Tibetan Studies has lacked such a repository of electronic, searchable texts. The automated recognition of printed texts, known as Optical Character Recognition (OCR), offers a solution to this problem; however, until recently, robust OCR systems for the Tibetan language have not been available. In this paper, we introduce one new system, called Namsel, which uses Optical Character Recognition (OCR) to support the production, review, and distribution of searchable Tibetan texts at a large scale. Namsel tackles a number of challenges unique to the recognition of complex scripts such as Tibean uchen and has been able to achieve high accuracy rates on a wide range of machine-printed works. In this paper, we discuss the details of Tibetan OCR, how Namsel works, and the problems it is able to solve. We also discuss the collaborative work between Namsel and its partner libraries aimed at building a comprehensive database of historical and modern Tibetan works—a database that consists of more than one million pages of texts spanning over a thousand years of literary production.
کپی لینک

پیشنهاد خرید
پیوستها
عنوان :
نام فایل :
نوع عام محتوا :
نوع ماده :
فرمت :
سایز :
عرض :
طول :
6d5781k5_42819.pdf
6d5781k5.pdf
مقاله لاتین
متن
application/pdf
549.70 KB
85
85
نظرسنجی
نظرسنجی منابع دیجیتال

1 - آیا از کیفیت منابع دیجیتال راضی هستید؟