رکورد قبلیرکورد بعدی

" Extending corpus annotation of Nepali: advances in tokenisation and lemmatisation "


Document Type : AL
Record Number : 939997
Doc. No : LA15t805x8
Language of Document : English
Main Entry : Hardie, Andrew; Lohani, Ram; Yadava, Yogendra
Title & Author : Extending corpus annotation of Nepali: advances in tokenisation and lemmatisation [Article]\ Hardie, Andrew; Lohani, Ram; Yadava, Yogendra
Title of Periodical : Himalayan Linguistics
Volume/ Issue Number : 10/1
Date : 2011
Abstract : The Nepali National Corpus (NNC) was, in the process of its creation, annotated with part-of-speech (POS) tags. This paper describes the extension of automated text and corpus annotation in Nepali from POS tags to lemmatisation, enabling a more complex set of corpus-based searches and analyses. This work also addresses certain practical compromises embodied in the initial tagging of the NNC. First, some particular aspects of Nepali morphology – in particular the complexity of the agglutinative verbal inflection system – necessitated improvements to the underlying tokenisation of the text before lemmatisation could be satisfactorily implemented. In practical terms, both the tokenisation and lemmatisation procedures require linguistic knowledge resources to operate successfully: a set of rules describing the default case, and a lexicon containing a list of individual exceptions: words whose form suggests a particular rule should apply to them, but where that rule in fact does not apply. These resources, particularly the lexicons of irregularities, were created by a strongly data-driven process working from analyses of the NNC itself. This approach to tokenisation and lemmatisation, and associated linguistic knowledge resources, may be illustrative and of use to researchers looking at other languages of the Himalayan region, most especially those that have similar morphological behaviour to Nepali.
کپی لینک

پیشنهاد خرید
پیوستها
عنوان :
نام فایل :
نوع عام محتوا :
نوع ماده :
فرمت :
سایز :
عرض :
طول :
15t805x8_42714.pdf
15t805x8.pdf
مقاله لاتین
متن
application/pdf
568.54 KB
85
85
نظرسنجی
نظرسنجی منابع دیجیتال

1 - آیا از کیفیت منابع دیجیتال راضی هستید؟