Named Entity Recognition (NER), search, classification and tagging of names name like frequent informational elements in texts, has become a standard extraction procedure for textual data. NER has been applied to many of texts and different types of entities: newspapers, fiction, historical, persons, locations, chemical compounds, protein families, animals etc. general a NER system's performance is genre and domain dependent and also entity categories vary (Nadeau and Sekine, 2007). The most general set of entities is usually some version of three partite categorization of, persons and organizations. In this paper we report first large scale and evaluation of NER with data out of a digitized Finnish historical collection Digi. Experiments, results and discussion of this research development of the Web collection of historical Finnish newspapers. Digi collection contains 1,960,921 pages of newspaper material from years1771-1910 both in Finnish and Swedish. We use only material of Finnish in our evaluation. The OCRed newspaper collection has lots of OCR; its estimated word level correctness is about 70-75 % (Kettunen and\\"a\\"akk\\"onen, 2016). Our principal NER tagger is a rule-based tagger of, FiNER, provided by the FIN-CLARIN consortium. We show also results of category semantic tagging with tools of the Semantic Computing Research (SeCo) of the Aalto University. Three other tools are also evaluated. This research reports first published large scale results of NER in a Finnish OCRed newspaper collection. Results of the research NER results of other languages with similar noisy data.
No takes yet. Share an insight, caveat, or question.
Kettunen et al. (2016) studied this question.