Exploiting Semantic Annotations in Math Information Retrieval

Sojka,  Petr

Exploiting Semantic Annotations in Math Information Retrieval

Varování

Publikace nespadá pod Ekonomicko-správní fakultu, ale pod Fakultu informatiky. Oficiální stránka publikace je na webu muni.cz.

Název česky	Využití sémantického značkování pro vyhledávání matematiky
Autoři	SOJKA Petr
Rok publikování	2012
Druh	Článek ve sborníku
Konference	Proceedings of ESAIR 2012
Fakulta / Pracoviště MU	Fakulta informatiky
Citace	SOJKA, Petr. Exploiting Semantic Annotations in Math Information Retrieval. In Jaap Kamps, Jussi Karlgren, Peter Mika, Vanessa Murdock. Proceedings of ESAIR 2012. Maui, USA: ACM, 2012, s. 15-16. ISBN 978-1-4503-1717-7. Dostupné z: https://dx.doi.org/10.1145/2390148.2390157.
www	DOI (ACM DL) workshop website preprint PDF poster
Doi	http://dx.doi.org/10.1145/2390148.2390157
Obor	Informatika
Klíčová slova	MIaS;MathML;indexing;search;canonical MathML;EuDML;digital libraries;information systems;information retrieval;mathematical content search;math indexing and retrieval;document ranking of math papers;text mining;DML-CZ;DML projects;semantics
Přiložené soubory	p15-sojka.pdf
Popis	This paper describes exploitation of semantic annotations in the design and architecture of MIaS (Math Indexer and Searcher) system for mathematics retrieval. Basing on the claim that navigational and research search are `killer' applications for digital library such as the European Digital Mathematics Library, EuDML, we argue for an approach based on Natural Language Processing techniques as used in corpus management systems such as the Sketch Engine, that will reach web scalability and avoid inference problems. The main ideas are 1) to augment surface texts (including math formulae) with additional linked representations (maps) bearing semantic information (expanded formulae as text, canonicalized text and subformulae) for indexing, including support for indexing structural information (expressed as Content MathML or other tree structures) and 2) use semantic user preferences to order found documents. The semantic enhancements of the MIaS system are being implemented as a math-aware search engine based on the state-of-the-art system Apache Lucene, with support for [MathML] tree indexing. Scalability issues have been checked against more than 400,000 arXiv documents.
Související projekty:	Účast ČR v European Research Consortium for Informatics and Mathematics The European Digital Mathematics Library