Unsupervised learning of part-of-speech guessing rules

ANDREI MIKHEEV

doi:10.1017/S1351324996001271

Unsupervised learning of part-of-speech guessing rules

Published online by Cambridge University Press: 01 June 1996

ANDREI MIKHEEV

Show author details

ANDREI MIKHEEV: Affiliation:
HCRC, Language Technology Group, University of Edinburgh, 2 Buccleuch Place, Edinburgh EH8 9LW, Scotland, UK. Email: Andrei.Mikheev@ed.ac.uk

Article contents

Abstract

Get access

Rights & Permissions

Abstract

Words unknown to the lexicon present a substantial problem to part-of-speech tagging. In this paper we present a technique for fully unsupervised acquisition of rules which guess possible parts of speech for unknown words. This technique does not require specially prepared training data, and uses instead the lexicon supplied with a tagger and word frequencies collected from a raw corpus. Three complimentary sets of word-guessing rules are statistically induced: prefix morphological rules, suffix morphological rules and ending guessing rules. The acquisition process is strongly associated with guessing-rule evaluation methodology which is solely dedicated to the performance of part-of-speech guessers. Using the proposed technique a guessing-rule induction experiment was performed on the Brown Corpus data and rule-sets, with a highly competitive performance, were produced and compared with the state-of-the-art. To evaluate the impact of the word-guessing component on the overall tagging performance, it was integrated into a stochastic and a rule-based tagger and applied to texts with unknown words.

Type: Research Article
Information: Natural Language Engineering , Volume 2 , Issue 2 , June 1996 , pp. 111 - 136

DOI: https://doi.org/10.1017/S1351324996001271 [Opens in a new window]

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

Article contents

Unsupervised learning of part-of-speech guessing rules

Abstract

Access options

Save article to Kindle

Save article to Dropbox

Save article to Google Drive

Reply to: Submit a response

Your details

You have entered the maximum number of contributors

Conflicting interests