Linguistic knowledge in statistical phrase-based word alignment

A. DE GISPERT; J. B. MARIÑO

doi:10.1017/S1351324905003931

Linguistic knowledge in statistical phrase-based word alignment

Published online by Cambridge University Press: 06 December 2005

A. DE GISPERT and

J. B. MARIÑO

Show author details

A. DE GISPERT: Affiliation:
TALP Research Center, Universitat Politècnica de Catalunya (UPC), Jordi Girona 1-3, Campus Nord D5, 08034 Barcelona, Spain e-mail: agispert@gps.tsc.upc.es, canton@gps.tsc.upc.es
J. B. MARIÑO: Affiliation:
TALP Research Center, Universitat Politècnica de Catalunya (UPC), Jordi Girona 1-3, Campus Nord D5, 08034 Barcelona, Spain e-mail: agispert@gps.tsc.upc.es, canton@gps.tsc.upc.es

Article contents

Abstract

Get access

Rights & Permissions

Abstract

In this paper, a novel phrase alignment strategy combining linguistic knowledge and cooccurrence measures extracted from bilingual corpora is presented. The algorithm is mainly divided into four steps, namely phrase selection and classification, phrase alignment, one-to-one word alignment and postprocessing. The first stage selects a linguistically-derived set of phrases that convey a unified meaning during translation and are therefore aligned together in parallel texts. These phrases include verb phrases, idiomatic expressions and date expressions. During the second stage, very high precision links between these selected phrases for both languages are produced. The third step performs a statistical word alignment using association measures and link probabilities with the remaining unaligned tokens, and finally the fourth stage takes final decisions on unaligned tokens based on linguistic knowledge. Experiments are reported for an English-Spanish parallel corpus, with a detailed description of the evaluation measure and manual reference used. Results show that phrase cooccurrence measures convey a complementary information to word cooccurrences and a stronger evidence of a correct alignment, successfully introducing linguistic knowledge in a statistical word alignment scheme. Precision, Recall and Alignment Error Rate (AER) results are presented, outperforming state-of-the-art alignment algorithms.

Type: Papers
Information: Natural Language Engineering , Volume 12 , Issue 1 , March 2006 , pp. 91 - 108

DOI: https://doi.org/10.1017/S1351324905003931 [Opens in a new window]

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

Article contents

Linguistic knowledge in statistical phrase-based word alignment

Abstract

Access options

Save article to Kindle

Save article to Dropbox

Save article to Google Drive

Reply to: Submit a response

Your details

You have entered the maximum number of contributors

Conflicting interests