Bi-text alignment lies at the heart of all data-driven machine-learning approaches to automatic translation and is an essential technique of analysis. This book provides a systematic, foundational introduction to the automatic alignment of parallel texts.
This book provides a systematic, foundational introduction to automatic alignment of parallel texts, a family of essential corpus analysis techniques for computing and learning the mappings between corresponding parts of the texts. Bitext alignment lies at the heart of all data-driven machine learning approaches to automatic translation, and the rapid research progress on alignment during the past two decades underlies the success of statistical machine translation approaches. Alignment is used across a wide range of resource acquisition applications including word sense disambiguation, terminology extraction, and grammar induction, as well as in translation memories and biconcordances for translators' assistants, bilingual lexicographers, and computer assisted language learners.
The book provides a systematic, foundational introduction to automatic alignment of parallel textsIt surveys a wide variety of fundamental alignment techniques including: IBM and HMM alignment models, techniques for aligning comparable corpora and learning of phrasal bilexicons, more recent alignment techniques such as greedy/competitive approaches and LTG modelsUseful for both practitioners and researchers in machine translation, natural language processing, bilingual lexicography, and computer assisted language learners
Dekai Wu
bilingual lexicography computer assisted language learning grammar induction machine learning statistical machine translation terminology extraction word sense disambiguation