Files Words Tokens Segments Arabic 31 110,690 141,058 7,102 Note: Word count is based on the untokenized Arabic source. Token count is based on the ATB-tokenized Arabic source.|
The purpose of the GALE word alignment task was to find correspondences between words, phrases or groups of words in a set of parallel texts. Arabic-English word alignment annotation consisted of the following tasks:
* Identifying different types of links: translated (correct or incorrect) and not translated (correct or incorrect)
* Identifying sentence segments not suitable for annotation, e.g., blank segments, incorrectly-segmented segments, segments with foreign languages
* Tagging unmatched words attached to other words or phrases
This release contains four types of files - raw, tokenized, treebank, and wa. The raw format contains the original Arabic and English sentences without any annotation. The tokenized format is the treebank tokenized version of the raw data which may contain Empty Category tokens (treebank leaves that have the POS label -NONE-). The treebank and wa files are treebank and word alignment annotations on the tokenized files.
Please view the following samples
* English Raw
* English Token
* English Treebank
* Arabic Raw
* Arabic Token
* Arabic Treebank
* Word Alignment
This work was supported in part by the Defense Advanced Research Projects Agency, GALE Program Grant No. HR0011-06-1-0003. The content of this publication does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.
None at this time.