|
Here is how I describe
alignment in my book
(under the auspicious heading "A Word of Caution on Alignment"):
In the context of
most TEnTs, alignment refers to the process of selecting file pairs in the
source and target language that were translated outside of a translation memory
environment, matching all the segments (sentences, headings, etc.), and
creating a translation memory database from those matches. The resulting
translation memory can then be applied to translate similar or identical texts.
[Virtually all TEnTs] contain alignment modules in some or all configurations.
At first glance, alignment seems like a great process that anyone starting to
use a translation environment tool should do to build up a nice translation
memory database.
True, alignment is
indeed a helpful process, but it's often misused. I've encountered many
situations where new users (both freelance and corporate) became enamored with
the idea of using alignment to "magically" turn their existing
translation materials into one large translation memory. They spent days or
weeks devoting their time to this task, and in the process they became so
frustrated with the use of their new tool that they stopped using it
altogether. The reason that alignment is often (and correctly) perceived as a
tedious process is its manual nature. Although each of the alignment modules .
. . applies well-chosen parameters to the alignment "suggestions,"
they all have to be verified, and--as anyone knows who has done alignment
before--often repaired. The parameters are typically punctuation and paragraph
markers, repetitions, and non-linguistic matches such as numbers and
abbreviations. This can go a long way toward making correct matches, but it
often requires user intervention. Typical cases where manual changes are
required are differences in sentence delimitation (one sentence in the source
becomes several in the target or the other way around), shifts in the order of
segments, different use and/or placement of footnotes, and index markers.
With all these
difficulties, why would alignment still be a helpful process? Alignment can be
a very powerful tool if you have specific sets of already-translated documents
that correspond closely to new documents which now have to be translated. The
amount of time you can save and the level of consistency and quality you can
achieve by aligning the existing documents and using that as the basis for your
translation can be immense, and there's simply no reason not to go that route.
But for other documents, unless you can hire someone else to do mass alignment
of existing materials (someone with the odd combination of being both cheap and
well-qualified. . .), I would strongly advise you to build up your translation
memory database by simply performing translation in the tool of your choice and
adding material to your translation memory segment by segment.
Now . . . there is
also a less "traditional" form of alignment, which in most situations
will be more successful.
Terminotix's AlignFactory
(see www.terminotix.com) offers an uncommonly
high accuracy of alignment because a) it uses the sophisticated alignment
engine developed by the Rali project at the University of Montréal and b) it
uses a number of filters that filter out any unlikely match (for instance,
based on differing lengths of segments). Furthermore, with AlignFactory
you can also select thousands of file pairs (including PDF files), have them
matched up (they have to follow certain naming conventions such as a language
identifier), and then have them aligned in one big swoosh. And it really is one
big swoosh: the speed of the alignment is mind-boggling. In fact, it's so fast
that I have repeatedly thought that something had gone wrong only to find that
it had already successfully completed the alignment. While it's not perfect, it
comes as close as I have seen in any tool.
Another Canadian
project is NoBabel's AutoAligner (see www.nobabel.com)
which performs the alignment of a multitude of files online. As of the writing
of this version of the Tool Box (6.1 -- January 2008), I have not had a chance
to test it.
Well, it seems to be
time for a new version of the Tool Box book, because I have now tested AutoAligner
and something tells me that I need to completely rework my chapter on alignment.
The cool thing about AutoAligner
is that nothing I said in the previous paragraphs about alignment is true for
this tool. It is not manual, it is not tedious, it does not look at sentences
like the other tools, and it does not work on a file-by-file basis. Oh, and it also
isn't perfect, just pretty darn close to it. But let me start from scratch.
First, unlike the other
tools, AutoAligner is a SaaS, meaning you don't have to install any
software on your computer. Instead, you upload files that need to be aligned to
an online server, and the software on the server does the alignment for you.
What? No manual corrections? That's right. The software is on its own. All you
need to do is upload files (for a test I uploaded about a thousand English,
German, and Chinese HTML files that did not necessarily correspond with each
other) without "telling" the AutoAligner anything about the
files, not even what language they are in. The files are analyzed on the server,
the language is correctly recognized, and the file pairs are matched up. Now, in
contrast to the other tools, the matching up is not performed on the basis of
the file name (onefile_de.html matches with onefile_en.html) but on the basis
of the actual content. So, according to a content analysis, the files are
matched up (in my case there were about 250 English> Chinese and English>
German file pairs matched up) and the remaining files are discarded. In the
next step, repositories for each of the language pairs are created in which all
the content of all the relevant files is read (of course, you don't see all
this; it's what happens in the background). The repository pairs are then
aligned according to a number of criteria, of which the order of sentences in
the files (which is the most important criterion for the other tools) is the
least important! The most important is a linguistic analysis, followed by file
origin. This means that not every possible pairing will end up in the translation
memory, but only those that are deemed appropriate (typically up to 5% are ignored).
But this also means that you can upload a list of terms in two different languages,
both sorted according to the language-specific sort order, and you will still
have correct alignment (try that with a traditional alignment tool!).
The resulting pairings
are then read into a TMX or Trados Workbench text format file and can be
downloaded by the user.
I spent some time with
the resulting TMX files (each of which had about 8000 translation units). I thoroughly
went through the first 4000 translation units of the English> German TM and
I found 3 (!) errors. I had only a cursory glance at the English> Chinese TM
and found no error.
Phew! Folks, this is
quite something -- in fact, it's unheard of!
Here are some of the
official specifics: The file types that are supported are text, Word,
WordPerfect, and HTML. (I pleaded with the developers to also include PDF which
is handled so beautifully by AlignFactory and they promised me to look
into it again.) The number of files you can upload is unlimited. Like I said, I
uploaded about 3000 files in three zip files. My first attempt to upload them
all in one zip file failed -- the system accepted only files up to 15 or so MB
at a time.
The languages that are
presently supported are English paired with French, Italian, German, Spanish,
Dutch, Portuguese, Russian, Polish, Chinese, Japanese, and Arabic. While the
limited number of language combinations is frustrating, it makes sense because
of the underlying linguistic data. It's only a matter of time until this number
grows and the need for English as source or target will vanish (essentially it's
a matter of funds to purchase the additional underlying dictionaries, etc.).
Since there are
linguistic processes underlying the analysis that this tool performs, it's not
super-fast. For the number of files that I had it took a few hours. But since
the processing did not happen on my computer, I could just log-off and check
back a few hours later to download the TM files.
Once the files are
generated and ready to download, you'll need to make a purchase decision. The
price is US$.02 per translation unit and it's up to you at that point to say yea
or nay (but when all that material is already available, who would decide at
that point not to take it?). So the price for my two databases would have been
approximately $160 (and the first $100 is free to any new user). By the way,
there are no duplicates in the files so you really are paying for individual
TUs.
What are the drawbacks?
I can think of only two. There truly is no context, which is a problem for
tools that have a feature which looks for preceding and succeeding matches to "guarantee"
a match (like Déjà Vu or MemoQ offer) or for tools that strongly
emphasize context (like Multitrans), and you will not get every single
match (see above). Even if it's only 2% that you miss out on, it still adds up
to 2% less repetition. So if you have to translate a manual of version 1.1 of a
product for which you already have version 1, it might make sense to align it
the good old-fashioned tedious way or with a tool like AlignFactory. However,
if you want to build TMs for matching and reference purposes and are
willing to invest some money, AutoAligner here we come.
|