ToolkitSmall

A computer newsletter for translation professionals

Issue 8-4-111
(the one hundred eleventh edition)
Contents
1. Alignment Afflictions, Adieu!
2. User Groups for TEnTs
3. Similis for Free
4. Translated Environment Tools
5. Terminology Tune-Up
6. Spell-Czech
The Last Word on the Tool Kit
Berlin Alexanderplatz

Last week during our little Easter vacation away from Berlin, I thoroughly enjoyed reading Alfred Döblin's Berlin Alexanderplatz, one of the most colorful and eccentric books about this city in the 1920s and about humanity as a whole. When I came to the passage where the protagonist Franz joins a band of robbers who--hmm, how should I say this--empty their bowels in the places they rob to complete the humiliation, I smugly grinned and thought to myself: Good that we're not in the 1920s anymore.

You may be able to guess where this is headed: when we came back to our temporary home in Berlin, we found that Franz and his gang had visited our apartment and had not changed their habits. Uggh.

My wife, whose computer was one of Franz's trophies, would like to share this urgent message with each of you: BACK UP! (And use a deadbolt!)

1. Alignment Afflictions, Adieu!

Here is how I describe alignment in my book (under the auspicious heading "A Word of Caution on Alignment"):

In the context of most TEnTs, alignment refers to the process of selecting file pairs in the source and target language that were translated outside of a translation memory environment, matching all the segments (sentences, headings, etc.), and creating a translation memory database from those matches. The resulting translation memory can then be applied to translate similar or identical texts. [Virtually all TEnTs] contain alignment modules in some or all configurations. At first glance, alignment seems like a great process that anyone starting to use a translation environment tool should do to build up a nice translation memory database.

True, alignment is indeed a helpful process, but it's often misused. I've encountered many situations where new users (both freelance and corporate) became enamored with the idea of using alignment to "magically" turn their existing translation materials into one large translation memory. They spent days or weeks devoting their time to this task, and in the process they became so frustrated with the use of their new tool that they stopped using it altogether. The reason that alignment is often (and correctly) perceived as a tedious process is its manual nature. Although each of the alignment modules . . . applies well-chosen parameters to the alignment "suggestions," they all have to be verified, and--as anyone knows who has done alignment before--often repaired. The parameters are typically punctuation and paragraph markers, repetitions, and non-linguistic matches such as numbers and abbreviations. This can go a long way toward making correct matches, but it often requires user intervention. Typical cases where manual changes are required are differences in sentence delimitation (one sentence in the source becomes several in the target or the other way around), shifts in the order of segments, different use and/or placement of footnotes, and index markers.

With all these difficulties, why would alignment still be a helpful process? Alignment can be a very powerful tool if you have specific sets of already-translated documents that correspond closely to new documents which now have to be translated. The amount of time you can save and the level of consistency and quality you can achieve by aligning the existing documents and using that as the basis for your translation can be immense, and there's simply no reason not to go that route. But for other documents, unless you can hire someone else to do mass alignment of existing materials (someone with the odd combination of being both cheap and well-qualified. . .), I would strongly advise you to build up your translation memory database by simply performing translation in the tool of your choice and adding material to your translation memory segment by segment.

Now . . . there is also a less "traditional" form of alignment, which in most situations will be more successful.

Terminotix's AlignFactory (see www.terminotix.com) offers an uncommonly high accuracy of alignment because a) it uses the sophisticated alignment engine developed by the Rali project at the University of Montréal and b) it uses a number of filters that filter out any unlikely match (for instance, based on differing lengths of segments). Furthermore, with AlignFactory you can also select thousands of file pairs (including PDF files), have them matched up (they have to follow certain naming conventions such as a language identifier), and then have them aligned in one big swoosh. And it really is one big swoosh: the speed of the alignment is mind-boggling. In fact, it's so fast that I have repeatedly thought that something had gone wrong only to find that it had already successfully completed the alignment. While it's not perfect, it comes as close as I have seen in any tool.

Another Canadian project is NoBabel's AutoAligner (see www.nobabel.com) which performs the alignment of a multitude of files online. As of the writing of this version of the Tool Box (6.1 -- January 2008), I have not had a chance to test it.

Well, it seems to be time for a new version of the Tool Box book, because I have now tested AutoAligner and something tells me that I need to completely rework my chapter on alignment.

The cool thing about AutoAligner is that nothing I said in the previous paragraphs about alignment is true for this tool. It is not manual, it is not tedious, it does not look at sentences like the other tools, and it does not work on a file-by-file basis. Oh, and it also isn't perfect, just pretty darn close to it. But let me start from scratch.

First, unlike the other tools, AutoAligner is a SaaS, meaning you don't have to install any software on your computer. Instead, you upload files that need to be aligned to an online server, and the software on the server does the alignment for you. What? No manual corrections? That's right. The software is on its own. All you need to do is upload files (for a test I uploaded about a thousand English, German, and Chinese HTML files that did not necessarily correspond with each other) without "telling" the AutoAligner anything about the files, not even what language they are in. The files are analyzed on the server, the language is correctly recognized, and the file pairs are matched up. Now, in contrast to the other tools, the matching up is not performed on the basis of the file name (onefile_de.html matches with onefile_en.html) but on the basis of the actual content. So, according to a content analysis, the files are matched up (in my case there were about 250 English> Chinese and English> German file pairs matched up) and the remaining files are discarded. In the next step, repositories for each of the language pairs are created in which all the content of all the relevant files is read (of course, you don't see all this; it's what happens in the background). The repository pairs are then aligned according to a number of criteria, of which the order of sentences in the files (which is the most important criterion for the other tools) is the least important! The most important is a linguistic analysis, followed by file origin. This means that not every possible pairing will end up in the translation memory, but only those that are deemed appropriate (typically up to 5% are ignored). But this also means that you can upload a list of terms in two different languages, both sorted according to the language-specific sort order, and you will still have correct alignment (try that with a traditional alignment tool!).

The resulting pairings are then read into a TMX or Trados Workbench text format file and can be downloaded by the user.

I spent some time with the resulting TMX files (each of which had about 8000 translation units). I thoroughly went through the first 4000 translation units of the English> German TM and I found 3 (!) errors. I had only a cursory glance at the English> Chinese TM and found no error.

Phew! Folks, this is quite something -- in fact, it's unheard of!

Here are some of the official specifics: The file types that are supported are text, Word, WordPerfect, and HTML. (I pleaded with the developers to also include PDF which is handled so beautifully by AlignFactory and they promised me to look into it again.) The number of files you can upload is unlimited. Like I said, I uploaded about 3000 files in three zip files. My first attempt to upload them all in one zip file failed -- the system accepted only files up to 15 or so MB at a time.

The languages that are presently supported are English paired with French, Italian, German, Spanish, Dutch, Portuguese, Russian, Polish, Chinese, Japanese, and Arabic. While the limited number of language combinations is frustrating, it makes sense because of the underlying linguistic data. It's only a matter of time until this number grows and the need for English as source or target will vanish (essentially it's a matter of funds to purchase the additional underlying dictionaries, etc.).

Since there are linguistic processes underlying the analysis that this tool performs, it's not super-fast. For the number of files that I had it took a few hours. But since the processing did not happen on my computer, I could just log-off and check back a few hours later to download the TM files.

Once the files are generated and ready to download, you'll need to make a purchase decision. The price is US$.02 per translation unit and it's up to you at that point to say yea or nay (but when all that material is already available, who would decide at that point not to take it?). So the price for my two databases would have been approximately $160 (and the first $100 is free to any new user). By the way, there are no duplicates in the files so you really are paying for individual TUs.

What are the drawbacks? I can think of only two. There truly is no context, which is a problem for tools that have a feature which looks for preceding and succeeding matches to "guarantee" a match (like Déjà Vu or MemoQ offer) or for tools that strongly emphasize context (like Multitrans), and you will not get every single match (see above). Even if it's only 2% that you miss out on, it still adds up to 2% less repetition. So if you have to translate a manual of version 1.1 of a product for which you already have version 1, it might make sense to align it the good old-fashioned tedious way or with a tool like AlignFactory. However, if you want to build TMs for matching and reference purposes and are willing to invest some money, AutoAligner here we come.

ADVERTISEMENT

SDL TRADOS Technologies, the world's leading translation tool provider, continues its commitment to the translation community.

It has been almost three years since SDL and Trados merged. What additional benefits has the combined SDL TRADOS business brought to the translation community?

Free service packs, educational web seminars, ideas.sdltrados.com and much more ... but how much more?

Click here to find out.

2. User Groups for TEnTs

I am often asked for addresses of user groups for the different TEnTs. I published a list awhile back, but it's now slightly outdated and incomplete. Here is a more complete list (please let me know what I've missed):

Déjà Vu: http://tech.groups.yahoo.com/group/dejavu-l/

Logiterm: http://tech.groups.yahoo.com/group/logiterm/

MemoQ: http://tech.groups.yahoo.com/group/MemoQ2/

MetaTexis: http://tech.groups.yahoo.com/group/MetaTexis/

MultiTrans: http://tech.groups.yahoo.com/group/multitrans/

OmegaT: http://tech.groups.yahoo.com/group/OmegaT/

SDLX: http://tech.groups.yahoo.com/group/sdlx/

Star Transit: http://tech.groups.yahoo.com/group/transit_termstar/

Swordfish: http://tech.groups.yahoo.com/group/swordfish_support/

Trados: http://tech.groups.yahoo.com/group/tw_users/

Wordfast: http://tech.groups.yahoo.com/group/wordfast/

Some of these groups are more active than others, but most of them will give you a fairly good idea of what the users think of their tool (and competing tools). In many cases this is the most efficient way to get help, and some of these groups also have non-English counterparts!
3. Similis for Free

Lingua et Machina, the maker of the TEnT Similis, is offering its tool for free in a limited yet functional version. Those who have read this newsletter for a while will remember that I have praised Similis in the past, in particular for its term extraction capability in the languages it supports (English, Dutch, German, Spanish, Italian, Portuguese, and French) where it achieves a much higher accuracy than any other tool that I am aware of. Of course, it is also a full-fledged translation tool with a translation memory and dictionary component. Since it comes with a good amount of linguistic material in the supported languages, be aware that it is by no means a small download and installation. Still, that shouldn't deter anyone who wants to have a good look at this tool from downloading and installing it.

The limitations of this free edition include a block on most features concerning data exchange (TMX, Trados, CSV) and you can process only the first 10,000 words in each document.
ADVERTISEMENT
across Personal Edition -- the next generation translation workbench!

FREELANCERS BENEFIT FROM
1. free license (worth 1450 USD / 980 euros)
2. free training and support
3. outstanding referral commission

Register now and join a community of more than 5,000 freelance translators!

4. Translated Environment Tools . . .

. . . is not what TEnT stands for, as Fernand Legros would very strongly argue. And he would be right. We all know it stands for Translation Environment Tool, but it is sadly true that many of these tools are only partially translated or not translated at all. Fernand particularly complains about Trados, which is officially translated into German, French, Spanish, and German (with a separate Japanese version), but often only on the surface. He points out that many components and manuals of the program are not localized, and says he feels "ridiculed and offended" by the fact that it is assumed "that ALL translators in the world are using the English interface of their software." He goes on:

Nobody would assert that a doctor is not entitled to medical care when he or she needs assistance because he or she is a medical care provider!

He's right, of course. However, I'd still like to defend Trados and the other TEnT makers a little (probably as wise an endeavor as sitting down in a bush of stinging nettles). Let's put it like this: TEnT makers tend to be aware of the need for translations of their software, and most have it translated into a number of languages. It's also true that most don't have all the components and/or documentation in all the languages that the main application is translated into, but . . . it ain't easy. On the one hand, we expect the tool vendors to implement changes quickly and frequently, necessitating new user interface components as well as documentation. On the other hand, we want those translated into a number of languages ASAP, which naturally would make the implementation of those features a lot slower.

I am often asked by colleagues whether I want to translate my Tool Box book or even this newsletter. So far the answer has always been no: considering how frequently I upgrade the content of both, it would create an unbearable administrative overhead to keep the translation coming. This may be neither a politically correct decision nor a very business-savvy one, but I think that many tool vendors are in the same leaky boat.

Some have decided to go a different route, most notably Star Transit which is available in Catalan, Czech, Chinese, French, German, Italian, Japanese, Spanish, and Swedish, all of which except Swedish with a complete help system in that language. Other tools such as Heartsome have chosen an interesting combo of languages with Polish, Japanese, and Chinese (not sure how Polish fits). Still others have relied on volunteers to translate their UI (I've been part of two of those teams in the past), a situation that also results in eclectic language combinations.

To me it seems (ooh, the stinging nettles are getting really thick now) that the tool vendors would do well to identify those languages whose speakers feel very strongly about having a complete product in their language--I assume this includes French and Japanese--and completely localize their products into those. Basic assistance can be provided to other language groups.

By the way, I mentioned awhile back that there is an ATA- and FIT-sponsored survey about translation tools. This particular issue is also being raised. Depending on the outcome of the survey, the tool vendors (and I) might change their minds. If you haven't had a chance to take it, I would encourage you to do so at www.surveymonkey.com/s.aspx?sm=Vestz_2fkHfp0eosNk_2bcXktA_3d_3d. It won't take more than 10 minutes.

5. Terminology Tune-Up

Rafael Guzmán asked me to mention his T-Manager again. Awhile back I recommended it as a great companion to an otherwise more automated terminology tool (or component of your TEnT), and that's really what it is. It's essentially a free (!) collection of Excel macros that allow you to:

  • flag duplicates/terminology inconsistencies/blacklisted terms,
  • generate a diagnosis on the status of the glossary,
  • compare two glossaries for missing inconsistently translated terms,
  • use terminology synchronization,
  • display the frequency in which terms occur,
  • calculate the number of words in each term,
  • generate reports and metrics on any variety of terminology-related things, and
. . . many other things. To describe it in one sentence, I would say it's a great, perhaps slightly geeky wellness center for terminology databases/glossaries.
6. Spell-Czech

I recently found two interesting items concerning spell-checking that might seem to cancel each other out.

The first is a nifty little tool that allows you to list and add up spelling errors in a Word document in one fell swoop. I was struck by how useful this tool would be for evaluating translations, if only it weren't for this second item I found:

Mary Maloof sent me a link to the site of University of Washington professor Sandeep Krishnamurthy (pronounced "Sandeep Krishnamurthy" as he mentions on the site) where he demonstrates "the Futility of Using Microsoft Word's Spelling and Grammar Check." This is great fun to read. He focuses mostly on the grammar checker, which he dismisses as (almost) completely useless unless you write well to start with:

As a result of my testing, I am convinced that this feature works well for good writers and not for bad ones. Good writers follow most of the rules and this feature can help them on the margins. If you are a bad writer with a poor understanding of the rules, this feature will not help you at all. This is, clearly, a problem. The feature does not help those who can most benefit from it.


It's good that most of us know how to write well in the language that we translate into. (That still doesn't mean that I use Word's Grammar Check, though.)
The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here is a reader who used that code on his website:

www.translab.gr

© 2008 International Writers' Group