ToolkitSmall

A computer newsletter for translation professionals

Issue 10-1-158
(the one hundred fifty-eighth edition)
Contents
1. Google Releases Google Transliteration IME
2. Get Those Terms Out of There! (Premium Edition)
3. Multilizer's Multilingual MultiMT Has Multicultural Multipurposes
4. Evaluation Criteria
5. Energizer-Bunnying Your Laptop
6. More Free Alignment, Foreva
The Last Word on the Tool Kit
Quality

You wanna hear something that will make you feel really uplifted in regard to the quality of your work?

Check out this quote from Leonardo da Vinci:

I have offended God and mankind because my work didn't reach the quality it should have.

Hmmm. How do you feel now about the quality of your last deliverable?

But then, in the brave new world of machine translation and crowdsourcing (now "sharecropping"), quality has been replaced by the concept of usefulness. Or has it? Or, phrased differently, how much difference does this distinction between quality and usefulness make for the projects that you and I have delivered last week?

 

Oh, Jeromobot? He's been thrown into a real identity crisis since being asked whether he is a mobot or a robot. Or something else entirely. We'll let you know. In the meanwhile, visit him on Twitter.

1. Γοογλε Ρελεασες Γοογλε Τρανσλιτερατιών ΗΜΕ / गूगल रेलेअसेस गूगल ट्रांस्लितेरातिओं इमे

(Google Releases Google Transliteration IME)

Just yesterday Google released IMEs (Input Method Editors) for 14 languages (Arabic, Bengali, Farsi, Greek, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Punjabi, Tamil, Telugu, and Urdu) that can be downloaded for free. Once installed, they become one of the keyboard languages you have available in Windows (just like the different languages you can install via Windows that are accessible through the small language icon on your taskbar). The concept is that you have easy access to those languages by being able to use a Western-language keyboard, type the pronunciation of the word in the respective language, and the IME will give you a number of auto-complete suggestions of alike-sounding real (or, in the case of my heading, unreal) words in that language with the language-specific characters you can choose from (very much like Windows IMEs for Chinese, Japanese, or Korean).

If you don't deal with these languages, it's fun to play with the tools. If you do deal with the languages, you might have more reasonable use than just play, but I would caution you a little. Though Google is usually very careful not to release something out of beta for many years, these tools may still be very beta-ish: I've crashed Outlook a number of times today while trying to use them in an email message.

The IMEs work for Windows 7, Vista, and XP.

2. Get Those Terms Out of There! (Premium Edition)
Now for those among you who are tired of hearing me talk about term extraction again and again when really this technology has still not found the breakthrough that it deserves: Give me a break. First of all, there are a whole slew of new subscribers who have not had the pleasure. And second, maybe there will be a breakthrough that will make this a much more widely used technology/practice. And if so, by the sheer logic of repetition, I will gladly take some of the credit!
The concept of term extraction is, of course, the ability to extract mono- or bilingual terminology from document(s) or databases to quickly create glossaries that will aid you with your translation projects. One reason that the termbase functionality is still so crudely underused with most tools is that it is tedious to use (and though it's actually no longer tedious to use in current versions of tools, it used to be tedious, and the user's mind still has it so classified). And even if it is easy to quickly enter terminology into your termbase or glossary as you translate, it means an interruption in the translation process, something that would be nice to avoid.
So, wouldn't it just be great if we could spend a couple of hours before a large project either harvesting terminology from existing projects of the same subject matter or quickly creating lists of source terms that are relevant to our project and translating those ahead of time? (Of course, this all becomes much more relevant and important when you are faced with multi-translator projects.)
Let's start with the no-brainer solution of extracting bilingual data from existing sets of translated documents or databases (typically in TMX format).
  • The most powerful application in the field of term extraction used to be the Xerox Terminology Suite (XTS), which was designed for the deep pockets of corporate users and was very powerful because it was based on preconfigured linguistic data in various languages. Today the suite is owned by TEMIS, where development has virtually (and literally) come to a halt. However, the translation environment tool Similis has integrated the XTS engine and therefore comes with a very high-level linguistic "knowledge" in seven EU languages (English, Dutch, German, Spanish, Italian, Portuguese, and French). Similis is able to apply a combination of linguistic and statistical rules to a number of processes, including automatic extraction of terms and phrases from translation memory content, with extremely high accuracy -- but unfortunately only in a handful of languages.
  • SDL offers two separate programs (MultiTerm Extract and SDL PhraseFinder -- they are sold as a bundle) that extract existing terminology and build up terminology databases or glossaries and present you with a proposed translated terminology list. MultiTerm Extract, the tool that originally comes from the Trados side of things, works on a purely mathematical level ("if word A always appears in sentences for which word B always appears in the translated sentence, then these words must form a word pair"). This means it supports all Windows-based languages. PhraseFinder, the former SDLX companion, works on a language-based level for English, French, German, Spanish, Dutch, and Portuguese. This means that overall all languages are supported, but users of the languages that are supported by PhraseFinder have drawn the longer straw since the recognition will be more accurate. (On the other hand, the PhraseFinder process is very resource-intensive and not particularly fond of large amounts of data.)
  • MultiCorpora's MultiTrans has always offered the extraction of monolingual term lists. With its latest version it added the WORDAlign feature that internally creates bilingual term lists to improve the accuracy of the alignment, but then can also be extracted as separate termbases.
  •  Another terminology extraction tool is SynchroTerm. SynchroTerm theoretically supports all languages; however, practically speaking there are different tiers of language support. In general, SynchroTerm relies on mathematical calculations to extract terminology pairs. For English, French, Spanish, German, Italian, Portuguese, Swedish, Russian, Greek, Polish, Turkish, Dutch, Hungarian, and Norwegian it also uses lists of stop words to filter those out automatically, and for English and French it also makes use of stemming rules, further improving the accuracy in those languages.
(As a side note: On the recommendation of Jan Dolejs I've run a few tests with another, very promising-looking tool called ParaConc for bi-, tri-, or quadrolingual term extraction, but I did not get very far -- maybe it's because I tested on Windows 7. If any of you have more luck let me know.)
These are the ways to semi-automatically create bilingual glossaries. Of course, if you start with just source documents, there need to be ways to extract just source terminology. You have two choices for that. You could use an integrated feature like those offered by Déjà Vu, Swordfish's Anchovy, and MultiTrans. These features create an index of all terms and phrases in the project and allow you to choose how long the phrases are to be, how many occurrences they are to have, etc. This is very helpful, and if at first it seems that there is a lot of useless material extracted, it's up to you to find good workflows to quickly locate the good stuff and delete the rest.
There are also external tools. These are called "concordancers." You can find a list of concordancers on the respective Wikipedia site or you can Bing for it (that was kind of fun to write!). You will quickly see that most of them come from academia -- clearly there is an interest for linguists to be able to analyze large corpora of text for the actual usage of terms and phrases. But there is also strong interest for us. Many of these concordancers are language-specific, which means that they come with information on what kind of terms or combinations of terms to ignore.
Yesterday I tested TermExtractor, a concordancer from Italy that curiously works on English texts (a German test that I ran failed spectacularly). This is what it says on their site about the tool:
The software helps a web community to extract and validate relevant domain terms in their interest domain, by submitting an archive of domain-related documents in any format. TermExtractor uses a novel method to extract terminology consensually referred in a specific application domain. The software takes as input a corpus of domain documents, parses the documents, and extracts a list of "syntactically plausible" terms (e.g. compounds, adjective-nouns, etc.). Two entropy-based measures, called Domain Relevance and Domain Consensus, are then used to select only the terms which are relevant to the domain of interest and consensually referred throughout the corpus documents. Domain Relevance is computed with reference to a set of contrastive terminologies from different domains. Finally, extracted terms are further filtered using Lexical Cohesion that measures the degree of association of all the words in a terminological string. Furthermore, TermExtractor is able to recognize in any type of document (txt, pdf, ps, dvi, tex, doc, rtf, ppt, xls, xml, html/htm, chm, wpd) the following text layouts: title, bold, italic, underlined, colored, capitalized, smallcaps, and to assign to terms with selected layouts a greater importance.
So, the extraction is a lot more intelligent than simply indexing all terms and phrases in the project, and I was indeed very impressed with the results of the file(s) that I uploaded and the results I received. I was a little concerned that it seems to focus exclusively on two-word phrases, but that may be fixed in later versions. Oh, yes, and it's free.
And as we are talking about terminology, CSOFT, the folks behind L10NWorks, a site I've written about in the past that collects all kinds of tools related to localization processes, has announced a new project: "TermWiki."
They use a lot of big words to describe what it is going to be. (They like big words: One of their discussion threads is introduced like this: "Ever heard of L10N, I18N, or G11N? Look like a bunch of gibberish? Well it's not." I happen to think it is.) Nonetheless, it looks like it's going to be a collaborative terminology management system with an easy-to-use user interface. We shall see. I would not even have mentioned it if not for the participation of Uwe Muegge, formerly of Medtronics and currently teaching in Monterey. With his participation, chances are that the results will be quite interesting.
ADVERTISEMENT
Wordfast, the world's leading provider of platform-independent TM software will bring you the NEW Wordfast Translation Studio in the coming weeks, which will include two powerful tools for one low price.
 
COMING SOON IN WORDFAST CLASSIC: The #1 MS Word-based translation memory tool  
  • Auto-complete feature
  • TXML, PDF, and TTX support 
COMING SOON IN WORDFAST PRO: The #1 multi-platform TM tool designed for translators and agencies alike  
  • Wordfast alignment tool
  • TTX and MIF support
  • Machine translation integration
  • and more... 
Learn more at www.wordfast.com
3. Multilizer's Multilingual MultiMT Has Multicultural Multipurposes

I was thinking of attempting to write a heading using only words that start with "multi" -- a plan that was foiled when I found out that there are more than 200! (Can you believe that? My favorite was "multimorbid," but I just couldn't find a way to make it fit.)

Anyway,

Multilizer, one of the smaller localization tools, introduced its new Multiple Machine Translation Technology -- or MultiMT -- technology. The idea is to pull machine translation from various sources and then automatically compare the different translations to "evaluate" their quality. It sounds really clever, but the principle behind it is rather simple:

MultiMT technology is based on this simple principle: whenever several machine translation engines return similar translations, those are likely to be appropriate.

To be sure this sounds a little simplistic to me, but if it works, hey, we'll take it. After all, the concept of translation memory is not that complex either, and most of us use it and benefit from it (most of the time). 

4. Evaluation Criteria
At the Interpreting the Future conference last October in Berlin, one of the speakers, Dino Azzano, presented his comparative research on various TEnTs (Across 4, Déjà Xu X, Heartsome 7, memoQ 3.2, MultiTrans 4.3, SDL Trados 2007, Transit 3.1, Wordfast 5.53). Astute readers will have quickly realized that in most cases these are not the most current versions, but his research is still interesting, if only to see what kind of criteria can be used to evaluate different tools. He looked at the treatment of non-textual elements in TEnTs ("Nichttextuale Elemente und ihre Bedeutung in CAT-Systemen").
The areas that he looked at included
  • Numbers
  • Dates
  • URLs and email addresses
  • Punctuation marks
  • Proper names
  • Inline graphics
  • Fields
  • Tags
    (The results for the last three categories were not in by the time he had to submit his paper for the proceedings so he does not say much about those.)
    Here are some of his results (as far as I know there is no online version available, so I cannot give you a link, but I am sure that Dino will pass that on to me once it becomes available.)
    • Numbers: All TEnTs recognize numbers in general and offer quality assurance features to verify the correct transfer of numbers between source and target. However, not all TEnTs offer an automatic conversion of decimal systems with numbers (my personal gripe: even Transit, which is an advanced system on so many levels, still does not do that). And even those TEnTs that do offer this feature don't work with a non-standard numbers (such as 1'000), which means that none offer a customizable system where the user could determine what needs to be converted to what. In cases of mixed numbers and letters (10KB), most tools also give up, with the laudable exception of Wordfast, which in almost all cases recognizes any number fragment as a number.
    • Dates: Numeric dates (10.1.2010) are recognized by most as dates, with the exception of Heartsome and memoQ, and alphanumeric dates (January 10, 2010) are recognized only by Trados.
    • Email addresses and URLs: If these are not part of a field (i.e., clickable), they are recognized only by Wordfast; otherwise, all tools recognize them. Only some, including Déjà Vu and Transit, also display the URL or email address in an editable manner. (I find this rather helpful, in particular when it comes to URLs, which often do have to be changed.)
    • Punctuation marks: These include all non-numeric and non-alphabetic characters. All TEnTs do fine with these, except when it comes to differences with spaces (protected vs. non-protected), where Heartsome, MultiTrans, and Trados flake out.
    • Proper names: Here the paper looks at names that are structurally different, such as in all caps (IEEE), mixed lower and upper case (JavaScript), alphanumeric sequences (A4), or with some special characters (%VERSION%). The two systems that do best with these are Trados, which at least recognizes words in all caps as proper names, and Wordfast, which tries to recognize all of the above.
    The paper then goes on to talk about how the different systems penalize differences in the above-mentioned categories (Transit is particularly rigid about penalizing for differences in end-of-the-segment punctuation) and convert them (memoQ, Wordfast, and MultiTrans and in particular Déjà Vu are good about stepping in for differences in punctuation) and on and on.
    In my opinion, this is helpful for suggesting some good parameters to look for when evaluating tools. Remember, some of the tools in this study are quite outdated, so don't throw tomatoes at Dino or me if you find that things have changed a bit.
    5. Energizer-Bunnying Your Laptop
    Here are some helpful insights about how to extend the life of a laptop's battery (I found these in one of Fred Langa's articles in the WindowsSecrets newsletter).
    Here are some of the highlights:
    • When your laptop is plugged in, it might be wise to remove the battery and store it in the fridge (make sure you wrap it tightly in a plastic bag before you do that). Also, if you run it down to about 40% charge before storing it, you'll get the best results.
    • Unlike other batteries, Li-ion batteries should not be run all the way down. They last longest when kept between 40% and 100% of their charge.
    • Check the date of manufacture before buying replacement batteries. The older the battery, the less life you'll get.
    Did you know all that? I sure didn't.
    6. More Free Alignment, Foreva

    The alignment tool YouAlign from the same company as the above-mentioned SynchroTerm has now been declared free for good (it was originally supposed to be for a limited time). It's a fantastic offer -- it uses the same engine as the commercial counterpart AlignFactory; it offers the same wide range of file formats, including PDF; and the only limitations are file size (1 MB) and that you can align only one file pair at a time (the latter limitation can easily be circumvented with file concatenators -- I use Twins File Merger).

    The Last Word on the Tool Kit

    If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

    Last week these readers added a link:

    www.languagestoday.com/translatorzone.html

     

    © 2010 International Writers' Group