Tool Box Newsletter Logo

 A computer newsletter for translation professionals


Issue 13-5-222
(the two hundred twenty-second edition)  
Contents
1.Getting Physical
2. Equipping Yourself
3. EasyMT? (Premium Edition)
4. Cross-checking
5. And Talking About Morphology . . .
6. Coming Up
7. New Password for the Tool Box Newsletter Archive
The Last Word on the Tool Box
Coming A Long Way

My apologies for the tardiness of this 222nd newsletter.  It's the first time I've been quite this late, but the good news is that it will now be only a couple of weeks until number 223!

I had hoped to write this while traveling during the last couple of weeks, but that didn't work out exactly the way I had planned. (You think it's hard to work on planes? Try doing it in a six-foot-seven-inch / 2-meter body!)

One place I visited was Kiev, where the first-ever Ukrainian translation conference was held with great gusto and 350 registered participants! An amazing number for a place without any existing organization on the ground. This would have been impossible without the indefatigable efforts of organizer Konstantin Dranch (a maker and shaker who is also the man behind translationrating.ru and Mozgorilla) and the pent-up excitement of the budding translation industry in Ukraine, an energy that made itself felt throughout the event.

Another conference that I attended was the ITI meeting in London. Though not as noisy and exuberant as its counterpart in Ukraine, it was also thoroughly enjoyable and had an impressive caliber of attendees. It was at the London conference that I shared this recent discovery.

Most of you know that many literary translators have long complained that publishers don't feel it's relevant to have the translators' names on the covers of books. In fact, this oversight led CEATL, the Conseil Européen des Associations de Traducteurs Littéraires, to offer a collection of book covers where translators are mentioned.

As I was rummaging through my in-laws' library the other day, I noticed a 1962 edition of Grass's The Tin Drum (pictured on the left). It stands in remarkable contrast to a 2009 edition of the same book (on the right):

In itself this probably does not qualify as a surprise, but here's the earlier version's imprint:

Translated by whom exactly?

(The translator, by the way, was Ralph Mannheim.)

There are many things we can criticize about the translator's position today, but based on this example, we seem to have come a long way, wouldn't you agree?

And how to propel us even further? Here's an idea: Give your clients, friends, and everyone you know a copy of Found in Translation.

Rafa Lombardino said it better on Twitter the other day than I ever could:

"It's about time I can recommend a book so people will know what it is that I do for a living! :-D"

Exactly!

1. Getting Physical 

My mother-in-law Donna is a remarkable woman. She is wise beyond her years, she is full of creative and unconventional ideas that defy her age, and she is very intelligent and highly practical. The quilt that she hand-made for our wedding many years ago is one of the most beautiful and meaningful creations I've ever seen.

However, when it comes to her computer, which would seem to present the perfect opportunity to use her plentiful creative and organizational skills, an impenetrable wall separates her from it. The digital divide that she experiences is more like a digital abyss, an abyss that seems to deepen rather than become more manageable every time she tries to bridge it. As a quilter, gardener, and pianist, if only she could actually physically get her hands around the applications she is working with, she would be great.

I think this lack of tactility with our computers is exactly what blocks many of us from success. It's what prevents us from being truly confident and efficient. We may have built ourselves tenuous, swaying bridges of vines to span the divide between the computer and ourselves, but few of us beyond the age of thirty are able to ignore the gap completely and walk straight into the digital world and its virtually tactile experience.

In recent workshops I've given for translators, I brought along an odd sculptural toy I've had since my youth, a network of chunky, interconnected wooden joints that can be twisted into unlimited patterns and forms. It really has no rhyme or reason, but I love to see how people are drawn to it, how they start playing with it almost unconsciously, trying to create shapes and taking pleasure from its fluid, ever-changing movement. In my workshops I've challenged attendees to recreate a certain figure that I love to make. There is no trick to making it; you simply need to gently force your will onto the toy until it gives way to that shape. I tell the attendees that's how we need to work with the programs in our computers. Don't be overwhelmed with the many different options and the apparent complexity of your translation environment tools or whatever you primarily use. Try the virtually tactile approach (and make sure to wipe those fingerprints of your screens afterward). 

ADVERTISEMENT

Terminotix acquires Portage, an accurate machine translation system.

Align your documents with the most powerful alignment tool: AlignFactory.

New resources added to the free Terminotix toolbar. Download it now!

Wonder why we call LogiTerm the Swiss Army knife of CAT tools? Ask for your free trial.

Don't miss our videos on YouTube on our products!

2. Equipping Yourself 

I can't wait to teach a three-day translation technology workshop at the University of Maryland on July 25-27. It's going to be very hands-on and practical and will cover all computer-related essentials that you'll need to be familiar with as a translation professional -- no matter whether you are a beginner or have been translating for a long time. The ATA has agreed to approve the workshop for 10 Continued Education points for participants, and the university has lowered the price to just a little more than $300.

The application deadline is July 1, but academia needs ample lead time for planning, so I know that they would like to see some early enrollments. Please sign up soon if you think this would be helpful for you and your career.

Two weeks prior to my course, Lynn Visson will also teach a three-day workshop on conference terminology and procedures 

3. EasyMT? (Premium Edition)

Some of you will remember Tony O'Dowd from his days at Alchemy Software, the company that develops tools such as the localization tool Alchemy Catalyst and a number of other lesser-known products.

Most who knew Tony were not too surprised when he unveiled KantanMT as his new project sometime after he left Alchemy. You see, he's a rather intense person, and imagining him not selling something just didn't feel right.

KantanMT is one of a number of products/companies -- including LetsMT, PangeaMT, tauyou, DoMT, or Asia Online -- that allow you to build a machine translation engine on the basis of your translation memory(s) and possibly some other data. Other services, such as SDL BeGlobal, not only build the engine but bring it right into your translation environment (see the 219th edition of the newsletter); still others, such as Microsoft Translator Hub, let you build machine translation engines in exchange for your data (see the 211th edition).

Like a number of its competitors, KantanMT is built with the open-source statistical machine translation engine Moses, but since it's primarily aimed at the midsize LSP market, it attempts to hide the dreaded technical aspects of that engine. So far, KantanMT boasts 400 clients in the trial phase due to end on June 1. (I meant to say "boast" since I talked to at least one LSP who was enraged to learn he was listed as an early adopter on the KantanMT website since he had dropped the product after giving it one test run.)

The idea behind this product is this: You upload your TM(s) (the typical size is 4-6 million source or target words, but if it's a highly specialized domain it can be as little as 1 million words) to the KantanMT site. The KantanMT engine -- which is hosted by Amazon Web Services -- performs some data cleaning and then processes the data to build a machine translation engine. This can be tagged on top of one of the "stock" engines that KantanMT might offer for that language combination and domain and is then used to translate files that you also upload to the site.

These files are translated by KantanMT in two passes -- first by leveraging it against the translation memory and then, if there is no fuzzy match above 85%, by machine translating the segments. Unfortunately, that threshold value can neither be changed nor is it apparent whether a segment was translated by translation memory or machine translation (well, chances are it will be quite apparent by the level of quality, but there is no formal way of filtering them). But that's supposed to change.

The file formats that are supported include bilingual formats (TMX, XLIFF, Wordfast TXML, and Trados TTX) and some monolingual formats such as Word and Excel documents, InDesign, HTML, and XML files, but the vast majority of files that have been translated are -- according to Tony -- TMX (translation memory exchange) files. This means, of course, that most users who have used the service so far are using it as an intermediary step before importing the resulting machine-translated TMX file into their internal translation workflow.

According to Tony, "clients know about the quality of their data" that they upload, so the focus of the whole process "is on quality of data, not quantity." Unlike him I would tend to say that midsize LSPs often do not know what the quality of the data in their translation memories is, but that's admittedly not a great message.

I don't mean to sound too negative about the product, though, because once you winnow through all the marketing fluff it is possible to build a relatively inexpensive machine translation engine for a tightly defined domain for certain projects and language pairs without having to be a technical expert. You will want to be an expert on your data, though. The 12 "data cleaning steps" that KantanMT performs on the translation memory might be good and fine, but they don't magically turn manure into gold -- so make sure that you start with an adequate data source.

When I inquired about other vendors that they were expecting to compete with, companies like SDL and Asia Online were waved off as "consultancy-led systems that come with a price." Only Microsoft's Translator Hub was named as something to be watched closely. Compared with SDL, I don't think the price is so much of a differentiator but it is true that the KantanMT team's energy will bring customized MT services to the forefront of more people's attention than ever before. 

ADVERTISEMENT

NEW! SDL ACADEMY

We understand that the modern translator needs to have a wide range of skills from business acumen to IT proficiency.

The NEW SDL ACADEMY has been developed as an online hub of information and learning resources for freelance translators, helping you focus on your professional development, to enhance your skills and help you find new ways to grow your business.

JOIN THE SDL ACADEMY TODAY

Are you looking to invest in translation memory software or considering an upgrade? Try the market-leading SDL Trados Studio 2011 free for 30-days. 

4. Cross-checking . . .

. . . could mean "to obstruct in ice hockey or lacrosse by thrusting one's stick held in both hands across an opponent's face or body." That sounds terrible! So what about: "to check (as data or reports) from various angles or sources to determine validity or accuracy"? Much better!

And it is indeed the latter that CrossCheck excels in. CrossCheck is a very clever and very free, completely browser-based quality assurance tool with rather amazing abilities. It was developed by Japanese translation company Idioma as part of the development of their own server- and browser-based translation environment tool iQube (which is not available for translators outside Idioma's own network of translators).

I had a chance to talk to Steen Carlsson, who runs Idioma's production center in Prague, about the genesis of his company's approach to quality assurance. In Japan, he said, many publishing houses have added translation services to make up for the near-collapse of the traditional publishing industry. Naturally, many of these entities did not have the necessary expertise in translation, so CrossCheck as a spun-off component of iQube was intended to make clients aware of some of their new competitors' lack of expertise (sneaky!) and at the same time allow clients to verify the quality of files translated by Idioma themselves (sneaky2).  In the process, the tool was opened up to the rest of the world. As a non-Idioma client, the only thing you will have to bear is the omnipresent offer of Idioma's paid services to fix the problems in the files that you have QA'ed  with their tool -- an offer you may politely refuse.

So, what exactly does this tool do? It allows you to upload two different types of files: bilingual files in the formats (SDL)XLIFF, TTX, and TMX that are to be quality-checked; and glossary files such as Excel files, SDL Trados termbase files, and TBX files that will serve as terminology resources to verify that correct terminology is being used.

The list of checks will sound familiar to you if you are familiar with what QA tools do. They include checks for untranslated text, empty translations, duplicated words or phrases, terminology adherence, spaces, capitalization, symbols and punctuation marks, numbers, brackets and parentheses (they call those "fences"), tag mismatch, languages-specific quotes, etc. -- you can find a complete list right here. Depending on what kind of checks you would like to perform, you can choose from a variety of profiles, and the results are collected in a number of ways.

You can download the results as a report in Excel or Word and make changes to the files in question on your computer, or you can view the errors right in the browser and actually perform edits in the browser. These edits will be reflected in the XLIFF, SDLXLIFF, TTX, or TMX files, which you can then download. I would not recommend the latter process because in the tests that I ran, the tool reformatted the TMX file that I tested with quite extensively.

The report option is very helpful, though, particularly because of the terminology adherence check. While most translation environment tools have a component that does something of that nature, the results are often useless because of the many false errors that are reported. Shockingly and frustratingly, most TEnTs are still flying blind when it comes to language-specific morphology. If the tool can't recognize that the plural form of a term might be the correct term even though you only have the singular form in your glossary, or that the genitive and nominative forms are just that -- different forms of one and the same term -- fuhgeddaboudit.

This is where CrossCheck shines, since it uses specifically developed morphology engines for the major Western European languages (sans languages like Maltese, Irish, etc.), plus Czech, Hungarian, Romanian, Russian, and Turkish. Cool, huh?

Now you should not expect this to always work flawlessly, but compared to the internal engines of Trados, memoQ, or Déjà Vu, it'll be great.

That brings us to languages. Internally at Idioma, CrossCheck is used for double-byte languages as well (after all, the majority of Idioma's clients are Japanese); for you and me, it's presently available only in Latin- and Cyrillic-based languages, but I was told that this is going to change soon.

Oh, I did notice a couple of things that I didn't like about the tool. One, the reformatted source files, I already mentioned. Another is the apparently easy-to-use user interface, which throws up some error messages that are anything but easy to understand -- such as "dehydrating value," the message I received to indicate I had uploaded files that did not perfectly align with the tool's very strict internal definitions of XLIFF or TMX files. Still, I know that I will run some of my projects through this tool from now on out. 

5. And Talking About Morphology . . . 

In the last newsletter I had a rather lengthy report on the free and open-source tool OmegaT. Shortly after that report was published, a completely new version of OmegaT was released (version 3.0). I will refrain from once again talking in great detail about OmegaT (you can find a good overview of the new features right here), but I would like to point out that in this new version, so-called tokenizers that provide better morphological recognition in termbase and TM recognition are now automatically installed. The supported languages include Arabic, Portuguese, Chinese, Japanese, Korean, Czech, Dutch, French, German, Greek, Persian, Russian, Thai, Danish, English, Finnish, Hungarian, Italian, Norwegian, Romanian, Spanish, Swedish, and Turkish.

Aside from OmegaT, to my knowledge there are only two other tools, XTM and GlobalSight, that also use the same technology. If anyone could help me understand why other tool vendors seem to not be interested in adding features like that, I would be obliged.

(Oh, and in case you think that I have an agenda here, you're right: I would love to see the technology that you and I use day-in and day-out become more intelligent.) 

ADVERTISEMENT

Measure and improve the quality of translations:  

memoQ 2013 will be released on 31 May, 2013

Linguistic Quality Assurance models, discussions, edit distance, fuzzy terminology lookup, improvements to Language Terminal, a set of file filter formats and much more are headed your way.

Register for the introductory memoQ 2013 webinars at http://kilgray.com/news-and-events/webinars. We look forward to e-meeting you!

6. Coming Up . . .

Here are some things that I would have loved to talk about this time but that did not quite make it in. But again, it's only a couple of weeks until the next newsletter. There you will read about:

  • The open-source alignment tool LF Aligner. Here's a glimpse of what it can do:
  • The new version of Kilgray's memoQ and its LanguageTerminal, which now also offers project management  functionality.
  • The "Swiss Post-Editing Score" for automated quality assessment of translations.
  • MultiQA, a Russian terminology management portal with support for inflectional languages (!) and integration with quality assurance.
I know! I can't wait either. 
7. New Password for the Tool Kit Archive

As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to May 2008.

You can access the archive right here. This month the user name is toolbox and the password is beachwalks.

New user names and passwords will be announced in future newsletters.

The Last Word on the Tool Box Newsletter

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box newsletter. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

If you are subscribed to this newsletter with more than one email address, it would be great if you could unsubscribe redundant addresses through the links Constant Contact offers below.

Should you be  interested in reprinting one of the articles in this newsletter for promotional purposes, please contact me for information about pricing.

© 2013 International Writers' Group