Tool Box Newsletter Logo

 A computer newsletter for translation professionals


Issue 13-10-227
(the two hundred twenty-seventh edition)  
Contents
1. Translating With the Big Guys (Premium Content)
2. The Veterans
3. MateCat (Premium Content)
4. New Password for the Tool Box Newsletter Archive
The Last Word on the Tool Box
Engagement

In my last few presentations, I've talked about translation technology (surprise!) and the effect it has on the world of translation. And a world it truly is. The focus of the last couple of years on technologies like machine translation has revealed that there really is no translation "industry." (And I'm the first to admit that I've used that term more than enough.)

Merriam-Webster defines an industry as "a group of businesses that provide a particular product or service." As an English-into-German translator, I do not provide the same service as translators of other language combinations, or even as other translators of my language combination with other fields of expertise or different levels of experience. And I most certainly don't provide the same service as a machine translation engine can provide, post-edited or not. Finally, if I compare my services to the services that most language service providers offer, I neither compete nor can I (or do I desire to) provide the same service.

Why is all that important? Well, for one thing, we like words and we should make sure that we get them right. But more important than that is our attitude toward the processes and people around us. The services provided by others who are also formally included in our "industry" might in fact be completely irrelevant to us and our translation work. Their work is definitely interesting, and we're well advised to take heed. But we should pay attention for the purpose of enriching ourselves rather than fretting about potential competition or threats to our livelihood.

I've recently been comparing us to a Chinese hand fan. Language, our common denominator, is like the pin at the bottom of the fan that holds together the bamboo segments (in our metaphor, this represents the different sectors). But from that one common point, the various parts diverge into many different directions, really without any overlap. 

Skeptical? Well, it works for me (especially -- and now I'm stretching it a little -- because the fan might just help us to cool down a bit from all the agonies of conflict).

Here are some things that all of us could actually do together, though:  Last week, on International Translation Day, Google did not have an appropriately themed Doodle (those fancily adorned Google logos to celebrate special occasions). Wouldn't it be great if there were one for next year's Translation Day? Anyone with great contacts at Google? If we all pull the strings we have, we can make our influence felt.

Also, Kevin Lossner contacted me last week about the woeful state of many translation-related Wikipedia articles. We've been tossing around a few ideas about how to improve these. One possibility would be forming a group representing the pin in the fan (i.e. the common denominator). If you have ideas, send them to me or tweet them with the hashtag #wikitrans.

And lastly, TAUS, the Translation Automation User Society, has launched a one-question survey asking whether translation is becoming a utility on the occasion of their conference themed "Translation is becoming a utility" next week in Portland, Oregon. I for one do not agree with that assessment, and I encourage you to make your voice heard, as well. You can find a link to the survey on the left-hand side of this page, or you can go directly to the survey right here. And if you choose to say more than just yes or no, send an email to translation.is.a@gmail.com. Alan Melby will represent the ATA at that meeting, and he will base his presentation partly on our responses.

1. Translating With the Big Guys (Premium Content)

Translation and language are very, very close to the center of the endeavors of companies like Google and Microsoft. In fact, Microsoft has recently released a flurry of announcements relating to translation technology.

One of the primary resources for translators who work in software translation, especially with products for the Windows platform, have been the so-called Microsoft glossaries. The "glossaries" really were large translation memories with the translation data of the user interface for many of Microsoft's software products. From 1994 through the summer of 2006 they were available for free on one of Microsoft's FTP sites. In 2006 these files were replaced with a multilingual glossary, which in 2009 was replaced by the Microsoft Language Portal. The portal offers an online search interface to Microsoft glossaries and translation memories, access to the Microsoft style guides for many languages, and the ability to download extensive glossaries in the termbase exchange (TBX) format.

A couple of days ago, the group within Microsoft that runs the Language Portal released the Terminology Service API, which essentially allows developers to integrate that service right into their tool. Now it can be only a matter of weeks until we see this show up in most translation environment tools. Be sure to say thank you to Microsoft's Palle Petersen for that.

The primary tool that this is already embedded in is Microsoft's simultaneously released Multilingual App Toolkit for Visual Studio (MAT) -- a very simple XLIFF editor for the translation of apps that connects to the MS's terminology and TM resources (you can find a good description of the editor here). Whoa, I thought, this is so cool, especially since I had had a conversation with a couple of folks just days prior to its release that a tool like that is sorely needed. But in my excitement I had forgotten how fickle XLIFF can be. In fact, not a single XLIFF file of the different flavors I tested could be opened in the MAT XLIFF editor. Most of the time the editor gave me some information why it couldn't open the files, but that still wouldn't qualify it for the easy-to-use all-around XLIFF editor that I was clamoring for.

(I had a conversation with Palle and two fellows from the Windows group who are responsible for the editor later that day, and they were kind enough to look at the XLIFF files that were generated by Trados Studio, memoQ, and Déjà Vu respectively and modify them so they could be edited. But they also noted -- understandably -- that their prime responsibility was to develop the editor to work with the Microsoft development environment and not to spend too many resources to build it so it would work for anyone and everyone.)

Still: if you are developing apps in Visual Studio or if you are working for someone who develops apps in Visual Studio, the MAT XLIFF editor is great. It's not a translation environment tool with all the bells and whistles -- you can leverage from other translated XLIFF files, but it has neither a private translation memory nor a termbase to which you could add data --  but it presents the translatable strings very handsomely and connects you to the terminology/TM resources I mentioned before (where it looks only for perfect matches). If nothing's found in there, it machine translates it with Microsoft Bing Translator (or other MT engines that you can connect). The developer can then choose to send the file -- including a download link for the tool -- to a translator or post-editor.

I asked the three guys on the phone whether they felt that this process had an impact on the translation quality of apps that were translated that way, assuming that many developers wouldn't worry about sending it off to a qualified linguist. Absolutely, they said, but in their estimate these issues in app development are handled quite differently than they used to be for large-scale programs. The developer typically will choose a crappy translation until he hears complaints by users. At that point he'll go back and put some more (human) effort into the translation.

Really?

(That was what I thought at first, too.)

But I bet they're right. It's a little scary from our point of view, but it's also exciting because it puts the end user in charge. And thanks to this tool, it's now easier for the developer to actually get it right... eventually.

 

As if this weren't already enough about Microsoft, here's some more. I wrote about the Bing Translator Widget about three years ago (then it was called the Website Translation Widget -- premium subscribers have access to this article in edition 162 of the newsletter archives). This project was finally raised out of its almost-eternal beta status and officially released last week. Remember what I said at the beginning of the newsletter: the world of translation has a lot of facets, most of which don't have much to do with each other, but they're still interesting. The Widget (a term almost as despicable as "dongle," by the way), a connector to the MS Bing Translator, sits on webpages so users can machine translate those pages. This is nothing new, and it's also not new that users can suggest better translations. What is new about it is that the website owner can use a collaborative model where he can invite certain users to contribute suggestions that will actually be used later on for that particular webpage (if approved by the owner). This works because the contributions are not only eventually fed back to the MT engine to improve it but also stored in a webpage-specific translation memory from which they are retrieved automatically. (When I talked to someone last time there were plans to make those translation memories available as TMX downloads to the website owner, but it seems that plan was unfortunately scratched.)

Here is Microsoft's description of that community translation model -- not too clearly written, I'm afraid, but who am I to talk?

So is this something that will ever have anything to do with any of us? I could actually imagine that it might be one way to sell web translation to clients so it's inexpensive to them and profitable to us. You introduce them to the Widget, they implement it and make you the preferred editor/translator, and that's it. Not the most elegant way for sure, but a lot better than unedited machine translation or no translation.And why is all this so relevant? Well, in a way these tools are the trailblazers of what we will all be doing in the not-too-distant future: working within our browsers. It's important to see how they manage it, especially in cases where some workarounds are needed. 

ADVERTISEMENT

TransitNXT - the ideal tool for translation and localization!

TransitNXT offers you an optimized translation memory and terminology management

solution that easily interfaces to all standard file formats, workflow and content management systems. Whether you use TransitNXT for translation, editing and proofreading, terminology

management, or project management, you will achieve your goals with its many intelligently designed and powerful features. All this and more in an efficient and user-friendly working environment.

To get a fully-functional trial version, please contact: transit@star-group.net

www.star-group.net

2. The Veterans

I met with two veteran tool makers in the last couple of weeks to receive a status report and explore their strategies.

 

I first talked with Star AG's CEO Josef Zibung about Star Transit. As many of you know, this is a very powerful translation environment tool, but sometimes it's not easy to find in the mix. I'm still waiting for some follow-up information that I requested, so I will hold off on reporting more about that conversation until the next newsletter, but one thing right away: Mr. Zibung assured me that he has had an awakening (that was not his actual word, but I think that's what he meant) in regard to marketing and publicity, and it will be a lot easier to find Star Transit from now on out.

 

Just yesterday I talked with Atril's new managing director Blandine Loze. Atril is the parent company of Déjà Vu, which, along with Star Transit, Trados, and IBM TM/2, was one of the first viable, commercially available translation environment tools. (Déjà Vu was actually the first Windows-based product.) Déjà Vu also was the first tool to really challenge Trados (at least in the minds of its users -- the "flame wars" between Trados and DV users have now become the stuff of folklore), and it came up with concepts that we assume as normal today (such as the tabular interface for translation) or are still cutting-edge (such as the automatic assembly of translated snippets into first draft translations).

Still, some of the excitement about Déjà Vu has cooled in recent years. There are still loyal users (including me for many projects), but many others have migrated to other tools because it seemed that Déjà Vu's speed of development and innovation could not keep up with its competitors, and many promises about release dates for new versions and features were not kept.

Now the new version of Déjà Vu X3 has been announced for release in December. Though I always welcome surprises, I wouldn't hold my breath that it actually will be released this year. Still, the new version is supposed to focus on a more ergonomic workflow, including a whole new user interface, inline spell-checking, preview, and more WYSIWYG formatting as you work on your translation.

Blandine may have been a little frustrated that I didn't fall off my chair when she mentioned those features.

Don't get me wrong: they're all nice, and I'm glad they will be in the new version, but I had hoped for something else entirely (or in addition to those features). It just so happens that Déjà Vu X2 has one of the most promising and innovative translation features on the market. Rather than trying to describe the feature from scratch, allow me to quote from a newsletter three years ago where I wrote about it:

To my knowledge, Déjà Vu X2 is the only tool out there that uses a feature called fuzzy match repair with machine translation. Fuzzy match repair has long been a specialty of Déjà Vu: Every time a fuzzy match is encountered (a fuzzy match is a translation memory match where the source segment in the text that needs to be translated is not quite the same as but very similar to a source segment in the TM), the internal processes of Déjà Vu try to identify what the differences are, whether it knows what the corresponding parts in the target segment are, and whether it can replace them with a translation that it knows.

In practice it works like this:

Imagine you have this segment that needs to be translated.

The TM contains:

Imagine you have this book that needs to be translated.

with the translation:

Stell dir vor, dass dieses Buch übersetzt werden muss.

That would be a fuzzy match. If the terminology database contains the entries

segment

and

book

Déjà Vu would automatically replace the translation of book with the translation of segment and come up with:

Stell dir vor, dass dieses Segment übersetzt werden muss.

and you end up with a repaired fuzzy match.

Naturally there are many variables that prevent this from always going so well (such as gender, etc.), but this is how it generally works. In many cases it does a great job, leaving nothing or very little for the translator to do.

This feature now has been enhanced with two processes that make it even more useful. The first is sub-segment matching, a process that virtually all translation environment tools use in some way or the other. What this would do in our example is that segment and book would not necessarily have to be separate entries anymore but, provided that they appear often enough in the translation memory, Déjà Vu would be able to extract that knowledge on its own.

The second feature is in a way a logical extension of that. Here only the old term needs to be known (in our case, book) so that it can be marked for replacement. If the new term cannot be found, Déjà Vu goes out to a machine translation engine of your choosing (the available ones presently are Google TranslateSYSTRANMicrosoft TranslatoriTranslate4, Asia Online, MyMemory and PROMT, provided you have the licenses), retrieves the translation for that term, and automatically uses it to replace just that part of the segment.

This works not only with single words but with phrases of any length as well.

Here is what I think: I don't want little improvements here and there in my tools. Sure it's great to see words underlined when they are misspelled, but I cannot imagine that I will make a decision for or against a tool based on that. What I want are features that have a major impact on my productivity -- like the one described above, only better and even more useful. Tools with a relatively small staff of developers cannot keep up with competitors like SDL, Kilgray, or Star that have significantly more developers working on their technology. So I wish and I hope that companies like Atril would focus on what makes them different, commit to staying different, and know that there will be enough users who appreciate that and put their money where their wishes are.

ADVERTISEMENT

memoQ: Join us for the conference marathon!

Meet our team at the Localization World conference in Santa Clara on 9-11 October!

Join us at tekom/tcworld Wiesbaden on 6-8 November or at the 54th ATA Conference in San Antonio, Texas on 6-9 November - where you will also be able to take part in the first-ever humanitarian Translate-a-Thon: Volunteer translators are needed on the spot for a large-scale, transcontinental charity project that Kilgray-memoQ organizes together with Translators without Borders for the International Federation of Red Cross and Red Crescent Societies.

Visit now and register!

Check out our event calendar for more events this autumn. 
3. MateCat (Premium Content)

As any cat owner knows (my very proud self included), mating cats ain't a pretty sight or sound. MateCat, however, might be much easier on your eyes or ears.

It's been more than a year since I talked about MateCat with Alessandro Cattelan and Marco Trombetti from Translated in Italy (the developers behind MyMemory). Back then it was still in the conceptual phase, and though it sounded interesting it didn't make complete sense to me. It was (and is) an EU-funded project that was supported by a large team of very impressive caliber, including Philipp Koehn, one of the leading developers of statistical machine translation and the open-source MT engine Moses.

The goal at that point was purely academic, as a tool to research how much time post-editing of machine translation takes. While there seem to be some kinds of standards emerging for how to compensate for post-editing machine translation, in many ways it's still a nebulous affair, so tools like this can be helpful.

Now the veil has been taken off MateCat, even though, as Alessandro emphasizes, it's still in alpha, meaning that it's still rough around the edges. Still, you might want to take a look even before it is officially launched in January.

Because it's EU-funded, the core of the project had to be open-source. This is what you find when you go to this site. Essentially it's the same as its commercial counterpart right here, with the one (important) difference that the open-source version does not provide for any file filters except XLIFF. So if you were to choose to use it, you would first have to convert your Word, InDesign, or XML files to XLIFF with something like the free localization champion Rainbow. On the other hand, right now there is really no reason to stubbornly use the open source vs. the commercial version since at this point both are free -- and, again, they are really early versions, so you might not want to use them in important projects anyway.

If you stop by the sites, you'll find a very easy and spartan interface with virtually no project management facilities except the ability to select your language combination, select a machine translation engine (or MyMemory, which contains both a large TM and MT), and create your personal TM (you don't have to do that, but if you don't all your translations end up in MyMemory, which might really frustrate your client).

Once that is done, you drag your file or files over to the page where they are analyzed. The analysis uses the concept of "equivalent words," something you might know as leveraged or weighted word count, i.e., word counts that take repetitions and matches into consideration.

The table-based translation view is also extremely simple (and a little fickle at this point) and essentially only offers to select from displayed translation memory matches and MT translations or to copy source to target and jump to the next segment. That's it.

There are several things happening in the background, though. First, you are of course being timed (remember this was the initial purpose of this tool to start with), and all kinds of statistical data is being collected. I'm not sure who has access to that kind of data, but you will eventually be able to install MateCat on your own computer or server, and then no one should be able to see it (you can view that data now by selecting the Editing Log link at the bottom of the screen). Also, the integrity of the individual segments is verified constantly to make sure that no tag or other element is missing. Lastly, and that's probably the most interesting part, an immediate learning of the associated statistical machine translation engine is happening. Now, statistical machine translation by default learns from the data you feed into it, but the updates only happen at certain times when the new data is computed and read into the engine (for instance, try to change Google Translate and assume that "your" correction will show up next time -- it's not gonna happen right away, but you might see it pop up in two weeks). MateCat is able to do this through a caching procedure rather than translation memory (you can find a description of it in this very technical article).

To come back to the file formats, the commercial version essentially supports all the file formats you want (all Office and OpenOffice formats, XML and HTML, TTX and even SDLX ITD, all the necessary DTP formats as well as a good number of software development formats) and of course XLIFF. What's great is that I took the same three XLIFF files that I had sent to the Microsoft developers earlier that their tool wasn't able to process and they worked beautifully in here.

So perhaps that's what this tool is right now and maybe even will be: the easy XLIFF editor (with some extras) that we discussed on Twitter the other day.

Of course, it can be a lot more than that. And since it's open-source, I wouldn't be surprised if quite a few language service providers might look at it as a potential alternative that they can configure to fit into their workflow. 

ADVERTISEMENT

Easier, Faster, Smarter - SDL Trados Studio 2014 is now here!

SDL Trados Studio 2014 has arrived, introducing a huge range of efficiency-improving features designed to make translating in a CAT tool more enjoyable.

Here are just a few of the new features available in Studio 2014: 

  • Merge files easily
  • AutoSave your open documents
  • Drag and drop translatable files straight into Studio
  • Automatic concordance search
  • New file format support
  • Enhanced comment handling

Learn more or buy a license today.  Tool Box readers receive an exclusive 15% discount!

"SDL Trados Studio is the Gold Standard for top-notch translation processes. There are no tools that can stand toe-to-toe with SDL Trados Studio 2014 and deliver the same productivity gains and quality." -- Stefan Gentz, tracom.de

Join the conversation on Twitter -- #Studio2014

4. New Password for the Tool Kit Archive

As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to 2007.

You can access the archive right here. This month the user name is toolbox and the password is shutdown.

New user names and passwords will be announced in future newsletters.

The Last Word on the Tool Box Newsletter

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box newsletter. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here are some webpages that proudly display the Tool Box logo:

glossarissimo.wordpress.com

www.hermetic.ch/wfc/wfc.htm 

www.hermetic.ch/wfc/wfcg.php

www.hermetic.ch/wfca/wfca.htm

If you are subscribed to this newsletter with more than one email address, it would be great if you could unsubscribe redundant addresses through the links Constant Contact offers below.

Should you be  interested in reprinting one of the articles in this newsletter for promotional purposes, please contact me for information about pricing.

© 2013 International Writers' Group