ToolkitSmall

A computer newsletter for translation professionals

Issue 9-6-142
(the one hundred forty-second edition)
Contents
1. The Google Translation Center That Was to Be (Premium Edition)
2. The Truth Is Out There, Part II
3. Some Dorky Licensing
4. Flexible Clay
The Last Word on the Tool Kit
Visionary Times

Sometimes when I'm writing this newsletter, I have to sit down and ponder what to write. Usually it's not hard to come up with something that I feel is relevant, but it might take a little bit of thinking. And then there are these other times, like -- you guessed it -- NOW, when things are just happening left and right.

Think about it: Within a span of a week and a half or so, SDL releases its new tool (and creates a lot of confusion out there in the process); the TAUS Data Association (TDA) releases its long-awaited translation memory sharing and term search facility (just after Linguee, the amazing new English<>German bitext portal was released -- I discussed that last week); and, oh yes, Google also (very silently, I might add) released its Google Translator Toolkit. And nothing against the importance of any of those other news items, but that final one probably tops them all. (And, yes, I am thinking about suing Google for the rights to the name! J)

On another note, the lovely Courtney Searls-Ridge sent me a memo last week after I mentioned the good life Alexander Pope enjoyed after his translation of Homer. Apparently he's also remembered for quoting his printer on the subject of translators (to the Earl of Burlington in November 1716):

Those are the saddest pack of rogues in the world. In a hungry fit, they'll swear they understand all the languages in the universe. I have known one of them to take down a Greek book and cry, 'Ay, this is Hebrew, I must read it from the latter end.'

By God, I can never be sure in these fellows, for I neither understand Greek, Latin, French, nor Italian myself. But this is my way: I agree with them for ten shillings per sheet, with a proviso, that I will have their doing corrected by whom I please; so by one or other they are led at least to the true sense of an author, my judgment giving the negative to all my translators.

Since Jeromobot's vision is a little fuzzy lately, he could not agree more!

1. The Google Translation Center That Was to Be (Premium Edition)

Remember about a year ago when news reached the public about a "Google Translation Center"? There was a lot of resulting hoopla from translators, tool vendors, and language service providers, and the responsible department at Google took a lot of heat. Back then I managed to talk to one of the folks in charge of the program, and it was one of the coolest conversations I ever had. In about ten minutes I asked all kinds of questions, and I always got the same answer: "Yeah, I can't really talk about that." It felt a little bit like talking to someone from the CIA: "Yeah, I could tell you, but then I'd have to kill you." (Though the Google guy was actually quite nice about it.)

Well, a lot of things have happened in the meantime, not least the radical change in the economy, and this may be one reason why Google's new Google Translator Toolkit (I'm going to call it GTT to avoid confusing it with an excellent newsletter I know of!) states this:

Google Translator Toolkit is free, but in the future, we plan to charge users whose translations exceed high-volume thresholds.

This joins many other things that are quite different from the original scenario.

But first things first: Here is what GTT is. It presents you with a rather well-designed front-end that allows you to do one of two things: You can either upload a file (in HTML, Word .doc [not .docx!], OpenOffice .odt, .txt, or RTF) or you can specify a URL and the corresponding HTML page will be uploaded.

When you select the file you need to select the language pair (the only available source language right now is English (!), but there is a choice of 48 available target languages, including double-byte and bi-directional languages). Then you need to choose whether this is a "shared" translation -- i.e., whether you are using and contributing to a large anonymous translation memory or whether you would like to upload your own translation memory in TMX format. Here you can also upload or define a glossary. Should you choose to upload a glossary, it needs to be in CSV format with a strictly defined pattern, but it can have fields for part-of-speech or definition aside from source and target.

When you have made those choices the upload happens and your original file is displayed on the left pane of a split window (unless you choose horizontal panels in the View menu). On the right side you can see a pretranslated version of the file. The pretranslated material comes first from the translation memory(ies). If nothing is found there -- and this is likely since at this point they don't seem to be particularly content-laden -- a machine translation with Google Translate is performed. (If you choose to forego the machine translation step, you can select Pre-fill with source text instead of machine translation under Settings.)

Before you start the actual translation, you should click the Show toolkit button on the top of the window. This will open a new pane at the bottom of the window that contains four different tabs: Translation Search Results (TM hits), Computer Translation (results from Google Translate), Glossary (hits from your glossary -- the number in parentheses indicates whether there are any and, if so, how many hits are being found), and Dictionary (this opens a search box in which you can enter single terms for lookup -- no idea where that content comes from). According to which translation segment you highlight in the Target pane, the different tabs will show the corresponding content -- if any is found.

The file is more or less displayed in WYSIWYG format, meaning that it is displayed the way you would see it in a browser or MS Word. Only when you select an individual translation segment (as in normal TEnTs, this typically corresponds to sentences) is the WYSIWYG view of that segment replaced with a text-only view in which inline codes are displayed as numeric codes with curly brackets ({1}) -- déjà vu, déjà vu! Unfortunately, it is not possible to enter the numeric codes with keyboard shortcuts. You will either have to use the ones that were automatically entered during the pretranslation and translate around them, or you can highlight the phrase that is surrounded by codes, select Insert HTML tags,select the ones you need, and they will be placed around the phrase (I'm not sure why it says "HTML tags" independent of the source format).

Since not all text is displayed in the normal WYSIWYG view of a document (think of keywords in HTML files or footnotes in Word files), there is a list of those at the bottom of the target pane under Hidden Text where they can be translated separately.

And while you are working on your translation (or before, or after), you can also invite others to participate in your translation/editing efforts by selecting Share> Invite people. All they need is some kind of Google account.

When you are finished with the work on your file, you can download the translated file in the original format (and formatting). And be happy ever after. Or not?

Let's look at the tool in more detail aside from its functional aspects.

The first thing I noticed when I started to look at it -- with the original Google Translation Center still burned in my memory -- is that it is much less project management- and process-oriented than what was originally planned. This new tool is aimed at the translator rather than the translation buyer. (Remember that glimpse of Google Translation Center that we all saw, where translation buyers could upload a document and then choose translators to work in it? None of that anymore.)

The official Google blog says this:

At Google, we consider translation a key part of making information universally accessible to everyone around the world. While we think Google Translate, our automatic translation system, is pretty neat, sometimes machine translation could use a human touch. Yesterday, we launched Google Translator Toolkit, a powerful but easy-to-use editor that enables translators to bring that human touch to machine translation.

The only argument that most of us would have with the statement that "sometimes machine translation could use a human touch" would be the "sometimes" (and that it's often more than just the "human touch" that's required). But the statement still tells us quite a bit about the immediate intent: to add the human expertise to Google's machine translation efforts. Remember, Google Translate, unlike engines like FreeTranslation or Babel Fish, uses a statistical machine translation engine that relies on good bilingual data -- lots of good bilingual data.

And I don't think that there is anything wrong with it if it's just that I am adding data from the translation that I am currently working on. What I did not like was this: As mentioned above, it is possible to upload TMX translation memories, and as you upload you can select whether you want to use it yourself or with others -- and that only seems fair. What I did not read before I uploaded a fairly large TM for testing purposes was this:

By submitting your content through the Service, you grant Google the permission to use your content permanently to promote, improve or offer the Services. If Google publicly displays any of the content you submitted through the Service, Google will display only portion(s) and not the entirety of the content at one time.

This means that even though I can now go ahead and delete my TM (so that I and possibly other users won't have access anymore), Google will continue to use it. That strikes me as odd -- to say the least.

Also, as you know, TMX stands for translation memory exchange, but it is not possible to get your TM (or your glossary) out of GTT once it's in there. The only thing you can get out is your translated file, though you can, of course, continue to use the TM and glossary content within GTT -- but only there.

To be fair, this was a problem in the first incarnation of Lingotek as well and they fixed it rather quickly, but we're talking about Google here, and I'm sure they don't do things without thinking them through. (Though other large companies do, and must also humbly concede defeat -- see below.)

Aside from the thorny intellectual properties issue, there are a number of other things that make me think that GTT in its current state is not for the professional translator (unlike what we saw -- or imagined seeing -- in the Google Translation Center):

  • The only source language (and UI) is English (though I imagine that this will change in no time).
  • There is no possibility to alter the segmentation.
  • There is no possibility for concordance searches.
  • There is no way to manage fuzziness and it's not very clear how fuzziness is decided on.
  • There are no QA tools and the only spell-checking is that of your browser.
  • The number of file formats is much too limited (no XML, Office 2007, InDesign, etc.).
  • You can only work on single files at a time.
  • There are no project management facilities.
  • There is no link to professional translation formats aside from TMX (such as TBX, TTX, or bilingual Word docs).

Where I think GTT will have a good response is on the semi-professional or non-professional translation market. The Wikipedia/Knol feature that allows you to take an English page, translate it, and publish it right to the server will certainly play a role in this (though I admit that I could not make this feature work!).

In its present form, though, this will not be a threat to small translation agencies or sites like ProZ or TranslatorsCafe as was feared with the first version. However, it is a "beta version," and my feeling is that if you were to ask Google where this is headed, they could tell you -- but they'd (regretfully and politely) have to kill you.
2. The Truth Is Out There, Part II

In the last issue we talked about Linguee, the German<>English corpus engine that I described as a "quantum step in search technology for translators."

To briefly recap some of the tool's cornerstones, it's a very large corpus of web-based translated materials from live online sources matched up with the help of a custom, user-generated dictionary and other web-based dictionaries. Once the data is found it is displayed in large chunks and you are offered links to the originating sites, giving you all the context and information about the data that you need. Though it covers only German and English at this point, other language versions are in the works and I would not be surprised to see something new in the next few weeks.

I also mentioned that the timing of Linguee's release was interesting as well, as it appears concurrently with the TAUS Data Association's release of the first version of its language corpus search engine. And this is what this article is about.

I have been aware of TDA's efforts for a while (and at their initial meeting suggested a model very similar to the one they are now using, if I may say so). While its benefits to translators might not seem too colossal at first glance, it's a very important undertaking that is only at the beginning of what it will eventually mean to our industry (and how it will compete with offerings like the one from Linguee or GTT).

The TAUS Data Association or TDA is, like its name says, an association of mostly large corporate translation buyers who originally came together to pool their translation memory data to - yes, you guessed it -- better train their machine translation engines. As mentioned above, to make statistical machine translation work, you need a lot of high-quality data, and even industry giants like Oracle, Microsoft, and Adobe do not have enough data on their own to get the results they hoped to achieve.

So, rather than semi-secretively pooling their data, they decided to open the data up to the public -- not as TM data, mind you, but as a terminology resource. If you want to get to the data as TM data, you can become a TDA member, contribute your own data, and download some other data for your own use.

But since this is financially out of reach for many of us, we can at least use it as a terminology resource.

After you register, you can search for single words or expressions in 80 different language pairs (they presently all originate from UK/US English, which are frustratingly treated as different languages that require separate searches) with a total -- as of today -- of 549,411,144 words. So far the data comes mainly from two major sources: material from the EU TMs (covering pretty much all the British English into other European languages within the "Legal Services" domain) and mostly from the domains of computer software, computer hardware, and professional services in US English into other languages. When you make a search it's relatively easy to see who has already contributed data (product names are not distorted), but it is not possible to know where an individual entry comes from (unless that very entry happens to have specific product names). To me this is the greatest weakness of this otherwise impressive offering, and an area where a tool like Linguee excels. For data to be truly valuable, we need to see where it comes from. Chances are we know that Term A could be translated with Term B, but we as translators need to know whether the specific client we work for uses Term B, C, or D. And very general domains like computer software or computer hardware simply are not specific enough to give adequate guidance.

Also, a tool like Linguee (and, no, I'm not getting paid to mention them so often -- I just think they have a very clever way of assembling and presenting data) uses the powerful combination of TMs, glossaries, and dictionaries to highlight source and target terms in the retrieved segments, something that TDA cannot do at this point.

So, overall it's a useful tool, but not a tool that will give you the one answer you're looking for; instead, it's more likely something that can give you a variety of terms you can then use to further narrow your search.

Also, it's important to remember that this tool was not developed primarily as a terminology research tool, but as a TM sharing tool. And I am eager to see how successful the implementation of that will be.
ADVERTISEMENT
3. Some Dorky Licensing . . .

. . . that's what many translators thought when they realized what the licensing of the newly released  Trados Studio 2009 version for upgraders entailed. While it did actually contain two licenses -- one for SDL Trados Studio 2009 and one for the SDL Trados 2007 Suite -- the 2007 license was time-limited and would cease to function after a period of 12 months. This is how SDL explained that move:

The reason for the time limited license is because we will be replacing this version of SDL Trados Studio 2009 within twelve months with one that will not require any of the legacy functions we use from SDL Trados 2007 Suite and to ensure that all upgrading users can continue to use their legacy product until they have had time to get used to all the features of the new one. At this point the upgrade will leave you only with SDL Trados Studio 2009. If we believe that an extension is required to these twelve months, we will make this extension available to all users.

The truth is that all this is really not that unreasonable -- it's relatively common to have the old license not officially be legal anymore after an upgrade has been purchased, but users typically don't care: Who wants to use the old software anyway? And even if you did, you just installed it and it ran just fine with the old licensing code. The problem here was that SDL's licensing codes are monitored online so they would not run fine after the license expired. In addition, and this is more important, the new Trados Studio really is more than just an update -- it's a whole new piece of software that may share some background coding but not much else that is visible to the user. And some features (alignment, TTX previews) are available only with the old version present.

So users panicked, lots and lots of emails were sent to forums and SDL Trados support, and within just a few hours the policy was dropped and replaced with this:

As you know, as part of the upgrade process, we asked you to return your old license first. After the licenses were returned you should have received your SDL Trados Studio 2009 license and a time-limited license of SDL Trados 2007 Suite. This policy has now been changed. Once you have returned your licenses we will now add to your account a permanent and unlimited license of SDL Trados 2007 Suite as well as a permanent license of SDL Trados Studio 2009.

Well, you certainly cannot blame SDL for not listening to users' responses, and you also can't say that user forums are useless.

Also, this comes following several changes to the licensing policy that had already taken place last week and that were initiated by user responses:

  • Originally it was planned to allow only one install per license. This has been changed to two (one for laptop, one for desktop) for upgraders. Folks who buy the new version from scratch have to pay an extra fee of 30 Euro to install two instances.
  • Originally it was planned to block running more than one freelance license on one network. This has also been changed so that it is now possible to run more than one on a workgroup network.
  • The AutoSuggest feature (which arguably is the most interesting new feature in the new version of Trados -- I reported on that a few weeks ago) was supposed to be limited to the more expensive Professional edition. For the time being it is now also part of the Freelance version.

I feel for the folks from SDL. They are trying to do the right thing, while at the same time knowing that releasing a completely reworked version of Trados is a huge gamble. As I said before, while the new version is superior in many ways to the old, users of the old Trados will have to sit down and learn some new tricks -- and those can be learned on the Trados Studio or one of its competitors.

ADVERTISEMENT

 Confused about translation technology?

TranslatorsTraining provides clarity. Save time and money by checking out our wide selection of comparative CAT tool video tutorials.

P.S. Sign up now and receive a free one-year subscription to the Premium Edition of the Tool Kit newsletter.

4. Flexible Clay

At the Localization World event in Berlin last week, the new version of Clay Tablet was released.

For those who are not sure anymore what Clay Tablet is: It offers ready-made connectors between commonly used content management systems (CMS) and a number of translation environment tools (TEnTs) and/or translation workflow systems so that text from the CMS can be monitored and easily extracted. It's a fairly high-end product with a corresponding price tag, but then it's also usually not purchased by the language service provider but by the end client on whose system it needs to be installed. Interestingly, last time I wrote about it (about 15 months ago) I said this: "Clay Tablet will offer its product with a very affordable SaaS -- Software as a Service -- model, essentially forestalling any huge expenditure for anyone." Today, about 80% of its sales are through the SaaS model. Things are changing quickly.

The new version (2.5) offers two new features -- InSync Master Asset Management Technology and MultiPoint Flexible Routing Technology. Quite a mouthful if you ask me, but still interesting. These new features give more flexibility in the routing of the text, so Clay Tablet is not only a text extraction and re-insertion service but it now offers a number of workflow management features that are aimed at helping to build higher quality TMs and less manual interaction.

Clay Tablet itself calls it a "game changer." I am not sure that I would agree with that on the whole, but from Clay Tablet's perspective it is, since it sort of redefines the tool and shows that its makers are on the ball to keep up the development, horizontally (by developing more and more interfaces) as well as vertically.
The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Last week this reader added a link:

www.asap-traduction.com

© 2009 International Writers' Group