ToolkitSmall

A computer newsletter for translation professionals

Issue 11-6-193
(the one hundred ninety third edition)  
Contents
1. 15 Words (Premium Edition)
2. Jamais Vu
3. Hybrid Hallucinations (Premium Edition)
4. Things That You Won't Read About This Time
5. A Simple Love Story (continued)
6. New Password for the Tool Kit Archive
The Last Word on the Tool Kit
Ygrjvslhv

Goodness, what a week or two! Not only has Twitter been aflitter with all kinds of messages on translation technology, but I know that many of you have probably been thinking about it more than usual. Naturally, I will address Google's withdrawal of the Google Translate API in this newsletter (and will try to put a slightly different spin on it than most other pundits), but I will also talk about the incredibly long-delayed new version of Déjà Vu X, and I'll try to clarify some things about machine translation. (As for the rest -- and there's been a lot -- we'll just have to wait another two weeks.)

I was also asked to mention the quickly approaching FIT World Congress in San Francisco in the beginning of August, and I'm happy to do so -- it should be a special event with great speakers and opportunities to meet other translating and interpreting folks from around the world.

The other day I was reminded of what may be the funniest audio clip I know. Those of you who live in the US probably know of Click and Clack the Tappet Brothers from NPR, two car mechanics from Boston who like to laugh and talk about cars and many other things. A while ago they had a little language-related feature that you can listen to right here. (And just in case anyone complains that this is biased of me to suggest this clip -- have you ever looked at my last name?!)

1. Deprecation (Premium Edition)

Let's talk about Google's announcement to deprecate the API for Google Translate (as well as a bunch of other APIs).

For the very few among you who haven't been following this story these last few days: Earlier this week, Google announced that it had decided to start immediately withdrawing the application programming interface (API) -- the part of the application that allows integration into other environments and tools -- for its machine translation engine Google Translate, with the final withdrawal happening in December of this year. Most of you know that this machine translation engine has made it into virtually all translation environment tools in the last couple of years as an optional alternative to the translation memory feature.

If you were to assume (as I had) that though it's there, it's really not used by professional translators, you would be wrong, as I learned six months or so ago when I made that assumption in writing and was subsequently lambasted with all kinds of emails like "YOU of ALL people SHOULD know that this technology can be helpful to translators."

So while some of us may not have used the integrated Google Translate components in tools like Heartsome, Trados, Déjà Vu X2, Wordfast, Fluency, Swordfish, Lingotek, memoQ, and Metatexis, many of our colleagues did, and it's been virtually a given that any new translation environment tool release would include them.

Many others have argued and written about the reasons for Google's actions, and while I think it's interesting to speculate about this, it's not what I'm really concerned with at this point, especially for this newsletter. The three questions that interest me most are these: a) What does this mean for the many translation tool vendors who have integrated Google Translate? b) What does this mean for Microsoft's Bing Translator, the most obvious competitor of Google Translate? and c) What does all this mean for Google Translator Toolkit, Google's own translation environment tool that most of us have tried to stay clear of but that most certainly will maintain the Google Translate feature?

Let's start with the last one first. I have to admit that for a split-second it actually rattled through my head that Google was shutting down the third-party TEnT access to Google Translate to promote its own tool. But then I shook my head in chagrin about such delusions of grandeur -- as if we were important enough for Google to even consider that! Taking down the API for Google Translate has little or nothing to do with the translation community and everything to do with Google's strategies for dealing with much bigger problems (read: Facebook, etc.) Still, as one person said whom I talked to this week, one of the many desirable side-effects for Google might be that Google Translator Toolkit could indeed become more attractive for those who frequently use Google Translate in their translation work. (Have you realized how often I've written "Google" in this article? "Google Translate" shall be GT from here on out!) Whether this really happens will depend on the response of other tool vendors (for instance, will they consider a business deal with Google to keep the next version of GT in their tools or will they make the Microsoft engine available?) and, of course, on Microsoft and/or other vendors of MT software (Systran, ProMT, Apptek, etc.).

I asked the translation environment tool vendors what they're going to do. Here are some of their responses:

István from Kilgray (memoQ):

I am personally not so concerned about this. I don't believe it will have an impact on translation tools. If nobody has Google Translate, there is no impact. We will be waiting until our customers ask for another engine.

Rodolfo from Maxprograms (Swordfish):

I'll add support for Bing. Support for Google's MT will be dropped unless Google offers a usable alternative, paid or free.

Carl from Western Standard (Fluency):

We currently have the option of switching from Google to Bing in Fluency. As long as Bing doesn't discontinue its support of the translation API, I think the changes in our tool will be minimal. In order to maintain our MT compare feature, we will be investigating other MT engines. While many translators understandably refuse to use MT in the translation process, we see the MT feature in CAT tools as another resource for the translator that, when used correctly, can be a powerful aid. Many have effectively incorporated it into their processes. I do think this opens the market up for other MT engines and might actually end up increasing MT quality now that the free Google lunch is off the menu.

 Daniel from Atril (Déjà Vu):

Now that Google has announced that they will phase out the Google Translate API, we will obviously have to switch to other MT systems -- but we were already planning on adding support for Bing Translator in one of the first monthly updates, and integration with commercial MT systems such as Systran, ProMT, etc. was already in the list of features to be added in the coming months.

And, lastly, Daniel from SDL (Trados):

Basically it will be very interesting to see if Google reaches out to the commercial vendors if and when they release version 2.0 of the Google Translate API which is currently in their labs, what the conditions will be, etc.

It will almost certainly be paying so we'll need to see what this means and how it impacts the MT  features in Studio. In the meantime, when we get closer to the shutdown of the current free service, we'll issue a hotfix for Studio where the plug-in will be disabled and/or a clear message will appear stating that this free service is no longer available. Hopefully by then we'll know what the future will be and what the shift to Google Translate 2.0 will look like.

I also asked Daniel from SDL whether there currently is a business deal with Google in place or even in the talks, and he confirmed that "there is no direct business deal at this point" -- thwarting some of the speculation swarming about on this.

So, to summarize, most vendors are a lot more relaxed about this than their users. The general consensus: "Yeah, we'll switch to Bing Translator and possibly some other solutions and otherwise we'll see what happens." Kind of refreshing, if you ask me.

I also contacted the Bing Translator team at Microsoft and asked them about their response. Here is what they said:

A large-scale translation service is certainly expensive to run, as Google correctly pointed out. It is also true that simply performing a translation doesn't teach us much besides the fact that we just did a translation. Our goal is to collaborate with our users and partners to continually improve the service that we can deliver to them. This is a model that has been working well for us so far.

The Microsoft Translator API is and remains open for any developer to use. Commercial use requires a commercial license agreement. We welcome localization and translation tool vendors to sign such a commercial license with us. We will provide translations free of charge. Instead of a payment, we are asking the tools authors to make effective use of the collaborative translation functionality provided in the Microsoft Translator API.

The collaborative functions in the Microsoft Translator API have been designed to provide a means for Microsoft and partners to collaborate on improving the quality of translations. A live sample of such an implementation is shown on our blog at http://blogs.msdn.com/translation. Simply put, use of the collaborative functionality allows one to authoritatively modify a Microsoft Translator suggestion for future translations within the same project, and makes it a collaborative alternative for that sentence to the broader community.

Ultimately, we want to be able to improve this service through partnerships and collaboration, while being transparent about availability, limitations and motivations. We have always been open to different types of partnerships that allow us to improve the service, so don't hesitate to contact us at mtlic@microsoft.com with your questions or proposals. We will also be publishing a white paper with additional information within the next two weeks.

Interesting, huh? Essentially, they say that they understand what Google is doing if it does not derive any direct benefit from it in the form of better translation quality of GT, so they will be looking at trying to do a better job with their engine. Awhile back I posted a sample of their MT/TM engine (still in beta) on my site for users to correct; you can correct the translations and it will be stored as a TM for that page and be used by the MT for further training. This is basically what they will be looking at while continuing to offering their APIs to tool vendors.

Oh, and to come back to usage of Google Translator Toolkit, the responses to the blog posting where Google announces the withdrawal of several of its APIs are very revealing. In effect, there are only two kinds of responses. The first is all about the GT API's importance for so many, and not just translators (hint: this is something that we should be able to capitalize on). The second is about Google's unreliability as a partner. Remember when Google stopped GOOG-411, its speech recognition service, awhile back? There was no need for it anymore from Google's point of view: it had collected all the data it needed.

To indirectly quote Carl from Western Standard: There is no free lunch. And since Google is no longer as concerned about its reputation, it has become particularly ruthless about sharing when it sees an immediate benefit and taking away if it doesn't. Not really the "technology partner" that I would like to work with.
ADVERTISEMENT

Join our Social Network for Translators: www.langmates.com

Automate your translation business: www.to3000.com

Start your own translation agency: www.projetex.com

Discover true word counts: www.anycount.com

Exclusive discount for Tool Kit Readers: http://special.translation3000.com/toolkit20/ 

2. Jamais Vu

So, Atril has finally released the latest version of Déjà Vu -- and a collective sigh has been heard among the still very loyal following of one of the most consistent workhorses in the world's TEnT stable. It's been a very long time coming, and of course the mockers were quick to point out that at first glance not much has changed -- the interface looks very similar, albeit a bit spruced up -- but behind the scenes there have been some changes.

I had a chance to look at the new version and also have a long email exchange with Daniel, the creative mind behind the tool.

Here are some of the new features in a nutshell:

  • Subsegment mining and matching through a feature called DeepMiner
  • Extension of the Assemble feature to subsegments and MT segments (still Google Translator, will be Bing Translator in the future)   
  • Predictive typing based on matches found in TM and termbase ("AutoWrite")
  • Use of project templates
  • Use of XLIFF as a format to exchange translation data at any point of the project (only available in the more expensive Workgroup version)
  • Storage and reuse of the rather tedious-to-construct but very powerful-to-use SQL statements that allow you to do virtually anything with your project and databases
  • Multi-file alignment
  • No more dongles!

There are also a lot of smaller things that the experienced user will quickly notice (ability to search on really short strings, faster processing, etc.), but these are not really showcase improvements.

The feature that has been garnering the most attention among early adopters has certainly been the AutoWrite. This is a huge help with productivity since there is no separate database necessary as in the case of Trados Studio, and the matches are automatically updated. In my opinion, the most promising feature is DeepMiner because it provides access to "stuff" that wasn't accessible before in the TMs -- at least not automatically. It's very much what memoQ has been doing for a while, and I think that it might still have to mature a little before it reaches full fruition. (Daniel mentioned in an email, "We have a number of improvements to the DeepMiner technology that we are still working on, which will increase speed and accuracy both for using DM to repair fuzzy matches and for AutoWrite." So there is certainly an awareness that things still need to be improved.)

When I spoke to an Atril representative last month in Vienna, she told me there was a high likelihood that Déjà Vu X2 (this is  the name of the new version) would be released in May. I just smiled at her, thinking to myself, "Yeah, right -- that'll go the way of many other promised release dates." But I was happily wrong and the new version did come out in May, so maybe the monthly update cycle they've announced will happen as well.

I mentioned some of the work that needs to happen on the DeepMiner. Among other observations, I noticed that the documentation was still at the level of the previous version, and there were some miscues in the features of the new and improved Project Explorer (formerly "Project Navigator") -- the pane that allows you to navigate through files in a project -- but those are relatively minor things that we should see fixed soon.

Overall, I've been playing with it for a few days now and have experienced no crashes or other calamities. (But then I also wasn't among the first-day wave of super-early adopters who got a bad build that crashed left and right. To Atril's credit, it was fixed quickly.)

The seasoned professionals among us still remember the "flame wars" between Déjà Vu and Trados users way back then when it seemed that those were really the only systems out there. Since then things have clearly changed, so I asked Daniel how he sees Déjà Vu's market position today, especially considering that it's relatively highly priced compared to the competition, and other tools have made a lot of inroads in the last few years.

He said, "Our intention is for DVX2 to help Atril regain its position in the market, both in terms of the competition (being the 'alternative' to SDL Trados), and in terms of regular releases (we intend to release new builds on a monthly basis, incorporating not only fixes but also new functionality which didn't make it into the first release of DVX2). I think that the new features in DVX2 (for which we've had some very encouraging feedback so far) will justify the slightly higher price tag."

I also asked him about the different editions (Standard, Professional, and Workgroup). The original plan was to cut the cheaper Standard edition: it didn't have many of the advanced features, causing some users to be frustrated with a cheaper but clearly less powerful version of the tool. But he pointed out that "feedback from certain customers revealed that there was indeed a place for a slightly limited Standard version at a very reduced priced -- but instead of having a drastically reduced feature set like in DVX1, we decided to include most of the automatic features, and limit the available filters to the file formats normally used by beginners." These missing filters he mentioned are for formats like InDesign, FrameMaker, XML, etc. -- but I gotta tell you, if you look into using Déjà Vu, don't mess with the lower-priced version. There are actually quite a few other features that are not available in it (you can find a comparison of the features right here).

(By the way, I tried to convince Atril to raffle off a copy of the new tool for you, my esteemed readers, but they were not to be convinced. There'll be giveaways for Trados and memoQ in some of the next few newsletters, though.)

I then asked Daniel to also talk about the future and topics like data sharing, cloud-based computing, and browser-based interfaces. He had some interesting things to say. He was cautious about data sharing, being very well aware of some of the legal hurdles. But he mentioned "short- and mid-term plans" to extend the TeaM Server product with project management features and a web-based interface for managing projects and for translation. Specifically about browser-based interfaces, he said

I was historically rather opposed to browser-based interfaces due to their limitations, but given the advances in client-side interfaces in the last couple of years I think it's perfectly feasible to construct a fully functional browser-based application that includes features like AutoSearch and AutoWrite.

I was glad to hear that, especially at a time when this clearly is something that most tools are moving toward. Kilgray's István Lengyel just exclaimed in a Twitter post earlier this week, "I am about to give up my belief that the web is useless for translation." We will see this first as an alternative interface with most tools, but -- you heard it here -- in four or five years you and I will be working almost exclusively through browser interfaces. And why not? If the browser is able to emulate all of the desktop's version of the interface, it puts an end to a lot of problems, not the least of which is of course a lack of compatibility for different operating systems.

To me the most exciting thing that Daniel said was in his comments about combining Déjà Vu's example-based machine translation approach -- which has long allowed for "fixing" fuzzy matches if the associated databases contained the necessary language data -- with the news subsegment "DeepMiner" feature and MT. He said that this "is probably what will have the largest impact on the market" and I agree wholeheartedly. In fact, this is the very thing I've been suggesting at a number of conferences lately (you can find the presentation of a talk I gave in Sin City -- I was a good boy, though -- a couple of weeks ago on Cetra's blog). What this means in practice would be that any fuzzy match would be analyzed on the basis of data in the TMs, termbases, and, now, MT engines, and the "bad" stuff (i.e., the parts that don't match) would be taken out and replaced with "good" stuff (the parts that are more likely matches). As I said, Déjà Vu has long been doing this, but due to the limited nature of termbases and TMs, there was only so much it could do. However, there certainly is a lot out there that can be done. This is where I see MT having the most profound impact as a productivity tool for the individual translator.

So, all in all, a good new version of Déjà Vu? I think so. If Atril can now keep up the monthly update schedule and work on some of the features that are in the pipeline, it might very well be a great new version.

3. Hybrid Hallucinations (Premium Edition) 

In the last newsletter I talked about hybrid MT systems in connection with Systran. I said that a hybrid MT system "would typically mean that all language processing is done with a combination of rules-based machine translation (RbMT) and statistical machine translation (SMT)." Looks like I have to backpedal a little with that.

"Hybrid" seems to be much more broadly used in this context. While it does refer to a combination of the two kinds of machine translation, it does not state to what degree they have to be used or what aspects need to be included. So, since Systran does build monolingual language models even with its desktop versions by analyzing existing content on your computer, and uses that in combination with its rules-based system, the claim to be hybrid would indeed be accurate (the Enterprise editions of Systran use both language and translation models in their SMT approach and combine that with the RbMT system).

Thanks to Laurie Gerber (you can read a slightly outdated article of hers where she explains hybridization) and to Mike Dillinger for helping me see the light on this.

And you say that's only semantics? And I say, yeah, but isn't that exactly what you and I do as translators? If we are not careful with our words, who will be?

ADVERTISEMENT

I started reading your book last night and I LOVE it. It has all the information I have always wanted to know. I will definitely be recommending it left and right.

Excellent job!!

Charlotte Brasler (English<->Danish)

The Translator's Tool Box: A Computer Primer for Translators

400 pages with (almost) everything you ever wanted to know about your computer  

4. Things That You Won't Read About This Time . . .

. . . but will give you so much to look forward to when Tool Kit #194 arrives at your proverbial doorstep:

  • The long-talked-about and finally released TM Repository by Kilgray, a translation memory management system
  • An effort to combine all kinds of open-source dictionaries into a really, really large dictionary in virtually all languages
  • An effort to translate the complete Tibetan Buddhist canon into English
  • A fascinating resource on African languages
  • A new tower of Babel in Argentina
So, all of this and -- of course -- much more next time.
5. A Simple Love Story (continued)

Here is another addition to my family of remarkable characters that I have been collecting over the years. These characters are simply striking.

If you have a Javascript-enabled browser, hold your cursor over the character for a definition.

6. New Password for the Tool Kit Archive

As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to May 2008.

You can access the archive right here. This month the user name is toolkit and the password is privilege.  

New user names and passwords will be announced in future newsletters.

The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here is a website that added the Tool Kit link this week:

www.leonhunter.com

© 2011 International Writers' Group