ToolkitSmall

A computer newsletter for translation professionals

Issue 10-5-166
(the one hundred sixty-sixth edition)
Contents
1. Lots and Lots of New (Versions of) Toys
2. Do Not Press This Button Again! (Premium Content)
3. Peedee'effs
The Last Word on the Tool Kit
Iron

Jeromobot and I just got back from a trip to the great city of Buenos Aires and a gathering of 1700 translators for the conference  of the Colegio de Traductores Públicos de la Ciudad de Buenos Aires. We were impressed with the high level of many sessions and the attendees' curiosity. (Jeromobot was also impressed with the excitement he garnered -- in the words of one attendee, "her colleagues threw themselves upon the e-books [and Jeromobot pins] like hungry translator-wolves.") But there was one event that really took my breath away: The conference was opened by none less than Argentine President Kirchner herself. OK, let me say that again. There was a translation conference in Argentina and it was personally opened by a major world leader, the president of Argentina.

Now, Kirchner's politics are not particularly popular with many of the attending translators (though everyone there seemed to succumb to her charisma) and what she actually said did not make much sense (but, again, there was the charisma), but I cannot even begin to say how much this meant for the Argentine translation community, and, dare I say, in extension, the rest of the worldwide translation community.

Folks, we are at a point where translation is more the focus of public attention than it may ever have been before. Not a week goes by without major stories about translation (or interpretation) in the major media. Agreed, often it concerns failures, such as the poor and stumbling interpreter for Mexican President Calderon when he met with Obama this week, or translation bloopers, like this series of articles in the New York Times, or, of course machine translation. But then there are also other stories, like this one from Amazon, which is planning to release a series of books consisting only of translated literature, and, well, the account of a state leader opening a translation conference.

What does that mean for us? It means we need to strike while the iron is hot and be vocal about ourselves and our profession. Write articles and try to publish them in major outlets, compose blog postings or other publications that are interesting not only to the translation community but to the general public as well, and present at conferences that are aimed beyond the boundaries of our industry. Because, believe it or not, without any unified industry-wide campaign, we are being noticed (again) -- so let's be proactive about it and work on strengthening our profession.

(And then we can also stop spamming each other to self-indulgently vote for each other's Language Blog of the Year!)

1. Lots and Lots of New (Versions of) Toys

There was a barrage of announcements of new versions of software-related tools this week (and it's not even Christmas -- but then, most of these tools are not exactly free either).

Among TEnTs (translation environment tools), Kilgray has (again) released a new version of memoQ (more on that later in this newsletter) and Alchemy Software has released a new version of Publisher. I've been trying to get a hold of some folks at Alchemy but failed to do so. I will talk to them before the next newsletter goes out and tell you more about where they see their product and its future.

In a nutshell, some of the more interesting new features now include TBX support, AutoComplete feature while you're typing, PDF support (we'll have to see what that means), and multi-file processing.

But, as I said, more on that in two weeks.

You will also be able to read more in two weeks about the two major workflow tools: LTC Worx and Plunet. Worx version 2 was just released (I would give you a link but their website is presently down . . .). I talked to one of the brains behind it, and we agreed it would be helpful to talk to a couple of their customers (and do the same for Plunet) and then contrast those views.

And lastly, Linguee, the highly useful German-English bilingual corpus search engine, has announced that it has officially ended the beta phase (whatever that means, especially considering that one of the co-founders used to work for beta-phase-happy Google). They have also added some changes that make its interface a little more streamlined and push it more toward the crowdsourcing concept. (And just to let you know: German translators are no longer part of the controversy whether one should say BCE/CE vs. BC/AD when referring to time. They just say BL and AL -- Before Linguee and After Linguee.)

But let's go one step back and remember what Linguee is. This is what I wrote almost exactly a year ago:

It's a very large corpus of English into German into English data (and, please, don't stop reading if those are not your languages!) of web-based translated materials. The two guys behind it have found ways to have web crawlers detect translated content online and match that up with the help of a 50,000+ entry dictionary and other web-based dictionaries. To look up a term or complete phrase, just enter it into the search box; the matches that are displayed are complete segment matches with the terms in question (both in source and target) highlighted. At first glance the data contains no metadata (origin, subject matter, etc.), but at second glance you will notice the little icons following the entries which actually are links to the originating sites, giving you all the metadata you could want. To search you don't have to register; as a registered user you can evaluate the translations and correct them, or you can add entries to the dictionary, which in turn are used to fine tune the matches.

When the founders started Linguee, translators were not even on their radar as potential and important clients; their main target was the "normal" English-as-a-foreign-language-user. I'm sure that many of the half-a-million daily search queries still come from that group, but many also come from professional translators.

The latest interface is rather streamlined -- exemplified in the thumbs-up and thumbs-down button following every entry that replaced the more cumbersome rating scenario from earlier versions -- but it also emulates the exact feel of most crowdsourcing applications. And indeed, another new feature is the ability to correct or suggest new translations -- and in return you'll get 20 advertisement-free searches (of course, you do have to be registered to suggest and correct data). This is an interesting model: rather than extracting payment, it promises less punishment (in the form of obnoxious ads), but to me it seems reasonable enough to work.

Oh, and the promise to add more language combinations later this year was given once more. I hope that Chinese, French, Spanish, and other translators will indeed be able to use that in 1 or maybe 2 AL.

And one more thing about the new Lionbridge Translation Workspace. While writing about it in the last couple of newsletters, I had to create a free trial account. In the process of creating a free account you'll need to give either credit card or PayPal information, and within a month the account becomes a paid account. If you do that, PLEASE make sure that you mark the day on your calendar -- there is no prior notification from Lionbridge that the account will be converted to a paid account. Not only will your credit card or PayPal account be debited for that amount, you will also have to pay a penalty to delete the account in the amount of three months' worth of subscription fee. Don't believe me? Go ahead and try it out. (By the way, Lionbridge did reimburse me at my request).

ADVERTISEMENT
WANTED: Loving Home for a Brand New iPad

Compete with linguists, translators, writers, and terminologists worldwide
to win a free iPad from TermWiki.com. Whoever enters the most terms wins.

All participants can download three times the amount of terms they submit.

More information can be found here: http://newsletter.termwiki.com/6/
2. Do Not Press This Button Again! (Premium Content)

Let me start with the very first thing I tried with the new version of memoQ (4.2): The Do not press this button button. Much to my dismay, this feature had been broken in the last few versions. Though I had launched numerous complaints, the folks at Kilgray honestly thought there were more important features to fix. How in the world?! You see, when you click this button, it tells you to "not press this button again" and sounds a cuckoo clock -- ingenious! (Fans will recognize the "Do not press this button again" from the Hitch Hikers Guide to the Galaxy. When I shared my insider recognition and how I loved reading the Guide with Kilgray's head developer Gábor, he almost started to cry -- who says that geeks can't be passionate?)

I just looked through past newsletters to see what I've said about memoQ (previously: MemoQ) over the last few years, and so much of what I've said in the past is still true, only now things feel so much more mature and well-rounded. A case in point is the opening screen: It makes you feel comfortable with cleverly designed images and a sense that you immediately know where to start. (I know the folks from Kilgray will REALLY hate me for saying this, but the opening screen and general conception is slightly reminiscent of Across and Trados Studio -- with perhaps a bit more elegance and subtlety.)

But you are not reading this to hear how something makes you feel. Let's cut to the chase: What are the new features?

Well, let's start with a quick general overview of what memoQ is.

Like many of its competitors, it comes in two different product lines: a stand-alone version that has all the features a stand-alone translator needs, and the memoQ Server version that allows larger organizations to post translation projects online (translation memories, terminology databases, and even the actual translation files if that is desired) and have translators work on those with their own version of memoQ or with mobile licenses that the client can give out to the translators for specific projects. Despite its (and their) young age, the folks from Kilgray were among the earliest to adopt server-based technology and have introduced technologies that have made this a reasonable alternative and attractive product for LSPs and end clients.

During the translation the translatable text is presented in a table format, the source on the left and target on the right (in this new version you can also have an editable horizontal field at the bottom), and matches from termbase and translation memory are displayed on the side. If you're reminded of Déjà Vu, you're right. A cursory look at the translation interface is indeed very reminiscent of that program. There are some important differences, though. Tags within sentences can be displayed as actual (write-protected) tags rather than numeric placeholders; the basic formatting (italics, bold, etc.) is displayed as such; there is a preview feature of the actual file for MS Office and HTML files (and for XML files, though I couldn't get this to work); and there is a squiggly-underline spell-checking feature (only with Hunspell dictionaries that you have to install under Tools> Options). Also, many of the underlying search algorithms (including a sub-segment search feature) and the structure of the databases are quite different.

You can mix and match any of the supported file formats in one project. (I won't list them all -- just imagine any of the file types that are typically supported by the major tools, including more exotic and helpful ones such as SVG graphic and Visio files and now, finally, the less exotic Office 2007 and 2010). You can then filter the views in a translation grid by any number or all files or any other criteria you like.

There are two main principles of how memoQ works that you might or might not like. It rarely uses wizards (dialogs where you make one or two choices and then click the Next button to make more choices on the next screen), and it does not use option-rich context menus (context menus are the context-specific menus that are displayed when you right-click somewhere). Both of these features are missing by design. I can live with the lack of wizards (they use very rich dialogs instead), but I would love to see more options in context menus -- to me this represents an easy way to "cheat" my way through a tool that I'm just getting to know. Once I know a tool, I don't use context menus but prefer keyboard shortcuts. In memoQ you will have to look for the necessary options on the toolbar and menus.

Speaking of menus, one way to familiarize yourself quickly with the features of a new tool is to look through the menus and open the "Options" dialog that you can find in almost any tool. In memoQ you can find it in the Tool menu, but you should also open a second dialog in the same menu, the so-called Resource Console. These two contain many of the features to fine-tune the tool. Some are easily understandable (such as AutoCorrect -- though I would love a feature here to easily import all AutoCorrect lists from Word or OpenOffice), whereas others require a certain understanding of things like regular expressions. (One example is segmentation rules, but the nice thing is that the segmentation exchange standard SRX has been well supported by memoQ for a long time, so if this looks too complicated you can easily create your tools somewhere else and just import them.) The Resource Console is interesting, and while it might not seem logical at first why some of the options are listed under Options and in the console, the idea is this: Each set of options can be exported as a resource from the console, sent to someone else, or simply transferred to another computer and imported right there. That's cool.

One interesting new feature that memoQ 4.2 now offers for anyone dealing with European languages is the integration of the large terminology resources of Eurotermbank (Kilgray was involved in the development of the underlying technology for that database). You can have hits shown up automatically in the default Translation results pane or include them in a manual lookup process (highlight a term and hit Ctrl+P for that). This second feature is indeed very helpful because it allows you to filter by subject matter. The first, automatic lookup, takes quite a bit of time to collect and display the terms (note that I tested this on the U.S. West Coast and it might be faster in Europe), and there does not seem to be a way to filter those to avoid getting lots of unhelpful stuff.

The other big achievement in this new version is the ability to export ongoing projects into bilingual formats that can then be processed outside of memoQ and reimported into the memoQ project at a later point. The formats include a Trados-segmented Word/RTF file, XLIFF, and a multi-column Word file that contains source, target, comments, and the status of the segment. Many translators have been waiting for this (often pointing to Déjà Vu which has offered this feature for a long time now) because it allows them to work with colleagues who might not own memoQ and because many prefer to do the editing in a different environment than the actual translation. For a company like Kilgray, a feature like this entails a certain amount of risk -- the risk that fewer copies of the software will be sold -- but I am convinced that it's a risk with a high reward: Openness is appreciated, especially in our community.

One new feature that memoQ does not have is an integrated machine translation -- a feature that most other tools at this point do offer. Quite honestly, I find this refreshing. I have a strong feeling that most tools -- really almost any of the major TEnTs -- have this "cheap" MT connection to Google Translate because it's an easily scored sales point (for the project manager or the company owner), but it really provides very little value to the translator. Don't misunderstand me: Machine translation can offer a lot of value, but so far not through the generic Google or Bing Translate.

ADVERTISEMENT
The Association for Machine Translation in Americas 9th biennial conference will be held in Denver, Colorado, October 31 - November 5, 2010, immediately following ATA.

Sessions aiming to bring together human translation and machine translation practitioners, including postediting demos and tutorials, and MT use case studies, will be held Sunday, October 31 and Monday, November 1.

For more information, please visit http://amta2010.amtaweb.org
or contact Laurie Gerber
3. Peedee'effs

In the workshop I gave in Buenos Aires, one of my slides (that I originally created several years ago) said this:

No TEnT really supports PDF files directly and no TEnT ever will.

Is this still a valid statement after various tools (Trados, Wordfast, memoQ, and apparently Alchemy Publisher) are now offering PDF support? It did not take me very long to decide to leave the slide untouched, and I still believe in the validity of that statement.

Let me explain. No, let my friend Tuomas Kostiainen explain (Tuomas and I will team up for the ATA tools workshop in Atlanta on June 5 and 6). Tuomas is one of the best-known Trados trainers here in the U.S. and he just sent the PDF translation feature of Trados Studio through its paces. You can read the results right here. These are his final thoughts:

1. Don't think that you can translate PDF files just like Word files even if you can open them in Studio.

2. If you open a PDF file in Studio, review the converted text in the Editor to see if there are problems with tags or hard returns. You can try to fix the problems by adjusting the PDF filter settings in Studio (and then opening/converting the file again) or by first saving the converted source file in Word format in Studio (File > Save Source As...), fixing the problems in Word and then finally opening the fixed Word file for translation in Studio. Remember that you can't edit the source segments in Studio, so the errors need to be fixed before you start translating the file.

3. If you need to convert PDF files frequently, consider buying a good conversion program, such as ABBYY Fine Reader (or PDF Converter) or Nuance OmniPage (or PDF Converter). They do the job much better, offer many more conversion options, allow editing and can be used for many other situations as well.

I am with him all the way.

You see, each of the TEnTs that offer PDF support use a third-party solution (Trados uses a converter by Solid Documents, Wordfast by BCL, and memoQ by Glyph & Cog) to convert the PDFs into a Word or text file and offer a translation interface for that. The problems that they encounter are similar to the ones that you encounter when converting PDFs, including the ones that Tuomas describes.

Please don't misunderstand -- I think it's great that there is some support for text within PDFs. But it's important to understand that this is all it is. It's not supporting PDFs as files like it would treat Word or FrameMaker documents, but as containers of texts that are placed into another format -- and sometimes the display and formatting within that format is similar to the originating PDF and sometimes it's not.

The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

This reader just added a link:

www.traducciones-montevideo.com

© 2010 International Writers' Group