ToolkitSmall

A computer newsletter for translation professionals

Issue 10-10-176
(the one hundred seventy-sixth edition)
Contents
1. PDFs in Translation Environment Tools (Premium Edition)
2. Updated Kids on the Block
2.1. Kid #1: Trados
2.2. Kid #2: memoQ
2.3. Kid #3: Fluency
2.4. Kid #4: Similis
3. Getting Even Smarter
The Last Word on the Tool Kit
Cathedrals Made of Fire

By now, most of you have probably read this article in the New York Times that likens the act of authoring to that of translation. The idea is nothing new in itself, but it's still heart-warming to see after the endless coverage of machine translation, particularly of the Google kind, in the popular media.

I looked at Twitter responses to this article and was not surprised to see that it garnered much more attention in the writing community than in the translation world. And I think that's great. After all, we've known about the similar nature of translation and writing -- both as acts of translation -- for some time, and it's good to see our friends, the writers, be reminded of that as well.

Speaking of Twitter, I had a couple of interesting experiences with using Twitter as a tool for applying some righteous pressure on corporations these past few weeks. With the conference season in full swing, Jeromobot and I have been all over the place, first in civilized Russia at a super-successful translation conference in Ekaterinburg (a city that otherwise distinguishes itself by being home to the only QWERT -- the Y is missing -- monument in the world; if you wonder why there is a Latin-based keyboard monument in Russia, don't -- the programmers who meet there every year on the Russian "Programmer's Day" for a contest in mouse-tossing, don't either), and then in less civilized San Francisco for a fun workshop. The only thing that was not so fun about San Francisco and distinguished it quite distinctly from a much safer Russia was this (and that's just the hand): BED BUGS! You may have been reading about these beasties in the papers and laughed it off as another media scare tactic, but unfortunately I can now attest that the scourge is for real! Now, to come back to Twitter. Expedia was giving me the runaround over reimbursement for the pitiful hotel with its menagerie of bugs (which, by the way, was the America's Best Value Inn & Suites at Union Square) until I posted it on Twitter. The reimbursement was made in a matter of minutes.

One more thing: I've been asking fellow translators to tell me what is important to communicate to the community of machine translation developers. I was asked to give a keynote lecture at the AMTA conference in Denver -- sort of as a representative of the translator community -- and this is a good opportunity to get things off your chest.

1. PDFs in Translation Environment Tools (Premium Edition)

I've mentioned the processing of PDFs in TEnTs plenty of times recently, but that's for good reason: within the last year or two, most vendors have recognized that there is a way to get to content in many PDFs (plus it sounds really good from a marketing perspective to say that PDFs are supported).

At this point, there are three different kinds of strategies when dealing with PDFs right inside the translation environment tool, and if PDFs are a common format for you, make sure to evaluate well which option suits you.

The "black-box" PDF-to-DOC method: This is how Trados, Alchemy Publisher, and Wordfast Pro approach the conversion. These tools use a third-party PDF-to-DOC conversion tool that they have integrated as a filter. It presents the results of the conversion as a Word document within their respective interfaces. This is clearly the easiest  method if everything goes according to plan: if the formats are maintained, line breaks are correctly placed (and not according to crazy PDF logic), and inline tags are placed only when necessary. Sometimes this happens, especially when the originating PDFs are short and go easy on the formatting. More often than not, however, this does not happen.

The user-controlled PDF-to-DOC method: This is the way that Fluency handles the conversion. Just like the other tools, it uses a third-party conversion tool, but rather than presenting the user with accomplished facts in the translation interface it adds an additional step where the user can edit the .doc file in an interface with the basic controls needed for that. Naturally this takes a little more time and effort, but there won't be any surprises when working with the converted PDF file in the translation interface. And since this particular tool also works in a WYSIWYG rather than an environment with inline tags, there won't be any big surprises there either. If this sounds too good to believe, it sort of is. If you do have a PDF file that presents a lot of conversion problems, the amount of time you'll have to invest into fixing the .doc file might just be too much to handle.

The PDF-to-text method: This is how memoQ and Similis handle the conversion. While this might sound rather crude -- all necessary formatting has to be applied after the file is translated -- in some cases this might be more time-efficient than doing it through a conversion that helps to maintain the formatting and creates a lot of work in the process. Of course, even with this method you won't easily get to text that is poorly organized in PDF files (such as text boxes or hard returns), and most certainly not to text within graphics.

This kind of text can only be accessed by a tool that combines PDF conversion with optical character recognition such as ABBYY Fine Reader (or PDF Transformer) or Nuance OmniPage (or PDF Converter). So, if the translation of PDFs is your bread and butter, that's the way you want to go: use a professional PDF conversion tool outside the translation environment, modify the resulting converted file as needed, and then bring it into the TEnT of your choice. (Of course, the other -- in many ways preferable -- option would be to stop accepting PDF jobs!)

2. Updated Kids on the Block

This past week (and the week ahead) has been accompanied by a whole slew of announcements of Service Packs, updates, and new versions of translation environment tools. This is not too surprising with the conference season in full swing, a time when tool vendors typically try to offer something new.

2.1. Kid #1: Trados

Most of you were probably not able to escape the announcement of Trados Studio with Service Pack (SP) 3. I had a chance to listen to an early presentation by one of the developers and have a look myself, and here is what's new and relevant (there are lots of smaller enhancements and bug fixes, but I won't bore you with those).

The upgrade went smoothly (note that the downloads for the updated versions of Trados Studio and MultiTerm consist of 600 MB, so you want to make sure to schedule this at an appropriate time). What bugged me was the annoying language selection process for the Freelance edition. I'm not sure why the old languages could not automatically be taken over, plus I had to do it twice because there was already an update to the update.

For the server-based versions of Studio, it is now possible to have more than two languages per translation memory (it has been a pet peeve of mine that this has never been well supported with Trados even though it was not explicitly disabled -- at least not in the pre-Studio Trados world). So this is a good thing, and it sure would be nice to have this feature available for desktop-based TMs as well. This fairly major structural change means that LSPs who use Studio SP3 should see that their vendors update as well to guarantee smooth interoperability between the different versions of Studio. The same is true for a combination of MultiTerm Server SP3, which also was restructured quite significantly and needs to be paired with Studio SP3.

Supported file types for all versions have not changed much in the new version of Studio: Office 2007/2010 has more stable support and the latest version of InDesign (CS5) is supported. Interestingly, this no longer goes through the .inx export format but exclusively through the more stable .idml files (I already mentioned this in my last newsletter). It sounds like the dawn of a new era of translatability for the ever-more-important InDesign format.

Also, with the recent SDL purchase of machine translation developer LanguageWeaver, that machine translation engine is now integrated into Trados Studio alongside the "old" SDL MT engine and Google Translate. Of course, these are all simply options that you can but don't have to choose to use. If you do choose to activate them, you might be confused about the choices that are presented to you (I was), but you'll eventually figure it out. The machine-translated segment will from then on out be displayed to you alongside possible TM hits (to start setting up the MT engines, you'll need to select Project> Project Settings> Language Pairs> Translation Memory and Automated Translation or Tools> Options> > Language Pairs> Translation Memory and Automated Translation). Most other tools that offer more than one MT engine integration have an either-or option; here you can choose to have all three different options displayed to you at the same time which might be helpful and if not that, it's at least interesting (if your language combination is supported by all three and your screen size allows for it).

Still, the most important addition to Trados Studio SP3 probably is the (once again) announced opening of SDL OpenExchange. This was heralded as a big step when the original version of Trados Studio was announced a year ago, but then nothing really happened. There was a website with a few relatively useless applications and that was it. Well, now there is a little bit more.

Let me first explain what this is. SDL essentially looked at the Apple App Store (and some other humbler examples) and decided to create something similar for the Trados Studio environment: a place where software developers can develop applications that work alongside Trados Studio and cater to specific needs of Studio users. These can then be sold (or downloaded for free) on a marketplace found on the OpenExchange site. The benefit for SDL is first, and most obviously, that developers have to pay an annual membership (€ 100), give 30% of the purchase price of each sale (if it's not a free download) to SDL, and are not allowed to market their applications elsewhere -- assuming they are accepted into the program, which is application-only. I know that many of you will now say, Aha! That's what it is -- SDL has even found a way to empty the pockets of developers! But I don't think that's the primary reason for OpenExchange.

From my vantage point, SDL is still struggling to convert the Trados world to its (no-longer-so-) new Trados Studio, and the competition is certainly not sitting still. So the idea is that a whole ecosystem of developers, applications, and buzz can only help to propel the product forward. No doubt there is also risk associated with that. What if no one takes them up on it (kind of like the widely broadcast idea to share AutoSuggest dictionaries that only a handful of folks responded to)? But this certainly has potential.

One thing that needs to happen is that all the application program interfaces (APIs) must be released (these are necessary for the developers to develop apps for specific aspects of the Studio programs); right now there are only a handful, with some important ones outstanding. And it sure would not hurt to have a couple of overnight-rags-to-riches success stories of some developers who created killer apps that we were willing to pay for. Since we all know that translators are so eager to spend money for essential tools, this should only be a matter of time, right? (See this heart-warming story on that topic.) But we'll see what happens, and truly I do think that this is a very interesting experiment.

Presently there are two interesting apps that can be downloaded (both are free and developed by SDL). One is called SDL TTX It! and allows you to quickly convert any supported file to the .ttx format without having to go through Trados Workbench or TagEditor. (For those who use third-party tools to translate .ttx files, a word of caution: since most of those tools require the .ttx file to be "pre-translated" within Trados 2007, this method might not work for you.) The other is called Bilingual Preview Generator and is really quite interesting. It's a tool that converts SDL XLIFF files (the format that Trados Studio uses as its interim translation format) into MS Excel, Word, or XML documents so that these can be reviewed outside the Trados environment and finally brought back into Trados. This is the route that Déjà Vu has gone since the dawn of time and has lately also been followed by memoQ, but either way, I'm glad that Trados is allowing for that as well -- not only because it gives a certain tool independence alongside the translation-editing chain, but also because it is sometimes very helpful to read text in a different kind of environment to find stylistic and typografical errors.

ADVERTISEMENT

Looking for a real on-line translation tool
that is powerful yet easy to use, flexible and economical
starting at less than €9 per month?

Discover the all new XTM Cloud!


Sign up for a free trial at www.xtm-intl.com/xtmcloud
2.2. Kid #2: memoQ

On to memoQ and Kilgray. Next week (or this week, depending on when I can finish this newsletter and send it out), Kilgray will release version 4.5 of memoQ. Last week I talked to István, one of the Kilgrayians, about whether Kilgray considered this to be a major or a minor release. He had to ponder an answer at first (he typically is very quick-witted, so this was a little surprising). Then he said that for all intents and purposes it is a major release, though they decided not to brand it as such ­-- partly because the last release, version 4.0, was a very major release for which the complete user interface and many of the internal workings had been overhauled and they did not want to go through anything like that anytime again soon.

I looked at the new version and I think it could indeed be justified as a major release -- particularly for two features, or should I say "concepts": the "resource concept" and the "desktop document concept" (and of course there are a host of other little new features and bug fixes as well).

Let's start with the second one first. If I were to point out one thing that differentiates Kilgray from all other vendors of translation environment tools, it would be this: from the very get-go, the earliest beginnings of the company, they did not focus on only one segment of the language industry but on all the segments -- translators, LSPs, and clients alike. This has resulted in a tool (set) that reflects the needs of all the groups equally and enjoys a good reputation with all of them (take that as a hint, new tool developer, whoever you might be). As a result, we can expect to see features in each of memoQ's new versions for each of the needs of the different groups.

The desktop document concept is directed toward LSPs and translation buyers with server-based installations. The idea is that server-based translation  projects that were completely server-based (including the translation files) can now be desktop-based, but these projects are in constant communication with the server and are updated accordingly. The benefit is that it is possible to spend time offline (such as when traveling or when your internet service provider goes down), and as soon as a connection is re-established the update will take place again.  Also, there is an interesting new feature in dealing with "floating licenses" (licenses that are temporarily given to translators for the duration of a project). Rather than assigning them separately, they are now packaged with and inside the actual project file, making the management of these licenses a lot less tedious.

According to István, the "resource concept" is a combination of the functions of MultiTrans, Star Transit, and the indexing tool dtSearch. Maybe this is reaching a little high (I could imagine that each of these vendors would have a thing or two to say about that), but it still describes some essential functions of this feature set.

The folks at Kilgray know all about the painful process of alignment, i.e., the conversion of a source and target file into a translation memory. Certainly, most of us would heartily agree. So rather than trying to perfect it, they just threw it out. Well, kind of. Instead of aligning file pairs and reading them into the translation memory, they now use something called LiveDocs. LiveDocs essentially stays separate from the translation memory; however, if you choose, it can be used as a resource in a manner that is very much like the translation memory. Clear as mud, right? Here's an attempt at clarification: rather than taking one source and one target file, matching them up, and then fixing it manually, you can now take any number of file pairs, align them on the fly and keep them as matched up file pairs for reference purposes. Aside from that you can also use bilingual files (such as XLIFF files) and monolingual files (for reference purposes). You can then immediately start using them as reference material (or "corpus"). In any given project, matches will show up just like TM matches, with the difference that you can see that they come from a LiveDoc rather than a TM and, if you choose to do that, have a lower ranking because of your ability to apply a penalty to presumably less safe LiveDoc matches. Should you find errors in the alignment of the files (and, trust me, you will, even though the alignment process has been improved by using termbase data, tags, and structural data such as codes in software formats), you can open the LiveDoc at the position of the error, fix it, and the alignment process is started from that point downward.

I am not sure what I think about this last step. It's certainly an interruption to the workflow, but then, so is adding terminology to the termbase as you translate -- a step I've gotten used to, use a lot, and whose results thrill me.

Since there are potentially many LiveDocs you can choose from, you can assign keywords to each LiveDoc, which are then used to give data from that particular LiveDoc preferential or deferential treatment. (For example, if you translate something for Adobe, preferential keywords might be "Adobe," "XML," or "software," and deferential terms "Microsoft," "iPod," or "database" if these terms are likely to occur in the LiveDocs.)

So, the file pair concept is similar to the idea of Star Transit, and the on-the-fly alignment corpus concept is similar to MultiTrans. But what about dtSearch? As mentioned above, you can upload any kind of monolingual file as a LiveDoc. This is then indexed and matches are displayed without translation, but you can open the reference document at that position to see the context of how that term is used -- sort of like a browser search, but in any kind of file format (including PDF) -- much like dtSearch or any other indexing tool.

All the data in the LiveDocs can be exported, but, interestingly, not into the translation memory exchange format TMX; instead, it can be exported into the translation file format XLIFF (one per file pair or bilingual document) so that the file associations with all the context can be preserved.

2.3. Kid #3: Fluency

The next tool that will roll out a new version is Fluency, the new tool that I mentioned recently in this newsletter. (The current edition of MultiLingual also carries my full-fledged review -- newsletter readers can get a free full year's digital subscription to MultiLingual by entering the code d90jst into www.multilingual.com/promo. I highly recommend it.) Version 2011 of Fluency will be released sometime before the ATA conference at the end of this month. I have a beta version here, but it's, well, a little too beta and unstable to really get a good idea of how well all the new features work, so I'm depending more or less on what the Fluency makers have told me is new.

This includes better processing of heavily formatted documents, an automatic suggest from terminology matches (I tried this and it works great) and an auto-suggest from TM matches (I could not get that to work), a number of newly supported file formats (WordPerfect and a number of software development formats), and there is a new and additional link to server-based OCR recognition by ABBYY, aside from other fixes and improvements.

ADVERTISEMENT

Tired of expensive tools that make you work harder for less?


Start saving time and money with Snowball!

 

Free 90-day trial: Lite (free), Freelance (€99), Pro (€199)

Download       Philosophy (Video)     Email

 

You translate. Snowball remembers.
2.3. Kid #4: Similis

And lastly, there is Similis. While there is no new version with additional features, its new pricing warrants a shout-out: it is now free for freelance translators. You can't beat that!

Here is a summary of what I wrote a couple of years ago about this tool:

Unlike most other tools, Similis comes with a very high-level linguistic "knowledge" in 7 EU languages (English, Dutch, German, Spanish, Italian, Portuguese, and French), which it derives from a powerful engine that was originally developed by Xerox for its XTS tool. This engine gives Similis the analytical power to apply linguistic rules to a number of processes, including alignment and automatic extraction of terms and phrases from translation memory content. Readers of my Tool Box book will remember that when I tested the Xerox (now Temis) XTS tool a few years ago, I was nothing short of awed by the accuracy of its terminology extraction (the ability to align translated documents and extract matching term and phrase pairs without much user intervention). There are a number of tools that offer that, but only on statistical rather than actual linguistic processes. Because Similis is able to use a combination of statistical and linguistic processes, the accuracy is extremely high -- so high, in fact, that I literally did not find a single error in the few tests I ran last week. Also, because of the integrated dictionaries and linguistic rules, the accuracy of alignment to create translation memories is extremely high (you can actually set the level of accuracy before you do the alignment; the only drawback at the highest level is the slow processing speed). And while it's not perfect, it's easy to correct errors, especially because the tool gives you a matching percentage alerting you to possibly problematic alignments. If you decide to use TMs that you have created in Similis in other tools, TMX export (and import) is supported.

Of course, all this is only good for creating translation memories and terminology databases. So what about the actual translation? It offers two different environments: a hybrid Word/Similis environment for the translation of all files directly compatible with Word (Word, RTF, text files, etc.) and PDFs (as a text extract), and a separate environment for HTML and XML files.

Both interfaces offer a split view between source and target, with the matches from the translation memory in-between and on the left matches from a set of integrated dictionaries. What makes the translation memory matches remarkable is the existence of "chunks," fragments of translation memory matches that the program was able to automatically extract from larger matches with the help of the XTS engine. And not only are those matches displayed, but you can also have them automatically inserted in the translation.


The general interface from where you control all different activities is very clean and intuitive (which it has to be because there is very little documentation aside from a few animated tutorials on the website and a French user guide).

Reading through this I have at least one question: Why in the world did I think things were intuitive back then? Nowadays I really don't think they are. Fortunately, there are tutorials (in French). Even if you don't speak French, they're helpful (and necessary) -- you can just follow the images and sort of get the hang of it.

Clearly the best feature of the tool is the alignment and, even better, term extraction. Many of you have heard me talk about Xerox's (failed) adventure with term extraction; the term extract feature in Similis represents the last remnants of this foray, and it works shockingly well. Of course, you have to work in the languages that are supported, but if you do you had better use this opportunity to grab it while it's available (though a Similis representative assured me that it's going to continue to be free).

After you download the tool you will have to apply for a one-year license to unlock most features -- the license will be sent back to you almost immediately by a server.

3. Getting Even Smarter

We've decided to push our efforts with TranslatorsTraining a bit. We envision TranslatorsTraining as the go-to site when you want to learn anything about translation environment tools -- whether because you are completely new to this kind of tool or because you want to branch out or even change your allegiance to a certain tool. There are a couple of other possibly less informed and costly alternatives out there at this point, but since our site offers videos that are done by the tool vendors themselves (according to our script and edited by us once we receive them), we know that we are the "gold standard" when it comes to comparing tools. We are currently looking at ways to also present the latest crop of tools, including recent newsletter features such as Lingo and XTM and previously covered tools such as Wordbee and Fluency.

The vast majority of TranslatorsTraining is now completely free (only the videos that show more specific processes of the covered TEnTs or localization tools still have to be paid for), and we have even asked the vendors to give our users significant discounts on their tools, including Trados, memoQ, and Déjà Vu, if you access their sites after you have watched the video.

If you decide to purchase one of the tools through our site, you will also receive my Tool Box ebook (value of $50) in addition to your discounted tool purchase.

It's a great deal. I encourage you to check it out.

The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

These readers just added a link:

www.greektranslator.gr

www.czechtranslation.com

© 2010 International Writers' Group