Tool Box Logo

 A computer journal for translation professionals


Issue 14-2-232
(the two hundred thirty second edition)  
Contents
1. Take the Lift! (Premium Edition)
2. Pssr, I've got a secret!
3. What's Wrong with Translation Technology?
4. This 'n' That
5. New Password for the Tool Box Archive
The Last Word on the Tool Box
Roller Coaster

First it puzzled me, then it pleased me, and now it just bugs me: Most of the translation professionals who bought the latest version of my Translator's Tool Box ebook after it was released a couple of weeks ago were already customers of earlier versions. It's flattering -- earlier users must appreciate the book or they wouldn't pay for an upgrade (which is only $25 rather than the $50 for first-time buyers). But it makes me wonder what I can do to convince new readers of the value of the book.

Maybe this will help: I wrote this book because I remembered how lost I was when I first started out in translation. I felt I was on good footing with my linguistic skills, but computers -- not so much! I had just finished a four-year stint writing a thesis on the history of Bible translation in China. As I said: not much technology and computers. Interestingly, I quickly found that the vast majority of my colleagues were in a similar position. Perhaps because of the painstaking and tedious (and fun) nature of the thesis-writing project, I decided to tackle the next painstaking project: the computer. I did not want to learn to write code (and I didn't!) or become a computer geek (I think I didn't), but I wanted to "get it," to understand the possibilities the computer had for me as a translator and how to access and use those possibilities. That's what this book is about -- essentially, it's my 17-year journey of understanding and harnessing the computer distilled into 400+ pages. Or, put differently: it's exactly what I needed at the beginning (and every other stage) of my career as a translator.

It makes me happy that quite a few universities understand that intention and are now using the book as a textbook. Maybe you should, too. You can find more information about it right here. 

1. Take the Lift! (Premium Edition)

One way to encourage technology development is to urge existing vendors to focus on areas that their users feel are underdeveloped -- that's what we did with the "what's missing webinars" (see below). Another avenue of development occurs when someone feels so strongly about a missing or underdeveloped feature that they set out to develop it on their own. This is what happened with Kevin Flanagan and his Lift project.

Kevin is a PhD student in computational linguistics at Swansea University and a French-to-English translator. As a translator, Kevin was frustrated with the way subsegmenting was implemented in translation environment tools -- or more accurately, how it was not implemented. Four years ago, when he came up with his proof of concept, that feature was essentially nonexistent in TEnTs.

At this point, most translation environment tools do offer some kind of subsegmenting (the recognition and automatic translation of parts of segments within a translation environment), but that doesn't mean that Kevin's efforts were fruitless. Quite the opposite, in fact. He really never looked for the kind of solution offered by the existing implementations but went about it quite differently.

At this point, tools like Déjà Vu, Trados, memoQ, Star Transit, and others have found ways to suggest subsegment matches by analyzing existing TM data. It's a logical thing to do -- if you have an existing translation memory. Computers are good at crunching numbers: if there's enough data on which probability analyses can be performed, and one subsegment in the source appears a number of times accompanied by the same subsegement in the target, then there's a great likelihood that it's a match. This is what's proposed to us as we translate, typically in the form of auto-complete suggestions.

But what if there's no existing translation memory data for a specific client or project? In that case, the feature does not work at all (for Trados Studio 2014, for instance, you need at least 10,000 translation units in a TM before you can even build an AutoSuggest dictionary) or not as reliably. So, why not use external data to evaluate each segment, regardless of whether there is any existing data in that particular TM, and come up with suggested matches? That's what Kevin's Lift product does.

After processing with a tokenizer and, if applicable, a lemmatizer, single words are sent to a combination of machine translation (for those single words only) and/or dictionaries. These dictionaries can be termbases, other locally installed dictionaries, or even online dictionaries; if there are few or no dictionaries for your language combination, a pivot language can be used. On the basis of that invisible process and its resulting feedback, internal proposals are made about which part of the target segment goes with which part of the source segment (as required by a new segment that contains a part of that segment). More internal processes further verify the likelihood of these proposals before a highly educated guess is presented to the translator. When I talked to Kevin, he said that he feels "it's acceptable if 10% of the time matches are not accurate." I quite agree because all of this is only an additional process to all the other TM, termbase, and already existing subsegmenting features you use anyway.

I'm not sure if I adequately explained what this is about (after all, most of us aren't computational linguists), but let's make it really simple: This process breaks up the mold of the data that we produce -- in the form of a TM -- and analyzes it with outside data to make it more effective. This means that we still use the -- presumably -- high-quality translation that came out of our own fingertips or was dictated into our microphones, but we get to analyze it at an even deeper level than has been possible so far.

Cool, huh?

Right here you can find a video that shows you how it works, with an add-on that Kevin created (for himself) for Trados Studio and an -- almost more interesting -- overview in the form of a poster he presented at a conference last year.

I'm not sure that this will make Kevin a rich man -- after all, not many have become rich with translation technology so far -- but I hope existing tool vendors will integrate his technology to our and his benefit and that this will be only the first step toward intelligently mining our own data through many tools.

ADVERTISEMENT

Even small translation teams can think BIG.

Kilgray recently launched memoQ cloud, a new server solution for small translation teams.

In the memoQ cloud, you can set up collaborative translation projects, share resources and manage translation projects for a monthly subscription fee.

Sign up and get started with a free cloud server trial now!

2. Pssr, I've got a secret!

A couple of weeks ago I realized that one of Microsoft's best-kept secrets is really still a secret, even to my journal and book readers.

I'm not sure whether Microsoft is really modest about communicating its strengths or whether it has such an antagonistic audience that it just doesn't get its point(s) across, but there are a couple of real jewels hidden deep inside Windows that shine all day long without (most of) us ever seeing it.

The Problems Steps Recorder of Windows 7 or Steps Recorder of Windows 8 is one such tool.

If you don't take your computer into a computer shop when you encounter problems, it can be very hard to explain what went wrong to someone on a phone help line. Of course, there are ways to share your computer with someone else -- the easiest might be join.me -- but another way is to record your problems and send the recording to someone else.

This is exactly what the (Problem) Steps Recorder does. This tool allows you to record everything on your screen. When the recording is done, it is not saved as a movie file but as an MHT HTML archive file and zipped up. Once unzipped, the MHT file can be opened with either Internet Explorer, Chrome, or Opera (or with Firefox or Safari with a special plug-in). It gives you a screen-by-screen description of what just happened on your computer, as well as a narration of the process and operating-specific information.

To start the recorder in Windows 7, click on the Windows button, type psr, and hit Enter.

In Windows 8, type steps in the Start screen and select the Steps Recorder.

Everything else is very self-explanatory.

It really is a freakishly good system and soo easy to use (once you know it's there). 

ADVERTISEMENT
Transit NXT -- the ideal tool for translation and localization!

Transit NXT offers you an optimized translation memory and terminology management solution that easily interfaces to all standard file formats, workflow and content management systems. Whether you use Transit NXT for translation, editing and proofreading, terminology management, or project management, you will achieve your goals with its many intelligently designed and powerful features. All this and more in an efficient and user-friendly working environment.

To get a fully-functional trial version, please contact: transit@star-group.net

Visit us at Localization World Conference in Bangkok!

www.star-group.net  

3. What's Wrong with Translation Technology?

In the last few editions of this journal I shared about a couple of webinars that I (and many of you) recently participated in. In case you missed out, here's the tale of their inception and results.

When Lucy Brooks from eCPD Webinars contacted me sometime last year to offer me a slot (or two) for a webinar, I tried to come up with something a little different.

Rather than talking about what I know about translation technology, I wanted to talk to other translators about where they felt we're missing out in translation technology. I planned to take those suggestions to the translation technology developers and get their responses -- whether they could see that an implementation could happen, whether they might already have something like that in place, or whether they might have other suggestions themselves.

Finally, two weeks after the first webinar I planned to do a follow-up webinar to report on the developers' responses and come up with an action plan of sorts.

I'm very happy to report that all of this happened just as planned and with very tangible results.

For the first webinar, titled "Translation Technology -- What's missing and what has gone wrong?", I had already polled you and my Twitter followers, so I was able to present the webinar attendees with some categories of missing or underdeveloped features -- voice recognition, access to external resources, termbases, translation memories/corpora, machine translation, general user-friendliness, exchange standards, and "other" -- that we then fleshed out during the webinar. You can find the specific criteria in a recent copy of the Tool Box Journal.

This long list of proposals and queries went out to more than 20 translation technology providers (essentially everyone I could think of who makes software relevant to freelance translators). The following providers responded, some with very comprehensive answers: Atril (Déjà Vu), KantanMT, Kilgray, Lingua et Machina (Similis), MemSource, Multicorpora (MultiTrans), SDL (Trados Studio), Star (Star Transit), Tauyou, Terminotix (LogiTerm), Wordbee, Wordfast, and XML-INTL (XTM Cloud). Among those companies that did not respond were some long-shots like Microsoft and companies like the Ukrainian AIT, whose owners probably have their minds on more elemental issues right now with recent events in Ukraine. On the other hand, the non-response of other disappointing MIAs such as Across, Heartsome, and Lionbridge, may reveal a certain dismissive attitude toward their users.

As I said, some of the responses were rather detailed, so it would go beyond my allotted space to give you all the responses, but I'm happy to share the compiled results as a large Excel spreadsheet -- just send me an email and I will send it to you.

In the meantime, though, here are some highlights.

Overall, our suggestions were not only welcomed but deeply appreciated by most vendors, who typically were very honest in their assessments. Take, for instance, Wordbee's introductory remark: "Our development plan is quite full. I would be lying if I said it's possible to turn it around now. Therefore I can only say 'we will do it' for a very small part of your list. But all points seem 'logical' and stuff to be done in the mid- or long-term." Great: if that's the result, we'll take it!

A number of tool vendors also used this opportunity to show off some of the features of their tools. Rather than this being annoying, however, it was helpful to see that in some categories we might already have made more progress than we (or I) thought we had. For instance, consider the automatic fixing of fuzzy TM matches and/or MT matches with termbase data or other materials. I was aware that Déjà Vu had been doing this for some time (this was easy to see in the wording of the proposal), but it turns out that Wordfast Pro, MultiTrans, and Star Transit are already using this important feature as well. And just as importantly, memoQ, XTM, and Wordbee are working on implementing it.

Interestingly, MemSource responded to that particular item with this: "None of our clients has asked us to implement fuzzy match repairs yet." If I had any doubts about the importance of this whole exercise, that response convinced me of its value. Far too often, technology (and other) vendors focus on their existing customers rather than assessing what other potential users would like to see in a technology they might want to later adopt. Now they all know, at least what freelancers would like to see.

This is exactly why the developers of relatively new tools were much more open to suggestions. If you look through the Excel spreadsheet, you will find that Wordbee, XTM, and MemSource much more frequently gave answers to the tune of "Great idea, we'll work on it" than companies like Star, Atril, and SDL. This is not to say that the more traditional companies ignored our suggestions. Star, for instance, pleaded for more suggestions of regular expression uses and usage examples. (We had suggested that the use of regular expressions may be powerful but also counter-productive because it essentially creates two different classes of users -- those who are not afraid of using computer language to control their tools and the rest of us).

Kilgray, Atril, and SDL showed themselves very open to creating user interfaces based on user profiles (something that we had suggested). MemSource, by the way, indicated that it wants to take that a step further by analyzing user actions over a period of time to create a user profile.

Overall, SDL had a slightly different strategy in responding to our requests. Rather than promising to do this or that, it consistently pointed to its OpenExchange app store, where third-party developers offer all kinds of apps -- many of which offer or could offer the very features we asked for. SDL's Daniel Brockmann even coined a new term: TEnP -- the Translation Environment Platform (as juxtaposed with TEnT -- Translation Environment Tool -- the term we use to refer to what were formerly called CAT tools). Daniel has a valid and interesting point. In many ways, SDL Trados Studio is (potentially) more able to respond to the needs of users by relying on the third-party developers who recognize that very need and develop solutions for it. The OpenExchange has been around for long enough to have gained some traction, and it will be interesting to see whether other developers (is able to) come up with something comparable.

A couple of other interesting tidbits: When it came to the topic of terminology management, the responses of the tool vendors made it very clear that there really are two very different approaches to terminology: the glossary-like approach assumes that a termbase essentially consists of a source term - target term terminology list, and the much more complex terminology approach satisfies not only the immediate need of the translator but also that of the terminologist. Clearly, tools like SDL MultiTerm, Star TermStar, Kilgray qTerm, and LogiTerm fall into the second category, while the various Wordfast products (Wordfast Pro, Classic, and Anywhere) fall into the glossary category.

The two (statistical) MT providers that responded (KantanMT and Tauyou) are now at least aware of requests that freelance translators have of their tools, in particular that there needs to be an immediate learning of the machine translation engine if the translator adds corrections in the post-editing process. (KantanMT already has an Instant Segment Retraining technology, whereas Tauyou is more cautious: "In some cases, [instant training] is possible, while not in others (or too dangerous).")

Beyond these and many other results (which you can find in the Excel spreadsheet), there were two more immediate action items.

Since there was a strong feeling among participants that it would be beneficial to have an exchange standard for keyboard shortcuts between different tools, I provided the tool vendors with a list of the 20 most important processes within a TEnT that would be helpful to be exchanged. I've also passed on a plea to revive the SRX standard, the standard that is concerned with exchanging segmentation rules between different tools. While the first of these requests is self-explanatory, the second was brought up in connection with the desire to have different sets of segmentation rules for different types of texts (and, of course, languages). It would be great if these kinds of rules didn't have to be developed for each tool but could be shared between the different technologies. SDL Trados Studio does not currently support SRX, and this would be a very helpful change.

Oh, and there is one more outcome: We'll have another comparable webinar in January 2015 to see where we've gotten with our requests and whether there are new ones that have risen to the top. Stay tuned for the announcements. 

ADVERTISEMENT

SDL OpenExchange App Store -- your complete translation toolkit

There are a host of great new apps on the SDL OpenExchange App Store to enable you to extend the functionality of SDL Trados Studio to suit your needs.

You will find over 90 apps available to choose from on the App Store, from converters for handling different file types to managing translation memories and machine translation widgets.

With no cost to download many of the apps, why not try it out today?

Visit Apps Store » 

4. This 'n' That

I'm pretty sure that there are no translators in European languages who are not aware of IATE, the large InterActive Terminology for Europe database. But I have to admit that though my languages are among those covered and my subject area is well-represented, I often used to strike out with IATE searches because it was difficult to know what to look for if the exact term was not available. This has changed now with a predictive typing feature, sort of like Google's suggestion feature, that lets you see suggestions as you enter your terms. It's been available for only a few weeks, but the number of successful searches that I have completed on IATE must have quadrupled during that time. Hope yours will, too!

 

I have mentioned the TAUS Translation Technology Showcase Webinars before. This last week was Across's term to present its system, and though I did not agree with the very concept of Christian Weih's presentation -- in a nutshell, that Across does not believe in exchange standards because it's not in the interest of their users -- I could only admire his chutzpah in unapologetically presenting his company's controversial stand on this topic.

You see, if it were up to technology providers with a good market share, we would not have exchange standards. But generally those companies don't have the guts to say it quite like that -- it's not a very popular thing to do. Christian, on the other hand, couldn't care less. The video is online now, and it's worth a watch.

 

The last time I mentioned the showcase webinars was in the last Tool Box Journal in connection with Kilgray, and I made an unfortunate error on the memoQ cloud offering. In an early edition I had said that "any user of the project manager edition of memoQ (which costs 1.500 euros) can purchase access to the cloud for a fee of 130 euros a month." That's not true. In fact, there are no ownership prerequisites to use the cloud product.

I was made aware of this error shortly after the journal went out and corrected it immediately on Twitter, which has become my go-to place to quickly communicate all kinds of important (and sometimes not-so-important) things. Make sure you subscribe to my (and Jeromobot's) Twitter feed.

 

I don't want to make this last item a whole new article, but I think it's still worth a mention. I was interested in how well the new Trados Studio 2014 aligner matches up with some specialized alignment tools. I downloaded an iPhone manual in English and German and aligned it in Trados Studio 2014, AlignFactory, and the open source LF Aligner (I reported on LF Aligner in this journal). After I tried to get the very unspecific Trados quality settings to produce the same number of matches as the default setting of AlignFactory (which ended up being a setting of 80%), these were the results: of the first 1,000 translation units, Trados misaligned 140, AlignFactory 103, and LF Aligner really struck out and produced way too many errors to even be useful -- which I think must have had to do with the fact that LF Aligner relies on dictionary data and did not find much to match Apple's iPhone-ish.

True, it's all very unscientific and non-representative, using an alignment pair with some inherent problems and a high error rate, but it's still kind of helpful (and the results were what I would have guessed as far as the difference between Trados and AlignFactory).

5. New Password for the Tool Box Archive

As a subscriber to the Premium version of this journal you have access to an archive of Premium journals going back to 2007.

You can access the archive right here. This month the user name is toolbox and the password is changingtimes.

New user names and passwords will be announced in future journals.

The Last Word on the Tool Box Journal

If you would like to promote this journal by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box Journal. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here is a website that mentioned the Tool Box last month:

http://quebec-japon.net/yokoso/ (cool site with a search engine for Japanese-French terms and a forum for French-speaking translators at http://www.quebec-japon.net/lerefuge/).

If you are subscribed to this journal with more than one email address, it would be great if you could unsubscribe redundant addresses through the links Constant Contact offers below.

Should you be  interested in reprinting one of the articles in this journal for promotional purposes, please contact me for information about pricing.

© 2014 International Writers' Group