Tool Box Newsletter Logo

 A computer newsletter for translation professionals


Issue 12-05-209
(the two hundred ninth edition)  
Contents
1.Lingueenee
2. The Better Kind of PDF (Premium Edition)
3. Corpulent Corpora
4. Open Minds, Open Source, Open Future (Premium Edition)
5. Managing Terminology in the Real World
6. MTV -- Machine Translation Views (Premium Edition)
7. Taming the Red Panda
8. New Password for the Tool Box Newsletter Archive
The Last Word on the Tool Kit
In Plain View

Last week I arrived for a lengthy appointment in a Eugene hospital, and I had a strange and wonderful experience. As the first of what would be many medical professionals entered my room, she introduced herself slowly and clearly and then pointed to a sharp-looking young man who introduced himself: "Und ich bin der Dolmetscher (And I am the interpreter)." My wife and I looked at each other in puzzlement. After all, my English is far from perfect, but I typically don't need an interpreter. Some over-eager beaver (I mean, duck -- you have to live in Oregon or really like US college sports to get that joke...) had seen my nationality on my chart and ordered an interpreter -- just in case. I ended up feeling sorry for the poor chap whom we had to send home, but I felt oddly blessed. It's too bad that a situation like this is far from ordinary.

To happen on a much more regular basis, there must be a great deal more awareness in the mind of the general public. (OK, there might have to be more finances as well, but stay with me for a second.)

Fortunately, there is a growing number of resources that might remedy this situation -- like the open call for Wikipedia pages profiling good translators, or the translation-specific URL shortener xl8.in. I love those ideas.

Another idea that I love -- really love, in fact - was hatched in Nataly Kelly's and my off-the-wall brains to celebrate our professions as translators and interpreters as well as the upcoming launch of our Found in Translation book. We are creating a video to shine a spotlight on our profession, and we need your help. All you need to do is say "I am an interpreter" (in your non-English language) or hold up a sign that reads "I am a translator" (in your non-English language) while making a short video of yourself. Send us the video by July 1 and you might very well see yourself in our compiled video that will virtually travel around the world and state loud and clear and proud who we are. You can find all the information about this on our website.

And just in case you missed it, here is the address again with the clever new URL shortener: xl8.in/video.

 

I find myself missing the more frequent editions of this newsletter: so much to talk about, so little space. Articles on Duolingo, The Big Wave and the new Déjà Vu TEAMserver had to go on the cutting floor for this edition, and there are still so many characters in the Characters with Character series I'd like to write about -- let's hope they make it into the next newsletter. Still, I do want to share an image that has been burned into my mind for several days now. One of the many 0 Through 9 works by Jasper Johns, this is an ever-moving depiction of the Arabic numerals that seem to shift once you feel you've settled your sight on just one. To me this speaks of many things, including language and translation: 

1 through 9 

1. Lingueenee

Although Linguee is available in only a handful of languages (English into and out of German, Spanish, French, and Portuguese), I would bet that in the world of translators it is much more widely known than that. That's not surprising to those of us who see an obvious value in what it offers for us and our work, but it was a surprise for the founders of Linguee. You see, they created the site for "average" folks who needed a quick way to find translations from real-life usage cases within their original contexts (even though most of them don't really understand the concept of context as you and I very much know.)

When I wrote about Linguee for the first time three years ago, I said this:

What is Linguee? Since I am a man of few words, I will simply say it's a quantum step in search technology for translators.

Really, I think that Linguee may just change the way that many of us look for information from now on. Here is what it is: It's a very large corpus of English into German into English data [with the addition now of the languages mentioned above] of web-based translated materials. The two guys behind it have found ways to have web crawlers detect translated content online and match that up with the help of a 50,000+ entry dictionary (the EN<>DE version can be downloaded right here) and other web-based dictionaries. To look up a term or complete phrase, just enter it into the search box; the matches that are displayed are complete segment matches with the terms in question (both in source and target) highlighted. At first glance the data contains no metadata (origin, subject matter, etc.), but at second glance you will notice the links to the originating sites, giving you all the metadata you could want. To search you don't have to register; as a registered user you can evaluate the translations and correct them, or you can add entries to the dictionary, which in turn are used to fine-tune the matches.

And in the following year I wrote:

And just to let you know: German translators are no longer part of the controversy whether one should say BCE/CE vs. BC/AD when referring to time. They just say BL and AL -- Before Linguee and After Linguee.

I'm aware that these are pretty strong endorsements, but they're mostly justified (there was a time when their corpus was full of machine-translate


d patents that didn't help very much, but they seem to have been kicked out). One of the founders, Gereon Frahling, worked for Google for a while, and this becomes apparent on many levels, including the site's visionary reach. While the Linguee team was surprised at the enthusiastic embrace by the translator community, their vision for the future goes far beyond that, and the numbers prove it: presently they have 7.6 million hits per month. That's a lot. To put this into perspective: it's even more hits than Jeromobot gets on YouTube!

One other area where the Google heritage becomes clear is an ongoing push for improvement and innovation. When I described the site three years ago, it was quite different and much less streamlined than it is today. What's particularly impressive today is its much stronger focus on terminology beyond the occurrence in the sample sentences. The supported languages all come with in-depth morphological support, so there is an immediate analysis of the grammatical form of the word you are looking for, including a listing of the respective infinitives and their various meanings.

The newest offering that was released just a few days ago and that I had a chance to discuss with Gereon encompasses two paid models, Linguee Premium and Linguee Professional. Currently they are in beta so they are still free, but from the end of June they will cost a monthly fee (4.99 or 9.99 euro per month if you buy a whole year of service).

My initial response when I heard about it was skeptical, but after having used it, I'm really impressed. You see, at first I thought this would just be another version of IntelliWebSearch in paid form -- that is, a way to quickly search for a term within any Windows application on Linguee. And that's true to some degree. You can search from almost any Windows application (more on that later) by simply holding the Ctrl key, clicking on a term, and executing a search within Linguee. What makes it really powerful, though, is that it not only searches for the one word you just clicked on but for the context of that term as well, so the matches that are displayed are not random but specifically geared to the kind of text that you are presently translating. This is very, very cool!

What's the difference between the two products? Essentially, one is created for the non-professional translator whereas the other -- yes, it's the more expensive one -- is for the pro. The difference between the versions is that only the Professional version is enabled to work within translation environment programs. It allows you to create a profile, and your services will be advertised a certain number of months on the site (in the future these will also be targeted to specific regions and areas of expertise). There will also be a search history (kind of like Google does to target search results more specifically), and you'll be able to search within certain preferred sources (though the number of sources is not particularly impressive).

I'm excited about the user profile and the targeted advertising since it offers a potentially great way to gain direct clients. And as far as the TEnT restrictions go, the Linguee team still has a little bit of work to do (but that's why the products are still in beta).

I ran some compatibility tests with a few tools. It worked great in Trados, Déjà Vu, Fluency, Transit, and Poedit. There were some problems with OmegaT, which seemed to be due to a faulty recognition of word boundaries. With Similis and memoQ it worked until I clicked on a word bordering a tag, which seemed to disable the system. It actually enabled the Premium version's lack of support for TEnTs, scrambling the search terms and spoiling the successful recognition within Linguee.

I passed those problems on to Linguee as well as a conflict with Dragon NaturallySpeaking that I encountered, and I'm sure they will be fixed quickly (or may already be fixed once this newsletter arrives in your inbox).

So far this offer is available only in the EN<>DE language directions, but the other supported languages will follow later this year after the German version is successfully launched.

ADVERTISEMENT

Empower your teams with TEAMserver by Atril, the most effective and affordable all-in-one server solution on the market!

Learn more here.

2. The Better Kind of PDF (Premium Edition)

I recently stumbled on a rather interesting twist to PDFs and the ability to process them directly in translation environment tools.

It has to do with the way that PDFs can be created and then processed in the free office suites OpenOffice.org and LibreOffice. (LibreOffice is an independent branch-off of OpenOffice.org, which happened in the fall of 2010 after Sun, the previous owner of OpenOffice.org, was sold to Oracle. Since that split, the development of LibreOffice has been more proactive, though OpenOffice.org, which has been donated to the Apache Software Foundation, just released a new version 3.4.)

LibreOffice (by default) and OpenOffice.org (through an extension that you can download right here) offer the option of creating a PDF that has a LibreOffice/OpenOffice.org file embedded, making this PDF completely editable within its originating application. In fact, when you open this PDF within LibreOffice/OpenOffice.org, it automatically opens in the word processing/spreadsheet/presentation component it was created in just like a normal document. If it's "only" a normal PDF that is not directly editable, it opens in the graphics component.

These kinds of PDF files are called hybrid PDFs, and you can create them by selecting File> Export as PDF> Create hybrid file (OpenOffice.org) or Embed OpenDocument file (LibreOffice).

I've always maintained that PDFs  will never be completely processable by translation environment tools, but I'm beginning to have the faint hope that this might just be a possibility with a feature like this for some PDFs. The feature just has to get used more widely, and Microsoft Office needs to add it as well.

 

Of course, there are plenty of translation environment tools that offer PDF compatibility, but in essence the tools have simply integrated one of the readily available PDF converters. For small and less heavily edited PDFs this often works fine, but for more complex PDFs you're usually better off converting the PDF using your own preferred method and then processing it in a TEnT.

The latest TEnT that now offers this PDF compatibility is the latest version of Déjà  Vu X2. It uses the same PDF converter as VisualTran Mate, Alchemy Publisher, and Wordfast Pro -- the BCL Technologies plugin -- but Déjà Vu X2 automatically also uses CodeZapper, a little utility that strips the resulting Word files of unnecessary codes, resulting in much cleaner files than some of its competitors.

3. Corpulent Corpora

Jason Christensen wrote to me the other day asking this:

Is it possible to use such online parallel corpora as Parasol (a Parallel Corpus of Slavic and Other Languages) with computer-assisted translation tools? Specifically, in translating from Russian to English, would it be possible to use translation/terminology management systems with these online parallel corpora in order to create a terminology bank?

This was my answer:

The easiest way to use a corpus in a translation environment tool would be to convert it to TMX and import it as a translation memory. The easiest way to convert to TMX typically is via the Xbench tool from ApSIC -- you only have to get the corpus into a delimited text format to get it into Xbench from which you can then export it as TMX.

I would be cautious about the usefulness of a corpus. For machine translation training purposes they are usually very useful, but for translation memory purposes it is usually more useful to have more targeted data.

I hope I did not frustrate Jason too much with my response, because it turns out that you have to actually apply for login credentials to search the database with a specific interface rather than downloading that particular corpus. (Parasol is an aligned corpus of translated and original belletristic texts in Slavic and some other languages. You can find a list of languages and the texts that are used right here.)

Still, what I mention in my response still stands for many other corpora that are downloadable. I remember when the EU first published its gigantic DGT corpus, which contained translated laws and regulations of the EU in its many languages. I spent a good day downloading, massaging, and importing the EN<>DE part into a translation memory and I'm ashamed to admit that I have used that particular TM very rarely. And the few times I did I think it would have been faster to Google what I needed to look for. If I specialized in translating EU law it might have proved a lot more useful, but if you're not specialized in the area that your TM excels in it's just a lot of data that makes processing a lot slower, and it might even make you err on the side of suggestions that you shouldn't use.

And yet it's good to know what's out there. As a matter of fact, that EU TM has just been rereleased and is three times its original size. You can find and download it right here. If you would like some more information about it, you can find it in this interesting (and very dense) article, which lists all the languages that can be found in this linguistic monster. Overall, this is really one of the easier-to-use corpora as it's delivered right in TMX, the TM exchange format that is compatible with virtually every tool.

Another not quite so easy-to-use corpus is the European Parliament Proceedings Parallel Corpus, which was also just rereleased in version 7. You can find that right here.

ADVERTISEMENT

Why do people work with memoQ?

The thing I really love about memoQ is that it always offers you (at least) two different ways to face your everyday issues. So you experience "safety" because you know you are relying on a tool that will not let you down. And it's "fun" because it feels like a video game with so many options and levels to discover!

Renato Renno, The Foreign Friend srl

To learn more, visit www.kilgray.com and download the 45 trial version. 

4. Open Minds, Open Source, Open Future (Premium Edition)

Olanto stands for Open Language Tools (not to be confused with the Oracle/Java Open Language Tools). It's a non-profit foundation in Switzerland that made its first foray into the public limelight at a recent conference in Luxembourg. After talking to Karim Benzineb, one of its founding members, I'm cautiously optimistic about what they are planning to do.

(The reason for my caution is that some other open-source projects, including OpenTM2 and in particular openTMS, have not really shown much of a track record despite great initial fanfare.)

It seems, though, that Karim and his team are fairly realistic about what they want to achieve, even though it's an awful lot.

Next month they're going to post two tools along with their source code: myCAT, a bi-text aligner and concordancer (a tool that allows you to take separate texts in source and target, align them into a bilingual corpus or translation memory, and search through that to locate translations); and myMT, a statistical machine translation tool based on the open-source MT engine Moses (more on Moses in the MTV article below.)

Both of these tools are fully developed (in fact, the development of myCAT was financed by none other than CERN), and the hold-up right now is due to an ongoing review of some legal documents in the Swiss bureaucracy.

At the end of this year, the website promises that these two tools will be followed by four more translation programs:

  • translation memory manager
  • terminology management system
  • cross-lingual search engine
  • content management system

How did Donald Trump put it so eloquently? "As long as you're going to think anyway, think big."

While these goals may sound ambitious, Karim, who started out as a translator himself and began building software at the request of his clients, seems very down to earth. His tools are already in use at a number of institutional clients, mostly located in Switzerland. There's a reason for that concentration -- namely, that the developers strongly suggest that their tools be installed and maintained only with adequate professional on-site assistance. Their website details what you can do to become a certified Olanto provider to support local customers with installation and support.

I have not yet tested the tools, which at this point are not designed for freelancers or very small companies -- though there is at least one slightly larger LSP in Switzerland that is already using some of the tools -- but I really hope that we'll be hearing more from Olanto. In fact, I wouldn't be at all surprised if some of Olanto's tools ended up in the tool arsenal of the European Union after its botched attempt to locate a tool provider to replace SDL Trados. Unfortunate experiences (from the EU's point of view) with Systran ended up in large court-ordered sums of money that had to be paid to Systran, so the EU is very open to some kind of open-source solution.

5. Managing Terminology in the Real World

Twice in the last couple of years I have asked you to respond to a survey by doctoral student Marta Gómez Palou Allard. This has now resulted in Marta's (excuse me: Dr. Allard's) thesis Managing Terminology for Translation Using Translation Environment Tools, which in its entirety is available for download. The fact that it's downloadable is no accident; in fact, Marta states in her Acknowledgments that she hopes to "be able to repay the community's generosity with the contribution this research brings to our field" -- and here she is talking about you.

The contribution that she makes is quite remarkable, and I think that we are repaid with much more than we ever invested ourselves.

Marta contends that there is a wide gap between how terminologists and translators view and use terminology. Accordingly, there's also a chasm between what is taught at universities and what translators actually do later on. But I'm most interested in her related contention that TEnTs (a term she thankfully uses throughout) are really geared more to the needs of terminologists than translators. She also looks at the percentage of the embedded terminology components that are actually systematically used by translators who use TEnTs; her not surprising but still frustrating answer of 40% stands in stark contrast to the obvious usefulness of terminology work.

Rather than proposing that translators should beef up on their terminology theory and change their ways, she actually proposes just the opposite: that terminologists should take real-life working scenarios into account when teaching classes or helping to build tools.

She also examines TBX, the terminology exchange standard, and comes to the conclusion that it could serve a very important purpose, but only in its TBX-Basic implementation, which has a much more simplified and more real-life approach to terminology work than its full-fledged cousin.

I love to see academic research that is modeled on real life and aims to improve it rather than theorize it, and this is a beautiful case in point. So, thank you, Tool Box readers, for responding to the surveys, and thank you, Marta, for making sense of the data.

ADVERTISEMENT

SDL Translation Games: Round 2

Thank you for attending the opening games. Thousands of you took part in the archery game, all testing your accuracy to be translation champion.

Our next game is based on consistency and takes part on the Javelin runway. Take part for a chance to win an iPad!

Join the games »

We have huge discounts of up to 35% off this month! So, maximize your consistency through the use of innovative features in SDL Trados Studio 2011 and save.

Discover our great offers » 

6. MTV -- Machine Translation Views (Premium Edition)

When we hear "machine translation," we obviously think of computers translating -- or trying to translate -- something from language A to language B. And that is indeed typically the case, unless you have something like the Microsoft Contextual Thesaurus, an MT system that translates English into English to thesaurize your phrases. (Not surprisingly, when I entered "thesaurize your phrases" it did not give me any alternate suggestions for "thesaurize" but plenty for "phrases" and even for "your".)

It's really quite a revelation, and I think you'll enjoy playing with it and very possibly actually using it. While it uses essentially the same logic as other MT systems, it's not a particularly fair comparison because it doesn't have to deal with many of the issues you have when going from one language into another (different syntax, morphology etc.). (So don't show it to your clients to illustrate the weakness of MT -- they might just quit working with you and rely on the Bing Translator instead!)

Microsoft has also published the API to the Contextual Thesaurus so that you can embed it into other applications. You can find it right here.

TAUS, the Translation Automation User Society, has released a report, authored by Achim Ruopp, about the current state of the open-source MT system Moses. This system allows you to create an MT system, provided that you have lots of high-quality data and are really computer-savvy. According to the report, that second condition is slightly changing (don't get any ideas, though, it's still not something that you or I should attempt) since there are more tools available to help with the process (I've reported on some of them in the past).

Overall, the report reaches this conclusion:

The major finding from the spring 2012 Moses User Survey is the fact that adoption of the open source package remains strong and the majority of users are now moving from early adoption to production use. This shift in the industry means changing priorities and could lead to tension with academic users. But thanks to the efforts by the Moses team, the community is healthy and active, and with renewed funding we believe any possible tensions can be overcome. What the industry should do, however, is to step up and give something back to the project so that the various gaps can be filled through mutual sharing. With continued contributions from both academia and industry, we believe that the project can continue its trajectory and become a tremendous success.

Taking Achim's last point to heart, Adobe has jumped into the fray by releasing the Adobe Moses Tools that they built for themselves in the process of working with Moses. You can find more information about these tools in this slideshow. I'm going to go out on a limb and predict that most of you will respond like me: yeah, right, I don't even know what these terms mean, let alone understand the technology behind it. 

7. Taming the Red Panda

It seems like forever now that Mozilla has been promising to plug some memory leaks within Firefox -- but many users still suffer from Firefox's large accumulated usage of memory. If you're among them and don't want to give up on Firefox (and why should you when it's the only browser you can dress up in Jeromobot style?), you might want to check out Firemin, a little application that keeps the memory usage of Firefox to what you want it to be. Mine has recently been around 800 KB (!) without any noticeable slowing in the response of the browser. 

8. New Password for the Tool Kit Archive

As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to May 2008.

You can access the archive right here. This month the user name is toolbox and the password is tuliptree.

New user names and passwords will be announced in future newsletters.

The Last Word on the Tool Box Newsletter

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box newsletter. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here are a couple of websites that mentioned the Tool Box last month:

asmarttranslatorsreunion.wordpress.com

steve-dyson.blogspot.com 

© 2012 International Writers' Group