Tool Box Newsletter Logo

 A computer newsletter for translation professionals


Issue 13-6-223
(the two hundred twenty-third edition)  
Contents
1.Getting There and Beyond (Premium Edition)
2. Last Call!
3. Swiss All the Way
4. The Next Step (Premium Edition)
5. New Password for the Tool Box Newsletter Archive
The Last Word on the Tool Box
You Do the Math

Imagine this: Two competing, relatively sophisticated applications to aid in, say, higher math are developed by two of the largest high-tech companies in the world. Each of these companies promotes its product heavily and proudly, proclaiming, in fact, that it will be one of the cornerstones of its future development plan. Not the development plan for the math application alone, mind you, but for virtually every other product they develop.

There is only one problem: about 8 out of 10 times the programs don't get their calculations right. They usually get pretty close to the actual results in their calculations, but just not quite there. Every once in a while they hit the bull's eye, but for every time that happens another calculation goes completely haywire. In those cases, the calculated result is not even pretty close; instead, it's on the very opposite side of the spectrum.

Well, yes, say the very confident representatives of the companies, the programs are not perfect, but it is so hard to get it right all the time, so it's amazing how well we do.

And, aside from some pretentious math geeks (who always seem to think that they're smarter than everyone anyway -- Gosh!), pretty much everyone agrees. In fact the leading media outlets regularly run  major stories on how wonderfully close many of those calculations come much of the time and marvel at the ingenious developers and the powerful computing processes that make all this possible. Some even speculate that one day -- ONE GLORIOUS  DAY! -- the programs will always get it right.

Well, virtually always.

 

I'm sure you've seen through my thinly veiled analogy. Of course, one could take issue with how closely a math program relates to a machine translation program, but sometimes odd comparisons allow us to view familiar things in a new and helpful context.

In fact, we should be proud to represent the profession that is performing what Google and Microsoft products unabashedly and proudly strive to achieve. It really is hard to get those translations just right, and because everyone knows it, no one dares to laugh about the imperfections of the results.

But aside from the weather forecast, can you think of any other area where this level of inexactitude is permissible? I can't, and we should remind ourselves of that fact when we lose perspective on how relevant our profession is.

It's like the quote that I found last week: Google Translate is "like a dog walking on its hind legs: although it is not done well, you are surprised to find it done at all."

(And before I receive a host of critical messages: no, the semi-free machine translation engines from Google and Microsoft are not the same as customized MT engines with appropriately integrated post-editing services, but that still doesn't make any of the foregoing irrelevant.)

1. Getting There and Beyond  (Premium Edition)

In many ways, memoQ has long since "arrived." It is the clear frontrunner among its competitors against SDL Trados in the TEnT (translation environment tool) market, and it is a known entity in all sectors of the language industry. But I realized only last week how much Kilgray, the maker of memoQ, really has established itself as a company when I talked with its CEO István Lengel. We naturally chatted about the new version of memoQ (more about that in a second), but we also talked about transitions within the company itself. It both took me by surprise and made me marvel at the maturity of the company to learn that much of the current leadership is going to step back to let some "fresher blood" take over. This doesn't mean that current CEO István or CTO Gábor Ugray or head of marketing Sandor Papp will completely resign. They will stay in the background in reduced roles with continued oversight, but the day-to-day business will be led by others (Gábor's shoes, for instance, will be filled by ex-Passolo/SDL developer Florian Sachse).

Not too shabby for guys who have barely reached their mid-thirties, wouldn't you think? I can't wait to see how this transition will work out, but my gut feeling is that it'll turn out well.

Let's talk about the tool.

Every tool on the market -- whether Trados Studio, Wordfast, Déjà Vu, or Multitrans -- has features that are distinguishing. memoQ is no exception, and as I reviewed them for this article, I was surprised to see how many there actually are.

Here is a list of distinguishing features that were implemented in previous editions and are unique to memoQ (and I mean "unique" in its language-stickler definition of "one of a kind"):

  • Version history: memoQ allows you to keep a full-blown version history of both source and target documents.
  • Translation memory driven segmentation: This feature matches the segmentation (how a translation file is separated into segments) to the associated translation memory. This is especially helpful if you have a translation memory from a different tool that may have used a different set of segmentation rules. By employing this feature, the new translation file will match its segmentation to translation units in the TM.
  • Content source: This allows you to select a content source so you don't have to add files manually. Instead, memoQ looks out for new or changed documents in that source and automatically imports the changes.
  • LanguageTerminal: This feature gives you the option to process unconverted InDesign files into a format that can be processed by virtually any translation environment tool, backup your memoQ file projects, exchange some memoQ resources, and (now in the latest edition) store some client- and project-specific information and perform some basic business functions. All this is cloud-based.

Other features, including the use of corpora ("LiveDocs") alongside the traditional TM, are not unique in the above-mentioned sense (for instance, MultiCorpora uses corpora as well) but still are distinguishing features in comparison to many tools.

What was left for the latest version, memoQ 2013 (6.5)? (By the way, I wish they had refrained from adopting the now oh-so-common convention of naming versions according to year. To me it always has the taste of cheap and under-handed pressure marketing. After all, who wants to admit to using a tool whose very name shows that it's two years old?)

As with previous major new versions, this also had a theme. This time around it's "quality," and that really is mostly represented in two new features: Linguistic Quality Assurance (LQA) and fuzzy terminology matching.

Neither sounds too impressive?  Let's see what you think after you read a little more about them.

Quality assurance in or around translation environment tools is mostly mechanical. Here's what I mean: The errors that the tools look for can be found by a computer because they compare source and target segments for things like differences in punctuation, length, or consistency. They also look for erroneous spaces or incorrect terminology on the basis of a specified glossary (more on that later). Linguistic QA à la memoQ brings in actual linguistic expertise that is represented by you and me. It's a system that sets up a customizable rating system where each kind of error that the editor detects in the process of editing gets a score depending on its severity. This provides very comprehensive reports as well as a pass or fail grade for the translation when the scores are summed up at the end of the editing (or proofreading or client review). Many of you who work as editors for larger LSPs have already worked with systems like this, only those were done manually in Excel spreadsheets or something similar rather than as part of the editing process.

memoQ uses a number of preconfigured QA models that can be modified, or you can use your own from the get-go. Here is the LISA model: 

I'd say that within, say, 6 months other tools will have something comparable. Long live the forces of the free market!

memoQ's new fuzzy terminology matching is a feature that most memoQ users will -- or at least should -- be using (though at this point, German translators will find it more useful than others -- danke, Florian!).

Isn't this what most other competing tools have as well, you might ask? Yes, most of them do. But the competitors apply the simple fuzzy TM algorithms that just compare sequences of characters (Star Transit and Across so far had been the exception in that for a small number of European languages). memoQ's fuzzy term matching is different.

For a good overview of this feature, be sure to read Kevin Lossner's blog. It comes down to this: rather than looking only for similarities in character sequence (which is the default concept of "fuzzy"), it also looks for things like language-specific character mappings (which denote typical changes of characters such as a -> ä, o -> ö or u -> ü in German for plural forms such as Kuss/Küsse) and character sequences in compound words. This last feature means that if you have a term like "Dateispeicherpfad" (file storage location), both "Datei" and "Speicher" would be recognized as term matches; "Pfad" would not be recognized because this feature only kicks in once a threshold of at least five characters is reached. (The compound recognition function is disabled in the terminology QA feature.)

So much for the good news. The bad news is that the character mapping has so far been implemented only for German, Hungarian, Italian, and Spanish, and the compound word recognition only for German. But is that really bad news?

Also, ich finde das geht in Ordnung!

Seriously, now: I think that tool vendors should resist the pressure to implement features only if they work equally well for many, most, or all languages just to make it equally marketable across the globe. Ultimately this all-or-none mentality impedes progress. While it's important for Kilgray to now expand the pool of applicable languages (though compound rules make sense only for languages with lovely words like Rindfleischetikettierungsüberwachungsaufgabenübertragungsgesetz), I'm glad that a first positive step has been taken.

Of course, the question is how far this kind of fuzzy matching will take us in comparison to term recognition that relies on a true morphological analysis. Kilgray has run some comparative tests of the new fuzzy model against the stemming of the Hunspell engine in German with the preliminary result that the fuzzy matching might actually deliver some better results with fewer false errors and missed recognitions. That would be great, though it remains to be seen whether this is true for other languages also. Ideally, though, we would have a morphological approach that uses lemmatization (thank you, readers, for that information), with actual word lists to find that "good" is the lemma for "better" or "bad" for "worse." Whether the new system will actually prove to be "better" or "worse" remains to be seen, but regardless there are a lot of exciting possibilities out there whose development will certainly make our tools smarter.

There are other new features as well, of course:

  • an API for the MT component (so that MT developers can develop their own connectors to memoQ if they need to);
  • a relatively sophisticated PO Gettext filter (the primary file format used in many open-source localization projects);
  • support for the new xliff:doc and TIPP package formats;
  • a web search feature that is comparable to (an anemic) IntelliWebSearch;
  • the concept of edit distance (this idea was first developed by the makers of MemSource and is an attempt to figure out how much effort is spent on the post-editing of machine-translated segments); and
  • improvements to the web-based WebTrans translation interface. (So far this is essentially a slightly feature-impoverished version of the desktop version of memoQ as a solution for LSPs whose translators and especially editors or proofreaders don't want to hassle with the desktop's version install. Note: I wish the team could get that environment to feature parity with the desktop version.)

Also, Kilgray assures us that not only is this a version that holds up "quality" as its own banner, it is quality-assured by much improved processes as no version before. (I did find some inconsistencies in testing the LQA feature, but I'll take their word for it.)

Without a doubt, this new version of memoQ is another step toward providing a continued strong contender to other tools, especially the market leader, and it's also a pointer to the future with some new features that we will surely see in other tools in some way or the other in the mid-term future. And that can only be in everyone's interest.

2. Last Call!

I've just been told by the University of Maryland that my three-day translation technology workshop still needs a bunch of sign-ups. You can read more about it by following the link above (you'll need to scroll down a little on that page). During those three days we'll be focusing on CAT tools (translation environment tools), complex file formats, data resources, machine translation, and the many smaller tools that are needed for the modern translator.

The dates were originally announced as July 25-27 but they are in fact July 26-28. As I mentioned before, it's going to be very hands-on and practical -- no matter whether you are a beginner or have been translating for a long time. Plus, the ATA has agreed to approve the workshop for 10 Continued Education points for participants.

And really important: the application deadline is July 1!

Hope to see you there.

ADVERTISEMENT

THE SDL ACADEMY & FREE WEBINARS

We understand that the modern translator needs to have a wide range of skills from business acumen to IT proficiency.

The NEW SDL ACADEMY has been developed as an online hub of information and learning resources for freelance translators, helping you focus on your professional development, to enhance your skills and help you find new ways to grow your business..

JOIN THE SDL ACADEMY TODAY

FREE WEBINARS to help you explore SDL Trados Studio 2011

3. Swiss All the Way

Most of you know that I like to mention where tool vendors come from -- partly to emphasize that we truly are an international bunch of folks, but also because it intrigues me when countries that would typically not be on my mental list of hotbeds of development are among the places where translation technology is being developed, including countries like Uruguay or Ukraine. But only rarely does a vendor whom I interview stress that his company is located in a certain country because of the positive associations with it -- especially for his particular product.

Well, that's what happened when I talked to (the native German) Mirko Plitt of Modulo Language about his Swiss Post-Editing Score product. A quality assurance product from the land of super-accurate watches! Can there be a better match?!

Now, you might ask why "Swiss" is mentioned in the name of the tool at all. It risks misleading potential non-Swiss clients, and it's surely not because "Swiss Post-Editing Score" is an easy-flowing and highly marketable name! Well, listen to what it does and you might agree that the particular Swiss brand of precision goes to the heart of what the tool does.

Consider this actual example (don't worry if your German is a little rusty -- it's OK if you don't really understand):  

Titelverteidiger Rafael Nadal hat erneut das Endspiel der French Open erreicht.

Nun trifft Nadal am Sonntag entweder auf den französischen Lokalmatador Jo-Wilfried Tsonga oder seinen Landsmann David Ferrer. Djokovic konnte sich damit nicht für die letztjährige Niederlage im Endspiel revanchieren. "Das ist ein sehr spezieller Sieg für mich", sagte Nadal nach seinem 20. Sieg im 35. Duell mit seinem Dauerrivalen: "Dieser Platz ist für mich etwas ganz Besonderes. Novak wird in einem anderen Jahr hier gewinnen, er ist ein großer Champion."

And then this:

Defending champion Rafael Nadal has again reached the final hell of the French Open.

Now Nadal meets on Sunday either on the French local hero Jo-Wilfried Tsonga or his compatriot experienced David Ferrer. Djokovic could not reciprocate for last year's defeat in staying the final. "This is a very special win for me," said ago Nadal after his 20th victory in the 35th duel with rival duration: "This place is for me something special. Novak will win in another year here, he's was a great champion."

Without a doubt you will stumble over the unidiomatic English -- that's not surprising since the English text was produced by machine translation. But you should really stumble over the surprising "final hell" of the French Open (unless you were just deeply engrossed in re-reading Dante and thought final hells could be found everywhere).

The "final hell" was inserted by Swiss Post-Editing Score. And it did that for no other reason than to be caught by the post-editor of this machine-translated text. The makers of SPES (sorry about the acronym, but my fingers are tired!) are not trying to evaluate the quality of machine translation; instead, they want to give MT users a way to evaluate post-editors. Companies that use MT will tell you it's hard to find good post-editors (if they find any at all), and there really are only very subjective ways to evaluate their quality. What about combining the well-proven ideas of sampling and error injection and merging them with measures of editing distance (how much a post-editor changes in the machine-translated text) to identify positive or negative outliers? This is what SPES does.

By injecting errors and automatically checking whether those have been corrected, it can come up with reports on the reliability of the individual translators. And if those number are also related to editing distance (and word count), it's possible to see whether the post-editor was just an (unnecessarily) eager beaver and corrected everything and anything anyway, or whether she focused on the "right kind of errors" (and, yes, dear passionate MT foe, I know, I know...).

You can see a sample "dashboard" report right here.

So far this product is in an alpha stage with only two LSPs using it. In fact, how they're using it at this point is less than sophisticated: they have to have the error insertion and the analysis done via intermediate XLIFF files. But when the product is launched in October it will be introduced as an API, allowing it to be directly integrated into any machine translation engine so the process will be automated. The price? It will be charged as a service, and will be approximately at the level of what Google charges for its Google Translate API, says Marko.

I'll let you know how this tool and concept progress.

Oh, Marko brought up something else that was interesting. As I mentioned above, it's very difficult to find qualified MT post-editors. How to eventually solve this? Let the laws of the market sort it out by significantly raising compensation. That's an idea! (And, yes again, dear MT foe, I know that this still does not mean that you'll touch it with a ten-foot pole.)

ADVERTISEMENT

Kilgray Translation Technologies released memoQ 2013

Adding significant functionality to the previous version, memoQ 2013 offers a wealth of productivity boosters for freelance translators, language service providers and enterprise customers alike.

Language Quality Assurance models, discussions, edit distance, fuzzy terminology lookup, improvements to Language Terminal, a set of file filter formats are just a few of the many new features that memoQ 2013 brings.

Attend Kilgray's free webinars, download the fully functional trial version of memoQ 2013 from www.kilgray.com/downloads, give it a try and join the ever-growing team of memoQ users!

4. The Next Step (Premium Edition)

The Muscovite company ITI ("International Translation and Informatics") was originally a software development company but quickly morphed into a localization and translation company. But it apparently never quite shook its development roots, because the product that it developed first for its internal use and has now opened up for the rest of us is both sophisticated and well thought-out.

One of the dilemmas that most LSPs face is that while it would be great to have one kind of translation environment technology for all projects (so that resources and expertise can be centralized), in reality they need to be able to work with the clients' requirements (provided they work with clients who have invested in translation technology themselves). This is even more true for single-language vendors (SLVs) who typically work for MLVs who do have some kind of technology in place. So, what to do about the resources that are being assembled in many different technology silos? Yes, there are exchange formats for translation memories, terminology databases, and translation files, but the reality is that the exchange is typically less than straightforward. And especially when it comes to terminology data where the many different TEnTs have such different approaches, ranging all the way from complex concept-based databases to simple glossaries, lots of data is lost in the transfer process. Plus, as our discussion on morphological recognition has shown, there really is a lot to do for any of the existing term databases as far as functionality goes.

This is where MultiQA enters. Even though its name suggests that it's primarily a quality assurance tool (and it does provide quality assurance -- more on that later), it really is a complex web-based terminology management tool with an attached desktop component.

Let's start with the desktop component. Rather than redeveloping a tool from scratch, ITI's developers forked (established a separate branch of development) off the open-source tool GoldenDict. GoldenDict is a tool that allows you to access a number of preconfigured and configurable websites in a search for terminology. It distinguishes itself from other tools that do similar things in that it uses the Hunspell stemmers for German, English, Spanish, French, Italian, Portuguese, and Russian so that inflected forms of words can be found on the web as well. Searches can be performed by highlighting a word and searching with a keyboard shortcut.

In the MultiQA version of GoldenDict (which can be downloaded on MultiQA's homepage), one of those searchable sites is, not surprisingly, your particular login to terminology resources stored on the MultiQA server. (In that respect it's similar to the SDL MultiTerm Widget.) But not only can you search for terms through that interface, you can also enter new terms into the termbase (the necessary dialog box is a little fickle for my taste, but it works once you know not to take your cursor off it).

To be able to use a "glossary" (it really is a termbase rather than a simple source-target word list) in MultiQA you'll need to set one up. And don't expect to log on and immediately know what to do -- it's rather complex. Once you understand the logic you will see its coherence, but only then.... (I spent about an hour with the team doing an interview about the tool and it still took me a couple of hours to "get" the actual workings of the system.)

But it's probably slightly unfair to criticize the complex interface: there really are very few limits to the complexity of the termbases you can set up -- and that requires a complex setup. (I would recommend that ITI produce a number of short videos explaining the individual steps that are necessary. While the integrated help system is helpful, it would be good to have another medium of instruction as well.)

During the setup of hierarchical termbases, you will set up not only the languages but also up to five different levels of user roles, a number of different statuses that any term entry can have (there is a special emphasis on the lifespan of terms), reports that project managers can run, glossary imports (and exports) through Excel and TBX, which fields are mandatory to enter, and on and on. You can find good descriptions of some of those features in this blog entry by Valerij Tomarenko.

All this might not be so different from other complex systems like SDL MultiTerm, TermStar, and qTerm, but there is one fundamental difference.

For English, Russian, German, Ukrainian, Kazakh, Chinese, Spanish, Polish, Slovak, Czech, and Norwegian (Bokmål), there are fully functional morphological engines implemented for nouns, adjectives, and prepositional phrases that work beautifully in the few tests that I ran. To actually see how they work you can call up the Automatic Parsing pane for each of the terms that you entered in the termbase (and that fall within these languages) and see what forms are automatically recognized. My German term "Indikator" was parsed beautifully (see below). Other more irregular terms did not fare that well, but it was still very impressive, and it's easy to manually correct and then save the corrections. 

  

Let's talk about the QA component that is part of the package as well. The QA feature checks bilingual files (TTX, bilingual RTF, Lionbridge's XLZ format, TMX, and XLIFF) for terminology, spelling, and inconsistent translations. Granted, this is less than other QA tools, but the morphological term recognition catches sentences that many others would not catch.

And the price for all of this? You can get a free account for 20 "glossaries" of up to 1,000 terms per glossary (without the QA checks), pay 5 euro per month/user for 200 glossaries of 10,000 terms per glossary, or 2,000 glossaries of 100,000 terms per glossary for 30 euros and six users.

Everything is hosted by ITI in Russia, and in my testing I ran into one stretch of very slow server responses, so you might want to encourage them to look into that.

To wrap it up, I admit I have always been in favor of having all my translation features under one roof (or "tent" -- see why I like the term "translation environment tool"?) so the various processes don't have to be separate from each other, and the translation memory can talk to the termbase, which also might check with a machine translation suggestion, etc. This is clearly not possible with a system like MultiQA. But in the real world, where many of us work with many TEnTs and continue to juggle separate terminology resources, this might be a great bet.

ADVERTISEMENT

Terminotix acquires Portage, an accurate machine translation system.

Align your documents with the most powerful alignment tool: AlignFactory.

New resources added to the free Terminotix toolbar. Download it now!

Wonder why we call LogiTerm the Swiss Army knife of CAT tools? Ask for your free trial.

Don't miss our videos on YouTube on our products!

5. New Password for the Tool Kit Archive

As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to May 2008.

You can access the archive right here. This month the user name is toolbox and the password is coolsummer.

New user names and passwords will be announced in future newsletters.

The Last Word on the Tool Box Newsletter

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box newsletter. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

If you are subscribed to this newsletter with more than one email address, it would be great if you could unsubscribe redundant addresses through the links Constant Contact offers below.

Should you be  interested in reprinting one of the articles in this newsletter for promotional purposes, please contact me for information about pricing.

© 2013 International Writers' Group