Tool Box Newsletter Logo

 A computer newsletter for translation professionals


Issue 12-09-213
(the two hundred thirteenth edition)  
Contents
1. Industry? Not Sure. Industrious? For Sure!
2. MT Post-Editing Made Easier (Ⱦ) (Premium Edition)
3. Lingotek, Take 23
4. Tracker Jackers (Premium Edition)
5. Creating Your Own Little World of Wonders
6. Two More Things
7. A Love Story (continued)
8. New Password for the Tool Box Newsletter Archive
The Last Word on the Tool Box
Oh, the Power

Four weeks ago I advised you to add a certain line to your webpage or online profile to raise your chances of being found by a Google search, because Google had set up a language-specific link on the results page of Google Translate. It seemed very generous, but why look a gift horse in the mouth? Then lo and behold, Google promptly took it off again..... Now, I'm not saying that this happened only because I mentioned it in the newsletter, but my requests to Google to provide an alternate explanation have not been answered. So I'm basking in a feeling of great power (however ill placed and misused), though I do apologize to all of you who may have changed your webpage for naught.

I would also like to empower you regarding this newsletter. I've had a number of folks mention this past week that they feel intimidated by the technical nature of some of the Tool Box articles. This was a blow, since opaque "technicalese" is exactly what I want to avoid. Naturally I promised to do better, but every once in a while it will be impossible to talk about some issues without getting a lot more technical. So from now on I'd like to use the forbidding looking Latin capital letter T with diagonal stroke (Ⱦ) to denote articles that are slightly more technical; I may even use two of them if an article's really technical (I won't use three because that would mean that I don't understand it myself!). I hope this will serve as a fair warning system that will allow you to skip the article in question if you prefer to do that.

Another kind of empowerment on a whole different level is Nataly Kelly's and my book, Found in Translation, that will be released by Penguin next month. I'm supposed to receive a copy of the first print run this week and I can't wait. I hope that you can't wait either and have already pre-ordered your copy at one of these places. In addition, a number of cities will be hosting book events, including Boston, Washington, Atlanta, San Diego, Monterey, Portland, Seattle, and Detroit. If you're close to one of these cities, please stop by and bring lots of clients and potential clients.

Why? Well, this brings us back to the empowerment: this book communicates clearly why translation is the very thread that holds together the fabric of modern society and the world around us, in a way that's both witty and serious without being guilt-producing. Sound like I'm tooting my own horn? Maybe. But I suspect there's a pretty good chance you'll agree.

1. Industry? Not Sure. Industrious? For Sure!

Sometimes I'll carry around an idea that I know just wants to come out and express itself in the form of an article, and right when I'm about to sit down and write it out, someone else comes out with something remarkably similar. This can leave me feeling one of two ways: dejected because someone was quicker than me, or elated because it confirms I was on the right track.

So I chose to be elated when Portuguese translator Fabio Said published his recent blog posting, "There Is No Such Thing as a 'Translation Industry'". Even though Fabio isn't exactly the first to discuss some of the points he addresses, they resonated more with me -- and I'm betting with many of you -- at this particular point in time.

To summarize his arguments, he writes (in the second part of the blog posting) that there is really not a single translation industry but many. Because there are so many kinds of translators with such different task sets, self-images, and technology needs, it's presumptuous to assume that my experience has much to do with that of someone who is not in exactly the same branch of the industry as I am.

Guilty as charged.

I know that I so often do just that -- assume that your needs are similar or identical to mine.

But I want to push Fabio's argument just a little bit further and have a look at all the people who form what we call the translation industry. In reality, the actual translators -- as diverse as they might be -- make up only one part of that industry, albeit the most crucial one. Other groups also include the following:

  • various folks who work for language service providers (LSPs), ranging from project managers to sales people to technical support staff;
  • people who work on the client's side, ranging from the mid-sizish company to the multi-national corporation, who manage projects, internationalize and localize products, and purchase services;

and of course

  • the other group that I often address: the technology providers who develop and market anything from word counting and text extraction tools to translation environment tools to machine translation solutions.

If Fabio's argument hits home when it's related "just" to translators, how much more relevant is it for all of the above? (I know that many of you don't like to see some of the groups I list above as part of the same industry -- for lack of a better term -- but for all means and purposes they are.)

So, yes, we need to be careful how we talk to each other. We need to be careful not to over-generalize our experience. But it would be wrong and ultimately self-defeating to focus only on the differences and stop talking to each other. After all, there are many more similarities than differences, especially in our shared focus on language -- in all its communicative glory and fallibility. If that were indeed our only remaining connector, we would still end up having a lot to say to each other, even beyond the technology that we use or the daily output we produce. 

ADVERTISEMENT

Empower your teams with TEAMserver by Atril, the most effective and affordable all-in-one server solution on the market!

Learn more here. 

2. MT Post-Editing Made Easier (Ⱦ) (Premium Edition)

As I was talking to some translation environment tool developers the other day, I asked whether their tool is also a good tool for post-editing machine translation output. They hesitated and then returned the question to me: "What exactly does a tool have to have to be a good MT post-editing tool?"

Ouch!

I didn't really quite know what to say. So I promised them I would think more about it and get back to them. Fortunately I had some Twitter friends, in particular @rubendelafuente and @javmallo, who put their thinking caps on as well.

Here are some of the things that we came up with.

First, just to make sure that we're all talking about the same thing, the goal is an editing environment that presents a productive method of post-editing raw machine translation output. I know many of you wouldn't touch that kind of work with a ten-foot pole -- and that's OK -- but others welcome it, whether for economic reasons or for the kind of challenge it presents. Challenge? Yes, there is a challenge to recognizing certain patterns of mistranslation and finding productive ways to fix them. It has sometimes been said that editing machine translation output is like editing fuzzy match output (fuzzy matches are translation memory matches that are similar but not identical), but I think this misses the point. If the base translation memory is of good quality, fuzzy matches are always inherently correct -- at least as a translation for another segment -- whereas MT matches are often not. This means that you will have to apply different strategies to fix those (or, all too often, translate them from scratch).

Still, some of the features of a good PE (post-editing) tool would be similar to existing translation environment tools (TEnTs). These include

  • translation memory so I can easily reference formerly fixed segments.
  • termbase so I can make sure that the prescribed terminology is followed.
  • QA checks so errors are immediately flagged.

But even these three features need to be modified to match the needs of post-editing:

  • Translation units (TUs) in the translation memory that come from formerly post-edited output need to have a different -- and lower -- status than TUs that come from "normally" translated projects. Why? Because often the specifications of post-editing are less stringent with quality and consistency. This in turn means that while I still want to be able to reference formerly post-edited TUs, TUs from other projects might be preferred.
  • The termbase needs to be much more deeply integrated into the editing process. This includes a more consistent use of white lists (must-have translations) and black lists (may-not-have translations) that automatically flag inconsistencies. Term recognition cannot just be on a pure fuzzy level as it is presently in most tools; instead, it needs to follow morphological patterns so that terms can truly be recognized in all forms (of course, this would be a lovely feature for any TEnT).
  • QA checks need to be more easily adjustable and should have different levels of severity. Some errors (such as end-of-segment punctuation or correct numbers) might have to be fixed for sure, while others, such as double spaces or differences in length of segment, should be flagged as optional fixes. Any QA check -- as is already the case in many TEnTs -- should be done in real-time, as I'm working on the segment.

Features that can be found in some TEnTs (though not all) but should definitely be part of our PE tool:

  • Auto-propagation: the automatic copying of fixes for one TU to all other perfect matches of that same segment (and flagging of fuzzy matches).
  • Batch file processing: the ability to process many files at once. This is really important since machine-translated projects are rarely small projects where only one file is processed, but typically tens, hundreds, or even thousands of files. The editor cannot afford to replicate changes in many files, nor does the one-file-at-a-time processing allow the tool to identify patterns (see below).

And then there are features that are presently not in TEnTs at all but should be:

  • Collection of indicators for work performed: These could include the number of keystrokes, time spent, or percentage of performed editing so that billing can be done in a reasonable manner (the tool MemSource already offers some of that).
  • Macro support: While this could be done with third-party tools, it would be very helpful to have an integrated macro recording tool for which the user doesn't have to have any knowledge of computer syntax -- just perform what you're doing, record it, save it, and then repeat it to your liking to fix similar problems.
  • Filtering tools: Most TEnTs already allow for filtering or sorting of segments in some way or the other -- but this is mostly based on perfect match criteria. For instance, you can filter on the occurrence of one term or phrase, or you can sort the source segments alphabetically. What we need in an ideal PE tool would be ways to filter according to fuzzy match criteria so that we can see all the segments that are similar and then quickly perform the changes that need to be performed.
  • Tag handling: Tags (the placeholders for formatting within segments) often present the greatest problems for machine translation engines. Often these engines assume that the sentence ends when they encounter a tag, so they restart the translation of a new sentence after the tag. You can imagine what kind of nonsense this produces. I don't have any practical solution, but I'm sure that PE tools can handle this better. One option would be to forego tags if the term that is enclosed by tags is contained in the termbase, and reapply the tags after the machine translation or the post-editing pass (if it's contained in the termbase, the tool should know which term the tags need to be applied to in the target segment).
  • Recording of patterns: The  PE should be able to record certain pattern of fixes and write reports on it. Why? So the machine translation engine could be modified accordingly to produce better results next time around

Any ideas of your own? Let me know and we'll continue this discussion. And rest assured: this is your megaphone to the TEnT developers, all of whom are reading along with you and hoping to integrate those features that will make their tool the best suited tool for post-editing. Your comments will be read by those who can make a difference. It's a promise! 

3. Lingotek, Take 23

Ask me about the product strategy for any translation-related product and I can probably come up with something (whether its developers would agree with my assessment is an entirely different matter!). But ask me about the product strategy for Lingotek and I'd have to shrug my shoulders. At the least I might be able to tell you that last year it was this and that, whereas the year before it was such and such. And now? Let me check and get back to you on that.

I've been writing about the translation environment tool Lingotek forever -- well, at least since it was launched in 2006 -- and I've been really fascinated by it. While it's probably fair to say that it's never really made it into the big leagues, it's been on the cutting edge when it comes to the employment of technology.

Just check out this list:

  • Long before Google Translator Toolkit, it was the first and at the time only tool that was completely web-based.
  • It was built on the central idea of sharing data in translation memories.
  • Aside from MultiTrans, it was the only tool that right from the get-go heavily promoted and supported sub-segment matching (though it was not particularly well-implemented in the first few editions).
  • Earlier than most other tools, it pushed for an integration of machine translation into its workflow (it presently supports Microsoft Bing Translator, Google Translate, SDL LanguageWeaver, SAIC Omnifluent, and Asia Online -- provided that you have the respective licenses).
  • Earlier than most other tools, it made translation crowdsourcing an integrated feature of its tool.
  • Earlier than any other tool, it began using a large-scale cloud solution -- Amazon's AWS -- to form its backbone.

Do some of these features annoy you? Well, I think that just proves my point. The makers of Lingotek started to work on those features long before others did, and the forerunners often bear the brunt of the "masses."

With those features in hand, the folks behind Lingotek have tried a number of different models -- with different rates of success. When they first started, they focused completely on the freelance translator. They offered a SaaS model with payments of something like $5 a month and strongly encouraged the sharing of data (you always had the choice to keep data to yourself as well). Not too many bought into that.

Then they focused heavily on the US government. Here they offered the same solution as for others, but it was installed and hosted on the premises of the client so that the confidential data could not escape the oversight of Big Brother (now they also offer access through AWS GovCloud). They were reasonably successful with that; in fact, about a third of their business today still comes from that source.

When crowdsourcing came up as a concept that could be used by enterprises and not just manga and game translators, they introduced some features to match those needs, including a voting system and more strictly defined  roles, and marketed it to large organizations. You can see how Adobe uses it right here and the Mormon church right here. (Well, you'd actually have to be a member of the LDS church to access that actual translation interface; for Adobe, on the other hand, you only need an "Adobe ID" and you can even get some free fonts if you translate some of their stuff -- nothing better than the good old barter system.) What has helped to facilitate this offering has been a strong push to integrate the system right into content management systems (including some from Alfresco, Drupal, Jive, Oracle, Microsoft, and Telligent).

Marketing to these large organizations has remained a key offering for Lingotek today and is now being supplemented with one additional target group, language service providers, as well as an attempt to sell translation or project management services alongside the technology. I'm not sure that this is really a great combination, and I anticipate it (continuing) to create conflicts of interest, so my guess is that they will not go very far with other LSPs. But they might very well be successful with their offer to combine technology with services. One of the latest features, the Translation Marketplace, might help in this regard. This directly pushes job offers from within Lingotek to networks like ProZ (Elance and oDesk are in the making -- by the way, did you know that 30% of all jobs on Elance are translation jobs?). The function of the service in that case? Calvin Scharff, the Lingotek-er I talked to, called it "Back Office" or "Management Broker."

I still have a login to Lingotek from its very first days, so I looked through the system again. Compared to its early days, it's a very comprehensive system. This is especially true for the project management component, where some training would be advisable before you can use it properly. While not as streamlined as the online translation interfaces of tools like XTM or Wordbee, the translation interface itself is still refreshingly easy to use. For the most basic features you can use re-assignable keyboard shortcuts, and any user presented with its interface intuitively knows how to use it.

Unfortunately, the file processing is still on a file-by-file basis (but there is an option for global search and replace), and the supported file types include Office 2003-2010, OpenDocument, RTF, TXT, Java Properties, HTML, XML, Adobe IDML, XLIFF, PO, and PDF. The last of these formats, PDF, should be used with caution (which must be the reason that it's still marked as beta). The test file that I ran through was imported very unsatisfactorily (line breaks were completely ignored) and could not be exported properly either.

What does this all mean for you? Well, so far the makers of Lingotek are still giving away free licenses for freelancers upon request, and they have no plans to change that. I asked them to make that a firm policy so that its users can rely on that for good -- so far with no luck. I'll let you know once that changes. 

ADVERTISEMENT

Focus on Translation Not Administration!

21% off all Efficiency Products from AITwww.translation3000.com

Meet the new version of AnyCount, now with multi-core processor support:  www.anycount.com

 Start your own translation agency: www.projetex.com   

4. Tracker Jackers (Premium Edition)

I had never spent much time with the TAUS Tracker, but after having a look I have to tell you that I really like it. It's a site that consists of three parts:

  • Translation Memory Systems (Why, oh, why, do they call them "TM Systems"? There is so much more to tools like Trados, memoQ, Wordfast, or any others than just the translation memory),
  • Machine Translation, and
  • Other Language Technologies. ("Other Language Technologies" consists mostly of tools geared toward the machine translation community -- an excellent example of that would be Olanto's newly released myCAT alignment tool that is actually not listed yet. For those who are more on the Ⱦechnical side of things, you should check it out right here.)

Within any of these categories you can search in freeform or with a search mask (click on Extended Search to see that) for the relevant categories. In the case of TM Systems, the categories include the distribution type (desktop, client-server, or web), operating system, supported file formats, whether or not there is an open application programming interface (API), what open standards are being supported, and the functionality during the pre-processing and the project set-up.

The search result will list the tool or the tools that match your criteria with a detailed description of the tool that was provided by the tool maker (limited to the categories mentioned above). It's really quite cool and can save you a lot of groundwork when you do your tool research. (For instance, I just found out that Wordbee natively supports Photoshop files -- who woulda known?)

5. Creating Your Own Little World of Wonders

Lisa John has discussed click.to on her excellent blog a number of times (this is her first post with links to additional posts at the bottom) and I was always a little skeptical. After taking the time now to try it out, I have (almost) become a convert.

Most of you who have read this newsletter for some time now are tired of me mentioning IntelliWebSearch, the free tool that allows you to search for any term or phrase in any kind of online or offline dictionary or browser that you have set up.

This is sort of what click.to does, only (surprise, surprise) at the click of a button rather than a keyboard shortcut and in an overall slightly slicker fashion. click.to comes with a lot of preconfigured targets that range from Google and Bing Translator to Facebook and Twitter to Outlook and PDF.

Say that again?

I know, it sounds a little overwhelming, but you are able to send any highlighted text or graphic or file to any of these "targets" simply by copying it (press Ctrl+C) from anywhere. So, you can read through an article online, see a quote that you like, highlight and copy it, and send it to your Twitter account by clicking on the Twitter button that magically appears when you have click.to installed and appropriately set up to do that. Or when you're done with a job and want to send the file to your client: simply copy it in Windows Explorer or your desktop, click on the Outlook or Gmail button, and just enter the address of your client into the email with the file already attached. Or -- and this might be the most relevant scenario for many of you -- highlight or copy a term and send it to your favorite dictionary to see the translation.

Still a little overwhelmed?

When you install click.to, it comes with 50 or so preconfigured click.to targets that you can choose to be presented to you every time you copy something. Some of those can become "Satellites"that are shown in the form of icons every time you copy something; all the others are stored in a dropdown menu that you can access at the end of your "Satellites."

Additionally, you can add any source that you think would be helpful beyond the preconfigured options. To do that, just right-click on the click.to icon in your System Tray, select Options, click Add, and then select Add a web action if you want to add a search in an online source or Add a Windows application call if you want to send something to an application on your computer. The instructions for how to add an online source or how to add a computer-based application are fairly easy to read, and you should be done setting up your personal wonder world within a snap.

The drawback? Can't be the price -- it's free -- but in my opinion it's just a little too intrusive. There are plenty of things to copy that you won't want to send with click.to, but the icon bar pops up whether you want to use it or not. And if you are using a TEnT that heavily utilizes the clipboard (such as the old version of Trados), you essentially end up seeing the icon bar all the time.

Still, I know that many of you will embrace it once you give it a go, so have a lot of fun. (And anyway, have you ever heard of an easier way to create a PDF than copying something and clicking a button?) 

ADVERTISEMENT

Translate apps and work up to 40% faster with Across Language Server v5.3!

You've likely heard that Across v5.3 is the fastest, most stable, and easiest-to-use Across yet.

But have you also heard that Across v5.3 includes the ability to translate mobile apps natively and easily?

This revolutionary ability sets the Across Language Server apart from the field completely!

  • Support of App Localization for Android, iOS (iPhone), and BlackBerry
  • Pretranslation up to 20% faster 
  • Building and exchange of crossTank Packages up to 100% faster

Interested?

Visit www.across.net/newversion  or download our 60-day free trial version from www.across.net/en/form-testversions-request.aspx.

6. Two More Things . . . 

Wesley Budd of SDL asked me to promote his bike run (bike as in bicycle rather than motor bike) through the Himalayas to raise money for the International Childcare Trust. You can fuel his efforts by supporting him at www.justgiving.com/wesridesnepal. Good luck, Buddy.

Also, I know that many of you will be coming to the ATA in San Diego this year. If you're not, you should. After all, Nataly's and my book will be officially launched there (I hear there are also some interesting sessions beyond that!). But before you finalize your itinerary, consider staying for a day or two longer to also visit the conference of the Association for Machine Translation in the Americas (AMTA) that follows right after the ATA. I mentioned previously that one of the keynote speakers on Monday will be Luis von Ahn of Captcha and DuoLingo and the other will be our own Caitilin Walsh. But more importantly, it would be great if the bridge building that started two years ago in Denver could be continued.

(Oh, and did I mention that you're invited to the Sunday evening welcome reception as an ATA member -- no registration necessary. And did I also say that their reception snacks are legendary?)

7. A Love Story (continued)

Here is another addition to my family of remarkable characters and writing systems that I have been collecting over the years. This is both valuable and diverse.

Hold your cursor over the characters for a definition. 

8. New Password for the Tool Kit Archive

As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to May 2008.

You can access the archive right here. This month the user name is toolbox and the password is greatthingstocome.

New user names and passwords will be announced in future newsletters.

The Last Word on the Tool Box Newsletter

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box newsletter. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here is a reader who added it last month:

www.ecpdwebinars.co.uk/events.html  

© 2012 International Writers' Group