Tool Box Newsletter Logo

 A computer newsletter for translation professionals


Issue 11-11-202
(the two hundred second edition)  
Contents
1. Conference Fatigue (Premium Edition)
2. Flexible Ideas
3. Paying for MT
4. This 'n' That
5. Eating Crow
6. Sneaky Stuff (Premium Edition)
7. Corporate Subscriptions
8. A Love Story (continued)
9. New Password for the Tool Box Newsletter Archive
The Last Word on the Tool Kit
Mad Gab

Are you familiar with the game called Mad Gab? My wife and I spent an hour last Sunday giggling hysterically with our 12-year-old as we intoned nonsense syllables and tried to break the code. Here's how it's played: One player shows a second player (the reader) a card with a phrase that may go something like this: "Aim Hunk Hear Inch." The reader must pronounce the phrase aloud with varying speeds and intonations until someone hears what it's supposed to be. In this case, it's "a monkey wrench." Another example would be "Ace Tray Taste Who Dent" that ends up being "a straight-A student."

Here are some customized Mad Gabs for you, my favorite language professionals:

  • Did Your Rome Owe Bought
  • Trends Lay Shun End Vie Run Meant You'll
  • He Brew Heir A Big Tea Yam
  • Gore Badge Inn Garb Edge Ow It
  • And my favorite: Trends Late Horse Art Gee Knee Yes Is

Have fun!

I was really pleased that the discussion on payment of MT post-editing that I referred to in the last newsletter was picked up by the very insightful blog of Kevin Lossner. And I was even more pleased to find out that the general consensus there seemed to be that the proposed model might actually be acceptable.

1. Conference Fatigue (Premium Edition) 

I think I've found a solution to conference fatigue: talk about something that surprises yourself!

The ATA -- which took place last week in Boston -- is one of the conferences that I usually really enjoy, but still, there have been a lot of conferences this year, and the year before, and the year before, and so on, ad infinitum.

To spice things up this year, I suggested several topics for conference talks that I thought were interesting and relevant but about which I really didn't have much to say when I suggested them. Fortunately, there are always a few months between suggesting a topic and giving a talk -- a time that I used to explore and attempt to truly understand those topics (and write about in this newsletter and elsewhere).

The topic of my ATA talk last week was How to Deal With More Data Than You Can Handle, but it mutated into a conversation about something bigger: how translation technology has changed in fairly small, incremental steps, but in the process has completely thrown out some of the old paradigms and replaced them with new ones.

Translation environment tools used to deal with data by lumping it into independent data silos that had little if anything to do with each other.

Silo #1, the translation memories, stayed largely untouched. Think about it. The chances of individual translation units within the TMs ever being touched or found again were extremely unlikely. I explored this idea in more depth in a recent article in the Translation Journal. Yes, manual concordance searches within the TMs were helpful, but they were tedious and time consuming.

Silo #2 contained the term- or phrase-based data in glossaries or termbases. This was a great resource. Unlike the manual concordance searches, these term matches came up automatically and were displayed and in some cases even entered into the target text. Typically it was an either-or situation, though: either you had perfect or fuzzy matches inserted from the TM or you had termbase matches. Only a couple of tools attempted to combine the content from these databases.

And finally, Silo #3 housed external data. These came in the form of dictionaries -- some paper, but mostly electronic -- as well as the many web-based sources of data and corpora that we talk about in this newsletter and machine translation engines such as Google Translate and Bing Translator.

Without a doubt, each receptacle contained highly valued resources, but they all sat silently side by side the other silos of data rather than "talking" to one another and combining forces.

In the new generation of tools, this has either already changed or is about to!

Because of subsegmenting features, TMs are no longer just collections of segments that may or, more likely, may not be used again in the future. Instead, they are vast collections of translation data that are continuously analyzed and presented to the user in chunks that are useful for the translation of the current segment. What does this mean? Aside from much more useful data, translation memory maintenance suddenly becomes a necessary reality. Those "Big Mama" TMs that collected everything you ever translated are a thing of the past. (You can read more about this in the Translation Journal article mentioned above.)

Let's talk about termbases. If the subsegmenting feature is working properly, there really is no need anymore for the simple source term - target term glossary (with the possible exception of the client-prescribed glossary). Why? Because terms are extracted from the TM automatically! Of course, this presumes that there is a reasonably large TM in the first place. But if it's there, the subsegmenting feature will not only extract terms, it will even extract the correct ones. Why again? Because it's a numbers game. If properly implemented, the subsegmenting feature extracts the translated terms and then also looks at the number of occurrences, presuming that even if there is a certain level of inconsistency among the translations, the one with the highest number of occurrences will be the first of the suggested translations and the statistically correct one.

So, glossaries are gone. What about more complex termbases? I think these have a bright future -- much brighter, in fact, than they enjoy right now. The next big frontier for translation environment tools will be to process actual linguistic data. Right now they look for similarities in sequences of characters on a term or segment level. And while some tools use some intelligence when it comes to things like plural endings for the large European languages, this is where it stops. What they -- and we -- need is a more in-depth knowledge of morphological rules that govern languages, all languages. As I mentioned before in this newsletter, this is not going to happen because the tool makers implement it -- they can't because it's simply too expensive. But I think that we will. I have a vision (forgive me for using such a grand term) of crowdsourcing our expertise for the languages that are relevant -- that is, all languages for which functional translations are being performed and translation environment tools are being used -- and collecting the data online so that the tool developers can "simply" open up their tools to process the data. (I put "simply" in quotation marks because I've already heard the complaints that this is not "as simple as it sounds.")

What does this mean in practice? Some tools already self-repair fuzzy matches, so a translation unit like "He read the book on the way to work (Er las das Buch auf dem Weg zur Arbeit)" can become a fuzzy-made-perfect match for the segment "He read the letters on the way to work (Er las die Briefe auf dem Weg zur Arbeit)" even though the gender and the number of the object are different in the target language. The tool would have to "know" the translation of "book" and "letter"; then, if the termbase contains information about grammatical categories, it can refer to the information in the online crowdsourced database and automatically change the noun's article and plural ending. Rather than specifying the different endings or word forms for each term in the termbase -- essentially an undoable task -- the individual users could simply map terms to categories that are created by all translators who take part in building up the online knowledge store.

If you think this sounds a lot like some principles of rules-based machine translation terminology processing, you're right. That's exactly what it does. The only difference is that it would not be limited to the small number of language combination that are currently available for rules-based MT processing, and, more importantly, you are completely in control of the terminology you want to have processed that way. (I know that I've mentioned this before, but in the beginning of December there is a meeting with a number of folks from the University of Illinois to see whether they can jump-start this project.)

To come back to the usefulness of termbases: yes, they will be very useful, but in a different way than they're currently used (if they're used at all) by most translators.

And what about external data? The march of external data right into our tools has already begun. Some tools offer live connections to online corpora or dictionaries, so if matches are found they are treated equally to data in TMs and termbases. Virtually all tools offer a machine translation component; in fact, at least one tool has now started to not only use it for gist translations but also for repairing fuzzy matches -- a potentially much more useful and effective use of Google Translate or Bing Translator.

Those separate data silos of old -- TMs, termbases, and external data? They're already extinct, or will be shortly -- not the data, but its separation.

Which tool exactly have I been talking about this whole time? Well, the single dream tool that combines all these features does not yet exist, but many tools incorporate some of the features. Here's a non-exhaustive list:

  • Subsegmenting: Multitrans, Lingotek, Trados Studio (via the AutoComplete dictionaries), memoQ, Déjà Vu X2
  • Automatic repairing of fuzzy matches: memoQ, Déjà Vu X(2)
  • Some morphological knowledge for some major languages: Across, Transit
  • Connection to morphological database: none yet (duh!)
  • Integration of outside resources: Fluency, Across, Lionbridge's Workspace, memoQ
  • MT integration: virtually all
  • Deeper MT integration to repair fuzzy matches: Déjà Vu X2 
ADVERTISEMENT

Save on Studio 2011 with the SDL Premier Packs!

Buy or upgrade to SDL Trados Studio 2011, the latest version of SDL's translation memory software, and save 20% plus receive a free online training session. For more information or to buy visit - www.sdl.com/promo/toolkit_Nov.

Offer ends 18th November.

2. Flexible Ideas

Awhile back I wrote about the über-cool Engkoo (英库), an add-on to the Chinese version of the Microsoft search engine Bing with machine translation, language learning, and much, much more. If you have any interest in Chinese and some time to spare, this is a good place to use that spare time.

It turns out that Microsoft has some other interesting language-specific projects, including Microsoft Afkar, which aims to enhance the web for Arabic users.

Afkar (أفكار) is Arabic for "ideas," and that's a pretty good common denominator. Here are some of the tools:

  • Maren (= مرن - Flexible) Reader: A browser tool (for IE, Firefox and Google Chrome) to transliterate phonetic, Latin-based Arabic into Arabic script.
  • Maren Transliteration: A browser-based, web-based, or desktop tool to type in Latin characters and convert into Arabic by simply pressing the Enter key.
  • Maren Autocomplete: A web-based auto-complete tool -- both on the word and phrase level -- for Arabic.
  • Maren Morph: An Arabic morphological analyzer that "analyzes a given Arabic word and displays possible analyses with corresponding diacritics for each analysis." Available only for Internet Explorer.

There are also various tools connected to the Arabic engine of Bing Translator, and then there's a tool that is interesting for more than just Arabic translators: WikiBhasha from Microsoft Research India. This tool is somewhat reminiscent of Google Translator Toolkit, but presently it's used only for Wikipedia pages. The idea is that you can expand Wikipedia pages in your language by translating more complete articles from other languages. You can translate as little or as much as you want on the basis of a machine-translated gisted version and compare your edits to those of other users. Once you are done you can submit right from within the tool to Wikipedia.

WikiBhasha is an interesting tool, but you'll need to be aware that the idea is to collect high-quality bilingual content in languages with little existing data -- one of the great challenges for machine translation developers. 

3. Paying for MT

Much -- maybe too much -- has been discussed about the new user-pay version of the Google Translate API, in contrast to the free Google Translate Toolkit or even the Google Translate website (which has seen some interesting changes lately, including some newly added virtual keyboards).

Some translation environment tools have already started using the new version of the Google API -- including memoQ and Wordfast -- and if you are intent on using Google Translate with one of those tools, you can find a helpful video on how to register and pay for it at Google right here.

If you use a tool like IntelliWebSearch rather than the directly integrated process of Google Translate, you can continue to use GT for free because IntelliWebSearch just retrieves data from the Google Translate webpage. You can find a description of how to do this right here.

And finally, if you want to switch from Google Translate to Bing Translator (for which you only have to pay for very high volumes), you can find information right here. 

ADVERTISEMENT

You follow your friends, your colleagues, translation-related events and many other things... Now follow your translation documents, and changes made to them, in your translation with memoQ 5.0!

Track changes, term extraction, cascading, TXML and Regex-based text filters, X-translate, Google v2 API integrated, and many other exciting features.

Download the 45 trial version from www.kilgray.com.

4. This 'n' That

After Snowball's developer moved from Denmark to the US, he not only changed his website address and design but also the licensing model. Rather than paying the modest licensing fee you had to pay previously, you now pay either $2.50 or $5 a month. The higher price is for a slightly enhanced "Pro" version (the only relevant difference that I can see is that it allows you to customize keyboard shortcuts).

Not familiar with Snowball?

I'm always "preaching" that no one should expect any translation environment tool to be easy to learn: translation is a complex process, so the main tool to support that process should naturally also be a complex tool. Well, Snowball might be the one exception to that rule. It is indeed very simple to use and to learn -- but it's also very limited in comparison to its larger competitors. Not only does it work only in Word, but it lacks many bells and whistles.

In short, it's a nice tool for someone who does a little translation on the side.

 

I had a chance to talk to Vahur Meus, the Estonian developer behind termbases.eu. Not surprisingly, it's a termbase tool whose free version allows you to share your termbases with others or keep them to yourself, and whose commercial version can be customized and integrated into other environments. Vahur seems to be a nice guy, but I'm not blown away by his tool -- and judging from the small number of publically available termbases, other users are not either. My main problem with a tool like this is that it's not readily integrated into any kind of translation environment or workflow -- this is where I see the main benefit of a terminology management tool. Still, the main building blocks are there, and it might prove to be useful to some.

 

And one last thing: one of Microsoft's power toys for Windows XP was the super helpful Image Resizer, which integrated right into Windows Explorer and allowed you to resize images with a simple right-mouse click. The project has been picked up for Windows 7 and you can find the download at this site. This is one of the tools I use all the time, and my friends love not receiving 5 MB photos from me! 

5. Eating Crow

The one tool that I've been encouraged to write about more than any other tool in the past is AutoHotkey, the ultimate (and free) interface for creating macros that work (almost) anywhere within Windows. The reason I've always refused is this: clearly I believe in making things as easy as possible on the computer -- this is partly what this newsletter is about. But I also believe in repeatability, meaning that things only really "work" if you can use them on pretty much any computer and/or any program. And that is especially true when it comes to creating keyboard shortcuts. I've often written about finger memory -- the difficulty of unlearning a shortcut once it's impressed into your finger brains -- and how much that can cripple you when changing computers.

All this said, there is no way I cannot write about AutoHotkey this time.

In Jeromobot's Twitter account, I tweeted the following query a couple of weeks ago:

Here is a utility I would like: add parentheses or custom quotes to word or phrase by highlighting it and pressing keyboard shortcut. Anyone?

A utility like that would be great, especially with all the many strange and often difficult-to-enter quotes -- here is a cool list -- that translators have to use, especially in translation environment tools. A number of followers almost immediately responded: use AutoHotkey, use AutoHotkey! And, yes, I admit I rolled my eyes just a tad. But when lovely Ignacio Hermo actually wrote an AutoHotkey script for me and sent it, I could not resist and finally installed it!

And it works amazingly well. All you need to do is download AutoHotkey under the link above, install it, save the following script in a text file, give the text file the extension .ahk. right-click it, and select Compile Script. A new little program icon will be created, and when it's double-clicked the magic starts.

Here is the script:

#p::; Windows+p to add parentheses to selection

Send,^c

Sleep, 001

Clipboard = (%clipboard%)

Send,^v

return

#q::; Windows+q to add quotes to selection

Send,^c

Sleep, 001

Clipboard = «%clipboard%»

Send,^v

return

Simply highlight the term or phrase and press WinKey+P to set parentheses or WinKey+Q to set quotes -- and either the shortcuts or the quotes and/or parentheses can obviously be adjusted. The only program that this has so far not worked in has been Star Transit XV -- everywhere else it's been like clockwork.

Now, you may not even be interested in this particular macro, but, again, you can automate virtually everything and anything with AutoHotkey. If programming a macro sounds too complicated to you -- here is a tutorial -- you might have to look for your own Ignacio!

ADVERTISEMENT

Fortis Revolution is the computer-assisted translation tool that will direct how you purchase and interact with translation software in the future.

Before Fortis Revolution, businesses had to piece together frustrating translation tools with confusing and hard-to-navigate product packages. Times have changed.

Fortis Revolution offers the simplicity of one product, one distribution and service channel, and one price, in an easy-to-use, yet effective tool to meet your individual translation needs.

One product, one place, one price -- the revolution has begun!

6. Sneaky Stuff (Premium Content)

Here's one piece of info that I know some of you will like. Some clients don't really care how a job gets done as long as it DOES get done, in time and with the desired quality. But other clients are not like that. They would feel slightly uncomfortable to find out that it was actually your friend who worked on a document, or that you spent only 20 minutes on a document for which you charged a whole hour. . . . (And, no, I'm not recommending either scenario.) However, you do need to be aware that Office documents in particular are keen to store all kinds of information about such things.

There are manual ways to remove this information, but these are in fact so manual that it's easy to forget or simply too tedious to do. Here is a quicker way: to remove personal information from a file, open your document and select the appropriate procedure for your software:

  • Office 2003 and earlier: Tools> Options> Security and check Remove personal information from file properties on save
  • Word 2007: Office button, select Word/PowerPoint/Excel Options, go to the Trust Center and select Privacy Options and Document Inspector.
  • Word 2010: File> Options> Trust Center> Trust Center Settings and Document Inspector.

While you're there (in all versions of Word), you can also find and select the option Warn before printing, saving, or sending a file that contains tracked changes or comments -- a helpful feature that may avoid some embarrassment caused by sending documents with change-tracked data that was not supposed to be seen by anyone but you, or some frustration when you print a document and forget to turn off tracked changes, thus making it virtually unreadable.

7. Corporate Subscriptions

With the introduction of new pricing for the premium edition of this newsletter at the beginning of this year, I also introduced "corporate" pricing, which refers to a subscription by folks in a company or organization who plan to share the newsletter with colleagues.

It's a laughable $150 a year for up to five users and $250 for more than five users. As I said: laughable. What makes it even better is that you also get a full 90 minutes of free tool consulting with me when you purchase one of these corporate options. This offer is valid for two months after the purchase.

For more information, I'll see you at http://www.internationalwriters.com/toolkit/. (And I look forward to consulting with you soon.)  

8. A Love Story (continued)

Here is another addition to my family of remarkable characters that I have been collecting over the years. These characters are not Unicode-compatible and probably never will be, but they are very colorful.

Hold your cursor over the characters for a definition. 

9. New Password for the Tool Kit Archive

As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to May 2008.

You can access the archive right here. This month the user name is toolbox and the password is cinquecento.

New user names and passwords will be announced in future newsletters.

The Last Word on the Tool Box Newsletter

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box newsletter. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here is a webpage that added the Tool Box link last month:

www.affordablelanguageservices.com

© 2011 International Writers' Group