ToolkitSmall

A computer newsletter for translation professionals

Issue 9-8-146
(the one hundred forty-sixth edition)
Contents
1. The EuroTermBankWord-AddinRelease
2. Using Search Engines (Premium Content)
3. Translated!
4. You align! No, you align! No, you align!
5. Clearing Up the DTP Conundrum
6. Type It, Baby
7. Getting the Most out of a Package (Premium Content)
The Last Word on the Tool Kit
Lucked Out

I received quite a bit of mail regarding the presumed quote by Blaise Pascal ("I have made this letter longer than usual, only because I have not had time to make it shorter."). Readers credited it to a variety of different authors, and I began to understand why when I received Ted Wozniak's email. He wrote:

About that quote about the "more time . . . shorter letter," I had heard that ascribed to William Gladstone the British Prime Minister. But when I went to confirm this on Google I found the following: [The author of the quote is] Saint Augustine, Pearl Buck, Winston Churchill, Marcus Tullius Cicero, Albert Einstein, Pliny the Elder, TS Eliot, Malcolm Forbes, Benjamin Franklin, Johann Wolfgang von Goethe, Ernest Hemingway, Oliver Wendell Holmes, Thomas Jefferson, Doctor Samuel Johnson, Abraham Lincoln, Jud Northbark, Blaise Pascal, Ezra Pound, Marcel Proust, Robert Sayre, Madame de Sevigne, Madame de Stael, Henry David Thoreau, Mark Twain, Voltaire, EB White, or Oscar Wilde.

I'm sure glad that Pascal is at least on the list of possible authors. . . .

1. The EuroTermBankWordAddinRelease

I get frustrated with Word's spell-checker sometimes. Why in the world would the longish term in the heading not be recognized as a correct word? (Just kidding.)

Anyway, there is a very interesting plug-in that was released earlier this week that allows a search through the EuroTermBank from within MS Word. Now, the EuroTermBank has been publicly available for the past two and a half years and it's a termbase with an "initial focus (...) on terminology collections from the 'new Europe,' including Estonian, Hungarian, Latvian, Lithuanian, and Polish terms and their equivalents in English, German, French, Russian and other languages (overall, almost 30 languages)."

I ran a number of tests with my (non-focus) language combination (EN> DE) and got surprisingly helpful results.

Of course, there are a host of on-line terminology resources of varying quality, but all things being equal, I tend to use those that allow access right from within my immediate work environment. That's why I like a tool (like IntelliWebSearch) that directly links me to the dictionaries and glossaries I want to go to, and that's also the strength of this plug-in.

The EuroTermBank Terminology Add-In for Microsoft Word, however, takes it a step further; it not only links to the desired sites, but it also brings the information right into the MS Word interface.

(The only real drawback I see is that MS Word is no longer the preferred environment of many translators, now where even the last two big tool vendors who swore by that environment -- Wordfast and Trados -- have gotten away from it. In fact, the only thing that I still use Word for on a regular basis is to write this newsletter. . . . )

But for those who still like Word as their main interface and who work in European language combinations, the add-in should be highly welcome. There are two different versions of it, one for Word 2003 and one for 2007. I mention this because there are some real functionality differences, and the 2007 version add-in is clearly superior.

But let's start with Word 2003. Once you have the add-in downloaded and installed, it simply adds "ETB Terminology" to all the other resource books (Thesaurus, Encarta, etc.) that you can access in the Research pane. This opens when you click on a term while holding the ALT key. You will find lists of translations of that term in all the languages that are found; there is no way, for instance, to set your desired language combination.

This is different in the Word 2007 version. There you have a large EuroTermBank button on the Review ribbon (which is essentially the only interesting ribbon for standard language-related activities) that you can click on when you have highlighted a word or any section within a text (you can achieve the same by pressing Ctrl +Shift +I or the right-click command Terminology). When you do that, a separate EuroTermBank pane appears on the right with all possible translations into all the languages, but you can quickly filter that down by selecting source and target language and subject area (these settings will be stored between Word sessions) to see a well-defined list of relevant terms. Also, you can send a whole sentence or paragraph using the same method, and the pane will show the terms for which it finds a translation by formatting them like hyperlinks. Clicking on them will lead you to the respective list of translated terms.

It's super easy to use, plus it's free and does not seem to interfere with anything else -- so you'd be silly not to use it.

2. Using Search Engines  (Premium Content)

I would just love to be more flexible with my use of search engines! In my column for the current ATA Chronicle I recall one of the rare times when Google was down in the last few weeks and the rigor that beset me -- I felt completely helpless and essentially stopped working for an hour or so. I did not even think about there naturally being plenty of other search engines that I could use just as well. It was one of those times when I felt really stupid!

Once I reflected on this I realized that once we become relatively expert in using one tool (Google, in this case) we automatically feel like we lose expertise in other tools. I love to use special search syntax like "define:<search term>", "filetype:XYZ <search term>", or "site: www.XYZ.com <search term>" in Google (I have written about that in the past) to quickly filter results -- but of course those parameters don't work in other search engines. It's entirely possible that there are comparable techniques in other search engines, but they are well-hidden.

Recently I found an interesting translation trick for Bing -- but what would really intrigue me is a "translation primer" from one search engine search syntax to another. Let me know if you're aware of something like that. (By the way, I do like Bing's travel comparison feature -- I found a great price for tickets for my trip to the BDÜ conference in Berlin next month.)

And while we're at browsers, LSP Globalization Partners International has "launched a new search engine for international business professionals, travelers, students, researchers and anyone else who needs to easily search the web by language, by country, and by search engine."

Glearch ("Global Search") allows you to search three different search engines (Google, Yahoo, and Bing -- though I was not able to see any results from Yahoo) simultaneously by language and country. It's very easy to use and returns interesting results -- but it also still has some bugs and I got a number of error messages when I tried exotic combinations. I am sure that these will be ironed out over time, so it might be good to bookmark it for later use. (Unfortunately, special syntax use like I mentioned above does not work with this tool.)

I was pointed to another, yet different multi-lingual search engine by Donna Parrish. 2Lingual is a simple but interesting implementation of Google machine translation and search technology. As you enter one term in the first search box, Google Translate produces the translation in the other search box and at the same time executes a search. I was a little distressed to see that "Jeromobot" got a good number of hits on the English side but very few on, say, the German side, but I think it's just a matter of time until he becomes a worldwide celebrity! (By the way, he has his eyes on a girl!)

So, to come back to 2Lingual: While the MT translation feature is kind of crude, I still can imagine some helpful uses for translators -- such as quickly locating websites on a certain topic across languages.

3. Translated!

This is the eureka cry of the translator after a completed job, but it's also the name of the company that claims to be offering the "world's largest translation memory."

Translated's MyMemory has been around for a while now, and I had been a little skeptical about it. But when I (and I'm sure many of you) received a note from them this week, I took the chance to look at it again and talk to its founder. And I was much more impressed.

First of all, here's what it is: MyMemory is a rather colossal translation memory of presently around 150 million segments that contains data from web alignments (app. 30% of the total data), corpora such as the EU corpus (app. 50%), and TMs that the mother company Translated contributed from its own work (after client approval) and contributions from other translators.

So far you can access data in a search mask that allows you to enter a term, phrase, or complete segment and then returns translations in your target language. A machine translated version is first offered, followed then by TM matches. Those TM matches can be used for your translation or you can choose to improve or delete them with very easy editing facilities that, according to CEO Marco Trombetti, are used at a rather impressive rate. The goal, of course, is to improve the quality of the data in a cooperative kind of way.

Besides making data available, the overall goal of the project is to use the data to feed Translated's own statistical machine translation engine as well as an effort to gain visibility in the translation market -- which, again according to Marco, works rather well.

The latest push is to encourage more translators to upload data (more on that later) and to introduce a new feature that allows you to upload a document and receive a TMX translation memory file for the translation of said document in the TEnT of your choice. This feature is in a rather early beta version right now, but they are making changes and improvements to it at an impressive rate (it has changed at least two or three times since I spoke to Marco yesterday). The downloaded TMX file will contain all the TM matches it finds, and for the remaining segments machine translation matches are being inserted. Personally I found the inclusion of MT matches to be a real show-stopper, but there will be an option to exclude that (you can already see that option on the respective dialog, though it's still grayed out). By the way, the MT engines that are used for the translation are different according to language combination. They range from Google Translate to Systran to their own systems.

When you contribute TMX files, you have the option to specify whether this is for your own private use (but I assume it will still be used for MT training) or for everyone, and you can choose whether you want to hide proper names such as company and product names. When I asked Marco about that last feature, he explained that he has developed algorithms which recognize these automatically by comparing source and target and essentially looking for identical terms that are all capitalized (for example) which are then assumed to be proper names. I can see this working in a combination with English as source or target, but I am not sure whether it would work in all language combinations. We shall see.

While this seems to be a good step toward providing some privacy for the data you upload, it's not clear what happens to the actual documents that you can upload to generate TMX files when this feature comes out of beta in October.  It seems that Translated would be well advised to explain the process by which the source documents are deleted once processed.

Overall, I think this is exciting, and I can see that this could be a model we might see more of in the future. The daily visitor rate of 13000 seems to confirm that notion.
ADVERTISEMENT

How do I fine-tune Windows so it works best for translation work? And where can I find info on freeware programs that allow me to operate more efficiently? Complex file formats--how do I translate those? Should I buy desktop publishing and graphic software? And, oh, what about translation tools?

Find answers to these questions and many, many more in the 360 pages of the classic computer primer for translators: The Translator's Tool Box.

4. You align! No, you align! No, you align!

This has the ring of my kids fighting about chores that need to be done -- only that in those cases the "align" is replaced with "do that."

I have known about YouAlign for quite a while now, but I was asked not to write about it until Terminotix, its owners, were ready to present a stable enough version. This has now happened, and even though I found a couple of bugs when I tested it, they were fixed at a moment's notice. I think many of you will be glad to see this (for the present time) free product.

Terminotix is a Canadian company that essentially offers three translation-related products: Logiterm, Synchroterm, and AlignFactory in various editions. In the past I have written about all of the products, but I have most consistently praised AlignFactory, which I think is the best product to turn alignment (the conversion of a source and target document into a translation memory) from a nightmare into a feasible and profitable part of the translation workflow. Now, AlignFactory is relatively highly priced and will be a hard purchase for some to justify, especially if you have not had proof of its power. So, Terminotix has decided to release a free online-based version of it that allows you to upload source and target documents -- including PDF files -- and receive a TMX file in return. (Just like Translated, Terminotix would also be well advised to explain how the original documents are being destroyed after the alignment process has finished.)

Those of you who have tested AlignFactory won't be surprised by its accuracy and speed. Everyone else should be in awe.

Presently there are only a few languages supported (Arabic, Chinese, German, English, Spanish, French, Japanese, and Russian), but Jean-François Richard, Terminotix's president, assured me that more will be added quickly.

There is also a limitation on file size of the source and target documents (1 MB each) and you can also not batch align many files at once (unlike in AlignFactory), but it's a great tool for smaller jobs and to whet your appetite for the full Monty.

5. Clearing Up the DTP Conundrum

For those of you who receive the Premium Edition of the newsletter, I published a table that displayed which TEnT can work with which DTP tool a few weeks ago (in the 144th edition, to be exact). I received a bit of feedback from TEnT vendors, so I have updated the table. Since the table was contained in a graphic stored on my server, you just need to pull up that newsletter, refresh, and you should see the new data.

6. Type It, Baby

John Yunker announced a little utility on his blog that might be helpful for entering text in a language you only rarely need to type in. In these cases it's a nuisance to look up special characters through the Character Map (under Start> Programs> Accessories> System) or to install a new keyboard. TypeIt is a shockingly simple online tool that allows you to enter text, with all its special characters, in Czech, Danish, Dutch, Finnish, French, German, Hungarian, Italian, Polish, Portuguese, Romanian, Russian, Spanish, Swedish, and Turkish. You simply open the respective language and enter text into a field in your browser. A special on-screen keyboard is displayed as well as keyboard shortcuts for the respective language. Once you've composed your text you can copy and paste it into the document you are working on.

Again, please don't use this as your main interface to translate, but if you need to enter the occasional Hungarian, Russian, or Turkish text, this might be a highly welcome quick-and-easy tool.

Interestingly, in a response to John's posting, someone pointed to a new tool by Google that's also very interesting: The Virtual Keyboard API. It says this on Google's blog:

It is often difficult for Internet users to input text in many non-Latin script-based languages for a variety of reasons. The correct keyboard layout may not be installed on the computer they're using -- sometimes such a layout may not be well developed or widely available. This poses a challenging problem for web developers because there is no way they can ensure that their users have access to this very basic input technology. Our Transliteration API can help, but requires that the user know multiple languages.

Right on the heels of introducing support for translating Persian (Farsi) [I mentioned this a few weeks ago -- Jost], we've added a new Virtual Keyboard API into the Google AJAX Language API to further assist with text input. With this, developers can help their users input text without relying on the right software being installed on the computer they happen to be using. (...) With this initial release, we are launching 5 language layouts. They are: Arabic, Hindi, Polish, Russian, and Thai. We plan to roll out support for more keyboard layouts in the future.

The procedures for entering the code to display virtual keyboards in those languages does indeed look very easy and will be a help to many web developers.

7. Getting the Most out of a Package (Premium Content)

Many TEnTs (translation environment tools) use packages created by the corporate versions that can then be edited by the freelance or even free "Lite" versions of the respective tools. Generally, these packages contain directory structures that are recognized by the Freelance or Lite version as a valid project directory, compressed in a zipped file. To the casual observer this is not very apparent at first because the extension of the package file is not .zip as we are used to with most compressed files, but it's an extension that is specific to the TEnT. To find out whether this is the case in your TEnT, open the file in a text editor (such as Notepad) and look for the first two letters in all the gobbledygook you'll find. If they are PK, you are in luck and you know this is a .zip file. (And then, please don't save the file in Notepad -- that would break it.)

The next thing you need to do now is to change the extension to .zip (just right-click on the file in Windows Explorer and select Rename). After you unzip it you will find the files that are contained within the zip file in a separate folder.

Now, some tool vendors have decided to internally encrypt these .zip files, thus building an artificial barrier to prevent the use of other tools with "their" files. I personally find this to be particularly poor judgment on their side because it goes against the spirit of standards and exchangeability of information that everyone on the outside is committing to. In fact, now, at a time when the most important exchange standards (TMX, TBX, and XLIFF) are supported by most tools, the two areas that are still not exchangeable are the workflows of translation management systems (such as the online access to translation memory, terminology, and project data of one particular tool) and these artificially encrypted packages. For the first area we will (have to) see APIs (interfaces that allow the access of third-party tools) in the near future; and for the second we just need to see the silly protection dropped (tool vendors, you know who you are, so just go ahead and do it!).

OK, after this tirade, let's come back to the issue at hand. Once you have unzipped the file, you end up with files you might be able to use or not, but chances are that there will be some part that you might find usable.

In the case of Star Transit, the Transit packages (with the extension .pxf) contain the standard SGML files that Transit uses as translation files (their extension is a three-letter language-specific abbreviation) and a subfolder that contains the necessary terminology in a .txe file. And that's what all this is about.

A client of mine had a problem the other day with just that file. After I told her how to get to it and also told her that this looks like standard MARTIF (the precursor of the terminology exchange standard TBX), the import as a MARTIF file into TermStar (Transit's terminology companion) failed. This procedure that was provided to her by Star's support did the trick:

Open the .txe file with a text editor, in the section <databaseDesc> delete all the information of the type DictProperty' id (the information of the type hyperlink separator and ExportedLangs can stay), and make sure that the tag pair <databaseDesc > ... </databaseDesc> stays in the file.

Now the .txe file can be imported as a MARTIF file into TermStar or other MARTIF-compliant tools. (And this, of course, would be a lovely extension of the otherwise powerful filter for Transit projects that MemoQ is offering . . . ).

The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Last week this reader added a link:

translationtimes.blogspot.com

© 2009 International Writers' Group