ToolkitSmall

A computer newsletter for translation professionals

Issue 9-2-133
(the one hundred thirty-third edition)
Contents
1. Information Super-Storage (Premium Edition)
2. And Let There Be Many TEnTs
3. 130 Pages of "Translation Software"
4. Mea Culpas
5. TEnT User Interface
6. Another Sad Demise?
7. A Love Story (continued)
The Last Word on the Tool Kit
Translation Is Dangerous

I was just made aware of this story, according to which an Afghan man who translated and distributed the Qur'an without printing the original Arabic alongside the translation is in danger of being sentenced to death. His court hearing is this coming Sunday. As a fellow translator -- who knows and cherishes the power of language -- I hope and pray that this will end well.

My own academic study was on the translation of the Bible into Chinese, but in the process I also dealt with Chinese Qur'anic translations -- some of which have the Arabic text alongside, and some of which don't. The issues of translating the Qur'an, which is seen as being a sacred revelation in its exact written Arabic form, are very different from those of translating the Bible, which is inherently translatable. But it seems to me that any translation becomes a related but completely distinct text, without acting as either a copy or a rival of the original. The original remains untouched and its sanctity unthreatened.

1. Information Super-Storage (Premium Edition)

With the content-centric Web 3.0 in the making, storage, retrieval, and interpretation of content is front and center to many technologies. But many of these technologies are concerned with large databases and are just a bit beyond the scope of the regular user, who's just looking for good ways to store and retrieve data efficiently. I recently looked at a number of tools designed to help with that, and I liked none quite as much as Evernote. Evernote comes in a number of flavors: for mobile devices, for a web-browser-based interface, and in a (Windows and Mac) desktop version. And it is this last edition that I had a look at. Evernote is a fairly large download -- a bit over 50 MB -- but it installs easily. One of the options during the install is for you to choose whether you would like to install handwriting recognition for German, English, French, and Russian.

This brings us to what this tool actually does: It allows you to collect snippets from any kind of application -- web browser, office application, Acrobat Reader, etc. -- in the form of text or image (!), uploads those to a server-based location -- so that you can access them from any computer -- and processes and indexes them. One of the processes that it applies is that of OCR (optical character recognition) as well as HWR (handwriting recognition). This means that if your images contain text in a decipherable manner, Evernote will also index those words and you will be able to search for them quickly.

It really is quite impressive. First I saved a couple of images that I knew would be hard to read, such as our logo to TranslatorsTraining with the turned-around r. Not surprisingly, that one could not be OCR'ed. But when I saved the logo of this newsletter (text: "The Tool Kit -- A computer newsletter for translators"), Evernote was able to do a beautiful job on it. I was able to "crack" the database that contains the index and this is what it contained for that image:

computer acomputer a computer acomputey a computey acompuler a compuler a coimputer aromputer a romputer ac omputer accipiter computed computes attributed reactivates newsletter for translators the tool toot kit hit kr kk kite kith kits kilt ten its tea tel tlc ltd ltv tem

When I entered any of these "words" into Evernote's interface, the graphic was called up and the presumed word highlighted. It's a very clever way of using OCR technology, I think. Rather than going for super-accuracy as is necessary when the OCR output needs to produce a document, here it uses its fuzzy abilities to make the application better -- making sure to get to the right word while allowing for typos during the search.

Once a snippet, a web page, a file, or a note (that you can manually enter) is saved, it becomes part of your extended memory that can be searched on the fly. And if there is no text in the item you saved, you can add tags manually that can be searched as well. Of course, any of the items is printable and can be saved elsewhere on your computer, and you can choose to organize your items in certain categories ("notebooks").

Evernote is free, unless you are a true packrat and want to save more than the 40 MB a month you are allowed to save with your free account. In that case there is a paid version ($5 a month) that gives you 500 MB a month.

It was interesting to see which languages are supported for plain text (this does not include the OCR or HWR ability). The only languages that I could not make work were Chinese, Japanese, and some Indic languages. Other languages, including Korean, Thai, Tibetan, Arabic, and Hebrew, worked fine.

So, what is this good for as a professional tool? Well, I can think of various uses. While working on a translation project -- especially a longer-term one -- there are many times when you find information during your research that might not have an immediate application. You end up discarding it only to kick yourself later when you need it. With a tool like Evernote, it is readily available. The same is true for the rare glossary that you might stumble on without having an immediate need for it. Sure, you could bookmark it, but you could just as well put it somewhere where it's easy to be found again. And on and on. . . .

ADVERTISEMENT

Albrecht Immigration Strategies, PC is a small law firm quietly doing big things for its clients.
  • Many other immigration lawyers only use their left brain hemisphere. I also use my right.
  • Many other immigration lawyers only see a picture of a hat when you show them a picture of a boa constrictor digesting an elephant. 
  • Many other immigration lawyers make simple things complicated. I make complicated things simple.
It should not be a surprise then that many other immigration lawyers actually hire me to help them win their own cases. Perhaps you too can hire me to help you win yours.

2. And Let There Be Many TEnTs

A few years ago, new translation environment tools (TEnTs) were showing up left and right, and it seemed that every other newsletter contained an announcement of a new tool. That certainly has slowed down somewhat. In the last few months, only AnyMem (see the ad in this newsletter) popped up as a completely new tool. However, now there is another new tool out there that also might be worth a second look: AidTrans Studio.

This tool from Krakow, Poland, seemed at first like a rather small tool, but it turned out to be a full-fledged application (that might still be a little wet behind the ears). I talked to developer Piotr Labuzek yesterday to find out more about what his tool is supposed to be and how he sees its position in the market.

He said that he "would like AidTrans Studio to become a 'popular tool,' affordable and relatively easy to use." The affordability he has certainly achieved -- right now all three versions (Basic, Professional, and Enterprise) are free in the beta testing phase until June, and from there on out the Basic edition will remain free. The Basic edition is a fully functional version without networking ability, batch processing, and "utilities" (regular expression test, encoding conversion, translation length verification, and custom tile format tags configuration).

The preconfigured supported file formats of AidTrans Studio are Microsoft Word 2003 (saved as .xml files), Word 2007 (saved as .xml or OpenOffice), PowerPoint and Excel 2007, XML, OpenOffice files, and Trados .ttx files. The fact that MS Word files are only supported through conversions is probably a real weakness. I asked Piotr about it and he said that he felt that XML is the way of the future and that's why he put his eggs in that basket. (Well, he may not have used those exact words, but that's what he meant!) Though that's true, we are still dealing with a lot of legacy documents that also need to be directly supported.

The interface of AidTrans Studio is very pleasant, a little reminiscent of the old version of Star Transit. This is also reflected in the underlying translation file structure, which consists of a text-based file for each source and target language (Transit uses SGML files). The database structure for the translation memory is similar to SDLX or Déjà Vu with the Microsoft Access Jet engine (but SQL Server for the Enterprise edition).

I did not find the workflow completely intuitive -- the creation of a project and the file import is a separated process and you have to independently initiate the "database environment" for the termbase and TM before they spring into action -- but those are things one could either get used to or which could be fixed in the post-beta phase.

I also think the tool could also benefit from a less technical look in certain areas (such as the configuration of regular expressions for a customized file filter -- which is basically nice but too complicated), but in general I was impressed by how complete and comprehensive the tool is.

Not surprisingly, this tool is not mentioned yet in the following compendium of translation software:

3. 130 Pages of "Translation Software"

John Hutchins, the great chronicler of translation software, and machine translation software in particular, has just released the 15th edition of his Compendium of Translation Software -- directory of commercial machine translation systems and computer-aided translation support tools.

It's really a very interesting document, if only to see how much software there actually is to support our work. One very practical application of the document is the index of language pairs for machine translation in the very back of the manual. I often receive questions about whether certain language pairs are supported by a particular system. Well, here are the many answers.

4. Mea Culpas

Since Barack Obama has shown us so impressively this week how to perform mea culpas, I cannot but contribute my own. And while I admit that erroneous information in this newsletter might be of a different magnitude than nominating ministers who have forgotten to pay taxes, here are mine nevertheless:

A number of readers noticed and alerted me to the fact that the InDesign for Translators manual does not apply to InDesign CS2 but to CS3 instead. My bad. However, the fact remains that it's sort of a shame that the manual does not cover the most current version: InDesign CS4.

On and beyond this subject, Arle Lommel sent me this information about CS4:

CS4 has introduced some profound changes relevant to translators (although some are not apparent yet in the UI). Perhaps the most significant one is the ability to export InDesign Markup Language (IDML) files, which are, in fact, a zip-compressed set of XML files (much like OpenOffice files are a bunch of compressed XML files). The really neat thing from a translator's perspective is that, if files are saved as IDML, you no longer need binary file-format filters for InDesign. Each story is saved as a separate XML file that can be manipulated with standard XML tools. If I were a TEnT tool developer, I would not bother with anything else now since making XML filters is far simpler than application-specific filters. This change also means that users of any tool that allows for customizable XML filtering can work with an InDesign-format document directly without the need to go to a tool that supports the InDesign binary format. (So there goes much of the marketing value of InDesign filters as a means to lock translators into a specific tool.)

Another new possibility would be the automatic replacement of graphics: since the entire content of the document is accessible, it would be a fairly simple process to have a set of build scripts (or even regular expressions) that could parse the XML files looking for things like graphic001_en.psd and replace it with graphic001_de.psd. (Or, if the content creator couldn't be bothered to use a naming convention like that, it would be easy to have a list of files that need to be replaced and go that way.) From experience I can say that managing replacement of links has been one area where most tools fall down (and where a lot of errors enter the localization process), but this would allow for some real improvements and for automatic error checking. While it wouldn't eliminate all manual work, it would eliminate the task of clicking through (and replacing) hundreds of links that we've all been through. And the nice thing is that anyone with a good text editor could do it, no special tools needed.

Now there are some other exciting things available in InDesign's engine that will be rolled out into the UI at some point. The most important one is that the compositing engine now supports right-to-left text. There is no UI to control it, but it can be scripted and the results for Arabic are impressive: it correctly uses the appropriate ligatures and positional forms and supports various kashida models and produces very nice-looking Arabic text, on par with text editors that advertise Arabic support as a strong point. While this ability is not ready for commercial deployment yet*, it does point to some major improvements that should surface in CS5. I also understand that there is also support for Indic scripts, but I haven't found any details on how to script them yet, so I have not tried that out yet.

(*After using it, I would definitely *not* recommend playing around with these features for someone who isn't comfortable with what may best be described as a pre-alpha experience -- [Those for whom this is relevant might want to try this -- Jost].)

I did some more research on what this IDML format is about (I was unfortunately not able to test it -- my own InDesign version is CS2). In said version CS2, Adobe introduced the InDesign-internal exchange format INX so that files between different versions of InDesign could be exchanged. While INX was XML-based, it was typically not possible to simply use a customizable XML filter of a TEnT for the processing of this format. Instead, the tool developers had to develop a specific XML-based filter for the INX format. In general these worked okay, but extremely complex files often suffered. One of my customers who exclusively translates Quark and InDesign files refuses to use the INX format in connection with his TEnT because of previous problems he has run into. Adobe itself says this (of course, only after the new version was released):

INX was difficult to read and manipulate because it was designed to be used by InDesign alone. Those who tried to manipulate INX encountered challenges with readability, robustness, extensibility, and compatibility with XML tools.

Now, the purpose of the new IDML format is specifically not just for InDesign-internal purposes, but instead to open up InDesign content to XML-enabled third-party applications -- among others, TEnTs.

This means that it is no problem to process these much more robust files with any of the (XML-enabled) TEnTs. All you need to do once you export the IDML file out of the original InDesign .indd file (File> Export) is to rename the .idml extension to .zip, unzip the file, locate the XML files that contain the story content -- the translatable text -- and import or open them with your TEnT. Of course, if your TEnT only processes files one by one, you might get slightly annoyed because you will have to deal with many files one after the other.

And while looking at all these things, I stumbled on a completely different solution for the translation of InDesign CS3 and CS4 files: StoryTweaker. Although this is not a solution that can (easily) be tied into a TEnT workflow and overall seems rather convoluted, it might be just the thing for someone who is looking for a cheap solution with some WYSIWYG support while translating InDesign files.

Back to the mea culpas.

I mentioned ways to avoid extreme finger-acrobatics for Trados user on a laptop, in particular with the shortcuts Alt+[Num+] and Alt+[Num*]. Well, it looks like there is a much easier way.

Trados always had alternative shortcuts (Ctrl+Alt+N for Set/Close Next Open/Get -- instead of Alt+[Num+] -- and Ctrl+Alt+Z for Translate to Fuzzy -- instead of Alt+[Num*]), but it never really publicized them much at all. That is, until it turned out that a bug with Trados and Word 2007 caused the Num-shortcuts not to work. Now the semi-official shortcuts are indeed the Ctr+Alt versions. (You can check them out in Article 2201 in the Trados knowledgebase -- thanks to Emma Goldsmith, Carmet Erez, and Amit Dharma for this tip.)

ADVERTISEMENT

AnyMem

New User-Friendly Translation Memory Tool

Happy 2009!

Hidden 35% discount on all AIT products:

5. TEnT User Interface

For the German-speaking readers of this newsletter, there is an interesting survey on the desired user interface for translation memory systems (TEnTs).

The author of the survey, who collects this survey data in the context of working on her translation diploma, writes: "The results of this survey are intended to help me and the developers of the open-source translation memory system OpenTMS to understand what a translator expects of the user interface of a TMS."

6. Another Sad Demise?

As has been discussed in various forums, Terminology Matters, the maker of two beloved tools -- the Word- and Trados-based QA tool Quintillian and, just as importantly, the TTX-to-Word-to-TTX conversion utility TTXpress -- has apparently shut down. I have tried to contact one of the former owners to a) know what happened and b) ask for permission to upload these utilities for download somewhere -- but I have not been successful. If anyone has any hints, I would be grateful.

7. A Love Story (continued)

I have been working on finding a good example for this one forever. Here it is.

If you have a Javascript-enabled browser, hold your cursor over the character for a definition.
The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

© 2009 International Writers' Group