ToolkitSmall

A computer newsletter for translation professionals

Issue 9-5-139
(the one hundred thirty-ninth edition)
Contents
1. Okapi in Java (Premium Content)
2. This'n'That
3. Looking for Glasses
4. Unding
5. XDXF
The Last Word on the Tool Kit
When Fate Strikes

I was sitting in a lovely French restaurant in Budapest last weekend with Jeromobot in the seat to my left. While I was engrossed in conversation with my neighbor to my right (Jeromobot can be a rather slow conversationalist), I failed to realize that SHE, in the guise of an aging flower girl, had approached, seen Jeromobot, grabbed him, and tried to get away. I tried to stop her by buying some flowers, but to no avail. She continued begging for Jeromobot until I had to give in and gave him to her. So this marks the end of Jeromobot's adventures. Or does it? (By the way, I am not making this up!)

Old Lady and Jeromobot

Fortunately, I had the chance to document some of Jeromobot's adventures in Budapest before SHE turned up, including some meetings with kings, horses, and chestnuts. Most importantly, perhaps, there was a crucial moment at the conference we attended -- MemoQFest, a very fun and warm user meeting -- where the management team of Kilgray, the company behind the TEnT MemoQ, bowed to popular pressure and officially re-pronounced MemoQ from \memok\ to \memo-q\. An official agreement was signed and you can see that Jeromobot was part of that, too.
1. Okapi in Java (Premium Content)

Lots of interesting news today. Yesterday, Yves Savourel, the developer of the Okapi tool set, sent me a note saying this:

Just a word to let you know that a new set of Java tools has been released today. The "old" .NET/Windows-based Okapi Framework has been completely redesigned and rewritten in Java. This allows the components to run on Windows, Linux and Macintosh. Moving to Java also made it possible for a few developers to join the new project and bring innovative ideas and more experience.

In this first milestone we have Rainbow v6, Ratel (an SRX editor), and a basic library. But the most interesting part is probably the common resource model all the components use, as well as the pipeline mechanism. This will be the backbone of the next releases. It will allow the users to really customize the tasks they need to run, and to create new tasks as well.

You can get the new packages from the Google Code site and you can get more information (including the beginning of a Developer's Guide) from the main site. We still have many things to develop, but hopefully this first milestone will give you a glimpse of what is coming.

All right, you might say, so what exactly is he talking about here? Well, Yves is the largely unsung hero who for years now has worked and released an open-source tool suite that deals with the often overlooked aspects of TEnT management, including conversions, maintaining TMs, and quality assurance. His tools include Olifant, the most advanced translation memory editor on the market -- this tool has not been ported to the Java platform, the new SRX editor that Yves mentions (SRX is the segmentation exchange standard that allows you to exchange the segmentation rules of your TEnT with another) - and, of course, the tool Rainbow. Rainbow is indeed an appropriate name for the tool: it seems to cover an amazingly wide spectrum of desirable features, including:

·         Encoding, line-break, URI, RTF, Byte-Order-Mark conversions

·         XSL transformation

Or for the less technical among us

·         Translation package creation

·         ID-based alignment

·         Translation comparison

·         TTX checker

Let's just look at the last four features.

The translation package creation allows you to create an XLIFF package or a package for OmegaT, Trados RTF, or Wordfast Classic from source documents that can be translated in the respective tools; the ID-based alignment allows you to align files with strings that have some kind of an ID (for instance Java .properties files) and comes with a full-fledged Alignment Verification editor; the translation comparison feature allows you to compare two versions of translated files to generate a score and an HTML or TMX report with the differences highlighted; and the TTX checker allows you to locate problem areas in Trados TTX files if they don't want to export properly. So, you can see, Rainbow is sort of a smorgasbord of things we often thought would be nice to have in commercial TEnTs.

I also really liked Ratel, the SRX editor. (I will need to ask Yves about that name. I'm sure there is some kind of clever explanation for it -- after all, "Okapi" relates to the "copacetic animals that are no strangers to innovation: they have blue tongues which they use to clean their ears. . . ." If that's not clever!) SRX is one of those standards that is only very slowly being adapted. The idea of SRX is great: The two large obstacles that TMX as a translation memory exchange format has to deal with are differences in segmentation and differences in codes within segments ("in-line codes"). SRX deals with the first issue by allowing tools to exchange an XML-based file that contains the necessary information about where a text is to be segmented and with what kind of exceptions. Unfortunately, the only tools that support SRX are MemoQ, Swordfish, Heartsome, SDLX, and XML-INTL.

And should you for some reason not like Yves' SRX editor, you can also try Maxprograms' SRX Editor, which also is available for free.
2. This'n'That

Here are a couple of things that I encountered during the last few weeks:

  • I have always been a very ardent defender of the OCR-based conversion of PDF into Word files offered by tools like PDF Transformer. And I still am. There are just too many advantages -- including the fact that image-based text will be converted as well as "normal" text -- but there is a new product that seems to do quite well with non-image-based PDFs (so, no PDFs that have been scanned). PDF-to-Word is a free online-based product that is surprisingly fast and accurate (in the few samples that I tested it with). Plus, it's always an advantage not to have to install anything. Of course, some of your clients may not be too thrilled about you uploading their documents into cyberspace. . . .
  • Someone on the Japanese Honyaku list mentioned the Chemistry Formatter Add-ins for Microsoft Word, Excel, and PowerPoint. Whoever said that academics have a good feel for funky names? I am not doing any chemistry translation but I could imagine that those of you who do will like this tool.
  • A couple of newsletters ago I mentioned the Resource Translation Toolkit by Mike Funduc, the developer of Search & Replace. It's a tool that was developed by Mike to support the translation of software resource files, including the RC and PO formats. Back then I said "according to the master himself, there most likely will also be more software file formats (such as Java .properties and .resx files) supported in future versions." Believe it or not, this has happened, along with a number of other enhancements. It's so fun to deal with developers who are passionate about their products.
  • Lastly, I wrote about the new Transit version awhile back and did so very positively. And rightly so -- it's a fabulous product on a lot of levels. I did not mention one thing, though, that really bugs me, and that is that they have thrown overboard a pricing strategy that I thought was very clever. Transit, along with SDLX, was one of the first TEnTs that offered a free version for translators. This version could not be used for standalone projects (i.e., you were not able to create projects on your own), but it was possible to translate projects that were sent by your client. This has been abandoned with the latest incarnation and has now been replaced with the Freelance version that "can be leased for an annual fee of 295 €." I think that's too bad. Transit has always been one of the more expensive products -- and many users would say rightly so -- but it seems to me that 300 Euro a year is an awful lot for a castrated version that does not even allow me to translate a simple Word file unless my client prepares it for me. But maybe that's just me. . . .
And to end the This'n'That on a positive note, Renato Reinau recommended WWW Search Interfaces for Translators -- a very clever interface to locate language-combination-specific glossaries. Cool stuff.

ADVERTISEMENT
MemoQ 3.5

"It's like playing a video game in permanent cheat mode"

To learn more, visit www.kilgray.com
3. Looking for Glasses

To me this is the silliest thing in the world: searching for glasses. Why do you have glasses in the first place? Because you can't see! So, what's the point of "looking" for them when you have misplaced them once again? (I have been trying to help my 20/20 wife understand this conundrum -- so far without much success.)

Same thing with help files! How can it be possible that I need help to open a help file that is supposed to help me?

Sadly (and understandably), some of us do need help with that. I wrote about this awhile back, but since I had several folks ask me about it this past week, here it is again:

Most of you know that I sell the Translator's Tool Box, my computer primer for translators, in parallel formats: a normal PDF file that lends itself to actual reading, and an HTML help (.chm) file that contains the same content but is more helpful when it comes to quick referencing. I have received a lot of emails from folks who bought the book/help file but could not open the help file on either a Macintosh computer or on a Windows Vista machine. Here's how to fix that.

For all versions of Windows, you might have to open the .chm file on your local computer; this means that you probably shouldn't have it on a network location. In Windows Vista, it is possible that you will be able to view the .chm file table of contents, but you might see the following message on the right-hand pane: "Navigation to this webpage was canceled." This error is caused by the stringent security requirements of Windows Vista and can be circumvented by right-clicking on the file, selecting Properties, and clicking on the Unblock button in the lower part of the dialog. You can find some more information on this procedure right here.

For Macintosh computers there are essentially two ways of going about this. You can either download and install the open-source application Chmox, which will allow you to open and use a .chm file, or you can install the CHM Reader Firefox add-on, which permits you to view .chm files inside Firefox. This add-on is also available as a Windows or Linux version.
4. Unding

This is one of my favorite German words: Unding. Literally "un-thing," I haven't arrived at a good translation for it -- "absurdity" might work, but that doesn't really quite cut it. (By the way, this was also was one of my mom's favorite words; unfortunately, she typically referred it to me in some way or other . . .).

Anyway, in about three weeks from now, on May 21, the second Localisation UnConference will be held, this time in Dublin. I mentioned the first about a year ago; it took place in Silicon Valley and was organized by Ultan O'Broin and Shawna Wolverton. This event is organized by Martin Wunderlich. Clearly, many conferences are struggling this year with attendance, but I wouldn't be surprised if this event would be well-attended. Here is what Martin writes about it in the conference blog:

Just to clarify, because I got a question on this: There is no registration fee for the un-conference. The event is free as in "free beer." Great news in times like these, isn't it? The sessions themselves then will hopefully be free as in "free speech."

And, yes, PowerPoint presentations are frowned upon.

5. XDXF

In a future newsletter, would you please tell us more about XDXF? I went to the site today, found a dictionary that I wanted, downloaded it, but it won't open. I get a message saying Microsoft doesn't recognize the format. Since I know nothing about XDXF (and the web site doesn't really explain anything... it seems to assume that you already know all about XDXF), I have no idea what to do. I didn't find anything in the Tool Box, so I'm stuck.

This is what Becky Blackley had to say. I had mentioned the XDXF project when I talked about the lonely developer who had developed MT2007. Like the name suggests, XDXF, the XML Dictionary eXchange Format, is an XML-based format that allows you to exchange dictionary content. The site offers many hundreds of dictionaries in all kinds of language combinations with a lot of interesting and good content. Since the file are XML files, you can simply open them in a text editor. But there is obviously a limited amount of use you would get from these files that way. At first I was sure that there had to be some kind of an exchange path to translation-related formats, but I am sad to report that after looking for an hour or so, aside from the proprietary path that is offered by MT2007, there does not seem to be one.

And then I looked again, and sort of understood. After all, a dictionary is different from a glossary or a terminology database. To make this truly usable in a TEnT, you would have to do more than a mere XML-based conversion to a format like TBX, TMX, or CSV.

Look at this random sample from the English-German dictionary that I downloaded:

<ar><k>(Little) Red Riding Hood</k>
Rotkäppchen {n}</ar>
<ar><k>(Sport) entry</k>
Nominierung {f}</ar>
<ar><k>(Textilien) shrinkage</k>
Einlaufen {n}</ar>
<ar><k>(Venetian) blind</k>
Jalousie {f}</ar>
<ar><k>(Viennese) tram</k>
Bim (österr.) (ugs.) {f}</ar>
<ar><k>(Wiener) schnitzel</k>
Wiener Schnitzel {n}</ar>
<ar><k>(a corporate name)</k>
BMW   Bayerische MotorenWerke</ar>

While these entries will be quite helpful from a dictionary perspective, they are hardly helpful from a TEnT perspective. The fact that the entries contain grammatical markers and explanations would be less than helpful in a translation environment. This data is not unhelpful, but it is helpful only in separate fields that I as the translator can refer to rather than enter into the translation together with the actual term.

Still, I see that there could be ways to eventually use this kind of data, especially the more specialized dictionaries (for which the above is not a good example), but someone would really have to sit down to build intelligent conversion parameters to TBX, the termbase exchange format.

The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

© 2009 International Writers' Group