ToolkitSmall

A computer newsletter for translation professionals

Issue 9-4-138
(the one hundred thirty-eighth edition)
Contents
1. Swordfish, AlignFactory, and openTMS are all in the Worx (Premium Content)
2. Interpreting URLs
3. Developers Are Lonely, Part II
The Last Word on the Tool Kit
Mi Chiamo Jost!

Those among us who live in the U.S. have been bombarded lately with a rather silly commercial by the language learning software Rosetta Stone. It shows a young, very farmer-like-looking farmer with the following text right beside it:

Farm boy dreaming on super model

Anyone with any kind of language background must have asked themselves the same question: What exactly does this software teach him to achieve this? Fortunately, a tongue-in-cheek reporter from the high-brow New Yorker magazine has found out.

Here are a couple of sentences from Lesson 1 (Beginning conversation):

During Milan's Fall Fashion Week, in which hotel will you be staying?
Durante la Settimana della Moda dell'Autunno di Milano, in quale hotel starà?

Could you please give me directions how to get to that hotel from western Illinois?
Per favore, mi potrebbe indicare la strada per quell'hotel partendo dall'Illinois occidentale?

No doubt that would impress her. But the article shows other useful ways to use his new Italian skills as well. My favorites are from "Lesson 7 (At the Fuel Co-Op)":

Good afternoon, Owney. I would like to buy two tanks of propane.
Buon pomeriggio, Owney. Vorrei comprare due serbatoi di propano, per favore.

I said, "I would like to buy some propane!"
Ho detto, "Vorrei comprare del propano!"

Of course you can't understand me. That is because I am talking in Italian.
Certo lei non può capirmi. Sarà perché sto parlando in italiano.

Laugh if you wish, Owney, but someday I will be having sex with a beautiful Italian supermodel in Milan, Italy, while you are still here sweeping fertilizer pellets off the floor.
Rida se vuole, Owney, ma un giorno farò sesso con una top model italiana a Milano, Italia, mentre lei sarà qui a spazzare via le palline di fertilizzante.

Well, goodbye, Owney. I will buy my propane another day.
Arrivederci, Owney. Comprerò il propano un altro giorno.

Speaking of impressive encounters, Jeromobot and I (and, yes, the rest of my family as well) went to the California redwoods the other day. Prepare yourself to be impressed.
1. Swordfish, AlignFactory, and openTMS are all in the Worx (Premium Content)

Yeah, kind of a lame little word play here . . . But there is news on so many fronts in the translation tool market that I decided to lump them all into one larger article.

Let's start with the last first. The project management application Worx is -- or rather has been -- the tool that the British LSP Language Technology Centre (LTC) has been marketing for about a year or so now. Worx is the successor to the LTC Organiser -- a tool that looks a little dumpy and old-fashioned now but that really was the trailblazer (Go Blazers!) in the PM tool market many years ago.

Yesterday, LTCannounced that it had split its technology division -- the one that supports and develops Worx and other tools -- into a separate company it now calls Agile.

Tobias Rinsche, now VP of Agile, had been trying to talk to me about this for a couple of weeks, but to be honest, I thought it just was not relevant enough to spend much time with it -- until I finally got to talk to him yesterday. The obvious thing that LTC wanted to achieve with this was a disentanglement of tools and services -- a status quo entanglement that much of our industry clearly suffers from. So far, so good. But what really excites Tobias -- and his excitement is contagious and well-founded -- is that Agile has a number of products in the pipeline that they have developed with and for LTC, products that are just waiting to be released. What's more, this whole new company (with 10 full-time developers) focused solely on technology and tool development will be good for their present and future customers.


Relative TEnT newcomer Swordfish has been quietly chipping away with the introduction of new features in the last few months. The more interesting ones are probably these:

  • Support for TXML (the Wordfast Pro XML format)
  • A plugin that allows for translation using Google's machine translation engine
  • Support for .ts (Qt Linguist) files

It's interesting and remarkable that Wordfast Pro, a new format of a tool that is hardly out of its beta phase, is already supported, but that just underlines the strong urge that exists to interchange TEnT-specific formats. The introduction of an actual machine translation component -- especially through the free Google Translate -- is becoming more and more common among tools. I still don't see a real-life use for it in my work -- not because it's machine translation in general, but because I don't think the data on Google Translate is specific enough to be of true value for a professional translator. But someone must be asking for it, or this wouldn't be introduced in so many tools.

 

I have written extensively about AlignFactory before, praising it as the best desktop-based alignment tool on the market. It achieves this with a sophisticated alignment engine developed by the University of Montréal and by using a number of filters that filter out any unlikely match.

Here are a few other highlights: It allows you to select thousands of file pairs at a time, allows you to process PDF files (!) for alignment, and does it all with breath-taking, neck-breaking, and gob-smacking speed. No, it's not perfect -- it still comes out with missed or wrongly executed alignments here and there -- but compared to the internal alignment tools of existing TEnTs, it's a no-brainer.

Last week they released a new version (2.0). It's much easier and more intuitive to change project settings, segmentation can now also be done on the basis of paragraphs -- making it easier for some bitext tools to work with it -- and you can now choose to retain some simple formatting in the TMX files (previously all formatting was stripped).

 

And lastly, the first full version of openTMS, the open source TEnT aimed at agencies and developed by FOLT, the German industry association, and Klemens Waldhör, who also was part of the team that originally developed Heartsome (and thus Swordfish), will be available on April 20. Naturally I have not had a chance to look at it, but it's very good news that things are happening so fast now after a long planning phase that had some people already writing this project off. It looks like it's happening after all. More on this after I have had a chance to look at it.

ADVERTISEMENT

Why pay a high price for a CAT tool and keep fighting with it every day?

Make your life simpler and easier by making the right choice! Heartsome Translation Studio is what you need and deserve:

  • based on open standards
  • cross-platform
  • supports all major database types
  • the most customizable CAT tool in the market
  • gives you the choice to deliver in XLIFF and TMX or tagged RTF and TXT TM formats

And as Jost puts it in the Tool Kit: "There are many obvious benefits to Heartsome -- platform-independence, a fairly small footprint on your system, a good choice of underlying database servers . . . and a very pleasant 'feel.' . . . It is certainly an interesting and powerful tool."

Enjoy the spirit of freedom. Make today your independence day! Buy now at www.heartsome.net.

2. Interpreting URLs

This might bore a lot of people who are more technical than the rest of us, but here it is anyway:

I have found it very helpful to learn to interpret URLs (web addresses) from a language/translation point of view. There are numerous parts of a URL that could identify the language of the webpage that it displays. And the cool thing is that if that is the case, chances are that the same webpage is also displayed in other languages (otherwise there is not much reason to note the language in the first place).

For instance, let's look at this long URL from the Microsoft help site:

http://windowshelp.microsoft.com/Windows/en-US/help/fe7ea80e-52a2-48d6-947a-05e02e78bc371033.mspx

Not really interesting you might say, but if you need to translate something for which you need that particular terminology, things might be different. Now, that URL has two language identifiers. One is very obvious -- en-US (in this case a mixture of the standards ISO 639-1 and ISO 3166). The other identifier may not be as obvious: the last four digits at the end are the widely used Microsoft Locale ID.

To change that page into, say, Japanese, you could just manually replace the URL in those two places with the appropriate codes:

http://windowshelp.microsoft.com/Windows/ja-JP/help/fe7ea80e-52a2-48d6-947a-05e02e78bc371041.mspx

and come out with the Japanese counterpart with all the Japanese terminology at your fingertips.

Or here is another one that I have been using a lot lately. There is a lovely English and German parallel SAP glossary at http://help.sap.com/saphelp_glossary/en/index.htm. To change the language in that case all you need to do is replace the en with de.

But since this is a glossary within an HTML frame, it's not quite as easy to get to specific entries. If you click on any of the actual English entries in the above page, the URL does not seem to change. However, Firefox has a sweet and easy way to let you open the page within the frame as a standalone page. Once you click on an entry and you have the English term and description displayed, right-click on that page and select This Frame> Open Frame in New Tab. This might open

http://help.sap.com/saphelp_glossary/en/3b/57a67b78608045852d629395c6844b/content.htm

And sure enough, just by changing it to

http://help.sap.com/saphelp_glossary/de/3b/57a67b78608045852d629395c6844b/content.htm

we get to the translated page.

Now I realize that this particular example is only good for the small handful of you who work in that language combination, but there are many other cases where this can be adjusted easily to other websites and language combinations.

Oh, and I felt really silly last week when I realized for the first time that Firefox offers a feature where you can simply highlight a word or phrase, right-click on it, and search for it in your default search engine. Sweet.

ADVERTISEMENT

Take the fast lane with AlignFactory

Align with speed and accuracy and increase the performance of your existing tools. Automate your document alignment process. Correct segments with the built-in alignment editor. Produce alignments in TMX format, with or without TM attributes, in LogiTerm bitexts format or in HTML without language tags.

sales@terminotix.com -- www.terminotix.com

3. Developers Are Lonely, Part II

Okay, here is the promised article on the second TEnT that was developed with a developer's mind -- and I am sure glad I held off on sending it out last week because it turned out that I really had seen only a small part of it.

So, let's start at the beginning. I didn't have to think much to come up with a title for this article, because the developer of MT2007 -- Andrey Uzbekov, aka Andrew Manson -- must be a very lonely man. Otherwise, I wouldn't understand why he put scantily clad women on the welcome screens of his software. (And if you think the woman in the main application is scandalous, check out the one in the alignment tool!) Seriously, I dithered back and forth on whether to write about this tool since it seemed just a little too much on the side of bad taste. Nonetheless, here it is.

It appears that the name -- MT2007 -- is an abbreviation of a reverse translation of Память переводов -- Memory Translation -- rather than Machine Translation as most of us would suspect. It is a Russian tool, and it is shockingly robust and feature-rich for being donation-ware. Aside from its rather appalling welcome screens, it's a very attractive program that provides a clear and easy-to-use interface . . . that is, if you read Russian or if you manage to switch to the (semi-)English interface (you'll need to run the LanguageSelector.exe in the bin folder for that). It comes with a very large amount of Russian-to-English glossary/translation memory data from the XDXF project (which makes it a very slow - and, in my case, often aborted -- 160-MB download, but I'm sure it's a very welcome resource for Russian translators). These glossaries, or I should probably say "dictionaries," are open-source material -- if you are not familiar with that project, have a look at the XDXF page for your language combination and you might just like what you find -- as are other components of this program, including the English WordNet dictionaries (also very much worth a look!) and most of the programming components. After you have downloaded the tool it does not require any installation; instead, just unzip it. 

(In case you are running a language version of an operating system that does not use commas as the decimal separator, you might not be able to actually use the program right away, since the font in the Translation Editor where the actual translation takes part will be obscenely large and unusable. This is because the font size is coded as a comma-separated value. To fix that, open MT2007.exe.config in the bin folder with a text editor and change the font size in all lines that contain "editor_font_size." And if you think that's frustrating, read the first two sentences of the manual: "As for now, MT2007 can be regarded as a perennial beta version. Which means that errors are guaranteed.")

While the interface of the tool is really simple, it might take some time to get used to it because the concepts (and terminology) are a bit different than what you might be used to. For instance, there is no real distinction between translation memories and terminology databases in this tool; instead, there is a feature that allows you to highlight subsegments, translate them, and send them to the "repository." The repository seems to be a collective term for linguistic materials including TM data -- which, just as in the upcoming version of Trados, uses SQLite as its underlying database system. Another core concept -- and here the developer's mind becomes only too apparent -- is the use of regular expressions. This is helpful if you want to build your own fuzzy search algorithms for languages with a high degree of flexion. That particular feature wouldn't be my cup of tea, but for those who are good with regex, more power to you.

The way you import TMX (translation memory exchange) data is exactly the same way you would import any other file you want to translate -- only with TMX you will send it to the repository after the import. Once the data is in there, automatic searches are relatively fast and there is a host of keyboard shortcuts (that are easily viewable by way of the little help icon in the Translation Editor) to switch between windows and enter data from the matches that are found. Like Déjà Vu, Heartsome, and Swordfish, MT2007 uses example-based machine translation (although I am still not convinced that this is the right term for what it actually does) by combining matches and "correcting" fuzzy matches through deleting certain parts of target segments and replacing others -- if the necessary material is available in the repositories or the other associated language resources.

Like many other tools at this point, this tool also just integrates true machine translation. PROMT, the Russian machine translation engine, is supported if it is installed along with -- surprise, surprise -- Google Translate.

The translation file formats it supports are MS Office 2007, OpenOffice, and MS Word 2003 saved as XML. Once the files are imported they are displayed in a tabular easy-to-work-with format.

To me it seems that this is a tool that would definitely warrant a second look by English<>Russian translators with a knack for regular expressions (and semi-naked girls), but it certainly has the potential to be more than just that. The author of the documentation for this tool, Konstantin Lakshin, has also written an interesting review of it in the publication of the Slavic division of the ATA that you should consult if you would like to dig deeper.

Overall I was impressed with the tool. I liked the concept -- just take a bunch of freely available components and make them into something new and useable. And while the tool might be primarily for one language combination, that does not have to be a bad thing. Too often we narrow our opportunities by trying to do something that fits everyone. And I like the very laid-back way of "marketing" the tool. Like its self-depreciating beginning, the manual ends by asking:

Before you move onto studying other features in MT2007, you should answer one question: "Do you personally find this approach to automation of translation interesting?"
The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

© 2009 International Writers' Group