ToolkitSmall

A computer newsletter for translation professionals

Issue 9-10-150
(the one hundred fiftieth edition)
Contents
1. O Canada -- Ô Canada (MultiTrans) (Premium Content)
2. Marketing to Translators
3. Fontastic!
4. Dealing with Embedded HTML Content in XML Files (Premium Content)
The Last Word on the Tool Kit
Misconceptions

What a couple of weeks! Translator Day on September 30 brought with it lots of surprises. The biggest, of course, was the release of the crowdsourcing application that Facebook had developed for its own site to all of its partners. I will spend more time in upcoming newsletters talking about the ins and outs of crowdsourcing.

Another surprise was the weird (and unknowledgeable) missive from President Obama's office announcing "automated, highly accurate, and real-time translations between the languages of the world -- greatly lowering the barriers to international commerce and collaboration." You would think they would have learned from Secretary of State Clinton's "Reset" button gaffe and taken that as a hint that machine translation is not quite where they apparently think it is. What struck me most was the similarity to IBM's announcement of its first public machine translation test in 1954:

"The potential value of this experiment for the national interest in defense or in peace is readily seen," Prof. Leon Dostert, Georgetown language scholar who originated the practical approach to the idea of electronic translation, declared to a group of scientists and United States government officials who witnessed the demonstration at IBM World Headquarters, 57th Street and Madison Avenue.

"Those in charge of this experiment now consider it to be definitely established that meaning conversion through electronic language translation is feasible."

Although he emphasized that it is not yet possible "to insert a Russian book at one end and come out with an English book at the other," Doctor Dostert predicted that "five, perhaps three years hence, interlingual meaning conversion by electronic process in important functional areas of several languages may well be an accomplished fact."

While ATA's Jiri Stejskal wrote an extremely well-crafted letter to Obama in response, Jeromobot (almost) proved IBM's vision to be true. You see, he and I visited Warsaw this past week for another very successful TM-Europe conference, and he decided to stay with a Polish/Irish family. Sebastian, the 11-year-old über-multilingual son, is now wondering how Jeromobot, "the little translation guy," does his translation. To quote his step-dad, "I think he thinks that you put the translation into him somehow and he translates while he drums." 

Before our painful separation, Jeromobot and I had some interesting encounters in Warsaw.
1. ᐆᑲᓇᑕ -- O Canada -- Ô Canada (MultiTrans) (Premium Content)

I had promised to write about MultiTrans a few editions ago, so I updated my long-expired trial license from earlier this year and reinstalled and tested and . . .  really did not get anywhere. Fortunately, I've found out in the meantime that much of this was my own fault. Still, I ended up asking to talk to them rather than continuing to meddle around unsuccessfully with their product. (By the way, a few years ago I actually used this tool for a couple of projects, but it had simply changed a little too much for me to be successful in trying out all its new features.)

So let's look at MultiTrans in general and the newer features in particular.

MultiTrans is a TEnT, a translation environment tool, that has a slightly different slant (but then, don't they all!).

Rather than relying on a traditional translation memory, it relies on something called bitext (or "corpus"). The difference between a TM and a bitext is essentially that a TM is made up of unrelated sentence pairs without any context, and a bitext consists of complete document pairs that naturally present you with all the context you would like. (A third database technology is by reference materials for which all your old file pairs have to be kept and you can utilize and index to automatically locate matches within these file pairs -- this is the technology that Star Transit uses.) Back to the bitexts. An interesting development in the differentiation between traditional database-based TMs and bitexts is that the lines have sort of blurred. At this point, most TEnTs use a feature called context matching, 101% matching, guaranteed matching, or perfect matching, which is essentially a way to assign a 100% match a higher match likelihood since it is surrounded by the same context. For this the context naturally has to be stored in the TMs, making them strikingly similar to bitexts. On the other hand, MultiTrans has lately also embraced some principles of traditional TM technology, for instance by storing matches of the current project in a temporary TM that will be kept until the project is finished (and the file pairs are at that point entered as bitext). Because of this mixture of technologies, MultiCorpora, MultiTrans' parent, has renamed its technology to "TextBase TM."

Another differentiator in which MultiTrans excelled was its subsegment leveraging. Through an intelligent combination of information for the terminology database and advanced indexing of the bitext, MultiTrans was early on able to locate subsegment matches, a technology that it is still trying to advance (in its present lingo MultiCorpora calls it ALTM -- advanced leveraging translation memory). Of course, with a number of other tools (including Trados and MemoQ) also supporting subsegement leveraging, this feature has encountered some competition.

The last big distinguisher was the alignment feature. Because MultiTrans uses the document-based approach rather than the segment-by-segment approach of the typical TEnT, the speed of the alignment is -- comparatively speaking -- lightning fast and the accuracy is, well, different than in other tools. Since there is no manual element to the alignment process, problems do occur, but these can be fixed once they are located during the translation. Remember, the materials are stored as complete documents, so it's no problem to quickly re-align a text on the fly once an error has been spotted and corrected. That said, WordAlign, one of the new features of MultiTrans, has greatly improved the accuracy of the alignment. This process creates and extracts glossaries during the alignment that it then further uses to assure that the ongoing alignment gets better and better (this is in addition to existing termbases you can select to enhance the alignment). No other TEnT can do anything like this, either in speed or in accuracy, and the only two tools that match MultiTrans' performance are (you, loyal newsletter reader, will already know what's coming) the fellow Canadian alignment tools AlignFactory and AutoAligner.

All right, maybe we should actually say something about the tool itself and not just its features.

MultiTrans is a complex tool with what I perceive to be a less-than-user-friendly workflow, but it's nothing that you can't get the hang of after playing with it a little bit. It works in various translation environments. MS Word is still its preferred third-party environment (alongside PowerPoint and WordPerfect), meaning that you actually sort of perform the translation within these tools. There is also a standalone "XLIFF Editor" which supports a whole range of additional formats (such as HTML, XML, FrameMaker, InDesign, and XLIFF), but it needs to be licensed independently.

When I say "sort of perform the translation within these tools," I mean that unlike the early versions, menus in Word, etc., are still used to connect to the MultiTrans data and processes, but now a little "Translation Agent" sits on top of Word, PowerPoint, WordPerfect, and (from the next release on) XLIFF Editor where you actually type the translation, search for data, and send data to, say, the terminology database. Once a segment is translated it is then entered into the Word or other document in the lower half where you can see everything in its full formatting (the only editor that does not allow for a WYSIWYG is the XLIFF Editor -- here the translation segments are displayed in a table format with only limited formatting on display).

If you need to work in the actual MultiTrans application to align texts ("TextBase Builder"), maintain bitexts or terminology data, use some of the available project management features, etc., you will first of all have to get used to a "different sense of logic." (For instance, if you want to align a large number of files, you will first have to open the completely separate ListBuilder application to match file names, a list which you then save and load into the main application -- yikes!) But you will also quickly find that there are a good number of power features (for instance, you can schedule alignment processing during non-working hours) that point to the most important client focus: large national and multinational organizations. Many of these organizations (and I am talking about institutions like the Canadian government, the African Union, or organizations of the UN) are set up with large in-house contingents of translators and a more limited number of freelancers. While the in-house translators naturally use the translation environment that MultiTrans provides, many of the freelancers don't (according to MultiCorpora's estimate, only 10-30% do). The rest use the bilingual Trados .doc/.rtf format which MultiTrans supports as an interim translation output format (and as a format that can be translated within the tool as well).

Don't misunderstand me: MultiTrans is set up to be used by all sectors of the language industry -- clients, LSPs, and freelancers -- but the latter two categories have just not been the main marketing focus so far. MultiCorpora assured me that this will change, but we shall see.

Other features that are helpful (but again, maybe more for larger organizations) are the integration into content management systems such as eDocs or Documentum, a close integration with the MT engine Systran (to be purchased as a standalone), and the close partnership with the workflow environment Plunet that I've mentioned in previous newsletters.

Overall this is a very, very powerful tool that in some areas is the most sophisticated of all the TEnTs -- or at least can match what its competitors have to offer. I look forward to a new focus on the freelance and LSP markets, which I hope will not only take the shape of marketing but will also include some usability studies and streamlining of some of the processes.

2. Marketing to Translators

Here is a completely underdeveloped field: the marketing of translated products to their translators. Actually, maybe this is not all that underdeveloped since it seems to work already. I have been doing a lot of translations for a watch manufacturer lately. And guess what my godson got for Christmas last year? I just could not resist after having translated tons of manuals and marketing lingo for that manufacturer. (On second thought, maybe the program that needs to be developed is one where translators are given employee discounts for products they have worked on -- that way I might be persuaded to buy one of those gas turbines for which I translated a spec sheet recently!)

And it works the same way with software. I always tend to use the software that I have just translated -- after all, I know all the tricks once the translation is finished. (Or I know to avoid it: my favorite story is that of QuarkXPress 5, which was so bad that Quark expressly forbade its translators to use it for the translation of its materials!)

Here are some things I recently learned that way about Google Chrome:

My new favorite feature is a way to create stand-alone applications of web-based applications in Chrome. This means that you can run any website not within the tabbed browser-interface but in an interface that has nothing but the actual application. (And if you really need to use the Back button or something like that, you can either use a keyboard shortcut or you can access that and other features by clicking on the application icon in the upper left-hand corner.)

I really like this because it prevents you from accidentally closing an important application that you're working in by closing your browser or browser tabs, and it lets you completely focus on your task. This is great for things like browser-based translation interfaces or many other important tasks for which it is not important to link continuously to other webpages.

Another likeable feature in Chrome is the ability to change interface languages on the fly (under Tools> Options> Under the Hood> Change font and language settings) or, of course, the versatility of its address bar by being able to use it as a search field as well.

And as far as the acclaimed speed improvements over other browsers goes: who cares? In my world they're all plenty fast. If my brain could move as fast as my browser, I'd be grateful.
ADVERTISEMENT

NOBABEL TRANSLATOR SUITE · IMPROVED
Easy to Use · More Language Pairs

ENHANCER - Best TOTAL Leveraging Tool
Adds New TUs to Small & Large TMs

AUTOALIGNER - Creates TMs Where None Existed Before
Simultaneously Aligns Multiple Documents in Multiple Languages · Aligns PDFs

No Upfront Cost · Save Money · Pay as You Go
FREE TRIAL NOW at www.nobabel.com

3. Fontastic!

David Pooley pointed me to a fabulous little freeware (donationware) program called NexusFont, a little utility with which you can display and easily manage all the fonts on your system. The only problem is that the help is available in Korean only. But no worries: even of you don't read Korean, the program is easy enough to use without the help files.

And while we're at fonts, could you read the heading for the MultiTrans article? If not, you don't have Code 2000 installed, which would give you the ability to read Inuktitut. Code2000 is arguably the most important font that a multilingual project manager should have on his or her system since it's the most comprehensive Unicode font with more than 60,000 glyphs and an impressive array of languages covered by Unicode. I don't want to bore you with endless lists of languages and scripts that you have never seen or heard of (and I can almost guarantee you that there are many among the ones that Code2000 covers) -- but let's just say, if you ever wanted to display Klingon alongside Inuktitut and Amharic, Code2000 is for you (and for anyone else dabbling in lots of languages).

The companion Code2001 supports some additional, mostly historical, languages.

Another interesting tool on the site is the Sample Unicode Test Pages, which will give you a good idea of what scripts you presently support, plus any number of interesting links.

And, by the way, Code2000 is shareware, so do the right things  and pay.

4. Dealing with Embedded HTML Content in XML Files (Premium Content)
Since ConstantContact (the application that I am sending this newsletter out with) gave me fits to publish this article with the proper coding, you can download it here as a PDF instead.
The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Last week this reader added a link:

www.lokanath.de

© 2009 International Writers' Group