ToolkitSmall

A computer newsletter for translation professionals

Issue 7-12-104
(the one hundred fourth edition)
Table of Contents
1. Imagine 2008
2. Synchronicity
3. Imaging in 2008
4. A Legal or Illegal Speech Recognition?
5. Do Something Worthwhile Between Now and the New Year!
The Last Word on the Tool Kit
Modern Times

Thanks for welcoming the new format. A small number of you (well, I guess that's "them" now) have unsubscribed, but the majority of you welcomed the new look and feel of the Tool Kit newsletter. I know there were some funky glitches with fonts and such, but we'll iron those out over time.

Also, some of you were worried that you received Premium Editions even though you weren't Premium subscribers. Not to worry: it was all done in true holiday gift-giving spirit.


Most importantly, as we enter 2008, I wish you a fun-filled and successful new year!
1. Imagine 2008

There is a fabulous book called Born on a Blue Day by Daniel Tammet, an autistic savant, in which he very eloquently describes how his mind associates shapes, numbers, and colors. Either we all have a bit of an autistic mind or maybe it's just me, but I have always associated numbers with colors. My 12-year old daughter does the same thing, and most of the time we agree. So I'm here to tell you that 2008 is a blue-green year! (I'll ask Lara when she comes home from school whether she agrees or not.)
 
Aside from being a blue-green year, what else is going to happen in 2008? Common Sense Advisory, with its primary focus on the business side of things, has just published its
predictions for 2008 for our industry with a surprisingly strong and accurate focus on technology and processes implemented by technology.

The one prediction that I hope will prove wrong is that "language industry standards still fall short on offering value."  Though I agree with CSA's assertion that "language technology standards like TMX and TBX will still be hobbled by the small-mindedness of vendors who focus more on switching users to their technology rather than openly sharing linguistic assets" -- this is true especially in regard to TBX -- I think that another standard -- XLIFF -- will gain more relevance this next year. Unlike TBX and TMX, standards that focus on the exchange of terminology databases and translation memories, XLIFF is a standard for the exchange of translation files, so it aims much more at the root level of the process. With some tools like Lingotek, Heartsome, and Idiom already using XLIFF in their core application, and an increasing realization among many translation vendors that the focus should be less on what tools are being used in the production chain and more on what processes can be supported, XLIFF will be a driving force behind this.

And TMX will also continue to play an important role. There are limitations to the standard. Many of us have painfully observed the loss of transferability of translation memory data between different tools (or different versions of the same tool!) due to differences in segmentation and dealing with inline codes (codes in the middle of a segment). While the differences in segmentation have been addressed with the additional SRX standard (which now only needs to be adopted by the different tools -- see CSA's comment above), the inline code problem will remain. But this does not mean that TMX is worthless. Instead, it will continue to serve as the preferred format to exchange TM content in its various applications. And here' another prediction: there will be more of those next year. (More on that in the next newsletter.)

I have already given some of my predictions in the 100th edition of this newsletter, but let me repeat one for its shock value and give another that I have only recently realized:

  •  2007 was the last year in which MS Word still played any significant role in the TEnT translation process. With Trados already having moved away from Word as its preferred translation platform, Multitrans and Wordfast on their way to doing the same thing, and Metatexis hoping to do the same, there really aren't that many left hanging on to Word.
  • That was a giveaway, but this prediction may be more interesting: SaaS! SaaS, or Software as a Service, has finally arrived in our industry. SaaS is the concept of not having to install the software on your local computer, but instead using it through a web browser, with most if not all of your language data being hosted by a server. It's by no means a new concept. Sometime in 2005 or 2006, the not-so-chic-anymore acronym ASP (Application Service Provider) got rebranded to SaaS, and now we are seeing it everywhere in the language industry. To be fair, there have been a number of applications working in that realm for a while, including most translation project management applications, even some TEnTs (Lingotek, Elanex, WebWordSystem, and Idiom), along with (ouch!) the many online machine translation engines.

When I first heard about server-based computing it sounded way too futuristic, and I resented the idea because it seemed to promise me less control. However, I've come to the conclusion that freedom (from software updates, computer problems, and backup worries) is not a bad thing, and even the traditional vendors will find ways to walk that plank (and I think that most of them will find out that they are pretty good swimmers).


But, really, it only makes sense. Many, many TMs and terminology databases have gone online anyway, so why would my TEnT not be all online as well? Imagine a world with no petty little conflicts on your local machine, no worries about backups or, for that matter, no differentiations between operating systems anymore. Well, welcome to 2008 (and 2009 and 2010 and . . .)!

2. Synchronicity

Remember that great Police album? By the way, I'm not a fan of all that Sting has done as a soloist, but I've enjoyed his latest album where he interprets songs by the 16th century composer John Dowland accompanied only by the lute.

Anyway . . .  Almost that long ago, Claude Simard and a couple of other readers asked me to write about the term extraction tool SynchroTerm. SynchroTerm has had a very interesting voyage: developed by the owner of BridgeTerm, a company in Montreal, for his translator dad, it was then given away for free to BridgeTerm's customers, and now it has finally joined forces with fellow-Montrealites Terminotix who are offering it as a professional (and paid) product to complete their offering of alignment, bitext, and now term extraction tools.

This partnership is an advantage because SynchroTerm can now employ the very, very powerful alignment technology of Logiterm and AlignFactory to align texts in their various formats to bitexts, extract terminology, and present long lists of terminology with reference information that can then be verified, annotated, and exported to a number of formats, including Logiterm, the two SDL MultiTerm formats, Excel, and the machine translation tool PROMT.

Aside from extracting terminology from aligned documents, it is also possible to import TMX translation memories or bitexts from Logiterm and have terminology extracted. There are a good number of settings that allow you to govern the extraction, including many fields you can automatically add to each term pair (so that your TEnT -- translation environment tool -- will be able to take that into consideration when processing the data and make suggestions to you).

When I tested the tool, I was impressed by its user friendliness and its speed. The tests that I ran were with English and German data. This is important to mention because different languages are treated differently in this tool. In general, SynchroTerm relies on mathematical calculations to extract terminology pairs. For English, French, Spanish, Italian, and Portuguese it also uses long lists of stop words to filter those out automatically, and for English and French it also makes use of stemming rules, further improving the accuracy in those languages. All other languages are not supported right out of the box, but can be added manually by adding stop lists for those languages to the file StopLists.txt which can be found in the installation directory. You can also find instructions for this in that file. (I would be cautious with non-Western languages, though. I imported a Russian TMX file which crashed the program.)

The best-supported languages are therefore English and French, the primary Canadian languages. (Canadian tools have a tendency to offer additional features for the English-French combination --  great for those translators, but what about us? It may be interesting in this context that the government of Nunavut has chosen Multitrans, another Canadian tool, as their major TEnT for the other official Canadian languages, Inuktitut and Inuinnaqtun.) So with German as a "second-tier language," my results did need some manual massaging, but they were still good enough to quickly generate a glossary.

Not surprisingly, SynchroTerm has some shortcomings, most annoyingly the fact that I could not resize the main window and working with it becomes tedious if your screen does not comply with the exact standard display. (I was promised that this would be fixed soon.) But overall it's an interesting tool with a very enthusiastic group of users.

By the way, the website says it costs CAN$1500, but according to Jean Francois Richard, its developer, that's the corporate price. For freelancers it costs CAN$350.

ADVERTISEMENT
AIT wishes you huge translation jobs and manageable deadlines for 2008!

Unbelievable 50%-OFF Sale of all the legendary AIT PM and word count products: AnyCount, Translation Office 3000, and Projetex! 

http://special.translation3000.com/happy2008

Grab our terminology products as well!

Have a Happy and Prosperous 2008!
3. Imaging in 2008

Now, just judging by its name, XnView does not sound like a trustworthy program, but let me tell you: it is. What does managing graphics have to do with translation? Depends on what you do, but I

use a graphics manager a lot. One of the drawbacks of TEnTs (Translation Environment Tools) is that, by principle, they don't worry about graphics. (WebBudget, a translation-memory-enabled tool for the translation of static websites, or the localization tools are the only exception to that -- they include graphics managers.) But if you need to write a quote on the translation of a website or any other project that contains graphics, you need to be able to scan the graphics quickly to see which contain translatable content and which don't.

Windows ME and XP offer relatively quick viewing of graphic files in the so-called Filmstrip view (this was unfortunately not taken over in Vista), but if you have to look through hundreds and hundreds of graphic files of the many different formats and manage them (i.e., delete, annotate, organize, or convert), Filmstrip doesn't cut it. For a long time I have been using ACDSee, a well-known product for managing graphics, but the marketing hyperbole that surrounds it, plus the fact that product development focused almost entirely on hobby photographers, really frustrated me. So I tried the oddly named XnView (which is free for personal use, but costs $26 for professional use) and I am blown away by it. It's awesome. I have yet to discover a format it does not support, it's a lot less heavy on memory than its competitors, and it's very intuitive. And the fact that it's multi-platform does not hurt either.

The only drawback is that it is not Unicode-enabled yet, but then I'm not sure what I would be using text for in that tool anyway.

4. A Legal or Illegal Speech Recognition?

In the last few weeks I have been working on some very large translation projects -- so large that the traditional TEnT-based typing approach just didn't cut it, particularly because I was having problems with one of my hands. Consequently I had to rely almost exclusively on speech recognition software. Since I had trained Dragon for the terminology of that particular client awhile back (I had imported large lists of terminology), I knew that I would fare very well with it. The combination of using TEnTs (in fact, two different ones for this project) and speech recognition made it possible to finish these projects that I would otherwise have had to reject.

Just as I was in the home stretch of that project, a colleague sent me this email from an agency that she has worked for in the past:


It has come to our attention both in the past and more recently that the usage of voice recognition software when translating can cause errors in the target text, errors that are not necessarily picked up by the proof reader or by spell checking. 

Whilst we understand that certain methods of translation are better suited to some translators than others, we would like to highlight the importance of proof reading, particularly when using voice recognition software.


Some may find using such software can speed up the translation process; others will find it hinders them.  Due to the nature of this software without thorough checking of final drafts, it has been found that errors can creep into the target text without the translator or proofreader being aware.


With voice recognition, what may not necessarily be grammatically incorrect or misspelled can be overlooked when proof checking and/or spell checking.  For example, "a legal" may be changed to "illegal" or vice versa.  Both are identified as correct, due to this fact neither will be flagged by our spellcheckers or even by ourselves when it comes to checking the translation prior to dispatch.  This potentially causes confusion to the reader of the translated text, i.e. our clients.


We do of course encourage you to adopt those methods which you find most effective and provide the best results, however we also find it important to bring this issue to your attention and urge you to take this into account when using or if you are considering using voice recognition software.

I think this is really interesting. In the past, translation agencies have often adopted technology first and have then encouraged (feel free to replace "encouraged" with a word of you choice) translators to use it as well. In this case it's sort of the other way around. Translators have adopted technology and agencies are noticing it by looking at the result -- in this case, unfortunately, not in a positive way. I don't think that this agency's nicely worded complaint is unfounded. In fact, I think it's right on the money in two areas.

No matter how well you work with speech recognition, it will introduce errors of the kind that are mentioned above, and unfortunately those are errors that are hard to spot. The redeeming quality of speech recognition is that there is no reason to look anywhere but on the screen where your text is being typed as you speak (no more glancing down at our keyboards for us poor typists). So you have a chance to put more emphasis on proofing as you write -- if you train yourself.

The other thing is that spell-checking often does much more then finding spelling errors. It can often be that last chance to change the one term that you were in doubt about or to spot other problems that may just be in the general vicinity of the spelling error. As the letter above indicates, dictating a text will produce virtually no spelling errors, robbing you of that last quality check before you send your project off to the client. I'm not sure that I have a good strategy to offer to counter this except the one that I mentioned above, but it's certainly something to keep in mind.

And then there is one other thing about speech recognition: it's easy to ramble. (No, in case you're wondering: this newsletter is being typed, not spoken!) But seriously, especially if you don't formulate your sentences in your mind before dictating, you may end with a lot of unnecessary fluff. This may be good if you are paid by the target word, but it's not good for the translation.

So, is speech recognition not what it promises to be? It absolutely can be -- if you adjust to it properly.

ADVERTISEMENT
Beetext Presents Beetext VirtualOffice

LSP inspired Project Management is now available to Freelancers through a flexible monthly online subscription. Offered as SaaS (Software as a Service), it's available from anywhere with an internet connection.

Imagine, managed archiving and daily backups wrapped in a PES (Project Efficiency Solution) package! And remember, with VirtualOffice there are no installation hassles or license registration.

Visit Beetext for more information.
5. Do Something Worthwhile Between Now and the New Year!

For those of you who are twiddling your thumbs after the holidays and before the New Year, here is a way to spend it: TranslatorsTraining.com has received very positive feedback from the first batch of users. It compares 13 different TEnTs with Flash-based videos and gives you a never-before-seen way of comparing and getting to know the tools in a hype-free and objective environment.

In addition, though we can't offer you a video link to Jay-Z like Common Sense Advisory does in the post mentioned above, our page does have a video by JZ!
 

Happy New Year!
The Last Word on the Tool Kit
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you can find under www.internationalwriters.com/sample.html into the HTML code of your webpage and the little icon that is displayed on that page with a link to my website will be displayed.

Here are two new websites that now contain the Tool Kit link:

www.beetext.com

© 2007 International Writers' Group