ToolkitSmall

A computer newsletter for translation professionals

Issue 9-11-155
(the one hundred fifty-fifth edition)
Contents
1. News of the (Translation) World
2. Athwart! (Premium Edition)
3. And Even More Windows 7 Tidbits (Premium Edition)
4. Happy New Year!
5. Watercooler
The Last Word on the Tool Kit
The Law

The Global Language Monitor awkwardly announced last month that

Twitter is the Top Word of 2009 in its annual global survey of the English language. Twittered was followed by Obama, H1N1, Stimulus, and Vampire. The near-ubiquitous suffix, 2.0, was No. 6, with Deficit, Hadron the object of study of CERN's new atom smasher, Healthcare, and Transparency rounded out the Top 10.

Now, I am not going to argue with that. The word Twitter truly has experienced a comet-like rise this year. What has been really fascinating to me, though, is that the founders of Twitter apparently assumed that they could coin two terms rather than just one -- and failed. We all know that English has no problems with turning a proper noun into a verb (just google that if you don't know what I mean . . .), and many other languages are more than happy to follow suit. "To twitter" would have worked just fine, but "to tweet" is asking just a bit too much. Even Twitter-savvy English-speaking folks have to really concentrate to say that they have just "tweeted" rather than "twittered" (see the "twittered" error in the announcement above). As far as I can tell, "tweet" has no chance of survival in non-English languages. Maybe it will come up in the official Twitter translations that are presently being crowdsourced, but I don't think it will really make any mentionable impact on actual language usage.

I may be wrong about this, but if not, this might be an interesting new "law": "Hope to coin one new term per product. That term may be contorted into all kinds of word forms, but don't be arrogant enough to think you can do that with two terms: the foreigners won't like it."

On another note, the Jewish festival of lights, Chanukah (also referred to as the festival of many spellings, such as "Hanukkah" or "Hanukah"), starts this year on December 12. For all to whom this relates, Happy Chanukah! And know that Jeromobot celebrates with you.
1. News of the (Translation) World

Again, there are a number of interesting releases and announcements that have been made in the last couple of weeks in and around the TEnT market.

  • The quality assurance tool QA Distiller has seen a new release with version 7. You can read about the new features right here. Incidentally I just worked on a project last week for which I had to work through a QA report by QA Distiller and I really liked it. It found a good number of consistency errors that the integrated QA tool (in this case, SDLX) had not been able to find.
  • I have written about Clay Tablet a number of times before. It's essentially a middleware tool that offers ready-made and to-be-designed connectors between content management systems and various TEnTs and workflow management systems. They announced a partnership with Translated.net, the company that publishes MyMemory, the ginormous online TM which consists of a mixture of aligned online resources, contributed translation memories, and user-corrected data. I am not sure that I completely buy into the promise of the press release -- in fact, I know that things like "everything that has already been translated doesn't have to be translated again" or "with over 210 million translated sentences in MyMemory, translation costs for many clients will be dramatically reduced" are wishy-washy marketing speak -- but it's still interesting. I would have expected an actual TEnT vendor to come up with a partnership with MyMemory rather than Clay Tablet, but there you go.
  • The TAUS Data Association (the industry association that is pooling large amounts of translation memory data primarily for machine translation training purposes but is also offering term lookups in its vast archives) is starting a program aimed at smaller LSPs and freelance translators. You can join the TDA for 75 Euro a year, and in exchange you can upload TMX data and download 10 times the amount you uploaded. I am not sure what data can be uploaded by LSPs or freelancers since the ownership issues might be a little unclear, especially when you join an organization in which the very clients you might have translated for are members. Still, this is an interesting step and we shall see how this will be embraced. The actual upload and download function will be enabled early next year, so you might have some time to think about whether it makes sense to join.
  • The translated Google web search got quite a bit of attention last week. A lot less attention was given to the almost simultaneous roll-out of the mono- and bilingual dictionaries at www.google.com/dictionary, which is probably more relevant to many of us. There is a total of 12 monolingual dictionaries (Chinese [Simplified and Traditional], Czech, Dutch, English, French, German, Italian, Korean, Portuguese, Russian, and Spanish) -- with the English dictionary by far the most concise -- and combinations between English and 36 different languages (Arabic, Bengali, Bulgarian, Chinese [Simplified and Traditional], Croatian, Czech, Finnish, French, German, Greek, Gujarati, Hebrew, Hindi, Italian, Kannada, Korean, Malayalam, Marathi, Portuguese, Russian, Serbian, Spanish, Tamil, Telugu, Thai). The appearance of the dictionary helped me with at least one thing: I finally knew where the odd dictionary entries in Google Translator Toolkit came from. Unfortunately, that did not help much, because we still don't really know where the data comes from. It looks like Google licensed data for some of the material from the dictionary publishing house Collins, but that only covers a few of the languages. Be that as it may, I am not sure how valuable the searches are for language combinations that are well supplied with good and specialized dictionaries (I am talking about the translator's point of view here), but for some of the other languages this might be a helpful resource. The one thing I did really like about it was that it also offers definitions from Wikipedia, Princeton, and other sources.
  • This is not new, but it was news to me: OPUS, an open-source parallel corpus. This is a very, very large resource of bilingual files in many language directions containing such varied materials as data from the European Medicines Agency, the European constitution, the  European Parliament Proceedings, the OpenOffice.org corpus, the opensubtitles.org corpus, and various open-source localization and software documentation files. The author of the site, Jörg Tiedemann, is a researcher working in natural language processing and machine translation, so the files are not especially made for translation memory -- most of them are in a text format -- but nothing that could not be converted to a TM-compatible format or even TMX (and the files for the European Medicine Agency are in fact in TMX). Very interesting stuff!
ADVERTISEMENT
Why pay a high price for a CAT tool and keep fighting with it every day? Make your life simpler and easier by making the right choice! Heartsome Translation Studio is what you need and deserve:
  •  based on open standards and cross-platform

  • the most customizable CAT tool in the market

  • gives you the choice to deliver in XLIFF and TMX or tagged RTF and TXT TM formats

And as Jost puts it in the Tool Kit: "There are many obvious benefits to Heartsome -- platform-independence, a fairly small footprint on your system, a good choice of underlying database servers . . . and a very pleasant 'feel.' . . . It is certainly an interesting and powerful tool."

Enjoy the spirit of freedom. Make today your independence day! Buy now at www.heartsome.net.

2. Athwart! (Premium Edition)

"Athwart" might be an archaic term for "across," but I realized once again this week that Across is not the name of an archaic product. In fact, it's very modern.

Across has been a stalwart contender in the TEnT market of Germany and other European countries, but it's only now gaining some ground in the U.S. and elsewhere. For the freelance translator, Across is unique in that it's offering its tool for free on the My Across site. Really! And it is in fact fully functional, not a "dumbed-down" version that can only process projects created by the enterprise version. You think that sounds suspicious? Well, Across figures that the money to be made is located not with the freelance translator but with the large translation buyers, and to some degree with corporate language service providers (a point that could be disputed). In order to successfully market to those segments, they need to be able to provide an infrastructure of translators who are able and willing to support the technology. And while Across does offer a component called crossWeb (get used to all the "cross" terms for all the program's modules with Across; admittedly thwartWeb would have sounded bad), which displays a browser-based translation interface that is surprisingly similar to the desktop interface, they know how hard it is to persuade translators to use web-based interfaces. (The numbers prove it: the Across representative I spoke with estimated that only 2 to 3 percent of all translators using Across use crossWeb.)

So, is Across a tool whose functions are primarily geared toward the enterprise user? In a certain way it is. It's very workflow-oriented, is supported by a rather large database system, and the ultimate goal is not to have users use the free Personal Edition as a standalone tool but as a means to connect to an enterprise server that holds all translation data (documents, TMs, and termbases). That said, there are some features in Across that I really like from a freelance translator's perspective. For instance, Across was the first TEnT to offer real-time spell-checking along the lines of MS Word. And in what may be its coolest feature, it also has a module called crossSearch that displays hits from internet, intranet, or local sources (such as glossaries, dictionaries, etc.). Naturally, all the settings for these resource settings are user-definable.

But these are things that have been with Across for a long time, and most of you have read about them.

Last week I looked at some of the new features of version 5 of Across, and here they are in a nutshell:

  • crossMining is a new module that allows you to extract terminology from translation memories for various purposes, including quality assurance checks of the TMs when those lists are compared to verified terminology databases (one of the strong points of Across term recognition is its morphological capabilities in various European languages) or the employment of these dictionaries as Auto-completion dictionaries for the translation process. The languages supported for this feature appear to be a number of European languages, including Turkish, and Arabic.
  • crossAnalytics is the new feature that tracks every record entered into the database in any number of ways so that it is very easy to track individual performance by translator, language, project, server, or even ROI (here: return on investment; Across also likes to use that abbreviation for "Rely on Independence" -- pointing to its own independence in comparison to some of its competitors). It's unfortunate that crossAnalytics can be used only from the moment it is installed, so any previously existing Across data is "invisible" to this module.
  • crossAutomate is a feature that -- surprise, surprise -- automates processes, slightly reminiscent of the way that Idiom works. While there is a visual workflow manager, the creation of customized workflows is presently only offered as a service, but it's in the works to also be created by the user.
  • And lastly there have been a number of changes to crossTerm, the terminology module, in particular the stricter separation of client resources. While in previous editions it was possible to separate terminology resources from different clients by filtering them accordingly, they are now stored separately and are thus easier to delete, export, or distinguish.

It won't surprise most of you that these new features -- with the exception of the refined terminology management -- are not available in the Personal Edition (the crossMining feature can be utilized by using Auto-completion dictionaries but you cannot build them yourself with the Personal Edition).

All in all, Across is a real powerhouse, which, not surprisingly, needs to be studied well before using. This is certainly true on the corporate level and to some degree also on the freelance level. Still, those slow and dark afternoons in December might be well spent with downloading and trying out this tool.
3. And Even More Windows 7 Tidbits (Premium Edition)

In the last newsletter I promised to write about the new life of the WinKey in Windows 7. More on that in a second, but how about this new little feature:

Windows 7's Problem Steps Recorder is a tool that lets you record everything on your screen (with the exception of text that you enter). Once the recording is done, it is not saved as a movie file but as an MHT HTML archive file and zipped up. Once unzipped you can open the MHT file with either Internet Explorer or Opera (or with Firefox or Safari with a special plug-in) and it will give you a screen-by-screen description of what just happened on your computer as well as a narration of it alongside operating-specific information.

To start the recorder, click on the Windows button, type psr, and hit Enter. Everything else is very self-explanatory. This is a great way of passing on information about problems you might be having with your machine to a technician or the like.

(The cynic, of course, will ask: Why would Microsoft have to include such a tool if everything is so hunky-dory with Windows 7? I know, I know. . . Fact is, I had a few smaller hiccups in the last three weeks with Windows 7, but certainly no crash, blue screen, or anything like that. So I am still happy.)

And here they are, the WinKey shortcuts. The WinKey has always been my little friend and I have written about it in the past, but its activities were more or less confined to

  • WinKey+E: Open the Start menu
  • WinKey+E: Display Windows Explorer
  • WinKey+R: Display the Run dialog box
  • WinKey+M: Display the desktop

There were a couple of other uses, but really those listed above were the only ones that were very useful. Now there is a whole plethora of new ones, such as:

  • WinKey+Up (Down): Maximize (Restore / Minimize) your current window
  • WinKey+Left (Right): Snap your current window to the left (right)
  • WinKey+T: Focus the first and then succeeding taskbar entries (WinKey+Shift+T cycles backward)
  • WinKey+Space: Peek at the desktop
  • WinKey+P: External display options (instead of the silly Fn+F key combinations)
  • WinKey+X: Open Mobility Center (access to things like brightness, volume, battery, wireless connectivity, external display, etc. -- this feature is also available on Vista)
  • WinKey+number key: Launch a new instance of the application in the Nth slot on the taskbar (same in Vista).

And here is my favorite:

  • WinKey + +: Zoom in to 200% (WinKey + - goes back to 100%).

And here is one other additional helpful Windows 7-specific shortcut:

  • Shift + Click (or Middle Click) on icon in taskbar: Open a new instance of the application

Happy shortcutting!

4. Happy New Year!

No, these are not premature congratulations but more an expression of anticipation of good things to come in the upcoming year. After all, MemoQ 4 will be released early next year and was introduced earlier this week in a webinar. Unlike the last couple releases that primarily were directed at the translator, the emphasis of this MemoQ release is going to be the language service provider and the corporate client, with some major new features for these groups. However, the freelance translator will also find some new features that will prove to be helpful (such as real-time spell-checks and an improved quality assurance checking).

Here are some of the "corporate" features:

There will be a completely new project management facility (humbly termed "Overview") which will allow the setup, pre- and post-processing, and managing of distributed multilingual projects. One feature within the Overview module has the potential to be a game-changer, I think. That is the so-called Post-translation analysis. The idea is that it would be a good idea to do analysis after the translation is finished. That way you would be able to see how many fuzzy matches (for example) you actually had, plus you can see how much one translator benefitted from a colleague working on the same project (assuming that one would start early in the process and another starting later would benefit more from the work of the early one). That is all fine and good. What makes me wonder a little is that the post-translation analysis also analyzes the percentage of fragment matches. I have talked about the potential savings through subsegment matching as one of the new battlegrounds of our industry, comparable to the price reduction conflicts that we first saw with perfect and fuzzy matches 10 or so years ago. With this kind of analysis, MemoQ is the first tool that potentially makes it possible to find new pricing models that would take subsegment matching into consideration. I'm not sure that this would be a good thing at this point. . . .

Another feature will be the treatment of any kind of setting as a "resource." This might not sound earth-shattering to you, but this will prove to be super-helpful because it means that settings can be compartmentalized, exported, and shared. And this is true for anything from keyboard shortcuts to TM and termbase settings to segmentation rules or QA settings. This will allow an easy exchange of the relevant settings between client and translator or between translators (or computers).

Projects for distribution to translators or colleagues will consist of zip files that exclusively contain files in exchange formats (XLIFF for translation files, TMX for TM files and CSV -- unfortunately not TBX -- for termbases). This of course means that you can work on those files in any tool that supports these formats, which at this point is almost every other TEnT. I really like and welcome this particular move. Other vendors have been a lot more quiet about what the "project files" contain, and it is a welcome change to make the exchangeability with other TEnTs so apparent and easy.

So, I (and I'm sure many others) am looking forward to seeing MemoQ 4 when it comes out early next year. (And maybe that subsegment analysis feature is not completely set in stone yet. . . .)

5. Watercooler

Watercooler, a place with "Tips, Tricks and Networking for Translators" is a great (virtual) place to hang out without the loud commercialism of other such places. I might just see you there.

The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Last week this reader added a link:

www.philiprand.it

© 2009 International Writers' Group