ToolkitSmall

A computer newsletter for translation professionals

Issue 10-3-162
(the one hundred sixty-second edition)
Contents
1. The End of the Five-Year-Rule (Premium Content)
2. @coöperate
3. Trados Starter
4. Searching, Zooming, and Wheeling
5. TermWiki
6. Error Recocknition (Premium Content)
7. A Love Story (continued)
The Last Word on the Tool Kit
Friede, Freude, Eierkuchen

I had some great talks with attendees of the very successful ATA Translation Company Division's conference in Scottsdale last week, and most of them seemed to center around relationships. Without wanting to moralize too much, these talks confirmed something that I have been mulling about for at least the last few weeks: What is the true nature of the language business, and what do we truly have (to offer)? Sure, the obvious answers are language and (hopefully) subject matter expertise. But those alone don't make a business.

The business is only possible because of the relationships we have created and maintain: relationships with our clients (again, that's the obvious one) but also with our peers and, for an LSP, vendors.  Aside from those two things -- skills and relationships -- there really isn't much more that we are made of. Yes, we may own some equipment and have access to technology and data, but those are meaningless without the other two.

While this might not sound too earth-shattering, my feeling is that two groups in our industry need to be reminded of this: Those who have strayed from the relationship focus need to re-embrace it as an essential value and cornerstone, and those who have maintained that focus may need to sharpen it and mine the unique strengths that come with relationships. I am particularly thinking of the lost art of cooperation (or as the New Yorker would put it, "coöperation").

And why exactly are we talking about this in a technical newsletter? Because technical know-how, expertise, and experience are exactly some of those areas that could be addressed by a stronger sense of community and cooperation. Why not have co-op-like constructs where we utilize each other's strengths, or where we support an infrastructure for companies that focus on offering technical services for our industry?

One area where this will become particularly relevant is an intelligent employment of machine translation. You can find some of the discussions on differences in the perception of machine translation in this report and more later in this newsletter.

1. The End of the Five-Year-Rule (Premium Content)

At the above-mentioned TCD conference there was a (semi-) official declaration that machine translation's five-year-cycle has ended. The five-year-cycle started in 1954 after the IBM 701 "dashed off its English translations [of a handful of Russian sentences] on an automatic printer at the breakneck speed of two and a half lines per second." The project leader Leon Dostert declared:

"Those in charge of this experiment now consider it to be definitely established that meaning conversion through electronic language translation is feasible."

And

although he emphasized that it is not yet possible "to insert a Russian book at one end and come out with an English book at the other," Doctor Dostert predicted that "five, perhaps three years hence, interlingual meaning conversion by electronic process in important functional areas of several languages may well be an accomplished fact."

We all know what happened afterward: This imagined five-year-goal was pushed out every five years for another five years, until we suddenly find ourselves in 2010. And the output of machine translation is mostly still not particularly "good," but it certainly is used. Since in many cases the concept of "good" has been replaced with "usable," it has become, well, usable (sort of). And therefore the five-year goal is a thing of the past. We now know that MT's limitations will stay (and gradually become "better"), and we have all accepted that there is a certain level of usability within those limitations.

And what we certainly do see is a lot of creative juices flowing to quickly collect more humanly translated data to train the machine translation engines to make them even more usable. Some of these efforts tend to go against our grain (for instance, see this blog post by someone with strong feelings on Google's efforts through the Google Translation Toolkit), while others may be viewed as more generous because there is more given in return.

A case in point is Microsoft's new Website Translation Widget (I was alerted to this by a blog posting by Dave Grunwald).

Like many before, this little MT application sits on a webpage and allows you to translate the content of that webpage into any of the two dozen offered languages. And the result is often, well, not very good. See for yourself right here. What's different about this is that it combines crowdsourcing (sorry, I know that many of you don't like this term!) and machine translation. Again, the result of the machine translation is less than stellar, but you as a visitor to the website (or as a collaborator or owner, for that matter) can click on any of the sentences and edit any of the segments. If in our example you choose German as the target language and click on the segments, you will see that the first three segments have been edited and approved (by me, as the website owner).

The difference between this and Google's machine translation correction efforts is that the human efforts do not (only) disappear into a black hole for the betterment of Google; instead, these corrected and approved translations will from now on always be used for your specific webpage. (The program is in beta right now, so sometimes the corrected translations will be listed as alternates rather than the actual translation. This will be changed and may even already have been changed when you read this.)

These translations will be specific to your webpage and will not be used as the standard translation for any other webpage, even if the original content of that webpage is the same.

Well, that's actually not quite true since they will be used by Microsoft to train its machine translation engine, so in a roundabout way they will be used by others as well.

So, yes, Microsoft will benefit by getting very high quality data, but it seems to me that the benefit given in exchange is tangible enough to make this an interesting barter (plus, at some point, the webpage owner can also download the translated data as a TMX file).
ADVERTISEMENT

SDL Trados Studio 2009 Service Pack 2 is here!

 

Celebrate the launch with a free 30 day trial for SDL Trados Studio 2009 SP2. Plus, enjoy exclusive discounts of up to 25% when you buy online. Hurry, offer ends 31st March.

 

To download your trial or buy online visit -- www.sdl.com/toolkit_03.


2. @coöperate

(You notice that I spared you from the all-too-common wordplays with "Twitter," "tweeting," and "tweets" for this heading? You're welcome.)

I would like to officially eat some crow and admit that I was wrong when I wrote some rather ditzy comments about my by-now favorite social network application (and really the only one I use): Twitter.

Twitter is useful.

If used right.

Here are the things that I -- well, actually Jeromobot and I -- use it for: We have about 200 people whom we follow -- all from within the language industry -- and we spend about 20 minutes or so per day to glance through the 500 or 600 tweets from these people and maybe add a couple of tweets ourselves. In those 20 minutes daily, we (OK, I'm going to drop the "we") have learned more about new developments, new tools, ongoing discussions, and outlooks on language and translation than I could otherwise have done in three or four times the amount of time.

It took me awhile to narrow down the people I follow to the ones who truly have something useful to say (not the "Oh-I-feel-like-taking-a-long-hot-bath-after-a-long-day-of-translation" kind of folks), and it has been useful to have learned how to use the Twitter tools.

The hash tag (#) is particularly useful when used properly. If it's placed before a word (note that the word can't contain special characters or spaces) it makes it searchable and hyperlinked and therefore relates it to other tweets that also have the same hash tag. Every couple of days I make a search for #tl8, the most commonly hash tagged word for translation-related tweets, to see whether I'm missing folks I should be following.

The hash tag also becomes important when you are at an event you are tweeting about and you want to make sure that other people see it, or so that you can see what other participants say about the event at the same time. Also, it makes it a fantastic tool to follow and interact at conferences and other events without actually being there. For instance, take a look at the discussions on last week's TCD conference at #tcd11. It will be very interesting to follow the ELIA conference in Istanbul -- #NDlst -- on April 15 to see what Jeff Chin and Michael Galvez will have to say about the Google Translation Toolkit. (By the way, I was promised that my legal fees would be paid if I sued Google for the copyright violation -- you can read all about it on Twitter!)

The @ symbol before a name denotes a specific person with a Twitter account, and using it within Twitter turns that person's name into a hyperlink that can be used to jump to that person's page where you can see all his or her tweets and subscribe (or not). @Jeromobot is my Twitter page, and if you look at my followers you get a pretty good idea of what I think is useful.

Oh, and if you don't want to sign up to Twitter but still receive tweets from a specific person, you can subscribe to the corresponding RSS feed, such as this one for Jeromobot, or for an event or a specific topic, such as this one for the ELIA conference.

Overall, Twitter is a good tool for building community, and since many within our industry have turned to Twitter, the community will only grow stronger. To reinforce the theme from the introduction: that is something we can only benefit from.

(An added benefit: There are also language jokes on Twitter. How about this one from @hansfens: "Why do the French only ever eat 1 egg? Because 1 egg is always un œuf. (Read aloud.)")
3. Trados Starter

In my last newsletter, I wrote about the new SP2 version of Trados Studio, and in particularly about the Starter Edition. I had been briefed by some SDL folks about it beforehand and thought I had a good impression of what it is and what it is supposed to do, but I was surprised to see some features described quite differently on the feature list on SDL's website. Cool, I thought, they changed that, and I wrote:

In some discussion groups, questions were raised about whether this move with the Starter Edition was a response to Lionbridge's upcoming Translation Workspace offering. But the more I looked into it, the more I am convinced that it's not. For instance, it is neither possible to connect to server-based TMs nor to open packages that clients might send you, so it has little in common with what the Lionbridge offering intends to do: be a cog in the larger workflow picture.

It turns out that SDL's website was erroneous (it's now corrected). So, yes, it is possible to "connect to server-based TMs or to open packages that clients might send you," which makes the Starter Edition, very much like Lionbridge's offering, a "cog in the larger workflow picture."

4. Searching, Zooming, and Wheeling

Here is a cool little trick for Firefox users. We all know that Firefox (and most of its competing browsers) has a little search bar in the upper right-hand corner. Presently the default search engine is Google (I would not be surprised if that would change anytime soon), but if you click on the little arrow to the left of the search it gives you a bunch of other choices of search engines or other searchable sites as well. You can add to this list or delete existing entries by selecting the Manage Search Engines option at the bottom of the list -- most other browsers have something similar.

I only recently learned that if websites use a certain code in combination with an XML code file on their server, Firefox automatically recognizes them as searchable. When it encounters that, the little down-arrow to the left of the search field gets a blue shading. If you open the search list, the currently opened website is already listed, waiting to be entered as one of your default search engines. Cool, huh?

Search options in Firefox

And here is another trick -- really, it's more of a bug fix. Many of us are used to being able to change the zoom in an MS Office document (and many other applications) by pressing the Ctrl key and turning the wheel on the mouse. It works great -- only not in some combinations of Windows Vista/7 and Word 2003/2007. The fix? Turn the mouse wheel REALLY hard and fast. Not very elegant, but it works.

5. TermWiki

Awhile back I mentioned the efforts of CSOFT on the development of TermWiki,which is now officially released. I took some time this week to talk to some folks at CSOFT in order to better understand the underlying concepts of the product, and I was impressed.

Why another terminology tool? I'm not sure this is the right question, and I imagine that the folks from CSOFT would respond: What do you mean another tool? Well, they may not quite say it like that, but they did look at the existing tool landscape and essentially found two different kinds of terminology tools within TEnTs: the first group essentially supports mere glossaries (i.e., bi- or multilingual lists of term pairs) and really frustrates terminologists because it essentially focuses only on the translation aspect rather than the whole document creation and terminology maintenance life cycle; the second group provides full-fledged terminology tools (think SDL MultiTerm or TermStar) that are so complex they are either not used at all or too little.

So the goal was to create a tool that provides real ease of use (to translators, authors, and terminologists) while providing enough complexity and customizability to make it useful for all parties involved, including the different levels of access rights that are necessary for the different roles within different groups. The platform that seemed most suitable was a Wiki: it's very customizable, easy to use, and role-centric, but at the same time allows for a great deal of collaboration. Thus the idea of TermWiki was born, and the development efforts were helped by a large CSOFT client whose commitment to a new terminology solution acted as a catalyst.

What is now offered are two different versions: an open-source version that is downloadable from CSOFT's site, and a paid version (with a per-user/annum fee) that includes certain additional features, including notifications as well as (the very important) import and export of data. Aside from that, the normal (un-technical) mortal might be hard-pressed to install and run the Linux-based open-source tool, so the paid version also comes in a hosted format (if that is desired).

So how does all this fit into a translation workflow? Well, it doesn't, and that's really what sort of convinced me. TermWiki (in its current state) is a platform that allows you to create and maintain corporate terminology with all the sophistication or simplicity you choose to have. And rather than bringing the TermWiki interface directly into the translation process, it allows you to export the necessary (and at any point very current) terminology data into a TBX or other kind of exchange file, which then can be imported into the terminology component of your TEnT and be as interactive with you as your tool allows it to be.

The LSP project managers' job suddenly becomes much easier because they receive only a link that allows them to export a data set that is project-specific rather than meddling with 50+MB large terminology that may or may not be current but for sure contain lots of unnecessary information. Translators might never actually see the TermWiki, but they benefit from the fresh data export. And the clients just need to know how to generate correct links (and, of course, how to maintain the data).

I think this sounds really clever, and as I am writing this I'm wondering whether this might not be something that TEnT vendors would be well advised to look at for their own use, in particular because of the open-source component.
ADVERTISEMENT

Having project management and tools interoperability in focus, memoQ 4.0 is just the perfect tool for language service providers.

Find out more on our new website at kilgray.com
6. Error Recocknition (Premium Content)

When I posted the title "Eror Recognition" for the first part of this article in an earlier newsletter, a reader wrote to ask whether I knew that I had it misspelled. I did.

The article was about our inability to see our own typos when editing our own work. I'm interested in this topic not only because I have that problem a lot in my day-to-day work, but also because I have long believed that we essentially read words very much like Japanese kanjis or Chinese characters. We don't decipher them stroke by stroke -- we read them as whole images, very much like we look at words. We experience this when we stare at a certain word and know intuitively that something is wrong -- typically we are right in our sense of wrongness, but we just can't see the forest for the trees. We need to force ourselves to look at it differently so the image dissolves and the letters appear.

One way to do this is to change fonts -- particularly if you change into a weird font that will first throw you off -- to force your mind to actually read. Another reader suggested that reading aloud is a really good way of finding those errors, or, for that matter, having the computer read to you with a text-to-speech program. (Someone else suggested reading backward -- I'm not sure how that would work. . . . ) But there are other ways to trick your brain within Word (some of these also work in other applications):

  • Switch to Draft font (make sure that you are in the Draft or Normal view of your document and select Tools> Options> View> Draft font in Word 2003 and below or Office button> Word Options> Use draft font in Draft and Outline views in Word 2007). This will give you a plain text view of your document without actually changing anything.

  • Change the color of the font by highlighting the text and using the appropriate button on the toolbar (Word 2003 and below) or the button on the Home ribbon (Word 2007).

  • Change the color of the background: You can either change that by printing the document out on colored paper -- but that's not very environmentally friendly -- or you can use the old DOS/WordPerfect look (blue background, white letters) by selecting Tools> Options> General> Blue background, white text. Unfortunately, this has been kicked out for Word 2007, but you can manually change the background color in Word 2007 under Page Layout> Page Color (you can then hover over the different colors and it will show you the different options immediately).

7. A Love Story (continued)

This is an intermittent series about characters I really like. I have a passion for exploring different writing systems, character sets, and fonts, and this one was particularly meaningful to me -- it made me feel at home.

If you have a Javascript-enabled browser, hold your cursor over the character for a definition.

The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Last week this reader added a link:

www.word4wordtranslation.com

© 2010 International Writers' Group