|
|
|
2011
| |
2011 is a year that's pregnant with meaning for translation in the English-speaking world because of the 400th anniversary of the publication of the "Authorized" or King James Version of the Bible -- a milestone perhaps comparable only to Kumārajīva's Buddhist translation into Chinese in the 5th century or Luther's translation of the Bible into German in 1534. And if you think all this "religious translation stuff" has nothing to do with us, not only will you have to take it up with Jeromobot, but the truth is that in many if not most cultures and languages, religious translations were the first to be done and often set the tone for the secular translations that followed.
I can't help but be touched by these words from the preface of the King James Version, which only at first glance seem antiquated:
Translation it is that openeth the window, to let in the light; that breaketh the shell, that we may eat the kernel; that putteth aside the curtaine, that we may looke into the most Holy place; that remooveth the cover of the well, that wee may come by the water, even as Jacob rolled away the stone from the mouth of the well, by which meanes the flockes of Laban were watered. Indeede without translation into the vulgar tongue, the unlearned are but like children at Jacobs well (which was deepe) without a bucket or some thing to draw with: or as that person mentioned by Esau, to whom when a sealed booke was delivered, with this motion, Reade this, I pray thee, hee was faine to make this answere, I cannot, for it is sealed.
And, swoosh, let's make a quick transition to 2011.
Two thoughts on the Tool Kit. First concerning the real one, the one you are reading right now. This is the first edition to appear in a Premium edition and a Basic instead of the previous "Standard" edition. The difference between the new Basic and the old Standard? The Basic's sho
Yes, it's shorter to remind you of all the good things that are to be had in the long and lovely Premium edition, which can be ordered right here.
In the past many of you asked me whether there was an archive of previous newsletters. Now there is one, going back to May of 2008 and available for Premium subscribers only. To access the archive, click here and enter the user name toolkit and the password amelie.
And here is the second thought, this one on the faux Toolkit, the one from Google. Roland Grefer reminded me a few weeks ago of the relevance of Google stopping GOOG-411, its speech recognition voice service. Why did it shut down? It had collected enough speech samples. What does that remind you of? Me, too! Once "we" have donated enough data to feed Google's hunger for translation memory data, Google Translator Toolkit might be shut down as well.
Will this actually happen? No, because Google's hunger for that data will not be satisfied for a very long time. And yet. Does this make you feel a bit like a pawn that is valuable as long as he produces? It sure does give me a creepy feeling, and I've decided once again that Google Translator Toolkit not only stole my name but is not a viable alternative to any of the commercial or open-source translation environment tools out there.
|
|
1. Hybrid Hubris? (Premium Edition)
| |
As machine translation gains a more significant foothold in some of our work lives, I've been writing more about it. In past newsletters you've seen me write about the two different kinds of machine translation technology that are presently used: rules-based machine translation (RbMT) and statistical machine translation (SMT).
Just to review, RbMT essentially refers to machine translation systems that are based on information about source and target languages consisting of dictionaries and grammars. On the basis of that information the input data is analyzed, transformed into an internal representation, and then used to generate the target language.
Statistical machine translation (SMT), on the other hand, is based on data -- lots of bilingual data (translation memories or aligned data) -- from which statistical models are generated, which in turn are used to translate data.
In principle, these are systems that are not easily combined, a state of affairs that might partly explain the historic hostility between the two camps. Just recently there have been some attempts, however, to combine the technologies, and some of these are even available in the commercial realm.
PROMT (depending on where you come from, you might pronounce this as "Pro MT" or "prompt") is a Russian company that was founded 20 years ago, very much like its Western counterparts with funding from the military. It has specialized in offering an RbMT solution for various user groups and commercial settings (the supported languages are English, German, French, Italian, Spanish, Russian, and Portuguese).
After RbMT market-leader Systran announced an RbMT-SMT hybrid solution in 2009 (I'll write about this in an upcoming newsletter), it was no surprise to have PROMT announce something similar late last year.
Last week I had a chance to talk about the new engine with the CEO of PROMT Americas, Olga Beregovaya, and I'll tell you the most frustrating news right off the bat: It's not available for you and me, even if "you" are a relatively large language service provider. Well, that's actually not quite accurate. It is available for a price starting at $40,000.
So with that out of the way, let's talk some nuts and bolts.
True, the two different concepts of machine translation are difficult to combine. However, the PROMT developers asked themselves what would happen if they combined them in a way where they evaluated each other rather than "interfering" with each other's processes. And this is essentially what happens in the "DeepHybrid" solution that is being offered for some of the server installations.
Rather than just generating one translation solution for each sentence, between 4 and 12 candidates (for sentences of 10-15 words) are being generated. While these are generated on the basis of RbMT, it's the SMT engine that then helps to decide which of these candidates is the best suited for the translation of the particular project (while at the same time passing on terminology to the rules-based engine). It's a multi-step process that seems cleverly structured, drawing on the strength of both engines or principles of machine translation.
Olga estimates that you need a translation memory or corpus of at least three million words per language pair and domain to adequately train the statistical machine translation engine. PROMT itself has run large-scale tests with data from the TAUS Data Association and other sources with relatively good success rates. But again, the rules-based engine for the supported languages comes with the system, but for the statistical engine you will have to supply your own material.
So this brings us back to the question of why this is only available in the corporate server editions and not for LSPs or freelance translators.
One reason is clearly the price point (mentioned above), but since this is obviously something the technology vendor decides on, there must be other reasons why this is priced in a range that is virtually inaccessible for language providers.
Clearly the large amount of data necessary is generally not available for language providers, calling into question whether it's reasonable for language providers to use a solution like this.
But maybe more importantly, there is the tension-filled history that you may be sensing once again right now between the machine translation community and the language industry, a history that may prevent the language industry from jumping on board and investing in every new MT development out there. So PROMT's idea is to sell this to large clients right now, see how it is accepted, and at some point in the mid-term offer a SaaS (software as a service) solution for language providers. (Olga had a classic quote on this: "You can't ask for a Mercedes for the price of a Yugo, but we will offer you a Mercedes rental facility at some point.")
So, now the most important question: What does this all mean? Is this the magic wand that MT developers have been looking for for the last 60 years? Umm, no. The tests that PROMT ran where they compared the output of translations from their customized RbMT system and the hybrid system had a difference of 5 points on the BLEU scale (BLEU is the commonly used method to evaluate the quality of machine translation), still far away from most human output and still very much in need of post-editing.
|
| ADVERTISEMENT |
Get Organized in 2011
Did you have that in you New Year's resolution? Get out of routine and chaos with TO3000 and Projetex. Reach new heights in 2011!
Focus on translation not administration! Get AIT products with 25% discount.
Click here!
|
| 2. The Zen of Translation Environment Tools | | I have no idea why I came up with that title, but it was a long time ago and I will try to make it as Zen as I possibly can. And what exactly is it? It's an ATA webinar next Tuesday (9 am PST, 12 pm EST, and 5 pm CET). It's open to ATA members and non-members alike, and I will talk about translation environment tools (duh!), what they are, where they are, and where they will be. I think it will be fun and I hope to see you there (there are still some seats available). Find more information right here. |
| 3. Tricky Word | |
Here are two things that I thought would be helpful for those of you who use MS Word.
The first one is relevant only to Word 2010 (as well as most other Office 2010 programs) -- and by the way, after having used Office 2010 for a few months now, the one (and only) feature that has crystallized itself as a real differentiator is the new search feature in Word (with all hits listed in a separate pane).
Anyway, if you have installed Office 2010, you will have noticed that most files which have been downloaded from the Internet or came as an email attachment are opened in "protected mode" -- a read-only mode for which you manually have to activate all editing features. This is a feature to reduce the potential risk that could come from those files. So if you want to be on the super safe side, go ahead and skip the next paragraph; otherwise, read on to find out how to disable this feature.
In MS Word/PowerPoint/Excel, etc. (you'll have to do this individually for each of these programs), open the File menu and select Options. On the left sidebar select Trust Center and then Trust Center Settings. On the ensuing Trust Center dialog, select Protected View. Here you will find all the relevant options that allow you to enable and disable the protected view for files from various origins. Make your choices and close the dialog boxes. From now on you can edit your documents right after (carefully) opening them.
Since the procedure described above takes a rather complicated route to get to the Options dialog -- which, as in the old versions of Office, is still the "command center" of sorts -- it's really helpful to ease access to that by adding a link for the Options dialog to the Quick Access toolbar in Office 2010 (and 2007) programs (the little icon bar to the right of the Office icon in the upper left-hand corner). To do this, select the down-arrow on the right-hand side of the Quick Access bar and select More Commands. In that dialog (which actually is the Options dialog), select All Commands under Choose commands from and then add Options to the right-hand pane by clicking on the Add button. This will make the Options dialog exactly one click away.
And lastly, here is something that I never knew about the Find and Replace feature in Word.
I have always found it annoying that it was not possible to search and replace something but leave the original text untouched. Doesn't make any sense? Well, here is a good example: Imagine someone who grew up using a typewriter and still adds those dreaded two spaces after periods in an English document. Before you process this with a translation environment tool, you need to take all those spaces out. So you could just make a search for two spaces and replace that with one space. That's easy. But let's imagine that this is a long and convoluted document where some of the double spaces that don't follow a period actually need to remain. So you could do a manual search and decide for every instance whether this needs to be replaced or not. But did I mention it's a long document?
So here is a (relatively) easy way to conduct a find-and-replace action for every instance where a period is followed by two spaces and a capital letter (as in the beginning of a sentence).
Press Ctrl+H to open the Find and Replace dialog in Word, select the More button to open up the extended options, and select Use Wildcards. Enter
(.) ([A-Z])
in the Find what field. This expression stands for "one period, followed by two spaces, followed by any capital letter." There are parentheses around the period and the [A-Z] expression to make them referable in the Replace with field. Because if we enter
\1 \2
into the Replace with field (first referable field, followed by one white space, followed by the second referable field), the second spaces will be removed but the periods and first letters of the following sentences will not.
Still doesn't make sense?
Imagine this: You are working on a table where names are listed with the family name first, followed by a comma, followed by the given name:
Smith, Roland
Doe, Jane
Kulongowski, Vladimir
Now your client wants you to change that for the translated version, and you need to sort this into family name following the given name. Here is how you can do this really easily.
Copy the table into a standalone Word document, press Ctrl+H, select Use wildcards, and enter
(<*>), (<*>)
(< = beginning of a word, * = 0 or more characters, > = end of a word, followed by a comma and a white space, followed by another beginning of a word, 0 or more characters, end of a word). Replace it with
\2 \1
The result will be this
Roland Smith
Jane Doe
Vladimir Kulongowski
The comma was taken out and the family and given name entries were switched.
If you still don't see any usefulness in this, file this topic under "fancy search and replace tricks that will come in handy one day at which time I will express my heartfelt gratitude to Jost for this great tip" and don't trouble your poor mind anymore.
|
| 4. CompeTEnT Updates (Premium Edition) | |
A few translation environment tool vendors have recently released upgrades for their tools, and though these were all "minor" upgrades (or as the geek likes to call it: "point releases"), they did contain a lot of goodies.
Let's start with Fluency, the tool that grows and changes at the speed of light. Here are some of the changes from the past couple of weeks:
- Added support for a number of software development formats, including .resx, .rc, and .properties files
- Added support for InDesign .idml files (I've reported previously that this will be the preferred InDesign exchange format for virtually all translation environment tools)
- Added more than 20 new languages for detection in its internal and updated optical character recognition application (for image-based files)
- Added concordance search in translation memories by target language (this is also something we are going to see more and more of)
- Released a new version (Fluency Enterprise) with real-time translation memory and terminology sharing and administration tools
Next comes memoQ with its latest release this week:
- Customizable metadata in the translation memory
- Added support for the OpenOffice Writer .odt format
- Added target language concordance search in translation memories (didn't I tell you?)
- Ability to export project-specific translation memory
- Ability to preview formatted XML files via XSLT style sheets
And lastly there is Déjà Vu. Most Déjà Vu users have been waiting for a completely new version for a very long time (and don't be fooled by Atril's website -- it's not here yet), but at least there is an update to the old Déjà Vu X version now with some new features:
- Support of the InDesign .idml format
- Support for XLIFF files (both of these features have been around for a little bit but have both been improved)
- Added support for Adobe InCopy .icml files, and Visio .vdx files
Of course, all of these new updates have long lists of fixes and performance enhancements, but I'm not going to bore you with those.
|
| ADVERTISEMENT |
Tired of wading through the marketing hype of tool vendors, downloading gigantic trial versions, and spending hours and hours of "trying" without getting anywhere? TranslatorsTraining.com offers free, video-based introductions into all leading translation environment tools, and now even special discounts on some of the tools. Save 20% off memoQ translator pro and Déjà Vu X Professional or 35% off SDL Trados Studio Freelance!
|
| 5. Second Thoughts (Premium Edition) | |
Marinus Vesseur wrote this response to my last newsletter (it's kind of longish, but a worthwhile read):
About your oft-recurring praise of IntelliWebSearch: this tool and the way you described the integration of a TAUS search made me think of something Douglas Adams once said: I am rarely happier than when spending an entire day programming my computer to perform automatically a task that it would otherwise take me a good ten seconds to do by hand.
Every so often -- when a reference like the one in your newsletter comes along, for example -- I feel like a complete idiot for still not using this wonderful and oh-so-simple tool. So once more I install the latest version, try one group search, get 10 meaningless non-results, half of which are by Google. (...)
I have a suspicion that I am not alone in my frustration with the somewhat more geeky tools and programs for translators. I have observed a basic incompatibility between technology and linguistics. Don't get me wrong, I love some of those CAT tools (which apparently are now TEnT for the L10N industry? What does that make translators? TEnT-acles? L8s? C3POs?) and use various more-or-less intuitive programs daily. I keep meeting translators (...) who are so scared of those tools, they don't even want to start using them. Not for the money it might cost, but for the complications that are to be expected and the tediousness of learning something so foreign to one's mindset. This is not a complaint about the article. The whole thing just triggered my thoughts on an underlying matter and my mail is just an attempt to point attention to what is perhaps the most important factor in the acceptance of any technology: its psychology. If the iPod had been complicated to use, it would not have become the success it now is. As long as those TEnT makers keep demanding an in-depth understanding of computer technology to be able to use [their tools], and at the same time are apparently not "digging" the translators' mind, they will not be very popular with translators, is my thesis.
Maybe I'm wrong and it is just a matter of getting used to a certain way of thinking that anyone can easily learn, such as working with a word processor or storing data in files and folders, but I doubt it. There is something mathematical about those TEnTs that goes against my grain. Something along the lines of hard facts versus made-up stories, exact science versus experience, a straight canal versus a meandering river, measured data versus intuition, you get my meaning. The way they force language into definable rules and cut up all documents into re-usable segments.
Which may seem to have very little to do with incorporating TAUS searches into IntelliWebSearch. I could have made it short and said that it's a tool for nerds, but I tried to phrase what has been bothering me about the iron grip with which the technology department of linguistics seems to want to submit language and its related profession to its diction, is all. Hope I made some sense. And maybe the tool IS in fact rather simple and very useful, if I only knew how, and you could spend a few lines on IntelliWebSearch for Dummies next time.
Man, I can feel his pain. He is completely right. There are indeed many in the language community who stand in front of this gap between what the modern language workplace requires and what the traditional concept of a translator has to offer.
Really, I think there are two ways to approach this gap if you are among those who experience it (and neither is without pain). One is to acknowledge it and say, "I'm going to not use the kind of technology that goes against my grain, but will search out a niche within our industry for which I don't have to use it." And, hey, there are many such niches. I recently talked to a well-known translator who charges three times the amount per word I do and refuses to use translation environment tools. She is making a great living (better, in fact, than I).
The other way is to acknowledge the gap and slowly try to overcome it piece by piece and step by step. I think that Marinus' description of the different natures of the beast(s) is very helpful. Yes, there is a psychological difference between a more technical approach to translation and a more traditional approach. The good thing is that this difference it very bridgeable. For folks in their 20s and 30s it might seem like an easy straddle, while the non-technology-natives will have to stretch a bit more.
Here are a couple pieces of advice from someone who spent a decade of his life in dusty archives reading letters of missionaries who argued back and forth on the translation of ONE word. Really! One word.
There is no reason to do it all at once. Try sitting down and making a list of things that you would like to understand one by one. Maybe it's Internet searches that you need to master first. Or how to compile glossaries. Or how to use glossaries. Or how to use translation memory. Or Microsoft Word and Excel. Or . . .
Once you've jotted down your list, make a plan on how you can achieve all this. There are actually a lot of opportunities out there. Start with training programs for specific tools (such as this one on IntelliWebSearch, Marinus), or the ones offered by other tool vendors or places like TranslatorsTraining.com, or of course through your local translation association.
Also, the newsgroups for many of the available tools (most are on Yahoo! Groups) are extremely supportive, plus they offer archives to access previous discussions.
Lastly, I would be remiss not to mention my ebook, which I (and many of its readers) think is a very helpful resource.
It's definitely doable. And while the goal might seem far off and maybe not even that attractive, I can attest to the enjoyment that you can derive from using your computer in a productive and creative manner.
And while we are talking about TDA searches with IntelliWebSearch (this was the topic that set Marinus off): if you need a list of language codes and industries you can use in your IntelliWebSearch expression, you can find them right here.
|
| 6. A Love Story (continued) | |
A long time ago I started collecting remarkable characters to share with you. This is the continuation after a long hiatus, and at the same time the solution to the "riddle" of the last newsletter.
If you have a Javascript-enabled browser, hold your cursor over the character for a definition.
|
| The Last Word on the Tool Kit | |
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
Here is a reader who recently loaded the code:
www.wordchameleon.com.au
© 2011 International Writers' Group
|
|