1. The Big Love Fest (Premium Edition)
| |
Here are some snapshots from the two conferences in Denver that may illustrate the historical import I've been touting (and hopefully convince you that I'm not completely mental in pronouncing the significance of these conferences).
On the last day of the ATA conference, the conference organizers had made sure to line up a successive string of sessions all dedicated to machine translation and its meaning to the translation profession. Since the program at the ATA has four or five simultaneous tracks, there were plenty of other sessions for the non-MT-inclined audience. Interestingly, though, it was the MT sessions that drew by far the largest audiences. Mike Dillinger's session on how to utilize machine translation as a translator and Laurie Gerber's session on post-editing of machine translation had full rooms of 200 participants each, and the closing panel with representatives from both the machine translation and the translation communities drew an audience of 500 or so. The numbers alone are remarkable -- in previous years maybe a dozen translators would have shown up, and then only to make strong statements against machine translation -- but even more interesting was the underlying spirit of the questions that was intended to learn more about MT rather than to insult. I can't overstate what a change this has been in comparison to previous years. While I'm not part of the "machine translation lobby," in the past I have had to soothe the fears of many proponents of this technology before they made presentations to translators by assuring them that we no longer believe in public executions.
Fast forward to the AMTA conference. The opening speakers of that conference were a very eloquent Nicolas Hartman, president of the ATA, and myself, both of whom tried to represent the translators' view of machine translation. And the same happened there, only in reverse. The reception was warm and decidedly non-critical and revealed a clear consensus that machine translation could only be taken to the next level with the translator rather than without.
One of the things that both Nick and I tried to communicate to the machine translation community was that, at least in the branch of statistical machine translation, there is a great indebtedness to human translators. For too long machine translation developers looked at data and corpora as if they had developed on their own rather than being the fruits of human translators' labor. We stressed the vital importance of acknowledging this and making sure that from now on translators play a much more active role in participating in ongoing developments.
Of course, this naturally means that we are also willing to engage rather than sit back and complain. But it certainly seems that these two conferences have created the infrastructure that will enable this to happen.
|
| 2. The Macro King At It Again | |
Many of you are familiar with CodeZapper -- in fact, I just mentioned it in the last newsletter. CodeZapper is a set of macros you can run for Word-based files that display a lot of tags after the import into a translation environment tool.
Now the creator of this tool, David Turner, has been at it again. His new tool is called PhraseMiner, and just like its predecessor it's a set of macros that you can run on any Word-based file or any other text copied into a Word document. But rather than deleting unwanted tags, this set of macros finds subsegment repetitions (or, alternatively, fuzzy matches, and -- in some languages -- extracts terms) and displays them in a newly opened Word document. While some tools such as memoQ, Multitrans, and Lingotek identify subsegment matches in the translation memory or corpus by default, many other tools do not. For these it is helpful to have subsegments (i.e., segments within a sentence) that will be repeated in the project identified, translated in advance, and then fed into the translation memory or terminology database.
I find it particularly interesting that it did not seem too difficult for David to write this set of Word macros in his free time. Without belittling his efforts at all, it seems odd that other translation environment tool developers with their dedicated staff of developers have not been able to come up with something equally good or better.
If your tool of choice presently does not support subsegment searches, you are well advised to look at David's solution. Mind you, it's not a perfect tool. Large documents take a long time to process, and David himself assumes the accuracy of identifying all subsegments, fuzzy matches, and -- in a limited number of languages -- terms, to be at about 70%. You can download a sample file that contains the macro in the files section of the Déjà Vu or memoQ user groups, and if you want to actually use the tool, let David know and he will send you the macro template and ask you for a donation of $20 for his efforts on his website. |
| ADVERTISEMENT |
Buy Wordfast before year's end and save 170 Euro on two industry-leading TM tools:
between now and December 31, Wordfast Classic + Wordfast Pro = 330 Euro.
Beginning January 1: Wordfast Classic OR Wordfast Pro individually = 350 Euro;
Wordfast Classic + Wordfast Pro = 500.
Learn more at www.wordfast.com/pricing.html
|
3. Communication Gorges (Premium Edition)
| |
When I first wrote about Lionbridge's Translation Workspace translation environment tool offering (see the ad in this newsletter), I mentioned the apparent contradiction of a service provider offering technology to competing service providers. Lionbridge strongly protested that notion, saying that "there is a clear separation between the services side of the organization and product business. This is why the GeoWorkz business unit was created." While I never really doubted the validity of this statement, last week it became apparent for all that it was indeed completely true.
Imagine this: the largest translation conference of the year is under way. The good folks from Lionbridge's GeoWorkz division have a very visible presence there as they try to introduce and explain their product. And at that precise moment, Lionbridge's VP World Wide Vendor & Supply Chain Management Didier Hélin sends out the now infamous "Lionbridge Vendor Management - IMPORTANT - DO NOT REPLY" message informing LB's vendors that they will have to cut their rates by 5%. (In case you were in Antarctica or on the moon during the last few days and did not hear about this, here is a good summary.) Now imagine yourself in the shoes of those good folks from GeoWorkz I mentioned. You think they still got to talk much about their product offering?
Should you have had any doubts whether there truly is a split between the service and technology divisions of Lionbridge, rest assured: it's more than a split. In fact, it's a very wide gap, at least as far as communication goes.
(And just to complete the whole picture, shortly before this newsletter was written, Lionbridge released grim-looking financial statements that some will say puts the forced price reduction into perspective. Others, myself included, would still argue that this was a communication gaffe par excellence.)
|
| 4. Machine Translation in Poetry | |
A few weeks ago I mentioned that Google had announced a new endeavor in the field of machine translationthat teaches the existing machine translation engines meter and rhyme rules to aid with the translation of poetry, or even to transfer prose into poetry. I thought this was a fascinating experiment, so I asked the team at Google to translate a piece of literature that is one of the most central passages for machine translation developers. I'm talking about the passage in the Hitchhiker's Guide to the Galaxy that contains the introduction to the Babel fish, the universal translator.
The original says this:
The Babel fish is small, yellow and leechlike, and probably the oddest thing in the Universe. (...) If you stick a Babel fish in your ear you can instantly understand anything said to you in any form of language.
I sent the "official" German translation of this passage to Dmitriy at Google (the engine is not open to the public yet) and asked him to pass it through his engine. Here is what he came out with:
The Babel fish is small and yellow, leech- like and perhaps the owner, all there is in universe. With Babel fish in his ears means at once all are told some in speech.
I used this passage during my talk at the AMTA conference (well knowing that it's sort of a no-no to present awkward-sounding machine translation passages to the machine translation community), but I've got to tell you that at this point I've read the Google engine rendering so often that I've almost started to like it.
What I found by re-reading this passage in the Hitchhiker's Guide, however, is either a lack of attention to detail or a mind-boggling failure to communicate among those who wish to claim the diminutive fish as their mascot. Clearly the Babel fish has been central in the imagination of machine translation developers -- after all, there's a reason why the first widely used free machine translation engine on the Internet was called Babelfish (you can find more on the history of machine translation in this article). And yet if you finish reading the same paragraph in the Hitchhiker's Guide that introduces the Babel fish, you will find the following statement:
Meanwhile, the poor Babel fish, by effectively removing all barriers to communication between different races and cultures, has caused more and bloodier wars than anything else in the history of creation
So much on that note.
Well, there is one more communication blunder in regard to machine translation that I find worthy of mention. SDL's Mark Tapling described the new BeGlobal MT service by SDL in an otherwise slickly produced video like this: "It's really quite a magical technology when you see it work." While I realize that he's trying to describe things with a marketing spiel, these are exactly the kind of communications that make it so hard to communicate with the translation community and completely misrepresent machine translation to the general public. One of the things that was touched on again and again during the ATA and AMTA conferences was the need to talk real. Can the appropriate use of machine translation be powerful? Absolutely. Is it magic? Absolutely not. |
| ADVERTISEMENT |
Special Offer from GeoWorkz - A division of Lionbridge
Thousands of Projects are being completed on Translation Workspace.
Signup for a 30-Day Free Trial!
You can use Translation Workspace, the leading on-demand translation productivity solution, for all your projects. Also, with your subscription you can add your profile to the GeoWorkz Directory - where you can promote your business to other subscribers and where other service providers can find you.
Get Started with a 30-Day Free Trial Translation Workspace Today!
|
| 5. High Time for Mac Users? | |
Office 2011 for the Mac has recently been released, and while I typically focus on tools for the Windows environment, this release seems important enough to mention because of the renewed support of Visual Basic. Not relevant to you? Let me guess, you're not using Wordfast Classic on your Macintosh computer. (And you might not even use a Macintosh computer.)
Before the advent of web-based and Java-based translation environment tools (XTM, Wordbee, Swordfish, or Heartsome), and long before a reasonable use of using a virtual Windows environment in the Mac, only a very small handful of translation environment tools were available for Macintosh users, with Wordfast Classic the most powerful among them. Not surprisingly, many Mac friends were dismayed to find that Word 2008 for Mac did not support the programming and macro language that Wordfast Classic uses to do virtually anything.
It would be nice to think that translators had a strong enough lobby to convince Microsoft to reinstate support for Visual Basic just for them, but I'm afraid that we were not the only ones who suffered from the withdrawal of that feature. So Microsoft reinstated it for the new version. Depending on which review of the new version of Office 2011 you read, you'll either find out that it's by far the best Office system that Microsoft has ever released or that it's a half-baked product that was released way too early. Be that as it may, Wordfast Classic Macintosh users will be delighted.
|
6. In Other News . . .
| |
Here are a couple of other things that ran across my way in the last couple of weeks:
The widely beloved terminology management/quality assurance/reference tool Xbench now offers spelling-checking for 16 languages through HunSpell plug-ins. Enough said. (But I would like you to realize that I did not use the much more magical "spell-checking.")
Two tools that I've run across in the last few days are ResourceBlender and Moses for Mere Mortals. I have not had a chance to try then out (and would love to hear from folks who have).
Here is what it says on ResourceBlender's website:
ResourceBlender is an open-source translation and internationalization application that offers an easy way to manage localized resources for inclusion with different applications. Available as an ASP.NET web application and a WPF desktop application, it makes localizing applications a breeze.
It's able to process .resx, .properties, .rdf, .po, and .php files and offers some kind of (rudimentary) translation/editing interface with some translation memory (and MT) functionality. As I said, I'd be interested to hear and report on your experience with it if you've had a chance to look at it.
Achim Ruopp brought to my attention Moses for Mere Mortals, a specially packaged and presumably much, much easier-to-use version of the open-source statistical MT Moses monster. So in theory, it's a free MT engine that is easy enough for the freelance translator or smallish LSP to use and feed with their own TM data and come out with some kind of results for specialized MT purposes. Again, I have no experience with it and would love some feedback for that as well.
And lastly, lots of folks responded to the articles on PDF processing and suggested their own favorite tools. Here is a sampling (and don't take my word for it -- these are recommendations by fellow newsletter subscribers):
- Eñaut Urrestarazu Aizpurua recommends the PDF converter, "which is for the moment the best one": Able2Doc.
- Melanie Farrugia recommends Free OCR: "I've been using it for about a year now and most of the times it does the trick. It all depends, however, on the resolution of the scanned document. The lower the resolution, the more "scrambled" the result will be, but for a free program it works quite well. You can even install additional languages -- by default it is English."
- Marinus Vesseur recommends NitroPDF, which gives him "better results than Omnipage."
- Radek Pletka likes "the free PDF XChange Viewer, which allows commenting and editing, copying text, stamping (even with your own signature), and which is compatible with latest PDF files."
- Renator Reinau commented on proofing PDFs, for which I had recommended Acrobat. His company uses Nuance's PDF Converter, the same tool that I mentioned for OCR and conversion purposes.
And lastly, Simon Sobrero came up with an idea that might have a lot of potential. He said:
. . . please would you put out a shout if anyone knows a service to convert non-editable PDFs into editable text (manually)? Surely any monolinguist could do this for a fraction of the price of a translator and make a good living out of it? I can't seem to find anyone in the yellow pages, and I am fed up doing it myself!
Anyone?
|
| The Last Word on the Tool Kit | |
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
© 2010 International Writers' Group
|
|