|
|
|
| Mi Chiamo Jost! | |
Those among us who live in the U.S. have been
bombarded lately with a rather silly commercial by the language learning
software Rosetta Stone. It shows a young, very farmer-like-looking
farmer with the following text right beside it: 
Anyone with any kind of language background
must have asked themselves the same question: What exactly does this software
teach him to achieve this? Fortunately, a tongue-in-cheek reporter from the high-brow
New Yorker magazine has found out.
Here are a couple of sentences from Lesson 1 (Beginning conversation):
During
Milan's Fall Fashion Week, in which hotel will you be staying? Durante
la Settimana della Moda dell'Autunno di Milano, in quale hotel starà?
Could
you please give me directions how to get to that hotel from western Illinois? Per
favore, mi potrebbe indicare la strada per quell'hotel partendo dall'Illinois
occidentale?
No doubt that would impress her. But the
article shows other useful ways to use his new Italian skills as well. My
favorites are from "Lesson
7 (At the Fuel Co-Op)":
Good afternoon,
Owney. I would like to buy two tanks of propane.
Buon pomeriggio,
Owney. Vorrei comprare due serbatoi di propano, per favore.
I said, "I would
like to buy some propane!" Ho detto, "Vorrei
comprare del propano!"
Of course you can't
understand me. That is because I am talking in Italian. Certo lei non può
capirmi. Sarà perché sto parlando in italiano.
Laugh if you wish,
Owney, but someday I will be having sex with a beautiful Italian supermodel in
Milan, Italy, while you are still here sweeping fertilizer pellets off the
floor. Rida se vuole,
Owney, ma un giorno farò sesso con una top model italiana a Milano, Italia,
mentre lei sarà qui a spazzare via le palline di fertilizzante.
Well, goodbye,
Owney. I will buy my propane another day. Arrivederci,
Owney. Comprerò il propano un altro giorno.
Speaking of impressive encounters, Jeromobot and I (and, yes, the rest
of my family as well) went to the California redwoods the other day. Prepare yourself to be impressed. |
|
1. Swordfish, AlignFactory, and openTMS are all in the Worx (Premium Content)
|
|
Yeah, kind of a lame little word play here . . . But there is news on so
many fronts in the translation tool market that I decided to lump them all into
one larger article.
Let's start with the last first. The project management application Worx
is -- or rather has been -- the tool that the British LSP Language
Technology Centre (LTC) has been marketing for about a year or so now. Worx
is the successor to the LTC Organiser -- a tool that looks a little
dumpy and old-fashioned now but that really was the trailblazer (Go Blazers!) in the PM tool market many
years ago.
Yesterday, LTCannounced that it had split its technology
division -- the one that supports and develops Worx and other tools --
into a separate company it now calls Agile.
Tobias Rinsche, now VP of Agile, had been trying to talk to me
about this for a couple of weeks, but to be honest, I thought it just was not relevant
enough to spend much time with it -- until I finally got to talk to him
yesterday. The obvious thing that LTC wanted to achieve with this was a
disentanglement of tools and services -- a status quo entanglement that much of
our industry clearly suffers from. So far, so good. But what really excites
Tobias -- and his excitement is contagious and well-founded -- is that Agile
has a number of products in the pipeline that they have developed with and for
LTC, products that are just waiting to be released. What's more, this whole new
company (with 10 full-time developers) focused solely on technology and tool
development will be good for their present and future customers.
Relative TEnT newcomer Swordfish has
been quietly chipping away with the introduction of new features in the last
few months. The more interesting ones are probably these:
- Support for
TXML (the Wordfast Pro XML format)
- A plugin that
allows for translation using Google's machine translation engine
- Support for .ts
(Qt Linguist)
files
It's interesting and remarkable that Wordfast Pro, a new format
of a tool that is hardly out of its beta phase, is already supported, but that
just underlines the strong urge that exists to interchange TEnT-specific formats.
The introduction of an actual machine translation component -- especially through
the free Google Translate -- is becoming more and more common
among tools. I still don't see a real-life use for it in my work -- not because
it's machine translation in general, but because I don't think the data on Google
Translate is specific enough to be of true value for a professional
translator. But someone must be asking for it, or this wouldn't be introduced
in so many tools.
I have written extensively about AlignFactory
before, praising it as the best desktop-based alignment tool on the market. It
achieves this with a sophisticated alignment engine developed by the University
of Montréal and by using a number of filters that filter out any unlikely
match.
Here are a few other highlights: It allows you to select thousands of
file pairs at a time, allows you to process PDF files (!) for alignment, and does
it all with breath-taking, neck-breaking, and gob-smacking speed. No, it's not
perfect -- it still comes out with missed or wrongly executed alignments here
and there -- but compared to the internal alignment tools of existing TEnTs,
it's a no-brainer.
Last week they released a new version (2.0). It's much easier and more
intuitive to change project settings, segmentation can now also be done on the
basis of paragraphs -- making it easier for some bitext tools to work with it
-- and you can now choose to retain some simple formatting in the TMX files
(previously all formatting was stripped).
And lastly, the first full version of openTMS,
the open source TEnT aimed at agencies and developed by FOLT, the German
industry association, and Klemens Waldhör, who also was part of the team that
originally developed Heartsome (and thus Swordfish), will
be available on April 20. Naturally I have not had a chance to look at it, but
it's very good news that things are happening so fast now after a long planning
phase that had some people already writing this project off. It looks like it's
happening after all. More on this after I have had a chance to look at it.
|
| ADVERTISEMENT |
Why pay a high price for a CAT tool and keep fighting with
it every day?
Make your life simpler and easier by making the right
choice! Heartsome Translation Studio is what you need and deserve:
- based on open standards
- cross-platform
- supports all major database types
- the most customizable CAT tool in the market
- gives you the choice to deliver in XLIFF and TMX or tagged RTF and TXT
TM formats
And as Jost puts it in the Tool Kit: "There are many
obvious benefits to Heartsome -- platform-independence, a fairly small
footprint on your system, a good choice of underlying database servers . . . and
a very pleasant 'feel.' . . . It is certainly an interesting and powerful
tool."
Enjoy the spirit of freedom.
Make today your independence day! Buy now at www.heartsome.net. |
2. Interpreting URLs
| |
This might bore a lot of people who are more technical than the rest of
us, but here it is anyway:
I have found it very helpful to learn to interpret URLs (web addresses)
from a language/translation point of view. There are numerous parts of a URL
that could identify the language of the webpage that it displays. And the cool
thing is that if that is the case, chances are that the same webpage is also
displayed in other languages (otherwise there is not much reason to note the
language in the first place).
For instance, let's look at this long URL from the Microsoft help site:
http://windowshelp.microsoft.com/Windows/en-US/help/fe7ea80e-52a2-48d6-947a-05e02e78bc371033.mspx
Not really interesting you might say, but if you need to translate
something for which you need that particular terminology, things might be
different. Now, that URL has two language identifiers. One is very obvious --
en-US (in this case a mixture of the standards ISO 639-1 and ISO 3166). The other identifier
may not be as obvious: the last four digits at the end are the widely used Microsoft Locale ID.
To change that page into, say, Japanese, you could just manually replace
the URL in those two places with the appropriate codes:
http://windowshelp.microsoft.com/Windows/ja-JP/help/fe7ea80e-52a2-48d6-947a-05e02e78bc371041.mspx
and come out with the Japanese counterpart with all the Japanese
terminology at your fingertips.
Or here is another one that I have been using a lot lately. There is a
lovely English and German parallel SAP glossary at http://help.sap.com/saphelp_glossary/en/index.htm.
To change the language in that case all you need to do is replace the en with de.
But since this is a glossary within an HTML frame, it's not quite as
easy to get to specific entries. If you click on any of the actual English
entries in the above page, the URL does not seem to change. However, Firefox
has a sweet and easy way to let you open the page within the frame as a
standalone page. Once you click on an entry and you have the English term and
description displayed, right-click on that page and select This Frame>
Open Frame in New Tab. This might open
http://help.sap.com/saphelp_glossary/en/3b/57a67b78608045852d629395c6844b/content.htm
And sure enough, just by changing it to
http://help.sap.com/saphelp_glossary/de/3b/57a67b78608045852d629395c6844b/content.htm
we get to the translated page.
Now I realize that this particular example is only good for the small
handful of you who work in that language combination, but there are many other
cases where this can be adjusted easily to other websites and language
combinations.
Oh, and I felt really silly last week when I
realized for the first time that Firefox offers a feature where you can
simply highlight a word or phrase, right-click on it, and search for it in your
default search engine. Sweet.
|
| ADVERTISEMENT |
Take the fast lane
with AlignFactory
Align
with speed and accuracy and increase the performance of your existing tools.
Automate your document alignment process. Correct segments with the built-in
alignment editor. Produce alignments in TMX format, with or without TM
attributes, in LogiTerm bitexts format or in HTML without language
tags.
sales@terminotix.com
-- www.terminotix.com |
3.
Developers Are
Lonely, Part II
| |
Okay, here is the promised article on the second TEnT that was developed
with a developer's mind -- and I am sure glad I held off on sending it out last
week because it turned out that I really had seen only a small part of it.
So, let's start at the beginning. I didn't have to think much to come up
with a title for this article, because the developer of MT2007 -- Andrey Uzbekov, aka Andrew
Manson -- must be a very lonely man. Otherwise, I wouldn't understand why he
put scantily clad women on the welcome screens of his software. (And if you
think the woman in the main application is scandalous, check out the one in the
alignment tool!) Seriously, I dithered back and forth on whether to write about
this tool since it seemed just a little too much on the side of bad taste.
Nonetheless, here it is.
It appears that the name
-- MT2007 -- is an abbreviation of a reverse translation of Память переводов
-- Memory Translation -- rather than Machine Translation as most of us would
suspect. It is a Russian tool, and it is shockingly robust and feature-rich for
being donation-ware. Aside from its rather appalling welcome screens, it's a
very attractive program that provides a clear and easy-to-use interface . . .
that is, if you read Russian or if you manage to switch to the (semi-)English
interface (you'll need to run the LanguageSelector.exe in the bin folder for that). It comes with a very
large amount of Russian-to-English glossary/translation memory data from the XDXF project (which makes it a very slow
- and, in my case, often aborted -- 160-MB download, but I'm sure it's a very
welcome resource for Russian translators). These glossaries, or I should
probably say "dictionaries," are open-source material -- if you are
not familiar with that project, have a look at the XDXF page for your language
combination and you might just like what you find -- as are other components of
this program, including the English WordNet
dictionaries (also very much worth a look!) and most of the programming
components. After you have downloaded the tool it does not require any installation;
instead, just unzip it.
(In case you are running
a language version of an operating system that does not use commas as the decimal
separator, you might not be able to actually use the program right away, since
the font in the Translation Editor where the actual translation takes
part will be obscenely large and unusable. This is because the font size is
coded as a comma-separated value. To fix that, open MT2007.exe.config in the bin
folder with a text editor and change the font size in all lines that contain "editor_font_size." And if you think that's frustrating,
read the first two sentences of the manual: "As for now, MT2007 can
be regarded as a perennial beta version. Which means that errors are
guaranteed.")
While the interface of the tool is really simple, it might take some
time to get used to it because the concepts (and terminology) are a bit
different than what you might be used to. For instance, there is no real
distinction between translation memories and terminology databases in this tool;
instead, there is a feature that allows you to highlight subsegments, translate
them, and send them to the "repository." The repository seems to be a
collective term for linguistic materials including TM data -- which, just as in
the upcoming version of Trados, uses SQLite as its underlying
database system. Another core concept -- and here the developer's mind becomes
only too apparent -- is the use of regular expressions. This is helpful if you
want to build your own fuzzy search algorithms for languages with a high degree
of flexion. That particular feature wouldn't be my cup of tea, but for those
who are good with regex,
more power to you.
The way you import TMX (translation memory exchange) data is exactly the
same way you would import any other file you want to translate -- only with TMX
you will send it to the repository after the import. Once the data is in there,
automatic searches are relatively fast and there is a host of keyboard
shortcuts (that are easily viewable by way of the little help icon in the Translation
Editor) to switch between windows and enter data from the matches that are
found. Like Déjà Vu, Heartsome, and Swordfish, MT2007
uses example-based machine translation (although I am still not
convinced that this is the right term for what it actually does) by combining
matches and "correcting" fuzzy matches through deleting certain parts
of target segments and replacing others -- if the necessary material is available
in the repositories or the other associated language resources.
Like many other tools at this point, this tool also just integrates true
machine translation. PROMT, the Russian machine translation engine, is
supported if it is installed along with -- surprise, surprise -- Google
Translate.
The translation file formats it supports are MS
Office 2007, OpenOffice, and MS Word 2003 saved as XML. Once the files are imported they are displayed
in a tabular easy-to-work-with format.
To me it seems that this is a tool that would definitely warrant a
second look by English<>Russian translators with a knack for regular
expressions (and semi-naked girls), but it certainly has the potential to be
more than just that. The author of the documentation for this tool, Konstantin
Lakshin, has also written an interesting review of it in the publication of the Slavic division of the ATA that you
should consult if you would like to dig deeper.
Overall I was impressed with the tool. I liked the concept -- just take
a bunch of freely available components and make them into something new and
useable. And while the tool might be primarily for one language combination,
that does not have to be a bad thing. Too often we narrow our opportunities by
trying to do something that fits everyone. And I like the very laid-back way of
"marketing" the tool. Like its self-depreciating beginning, the
manual ends by asking:
Before you move onto studying other features in MT2007,
you should answer one question: "Do you personally find this approach to
automation of translation interesting?"
|
The Last Word on the Tool Kit
|
|
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed. © 2009 International Writers' Group | |