|
|
|
| Visionary Times | |
Sometimes when I'm writing this newsletter, I have to sit down and ponder
what to write. Usually it's not hard to come up with something that I feel is
relevant, but it might take a little bit of thinking. And then there are these
other times, like -- you guessed it -- NOW, when things are just happening left
and right.
Think about it: Within a span of a week and a half or so, SDL releases
its new tool (and creates a lot of confusion out there in the process); the TAUS
Data Association (TDA) releases its long-awaited translation memory
sharing and term search facility (just after Linguee, the amazing new
English<>German bitext portal was released -- I discussed that last week);
and, oh yes, Google also (very silently, I might add) released its Google
Translator Toolkit. And nothing against the importance of any of those
other news items, but that final one probably tops them all. (And, yes, I am
thinking about suing Google for the rights to the name! J)
On another note, the lovely Courtney Searls-Ridge sent me a memo last
week after I mentioned the good life Alexander Pope enjoyed after his
translation of Homer. Apparently he's also remembered for quoting his printer
on the subject of translators (to the Earl of Burlington in November 1716):
Those are the saddest
pack of rogues in the world. In a hungry fit, they'll swear they understand all
the languages in the universe. I have known one of them to take down a Greek book
and cry, 'Ay, this is Hebrew, I must read it from the latter end.'
By God, I can never
be sure in these fellows, for I neither understand Greek, Latin, French, nor Italian
myself. But this is my way: I agree with them for ten shillings per sheet, with
a proviso, that I will have their doing corrected by whom I please; so by one
or other they are led at least to the true sense of an author, my judgment
giving the negative to all my translators.
Since Jeromobot's vision is a little fuzzy lately, he could not agree more!
|
|
1. The Google
Translation Center That Was to Be (Premium Edition)
| |
Remember about a year ago when news reached the public about a "Google
Translation Center"? There was a lot of resulting hoopla from
translators, tool vendors, and language service providers, and the responsible department
at Google took a lot of heat. Back then I managed to talk to one of the
folks in charge of the program, and it was one of the coolest conversations I
ever had. In about ten minutes I asked all kinds of questions, and I always got
the same answer: "Yeah, I can't really talk about that." It felt a
little bit like talking to someone from the CIA: "Yeah, I could tell you,
but then I'd have to kill you." (Though the Google guy was actually
quite nice about it.)
Well, a lot of things have happened in the meantime, not least the
radical change in the economy, and this may be one reason why Google's new Google Translator Toolkit (I'm
going to call it GTT to avoid confusing it with an excellent newsletter I
know of!) states this:
Google Translator Toolkit
is free, but in the future, we plan to charge users whose translations exceed
high-volume thresholds.
This joins many other things that are quite different from the original
scenario.
But first things first: Here is what GTT is. It presents you with
a rather well-designed front-end that allows you to do one of two things: You can
either upload a file (in HTML, Word .doc [not .docx!], OpenOffice .odt, .txt,
or RTF) or you can specify a URL and the corresponding HTML page will be
uploaded.
When you select the file you need to select the language pair (the only
available source language right now is English (!), but there is a choice of 48
available target languages, including double-byte and bi-directional languages).
Then you need to choose whether this is a "shared" translation -- i.e.,
whether you are using and contributing to a large anonymous translation memory
or whether you would like to upload your own translation memory in TMX format.
Here you can also upload or define a glossary. Should you choose to upload a
glossary, it needs to be in CSV format with a strictly defined pattern, but it can have fields for part-of-speech or definition
aside from source and target.
When you have made those choices the upload happens and your original
file is displayed on the left pane of a split window (unless you choose horizontal
panels in the View menu). On the right side you can see a pretranslated
version of the file. The pretranslated material comes first from the
translation memory(ies). If nothing is found there -- and this is likely since
at this point they don't seem to be particularly content-laden -- a machine
translation with Google Translate is performed. (If you choose to forego
the machine translation step, you can select Pre-fill with source text
instead of machine translation under Settings.)
Before you start the
actual translation, you should click the Show toolkit button on the top
of the window. This will open a new pane at the bottom of the window that contains
four different tabs: Translation Search Results (TM hits), Computer
Translation (results from Google Translate), Glossary (hits
from your glossary -- the number in parentheses indicates whether there are any
and, if so, how many hits are being found), and Dictionary (this
opens a search box in which you can enter single terms for lookup -- no idea
where that content comes from). According to which translation segment you
highlight in the Target pane, the different tabs will show the
corresponding content -- if any is found.
The file is more or less displayed in WYSIWYG format,
meaning that it is displayed the way you would see it in a browser or MS Word.
Only when you select an individual translation segment (as in normal TEnTs, this
typically corresponds to sentences) is the WYSIWYG view of that segment replaced
with a text-only view in which inline codes are displayed as numeric codes with
curly brackets ({1}) -- déjà vu, déjà vu! Unfortunately, it is not possible to
enter the numeric codes with keyboard shortcuts. You will either have to use
the ones that were automatically entered during the pretranslation and
translate around them, or you can highlight the phrase that is surrounded by
codes, select Insert HTML tags,select the ones you need, and
they will be placed around the phrase (I'm not sure why it says "HTML
tags" independent of the source format).
Since not all text is displayed in the normal WYSIWYG
view of a document (think of keywords in HTML files or footnotes in Word
files), there is a list of those at the bottom of the target pane under Hidden
Text where they can be translated separately.
And while you are working on your translation (or
before, or after), you can also invite others to participate in your
translation/editing efforts by selecting Share> Invite people.
All they need is some kind of Google account.
When you are finished with the work on your file, you
can download the translated file in the original format (and formatting). And
be happy ever after. Or not?
Let's look at the tool in more detail aside from its functional
aspects.
The first thing I noticed when I started to look at it
-- with the original Google Translation Center still burned in my memory
-- is that it is much less project management- and process-oriented than what was
originally planned. This new tool is aimed at the translator rather than the
translation buyer. (Remember that glimpse of Google Translation Center
that we all saw, where translation buyers could upload a document and then
choose translators to work in it? None of that anymore.)
The official Google blog says this:
At Google, we
consider translation a key part of making information universally accessible to
everyone around the world. While we think Google Translate, our
automatic translation system, is pretty neat, sometimes machine translation
could use a human touch. Yesterday, we launched Google Translator Toolkit,
a powerful but easy-to-use editor that enables translators to bring that human
touch to machine translation.
The only argument that
most of us would have with the statement that "sometimes machine
translation could use a human touch" would be the "sometimes" (and
that it's often more than just the "human touch" that's
required). But the statement still tells us quite a bit about the immediate intent:
to add the human expertise to Google's machine translation efforts.
Remember, Google Translate, unlike engines like FreeTranslation or
Babel Fish, uses a statistical machine translation engine that relies on
good bilingual data -- lots of good bilingual data.
And I don't think that
there is anything wrong with it if it's just that I am adding data from the
translation that I am currently working on. What I did not like was this: As mentioned
above, it is possible to upload TMX translation memories, and as you upload you
can select whether you want to use it yourself or with others -- and that only
seems fair. What I did not read before I uploaded a fairly large TM for testing
purposes was this:
By submitting your
content through the Service, you grant Google the permission to use your
content permanently to promote, improve or offer the Services. If Google
publicly displays any of the content you submitted through the Service, Google
will display only portion(s) and not the entirety of the content at one time.
This means that even though I can now go ahead and delete my TM (so that
I and possibly other users won't have access anymore), Google will
continue to use it. That strikes me as odd -- to say the least.
Also, as you know, TMX stands for translation memory exchange,
but it is not possible to get your TM (or your glossary) out of GTT once
it's in there. The only thing you can get out is your translated file, though you
can, of course, continue to use the TM and glossary content within GTT
-- but only there.
To be fair, this was a problem in the first incarnation of Lingotek as well and
they fixed it rather quickly, but we're talking about Google here, and I'm
sure they don't do things without thinking them through. (Though other large
companies do, and must also humbly concede defeat -- see below.)
Aside from the thorny intellectual properties issue, there are a number
of other things that make me think that GTT in its current state is not
for the professional translator (unlike what we saw -- or imagined seeing -- in
the Google Translation Center):
- The only source
language (and UI) is English (though I imagine that this will change in no
time).
- There is no
possibility to alter the segmentation.
- There is no
possibility for concordance searches.
- There is no way
to manage fuzziness and it's not very clear how fuzziness is decided on.
- There are no QA
tools and the only spell-checking is that of your browser.
- The number of
file formats is much too limited (no XML, Office 2007, InDesign, etc.).
- You can only
work on single files at a time.
- There are no
project management facilities.
- There is no
link to professional translation formats aside from TMX (such as TBX, TTX,
or bilingual Word docs).
Where I think GTT will have a good response is on the
semi-professional or non-professional translation market. The Wikipedia/Knol
feature that allows you to take an English page, translate it, and publish it
right to the server will certainly play a role in this (though I admit that I
could not make this feature work!).
In its present form, though, this will not be a
threat to small translation agencies or sites like ProZ or TranslatorsCafe
as was feared with the first version. However, it is a "beta version,"
and my feeling is that if you were to ask Google where this is headed, they
could tell you -- but they'd (regretfully and politely) have to kill you.
|
2. The Truth Is Out
There, Part II
| |
In the last issue we talked about Linguee,
the German<>English corpus engine that I described
as a "quantum step in search
technology for translators."
To briefly recap some of the tool's cornerstones, it's a very large
corpus of web-based translated materials from live online sources matched up
with the help of a custom, user-generated dictionary and other web-based
dictionaries. Once the data is found it is displayed in large chunks and you
are offered links to the originating sites, giving you all the context and
information about the data that you need. Though it covers only German and
English at this point, other language versions are in the works and I would not
be surprised to see something new in the next few weeks.
I also mentioned that the timing of Linguee's release was interesting
as well, as it appears concurrently with the TAUS Data Association's
release of the first version of its language corpus search engine. And this is
what this article is about.
I have been aware of TDA's efforts for a while (and at their
initial meeting suggested a model very similar to the one they are now using,
if I may say so). While its benefits to translators might not seem too colossal
at first glance, it's a very important undertaking that is only at the
beginning of what it will eventually mean to our industry (and how it will
compete with offerings like the one from Linguee or GTT).
The TAUS Data Association or TDA is, like its name says,
an association of mostly large corporate translation buyers who originally came
together to pool their translation memory data to - yes, you guessed it --
better train their machine translation engines. As mentioned above, to make statistical
machine translation work, you need a lot of high-quality data, and even
industry giants like Oracle, Microsoft, and Adobe do not
have enough data on their own to get the results they hoped to achieve.
So, rather than semi-secretively pooling their data, they decided to
open the data up to the public -- not as TM data, mind you, but as a terminology
resource. If you want to get to the data as TM data, you can become a TDA
member, contribute your own data, and download some other data for your own
use.
But since this is financially out of reach for many of us, we can at
least use it as a terminology resource.
After you register, you can search for single words or expressions in 80
different language pairs (they presently all originate from UK/US English,
which are frustratingly treated as different languages that require separate searches)
with a total -- as of today -- of 549,411,144 words. So far the data comes mainly from two
major sources: material from the EU TMs (covering pretty much all the British English into other European
languages within the "Legal Services" domain) and mostly from the
domains of computer software, computer hardware, and professional services in
US English into other languages. When you make a search it's relatively easy to
see who has already contributed data (product names are not distorted), but it
is not possible to know where an individual entry comes from (unless that very
entry happens to have specific product names). To me this is the greatest
weakness of this otherwise impressive offering, and an area where a tool like Linguee
excels. For data to be truly valuable, we need to see where it comes from.
Chances are we know that Term A could be translated with Term B, but we as
translators need to know whether the specific client we work for uses Term B,
C, or D. And very general domains like computer software or computer hardware
simply are not specific enough to give adequate guidance.
Also, a tool like Linguee (and, no, I'm not getting paid to
mention them so often -- I just think they have a very clever way of assembling
and presenting data) uses the powerful combination of TMs, glossaries, and
dictionaries to highlight source and target terms in the retrieved segments,
something that TDA cannot do at this point.
So, overall it's a useful tool, but not a tool that will give you the
one answer you're looking for; instead, it's more likely something that can
give you a variety of terms you can then use to further narrow your search.
Also, it's important to remember that this tool
was not developed primarily as a terminology research tool, but as a TM sharing
tool. And I am eager to see how successful the implementation of that will be. |
3. Some Dorky Licensing . . .
| |
. . . that's what many translators thought when they realized what the
licensing of the newly released Trados Studio 2009 version for upgraders entailed. While it did actually contain
two licenses -- one for SDL Trados Studio 2009 and one for the SDL
Trados 2007 Suite -- the 2007 license was time-limited and would
cease to function after a period of 12 months. This is how SDL explained that move:
The reason for the time
limited license is because we will be replacing this version of SDL Trados
Studio 2009 within twelve months with one that will not require any of the
legacy functions we use from SDL Trados 2007 Suite and to ensure that all
upgrading users can continue to use their legacy product until they have had
time to get used to all the features of the new one. At this point the upgrade
will leave you only with SDL Trados Studio 2009. If we believe that an
extension is required to these twelve months, we will make this extension
available to all users.
The truth is that all this is really not that unreasonable -- it's relatively
common to have the old license not officially be legal anymore after an upgrade
has been purchased, but users typically don't care: Who wants to use the old
software anyway? And even if you did, you just installed it and it ran just
fine with the old licensing code. The problem here was that SDL's licensing
codes are monitored online so they would not run fine after the license expired.
In addition, and this is more important, the new Trados Studio really is
more than just an update -- it's a whole new piece of software that may share
some background coding but not much else that is visible to the user. And some
features (alignment, TTX previews) are available only with the old version
present.
So users panicked, lots and lots of emails were sent to forums and SDL
Trados support, and within just a few hours the policy was dropped and replaced with this:
As you know, as part of
the upgrade process, we asked you to return your old license first. After the
licenses were returned you should have received your SDL Trados Studio 2009
license and a time-limited license of SDL Trados 2007 Suite. This policy
has now been changed. Once you have returned your licenses we will now add to
your account a permanent and unlimited license of SDL Trados 2007 Suite
as well as a permanent license of SDL Trados Studio 2009.
Well, you certainly cannot blame SDL for not listening to users'
responses, and you also can't say that user forums are useless.
Also, this comes following several changes to the licensing policy that
had already taken place last week and that were initiated by user responses:
- Originally it
was planned to allow only one install per license. This has been changed
to two (one for laptop, one for desktop) for upgraders. Folks who buy the
new version from scratch have to pay an extra fee of 30 Euro to install
two instances.
- Originally it
was planned to block running more than one freelance license on one
network. This has also been changed so that it is now possible to run more
than one on a workgroup network.
- The AutoSuggest
feature (which arguably is the most interesting new feature in the new
version of Trados -- I reported on that a few weeks ago) was supposed
to be limited to the more expensive Professional edition. For the time being
it is now also part of the Freelance version.
I feel for the folks from SDL. They are trying to do the right thing,
while at the same time knowing that releasing a completely reworked version of Trados
is a huge gamble. As I said before, while the new version is superior in many
ways to the old, users of the old Trados will have to sit down and learn
some new tricks -- and those can be learned on the Trados Studio or one
of its competitors.
|
4.
Flexible Clay
| |
At the Localization World event in Berlin last week,
the new version of Clay Tablet
was released.
For those who are not sure anymore what Clay Tablet
is: It offers ready-made connectors between commonly used content management
systems (CMS) and a number of translation environment tools (TEnTs) and/or translation
workflow systems so that text from the CMS can be monitored and easily
extracted. It's a fairly high-end product with a corresponding price tag, but
then it's also usually not purchased by the language service provider but by
the end client on whose system it needs to be installed. Interestingly, last
time I wrote about it (about 15 months ago) I said this: "Clay Tablet
will offer its product with a very affordable SaaS -- Software as a Service --
model, essentially forestalling any huge expenditure for anyone." Today,
about 80% of its sales are through the SaaS model. Things are changing quickly.
The new version (2.5) offers two new features -- InSync
Master Asset Management Technology and MultiPoint Flexible Routing
Technology. Quite a mouthful if you ask me, but still interesting. These
new features give more flexibility in the routing of the text, so Clay
Tablet is not only a text extraction and re-insertion service but it now
offers a number of workflow management features that are aimed at helping to
build higher quality TMs and less manual interaction.
Clay Tablet itself calls it a "game changer." I am not
sure that I would agree with that on the whole, but from Clay Tablet's
perspective it is, since it sort of redefines the tool and shows that its
makers are on the ball to keep up the development, horizontally (by developing
more and more interfaces) as well as vertically.
|
The Last Word on the Tool Kit
|
|
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
Last week this reader added a link:
www.asap-traduction.com
© 2009 International Writers' Group | |