1. The EuroTermBankWordAddinRelease
| |
I get frustrated with Word's spell-checker
sometimes. Why in the world would the longish term in the heading
not be recognized as a correct word? (Just kidding.)
Anyway, there is a very interesting plug-in that was
released earlier this week that allows a search through the EuroTermBank from within MS Word.
Now, the EuroTermBank has been publicly available for the past two and a half
years and it's a termbase with an "initial focus (...) on terminology
collections from the 'new Europe,' including Estonian, Hungarian, Latvian,
Lithuanian, and Polish terms and their equivalents in English, German, French,
Russian and other languages (overall, almost 30 languages)."
I ran a number of tests with my (non-focus) language
combination (EN> DE) and got surprisingly helpful results.
Of course, there are a host of on-line terminology
resources of varying quality, but all things being equal, I tend to use those
that allow access right from within my immediate work environment. That's why I
like a tool (like IntelliWebSearch) that directly links me to the
dictionaries and glossaries I want to go to, and that's also the strength of
this plug-in.
The EuroTermBank Terminology Add-In for Microsoft Word, however, takes it a step further; it not
only links to the desired sites, but it also brings the information right into
the MS Word interface.
(The only real drawback I see is that MS Word
is no longer the preferred environment of many translators, now where even the
last two big tool vendors who swore by that environment -- Wordfast and Trados
-- have gotten away from it. In fact, the only thing that I still use Word
for on a regular basis is to write this newsletter. . . . )
But for those who still like Word as their main
interface and who work in European language combinations, the add-in should be highly
welcome. There are two different versions of it, one for Word 2003 and
one for 2007. I mention this because there are some real functionality
differences, and the 2007 version add-in is clearly superior.
But let's start with Word 2003. Once you
have the add-in downloaded and installed, it simply adds "ETB
Terminology" to all the other resource books (Thesaurus, Encarta, etc.)
that you can access in the Research pane. This opens when you click on a
term while holding the ALT key. You will find lists of translations of that
term in all the languages that are found; there is no way, for instance, to set
your desired language combination.
This is different in the Word 2007 version.
There you have a large EuroTermBank button on the Review ribbon
(which is essentially the only interesting ribbon for standard language-related
activities) that you can click on when you have highlighted a word or any
section within a text (you can achieve the same by pressing Ctrl +Shift +I or the right-click
command Terminology). When you do that, a separate EuroTermBank
pane appears on the right with all possible translations into all the
languages, but you can quickly filter that down by selecting source and target
language and subject area (these settings will be stored between Word
sessions) to see a well-defined list of relevant terms. Also, you can send a
whole sentence or paragraph using the same method, and the pane will show the
terms for which it finds a translation by formatting them like hyperlinks.
Clicking on them will lead you to the respective list of translated terms.
It's super easy to use, plus it's free and does not
seem to interfere with anything else -- so you'd be silly not to use it.
|
2.
Using Search Engines (Premium Content)
| |
I would just love to be more flexible with my use of
search engines! In my column for the current ATA Chronicle I recall one
of the rare times when Google was down in the last few weeks and the
rigor that beset me -- I felt completely helpless and essentially stopped
working for an hour or so. I did not even think about there naturally being plenty
of other search engines that I could use just as well. It was one of those
times when I felt really stupid!
Once I reflected on this I realized that once we
become relatively expert in using one tool (Google, in this case) we
automatically feel like we lose expertise in other tools. I love to use special
search syntax like "define:<search term>",
"filetype:XYZ
<search term>", or "site: www.XYZ.com <search term>"
in Google (I have written about that in the past) to quickly filter
results -- but of course those parameters don't work in other search engines.
It's entirely possible that there are comparable techniques in other search
engines, but they are well-hidden.
Recently I found an interesting translation trick for Bing -- but what would really intrigue me is a "translation
primer" from one search engine search syntax to another. Let me know if
you're aware of something like that. (By the way, I do like Bing's
travel comparison feature -- I found a great price for tickets for my trip to
the BDÜ conference in Berlin next month.)
And while we're at browsers, LSP Globalization
Partners International has "launched a new search engine for
international business professionals, travelers, students, researchers and
anyone else who needs to easily search the web by language, by country, and by
search engine."
Glearch ("Global
Search") allows you to search three different search engines (Google,
Yahoo, and Bing -- though I was not able to see any results from Yahoo)
simultaneously by language and country. It's very easy to use and returns
interesting results -- but it also still has some bugs and I got a number of
error messages when I tried exotic combinations. I am sure that these will be
ironed out over time, so it might be good to bookmark it for later use.
(Unfortunately, special syntax use like I mentioned above does not work with
this tool.)
I was pointed to another, yet different multi-lingual
search engine by Donna Parrish. 2Lingual is a simple but interesting
implementation of Google machine translation and search technology. As
you enter one term in the first search box, Google Translate produces
the translation in the other search box and at the same time executes a search.
I was a little distressed to see that "Jeromobot" got a good number
of hits on the English side but very few on, say, the German side, but I think it's
just a matter of time until he becomes a worldwide celebrity! (By the way, he has his eyes on a girl!)
So, to come back to 2Lingual: While the MT
translation feature is kind of crude, I still can imagine some helpful uses for
translators -- such as quickly locating websites on a certain topic across
languages.
|
3. Translated!
| |
This is the eureka cry of the translator after a completed
job, but it's also the name of the company that claims to be offering the
"world's largest translation memory."
Translated's MyMemory has been around for a
while now, and I had been a little skeptical about it. But when I (and I'm sure
many of you) received a note from them this week, I took the chance to look at
it again and talk to its founder. And I was much more impressed.
First of all, here's what it is: MyMemory is a
rather colossal translation memory of presently around 150 million segments that
contains data from web alignments (app. 30% of the total data), corpora such as
the EU corpus (app. 50%), and TMs that the mother company Translated
contributed from its own work (after client approval) and contributions from
other translators.
So far you can access data in a search mask that
allows you to enter a term, phrase, or complete segment and then returns
translations in your target language. A machine translated version is first offered,
followed then by TM matches. Those TM matches can be used for your translation or
you can choose to improve or delete them with very easy editing facilities that,
according to CEO Marco Trombetti, are used at a rather impressive rate. The
goal, of course, is to improve the quality of the data in a cooperative kind of
way.
Besides making data available, the overall goal of the
project is to use the data to feed Translated's own statistical machine
translation engine as well as an effort to gain visibility in the translation
market -- which, again according to Marco, works rather well.
The latest push is to encourage more translators to
upload data (more on that later) and to introduce a new feature that allows you
to upload a document and receive a TMX translation memory file for the
translation of said document in the TEnT of your choice. This feature is in a
rather early beta version right now, but they are making changes and
improvements to it at an impressive rate (it has changed at least two or three
times since I spoke to Marco yesterday). The downloaded TMX file will contain
all the TM matches it finds, and for the remaining segments machine translation
matches are being inserted. Personally I found the inclusion of MT matches to
be a real show-stopper, but there will be an option to exclude that (you can
already see that option on the respective dialog, though it's still grayed
out). By the way, the MT engines that are used for the translation are
different according to language combination. They range from Google Translate
to Systran to their own systems.
When you contribute TMX files, you have the option to
specify whether this is for your own private use (but I assume it will still be
used for MT training) or for everyone, and you can choose whether you want to
hide proper names such as company and product names. When I asked Marco about
that last feature, he explained that he has developed algorithms which recognize
these automatically by comparing source and target and essentially looking for
identical terms that are all capitalized (for example) which are then assumed
to be proper names. I can see this working in a combination with English as
source or target, but I am not sure whether it would work in all language
combinations. We shall see.
While this seems to be a good step toward providing some
privacy for the data you upload, it's not clear what happens to the actual
documents that you can upload to generate TMX files when this feature comes out
of beta in October. It seems that Translated
would be well advised to explain the process by which the source documents are
deleted once processed.
Overall, I think this is exciting, and I can see
that this could be a model we might see more of in the future. The daily
visitor rate of 13000 seems to confirm that notion.
|
| ADVERTISEMENT |
How do I fine-tune Windows so it works best for translation work? And where can I find info on freeware programs that allow me to operate more efficiently? Complex file formats--how do I translate those? Should I buy desktop publishing and graphic software? And, oh, what about translation tools? Find answers to these questions and many, many more in the 360 pages of the classic computer primer for translators: The Translator's Tool Box.
|
4.
You align! No, you
align! No, you align!
| |
This has the ring of my kids fighting about chores
that need to be done -- only that in those cases the "align" is
replaced with "do that."
I have known about YouAlign for quite a while now, but I was
asked not to write about it until Terminotix, its owners, were ready to
present a stable enough version. This has now happened, and even though I found
a couple of bugs when I tested it, they were fixed at a moment's notice. I
think many of you will be glad to see this (for the present time) free product.
Terminotix is a
Canadian company that essentially offers three translation-related products: Logiterm,
Synchroterm, and AlignFactory in various editions. In the past I
have written about all of the products, but I have most consistently praised AlignFactory,
which I think is the best product to turn alignment (the conversion of a source
and target document into a translation memory) from a nightmare into a feasible
and profitable part of the translation workflow. Now, AlignFactory is
relatively highly priced and will be a hard purchase for some to justify,
especially if you have not had proof of its power. So, Terminotix has decided
to release a free online-based version of it that allows you to upload source
and target documents -- including PDF files -- and receive a TMX file in return.
(Just like Translated, Terminotix would also be well advised to
explain how the original documents are being destroyed after the alignment
process has finished.)
Those of you who have tested AlignFactory won't
be surprised by its accuracy and speed. Everyone else should be in awe.
Presently there are only a few languages supported
(Arabic, Chinese, German, English, Spanish, French, Japanese, and Russian), but
Jean-François Richard, Terminotix's president, assured me that more will
be added quickly.
There is also a limitation on file size of the source
and target documents (1 MB each) and you can also not batch align many files at
once (unlike in AlignFactory), but it's a great tool for smaller jobs
and to whet your appetite for the full Monty.
|
5.
Clearing Up the DTP
Conundrum
| |
For those of you who receive the Premium Edition of
the newsletter, I published a table that displayed which TEnT can work with which
DTP tool a few weeks ago (in the 144th edition, to be exact). I
received a bit of feedback from TEnT vendors, so I have updated the table.
Since the table was contained in a graphic stored on my server, you just need
to pull up that newsletter, refresh, and you should see the new data.
|
6.
Type It, Baby
| |
John Yunker announced a little utility on his blog that might be helpful for entering text in a language
you only rarely need to type in. In these cases it's a nuisance to look up
special characters through the Character Map (under Start>
Programs> Accessories> System) or to install a new keyboard. TypeIt is a shockingly simple online tool
that allows you to enter text, with all its special characters, in Czech,
Danish, Dutch, Finnish, French, German, Hungarian, Italian, Polish, Portuguese,
Romanian, Russian, Spanish, Swedish, and Turkish. You simply open the
respective language and enter text into a field in your browser. A special
on-screen keyboard is displayed as well as keyboard shortcuts for the
respective language. Once you've composed your text you can copy and paste it
into the document you are working on.
Again, please don't use this as your main interface to
translate, but if you need to enter the occasional Hungarian, Russian, or
Turkish text, this might be a highly welcome quick-and-easy tool.
Interestingly, in a response to John's posting,
someone pointed to a new tool by Google that's also very interesting: The Virtual Keyboard API.
It says this on Google's blog:
It is often difficult for Internet
users to input text in many non-Latin script-based languages for a variety of
reasons. The correct keyboard layout may not be installed on the computer
they're using -- sometimes such a layout may not be well developed or widely
available. This poses a challenging problem for web developers because there is
no way they can ensure that their users have access to this very basic input
technology. Our Transliteration API can help, but requires that the user
know multiple languages.
Right on the heels of introducing
support for translating Persian (Farsi) [I mentioned this a few weeks ago -- Jost], we've
added a new Virtual Keyboard API into the Google AJAX Language API to
further assist with text input. With this, developers can help their users
input text without relying on the right software being installed on the
computer they happen to be using. (...) With this initial release, we are launching 5 language
layouts. They are: Arabic, Hindi, Polish, Russian, and Thai. We plan to roll
out support for more keyboard layouts in the future.
The procedures for entering the code to display
virtual keyboards in those languages does indeed look very easy and will be a
help to many web developers.
|
7.
Getting the Most out of
a Package (Premium Content)
| |
Many TEnTs (translation environment tools) use
packages created by the corporate versions that can then be edited by the
freelance or even free "Lite" versions of the respective tools. Generally,
these packages contain directory structures that are recognized by the Freelance
or Lite version as a valid project directory, compressed in a zipped file. To
the casual observer this is not very apparent at first because the extension of
the package file is not .zip as we are used to with most compressed files, but
it's an extension that is specific to the TEnT. To find out whether this is the
case in your TEnT, open the file in a text editor (such as Notepad) and
look for the first two letters in all the gobbledygook you'll find. If they are
PK, you are in luck and you know this is a .zip file. (And then, please don't
save the file in Notepad -- that would break it.)
The next thing you need to do now is to change the
extension to .zip (just right-click on the file in Windows Explorer and
select Rename). After you unzip it you will find the files that are
contained within the zip file in a separate folder.
Now, some tool vendors have decided to internally
encrypt these .zip files, thus building an artificial barrier to prevent the
use of other tools with "their" files. I personally find this to be particularly
poor judgment on their side because it goes against the spirit of standards and
exchangeability of information that everyone on the outside is committing to.
In fact, now, at a time when the most important exchange standards (TMX, TBX,
and XLIFF) are supported by most tools, the two areas that are still not
exchangeable are the workflows of translation management systems (such as the
online access to translation memory, terminology, and project data of one
particular tool) and these artificially encrypted packages. For the first area we
will (have to) see APIs (interfaces that allow the access of third-party tools)
in the near future; and for the second we just need to see the silly protection
dropped (tool vendors, you know who you are, so just go ahead and do it!).
OK, after this tirade, let's come back to the issue at
hand. Once you have unzipped the file, you end up with files you might be able
to use or not, but chances are that there will be some part that you might find
usable.
In the case of Star Transit, the Transit
packages (with the extension .pxf) contain the standard SGML files that Transit
uses as translation files (their extension is a three-letter language-specific
abbreviation) and a subfolder that contains the necessary terminology in a .txe
file. And that's what all this is about.
A client of mine had a problem the other day with just
that file. After I told her how to get to it and also told her that this looks
like standard MARTIF (the precursor of the terminology exchange standard TBX),
the import as a MARTIF file into TermStar (Transit's terminology
companion) failed. This procedure that was provided to her by Star's
support did the trick:
Open the .txe file with
a text editor, in the section <databaseDesc> delete all the information of the type DictProperty' id (the information of the type hyperlink separator and ExportedLangs can stay), and make sure that the tag pair <databaseDesc > ... </databaseDesc> stays in the file.
Now the .txe file can be
imported as a MARTIF file into TermStar or other MARTIF-compliant tools.
(And this, of course, would be a lovely extension of the otherwise powerful
filter for Transit projects that MemoQ is offering . . . ).
|
The Last Word on the Tool Kit
|
|
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
Last week this
reader added a link:
translationtimes.blogspot.com
© 2009 International Writers' Group | |