|
1. Information Super-Storage
(Premium Edition)
|
With the content-centric Web 3.0 in the making, storage, retrieval, and
interpretation of content is front and center to many technologies. But many of
these technologies are concerned with large databases and are just a bit beyond
the scope of the regular user, who's just looking for good ways to store and
retrieve data efficiently. I recently looked at a number of tools designed to help
with that, and I liked none quite as much as Evernote. Evernote comes in a
number of flavors: for mobile devices, for a web-browser-based interface, and
in a (Windows and Mac) desktop version. And it is this last
edition that I had a look at. Evernote is a fairly large download -- a
bit over 50 MB -- but it installs easily. One of the options during the install
is for you to choose whether you would like to install handwriting recognition
for German, English, French, and Russian.
This brings us to what this tool actually does: It allows you to collect
snippets from any kind of application -- web browser, office application, Acrobat
Reader, etc. -- in the form of text or image (!), uploads those to a
server-based location -- so that you can access them from any computer -- and
processes and indexes them. One of the processes that it applies is that of OCR
(optical character recognition) as well as HWR (handwriting recognition). This
means that if your images contain text in a decipherable manner, Evernote
will also index those words and you will be able to search for them quickly.
It really is quite impressive. First I saved a couple of images that I knew
would be hard to read, such as our logo to TranslatorsTraining with
the turned-around r. Not surprisingly, that one could not be OCR'ed. But when I
saved the logo of this newsletter (text: "The Tool Kit -- A computer newsletter for translators"), Evernote was able to do a
beautiful job on it. I was able to "crack" the database that contains
the index and this is what it contained for that image:
computer acomputer a computer acomputey a computey
acompuler a compuler a coimputer aromputer a romputer ac omputer accipiter
computed computes attributed reactivates newsletter for translators the tool
toot kit hit kr kk kite kith kits kilt ten its tea tel tlc ltd ltv tem
When I entered any of these "words" into Evernote's
interface, the graphic was called up and the presumed word highlighted. It's a
very clever way of using OCR technology, I think. Rather than going for
super-accuracy as is necessary when the OCR output needs to produce a document,
here it uses its fuzzy abilities to make the application better -- making sure to
get to the right word while allowing for typos during the search.
Once a snippet, a web page, a file, or a note (that you can manually
enter) is saved, it becomes part of your extended memory that can be searched
on the fly. And if there is no text in the item you saved, you can add tags
manually that can be searched as well. Of course, any of the items is printable
and can be saved elsewhere on your computer, and you can choose to organize
your items in certain categories ("notebooks").
Evernote is free,
unless you are a true packrat and want to save more than the 40 MB a month you
are allowed to save with your free account. In that case there is a paid
version ($5 a month) that gives you 500 MB a month.
It was interesting to see which languages are supported for plain text (this
does not include the OCR or HWR ability). The only languages that I could not
make work were Chinese, Japanese, and some Indic languages. Other languages,
including Korean, Thai, Tibetan, Arabic, and Hebrew, worked fine.
So, what is this good for as a professional tool? Well, I can think of
various uses. While working on a translation project -- especially a
longer-term one -- there are many times when you find information during your
research that might not have an immediate application. You end up discarding it
only to kick yourself later when you need it. With a tool like Evernote,
it is readily available. The same is true for the rare glossary that you might
stumble on without having an immediate need for it. Sure, you could bookmark
it, but you could just as well put it somewhere where it's easy to be found
again. And on and on. . . .
|
| ADVERTISEMENT |
-
Many other immigration lawyers only use their
left brain hemisphere. I also use my right.
-
Many other
immigration lawyers only see a picture of a hat when you show them a picture of
a boa constrictor digesting an elephant.
- Many other
immigration lawyers make simple things complicated. I make complicated things
simple.
It should not be a
surprise then that many other immigration lawyers actually hire me to help them
win their own cases. Perhaps you too can hire me to help you win
yours.
|
| 2. And Let There Be Many
TEnTs | |
A few years ago, new
translation environment tools (TEnTs) were showing up left and right, and it
seemed that every other newsletter contained an announcement of a new tool.
That certainly has slowed down somewhat. In the last few months, only AnyMem
(see the ad in this newsletter) popped up as a completely new tool. However,
now there is another new tool out there that also might be worth a second look:
AidTrans Studio.
This tool from Krakow, Poland, seemed at first like a rather small tool,
but it turned out to be a full-fledged application (that might still be a
little wet behind the ears). I talked to developer Piotr Labuzek yesterday to
find out more about what his tool is supposed to be and how he sees its
position in the market.
He said that he "would like AidTrans Studio to become a
'popular tool,' affordable and relatively easy to use." The affordability
he has certainly achieved -- right now all three versions (Basic, Professional,
and Enterprise) are free in the beta testing phase until June, and from
there on out the Basic edition will remain free. The Basic edition
is a fully functional version without networking ability, batch processing, and
"utilities" (regular expression test, encoding conversion,
translation length verification, and custom tile format tags configuration).
The preconfigured supported file formats of AidTrans Studio are
Microsoft Word 2003 (saved as .xml files), Word 2007 (saved as
.xml or OpenOffice), PowerPoint and Excel 2007, XML, OpenOffice
files, and Trados .ttx files. The fact that MS Word files are
only supported through conversions is probably a real weakness. I asked Piotr
about it and he said that he felt that XML is the way of the future and that's
why he put his eggs in that basket. (Well, he may not have used those exact
words, but that's what he meant!) Though that's true, we are still dealing with
a lot of legacy documents that also need to be directly supported.
The interface of AidTrans Studio is very pleasant, a little reminiscent
of the old version of Star Transit. This is also reflected in the
underlying translation file structure, which consists of a text-based file for
each source and target language (Transit uses SGML files). The database
structure for the translation memory is similar to SDLX or Déjà Vu
with the Microsoft Access Jet engine (but SQL Server for the Enterprise
edition).
I did not find the workflow completely intuitive -- the creation of a
project and the file import is a separated process and you have to
independently initiate the "database environment" for the termbase
and TM before they spring into action -- but those are things one could either
get used to or which could be fixed in the post-beta phase.
I also think the tool could also benefit from a less technical look in
certain areas (such as the configuration of regular expressions for a
customized file filter -- which is basically nice but too complicated), but in
general I was impressed by how complete and comprehensive the tool is.
Not surprisingly, this tool is not mentioned yet in the following
compendium of translation software:
|
3.
130 Pages of
"Translation Software"
| |
John Hutchins, the great chronicler of translation software, and machine
translation software in particular, has just released the 15th
edition of his Compendium of Translation Software -- directory of commercial machine translation
systems and computer-aided translation support tools.
It's really a very interesting document, if only to see how much
software there actually is to support our work. One very practical application
of the document is the index of language pairs for machine translation in the
very back of the manual. I often receive questions about whether certain
language pairs are supported by a particular system. Well, here are the many
answers.
|
| 4. Mea Culpas |
Since Barack Obama has shown us so impressively this week how to perform
mea culpas, I cannot but contribute my own. And while I admit that erroneous
information in this newsletter might be of a different magnitude than
nominating ministers who have forgotten to pay taxes, here are mine
nevertheless:
A number of readers noticed and alerted me to the fact that the InDesign for Translators manual
does not apply to InDesign CS2 but to CS3 instead. My bad. However,
the fact remains that it's sort of a shame that the manual does not cover the
most current version: InDesign CS4.
On and beyond this subject, Arle Lommel sent me this information about CS4:
CS4
has introduced some profound changes relevant to translators (although some are not
apparent yet in the UI). Perhaps the most significant one is the ability to
export InDesign Markup Language (IDML) files, which are, in fact, a zip-compressed
set of XML files (much like OpenOffice files are a bunch of compressed
XML files). The really neat thing from a translator's perspective is that, if
files are saved as IDML, you no longer need binary file-format filters for
InDesign. Each story is saved as a separate XML file that can be manipulated
with standard XML tools. If I were a TEnT tool developer, I would not bother
with anything else now since making XML filters is far simpler than application-specific
filters. This change also means that users of any tool that allows for
customizable XML filtering can work with an InDesign-format document
directly without the need to go to a tool that supports the InDesign binary
format. (So there goes much of the marketing value of InDesign filters as
a means to lock translators into a specific tool.)
Another new possibility
would be the automatic replacement of graphics: since the entire content of the
document is accessible, it would be a fairly simple process to have a set of
build scripts (or even regular expressions) that could parse the XML files
looking for things like graphic001_en.psd and replace it with
graphic001_de.psd. (Or, if the content creator couldn't be bothered to use a naming
convention like that, it would be easy to have a list of files that need to be
replaced and go that way.) From experience I can say that managing replacement
of links has been one area where most tools fall down (and where a lot of
errors enter the localization process), but this would allow for some real
improvements and for automatic error checking. While it wouldn't eliminate all
manual work, it would eliminate the task of clicking through (and replacing)
hundreds of links that we've all been through. And the nice thing is that
anyone with a good text editor could do it, no special tools needed.
Now there are some
other exciting things available in InDesign's engine that will be rolled
out into the UI at some point. The most important one is that the compositing
engine now supports right-to-left text. There is no UI to control it, but it
can be scripted and the results for Arabic are impressive: it correctly uses
the appropriate ligatures and positional forms and supports various kashida
models and produces very nice-looking Arabic text, on par with text editors
that advertise Arabic support as a strong point. While this ability is not
ready for commercial deployment yet*, it does point to some major improvements
that should surface in CS5. I also understand that there is also support
for Indic scripts, but I haven't found any details on how to script them yet,
so I have not tried that out yet.
(*After using it, I
would definitely *not* recommend playing around with these features for someone
who isn't comfortable with what may best be described as a pre-alpha experience
-- [Those for whom this is relevant might want to try this -- Jost].)
I did some more research on what this IDML format is about (I was unfortunately
not able to test it -- my own InDesign version is CS2). In said
version CS2, Adobe introduced the InDesign-internal exchange
format INX so that files between different versions of InDesign could be
exchanged. While INX was XML-based, it was typically not possible to simply use
a customizable XML filter of a TEnT for the processing of this format. Instead,
the tool developers had to develop a specific XML-based filter for the INX
format. In general these worked okay, but extremely complex files often
suffered. One of my customers who exclusively translates Quark and InDesign
files refuses to use the INX format in connection with his TEnT because of
previous problems he has run into. Adobe itself says this (of course,
only after the new version was released):
INX was difficult to
read and manipulate because it was designed to be used by InDesign
alone. Those who tried to manipulate INX encountered challenges with
readability, robustness, extensibility, and compatibility with XML tools.
Now, the purpose of the new IDML format is specifically not just for InDesign-internal
purposes, but instead to open up InDesign content to XML-enabled
third-party applications -- among others, TEnTs.
This means that it is no problem to process these much more robust files
with any of the (XML-enabled) TEnTs. All you need to do once you export the
IDML file out of the original InDesign .indd file (File> Export)
is to rename the .idml extension to .zip, unzip the file, locate the XML files
that contain the story content -- the translatable text -- and import or open
them with your TEnT. Of course, if your TEnT only processes files one by one,
you might get slightly annoyed because you will have to deal with many files
one after the other.
And while looking at all these things, I stumbled on a completely
different solution for the translation of InDesign CS3 and CS4
files: StoryTweaker.
Although this is not a solution that can (easily) be tied into a TEnT workflow
and overall seems rather convoluted, it might be just the thing for someone who
is looking for a cheap solution with some WYSIWYG support while translating InDesign
files.
Back to the mea culpas.
I mentioned ways to avoid extreme finger-acrobatics for Trados
user on a laptop, in particular with the shortcuts Alt+[Num+] and Alt+[Num*].
Well, it looks like there is a much easier way.
Trados always had
alternative shortcuts (Ctrl+Alt+N
for Set/Close Next Open/Get -- instead of Alt+[Num+] -- and Ctrl+Alt+Z
for Translate to Fuzzy -- instead of Alt+[Num*]),
but it never really publicized them much at all. That is, until it turned out
that a bug with Trados and Word 2007 caused the Num-shortcuts not to work. Now the
semi-official shortcuts are indeed the Ctr+Alt
versions. (You can check them out in Article 2201 in the Trados
knowledgebase -- thanks to Emma Goldsmith, Carmet Erez, and Amit
Dharma for this tip.)
|
| ADVERTISEMENT |
AnyMem
New User-Friendly Translation Memory Tool
Happy 2009!
Hidden 35% discount on all AIT products:
|
| 5. TEnT User Interface | |
For the German-speaking readers of this newsletter, there is an interesting survey on the desired user interface for translation memory systems (TEnTs).
The author of the survey, who collects this survey data in the context
of working on her translation diploma, writes: "The results of this survey
are intended to help me and the developers of the open-source translation
memory system OpenTMS to
understand what a translator expects of the user interface of a TMS."
|
| 6. Another Sad Demise? |
As has been discussed in various forums, Terminology Matters, the maker
of two beloved tools -- the Word- and Trados-based QA tool Quintillian
and, just as importantly, the TTX-to-Word-to-TTX conversion utility TTXpress
-- has apparently shut down. I have tried to contact one of the former owners
to a) know what happened and b) ask for permission to upload these utilities
for download somewhere -- but I have not been successful. If anyone has any
hints, I would be grateful.
|
| The Last Word on the Tool Kit |
|
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed. © 2009 International Writers' Group | |