1. Analyze This or
Analyze That!
| |
Gabi Ortiz pointed me to a site -- www.textalyser.net -- that
performs all kinds of text analysis. Like Gabi, I'm not sure how helpful this
really is for translators -- but it's worth a glance anyway. The idea is that
you can copy and paste text into a provided text field and then perform
analyses that you did not even know were possible.
Here is a sampling of the kinds of numbers you can
get:
·
Number of different words
·
Complexity factor (lexical density)
·
Readability (Gunning-Fog Index)
·
Total number of characters
·
Number of characters without spaces
·
Average syllables per word
·
Sentence count
·
Average sentence length (words)
·
Max/min sentence length (words) ·
Readability (100-easy, 20-hard,
optimal 60-70)
I couldn't find information on what makes up a
word in the minds of Textalyser's makers, but this is still interesting
data -- albeit primarily geared toward French and English. The site looks a
little rough around the edges, but on short texts the analysis is indeed
performed blazingly fast. Longer texts -- I (very unfairly) uploaded the
complete text of War and Peace -- make the site, well, crash. But then, I am not sure you
would perform this kind of analysis before your next War and Peace translation
project. . . .
|
2.
Context Menus in Word
and Other Problems (Premium Edition)
| |
I recently installed the dictionary application Babylon
for testing purposes. Once I was done with my tests I immediately uninstalled
it again, but to my great frustration the context menu items that were
installed in many, many MS Word context menus (the menus that appear
when you right-click on any item within the Word interface) were still
there. Really annoying.
One way to deal with that would have been to delete
the normal.dot
template and let Word recreate one when it restarts. That's easy to do
and the unnecessary menu items would have been gone, but all the other
customizations to the way that Word behaves and looks that I have added
over the years would have been gone, too. So, that's not good.
But then I remembered a trick I had used years before,
and I thought it might be helpful to share: Most of you know that it's easy to alter
the toolbars in Word 2003 and before (and
most other programs). You just need to right-click on the toolbar area, select Customize,
and then drag and drop commands from and to the toolbar. Again, these
alterations are only stored in the normal.dot -- but if you are careful you can
keep that for a long time.
Now, the same dialog also gives you access to the
context menus ("shortcut menus" in MS speech). Once you are on the
first tab of the Customize dialog (again, you'll need to select Customize
after right-clicking on the toolbar area or you could also select Tools>
Customize), the Toolbars tab, you can find the Shortcut Menu
item. If you double-click on it, a new "toolbar"-sort-of-thing will
show up on your screen with three menu items: Text, Table, and Draw.
Clicking on each of those will reveal long lists of "situations" for
which context menus might be displayed. If you select any of these, you can see
what respective commands the context menus in these situations contain. You can
change them by simply dragging new commands from the Commands tab of the
Customize dialog to the list of commands, or by dragging them out of the
lists if you want to delete them. Not bad, huh?
Of course it's rather involved, but it might be worth
it if you are terribly annoyed by that one silly command that you find every time
you right-click on something, or you always long for one command that's just
not there.
Now, in Word 2007, the feature is gone. The
only way to mess with the context menus there is through some heavy-duty programming, so you might be stuck with what you have. (Good thing the Babylon
commands were not as nasty there as they were in Word 2003 when I tested
it.)
Oh, and while we're at Word, I encountered a
terrible nuisance with a copy of Word 2007 on Vista recently, and
from what I've seen it looks like it's a rather common problem.
One day while I was doodling around in Word 2007
(you know how much time translators have to doodle around . . .) Word
stopped responding to my mouse. I couldn't select text or open menus or
anything like that with my mouse. And when I closed Word I received a
message informing me that Word had just crashed.
Here is what I found for a fix:
The underlying problem is a damaged Word Data key in the Windows registry. You will
need to open the Registry Editor by typing regedit into the Search field in the Windows
Vista Button menu.
(Since the registry is the Achilles heel of Windows you should
now, once the Registry Editor is open, make sure to back it up by
selecting File> Export and selecting All under Export Range.)
Then you will need to navigate to
\\HKEY_CURRENT_USER\Software\Microsoft\Office\12.0\Word\Data
and either rename the Data
key or delete it (Word 2007 will create a new Data
key using the default settings when it restarts.) That did it for me. It recurred
once after a few days of using Word properly, so I fixed it again and
this time it has not popped up again.
|
| ADVERTISEMENT |
SDL
Trados Studio 2009 - 25th Anniversary Offer!
SDL Trados Studio 2009 Translate 30% faster with
AutoSuggest - One reason to choose or upgrade to SDL Trados Studio
2009 If one is not enough ... check our Flash Demonstration with all the new features!
To find
out more please visit www.sdl.com/toolkit09 Valid until 31/09/09
|
3. The Alignment
Progress Report
| |
In the last issue of this newsletter I wrote about the
alignment tool(s) offered by Terminotix. Every time I do that I hear from Ilia
Kaufman or someone else from KCSL and NoBabel.
But that's OK since this seems to be a good way of learning about the progress
being made in terms of product development.
For those who don't remember: NoBabel is a
company that offers a number of products for translation memory optimization
and generation, among them AutoAligner. Unlike Terminotix's tools, AutoAligner
is a SaaS (Software as a Service), meaning you don't have to install any
software on your computer. Instead, you upload files that need to be aligned to
an online server, and the software on the server does the alignment for you
(and you pay for every matched sentence pair). In fact, you don't even have to
"tell" the AutoAligner anything about the files, not even what
language they are in. The files are analyzed on the server, the language is correctly
recognized, and the file pairs are matched up on the basis of the actual
content.
Once the files are matched up, the actual segment
matching is performed on the basis of a linguistic analysis. This means that
not every possible pairing will end up in the translation memory, but only
those that are deemed appropriate (typically up to 5% are ignored). But this
also means that you can upload a list of terms in two different languages, both
sorted according to the language-specific sort order, and you will still have
the correct alignment.
So far so good, and the match accuracy is amazingly
high. But, again, if you are bound to get every possible segment pair out of
the alignment, this might not be the right tool (the same, by the way, is true
for Terminotix's products).
When I last reported on AutoAligner, I
had two main points of criticism: the limited number of file formats it
supported (text, Word, WordPerfect, and HTML) and its limited
number of possible language combinations. Both of these have changed in the
latest version. As far as file formats go, PDF is now also supported, which
obviously is extremely important since so much web-based and other alignable
content sits in PDF files. And the number of languages has been dramatically
increased: it is now possible to do English <> French, Italian, German, Spanish,
Dutch, Polish, Portuguese, Russian, Arabic, Chinese, Japanese, and Korean as
well as German <> French and Spanish.
|
4.
EuroTermBank, the
Other Option
| |
In the last newsletter I reported on the clever EuroTermBank Terminology Add-In for Microsoft Word which brings the content
of that helpful termbase right into MS Word. I also mentioned that,
while it's extremely helpful, it would be nicer to have the possibility of having
the terminology in all kinds of applications and not just Word through a
tool like IntelliWebSearch.
Mike Farrell, IWS's creator, was kind enough to
write the search mask for it. Here it is:
Label=EuroTermBank Start=http://www.eurotermbank.com/Search.aspx?text= Finish=&langfrom=en&langto=it&subject= Notes=Change ISO codes in Finish string according to your languages Quotes Off=No Pluses Off=No Encoding=UTF-8 Case=1
All you need to do is change the
third line to the language combination of your desire, copy the above code to
your clipboard, open IWS's Search Settings and select Share>
Import from clipboard. This will produce an already filled-out Import
dialog that you can then edit if you need to and save. From then on out, you
can search the EuroTermBank from anywhere in Windows.
Cool.
|
5.
Blogging All Over the
World
| |
When I was 12, I brought home that (admittedly
dreadful) Status Quo single "Rocking All Over the World" and put
it on my newly acquired turntable. My dad just looked at me and then uttered
the hurtful (and what seemed to me very ignorant) words: "Oh, this isn't
music! All I hear is that monotonous beat!"
It's a good thing that blogs aren't monotonous,
especially when they come from translators! The ATA has taken it upon itself to
maintain a little list of translation-related blogs. Good stuff when you have
time to doodle around.
|
| ADVERTISEMENT |
How do I fine-tune Windows so it works best for translation work? And where can I find info on freeware programs that allow me to operate more efficiently? Complex file formats--how do I translate those? Should I buy desktop publishing and graphic software? And, oh, what about translation tools? Find answers to these questions and many, many more in the 360 pages of the classic computer primer for translators: The Translator's Tool Box.
|
6.
Not So Fast!
| |
I mentioned in my last newsletter the following in
regard to Microsoft Word and its usefulness as a translation interface: "Even
the last two big tool vendors who swore by that environment -- Wordfast
and Trados -- have gotten away from it."
Well, Monsieur Champollion, the maker of Wordfast,
was not happy about that comment, and rightly so. I was glad he contacted me
because it gave me a chance to spend a little time talking with him (in the
past, Wordfast had not been particularly conversational).
So, Wordfast has released Wordfast 6 (see, I almost said it myself!).
They have, of course, released Wordfast Pro, which -- and this is the
message that Yves would like to convey -- is not an update of Wordfast 5.5 (darn it, again!) Wordfast
Classic, but a separate version that eventually might or might not be
accepted as the standard Wordfast version. While Yves is convinced that
this will happen at some point, he is committed to fully supporting the Word-based
Wordfast Classic until, well, until forever. It no longer works on Macs
with the latest version of MS Office, but it does work (according to some
preliminary tests) on the upcoming Word 2010 for Windows,so
there can be some continuity there.
I find this whole story quite fascinating and something
that could serve as a great lesson to other tool vendors as well. Wordfast has
been a very popular tool among translators for a number of years now. It was a
lot cheaper than most of its competitors, particularly Trados, and while
it may seem a bit small, many have found that there is surprising power under
its hood. It was the first TEnT to introduce the QA concept, it has offered the
Very Large Translation Memory (VLTM), an online shared
translation memory for many language pairs, and has maintained a very loyal
user base through outstanding support by Yves. The limitations had been clear
for a while, though: Because Wordfast (now Wordfast Classic) is Word-based,
it could process only those files that are either directly or indirectly
processable in Word, which is great for folks who mostly work in these
file formats, but not so great for others.
So it was clear at some point that the development of
an independent interface not bound to any third-party tool or format would be
extremely desirable, and this was the Java-based Wordfast Pro
(which, to Yves's great regret, was first called Wordfast 6). Alas, it
turned out that it's a formidable task to create a TEnT from scratch,
considering the many formats that it should be supporting (presently MS
Office formats, HTML, and InDesign are supported) and considering
how conservative translators are! Yes, I said it! Translators are a
slow-moving, conservative bunch, hard to convince to move away from something
that, well, sort of worked.
I'm glad that Yves sounded upbeat about the new
version of Wordfast Pro and its good chances of being simply Wordfast
again at some point. But I'm afraid that many, many of his customers have yet
to be convinced.
(And if you think that this is a bit reminiscent
of SDL's situation with the new Trados, you are spot on.)
|
7.
Toxic Stuff
| |
I regret that I haven't written more often about OmegaT,
the open-source TEnT out there with the greatest following. It's not that I
don't want to; it's more that I just don't hear it when something new happens.
I do receive mailings from OmegaT's alive and active mailing list,
but then I receive many other lists' mailings as well.
But last week I stumbled onto something that I
really liked. Marc Prior has developed a little utility with the funky name Toxic (Trados-OmegaT-eXchange)
that converts Trados TTX files and Wordfast Pro TXML files to an OmegaT
format and, once the translation is done, back to the original format. I
performed some trial conversions and compared the resulting files and it looked
good. According to the list there apparently are some problems here and there,
but nothing that a friendly bunch of co-open-sourcers couldn't and wouldn't
help with.
|
8.
I'm Just Sure I've
Seen Those Tips and Tricks Before! (Déjà Vu, Déjà Vu), Part II (Premium
Edition)
| |
I just finished a very smooth and fun project in Déjà Vu, and here are some more of the things that I appreciate about that tool.
Last time I did this I mentioned interesting uses of
the "lexicon":
The lexicon, that
ominous third project-internal database, is good for many things -- as the
preferred terminology container, as a temporary holding place for data that
still needs to be verified before being sent to the terminology database, or as
a means to extract terminology before starting a project to get some terminology
research out of the way before diving into the midst of things. But aside from
that, I use it increasingly often as a means to import data quickly into the
translation memory or terminology database. Most (advanced) Déjà Vu
users are familiar with the slow performance when it comes to database
maintenance. It can take hours to import large external databases into an
existing database (and just as long to export it again). Amazingly, importing
into the simple lexicon (via File> Import> Lexicon) is done in a
snap, and sending the lexicon to either a translation memory or a terminology
database (via Lexicon> Send Lexicon to ...) is also done in a
heartbeat. Once you are done you can delete your provisional lexicon from your
project by simply right-clicking on it and selecting Delete. The
additional benefit of this is that even the project-specific data on subject
and client are sent as well, which means you don't have to fiddle with SQL
commands to add that retrospectively.
I also spoke of surprisingly easy ways to use SQL
(Structured Query Language) with Déjà Vu, good ways of sorting and
filtering data in translation projects, and the fact that it also supports SDLX
.itd projects (by the way, since then I have run into a good number of .itd
files where an import into Déjà Vu was not possible).
Here are some other things that I really like. The
feature that I probably have used most frequently is the Filter on Selection
function. This allows you to highlight any term or phrase in the source or
target column, right-click, select Filter on Selection, and only rows
that contain that very term or phrase will be displayed. This is a very cool
feature to quickly check consistency, and it's much faster than scanning for a
specific entry in the translation memory of something you have translated
before. The only two caveats to this feature are that it's easy to forget you
filtered your project and erroneously think you are done (you need to right-click
again and select Unfilter to reset the filter), and there are no
keyboard shortcuts assigned to it. The latter problem is easily solved, though.
Just like in the Word article above, right-click on the toolbar section,
select Customize, click on Keyboard,and locate the three
respective commands under View (the first occurrence of Filter on
Selection is for the source, the second for target). Assign keyboard
shortcuts (I have Alt+Ctrl+F for
filter on source, Alt+Ctrl+G for
filter on target, and Alt+Ctrl+U
for unfilter), and you are good to go. As far as I am aware, MemoQ is
the only other tool that offers filtering features as powerful as this.
Another very helpful feature is the External View
feature. Under File> Export> External View you can create an RTF or
HTML document in a multi-column format or a bilingual Trados document
(the latter only in the Déjà Vu X Workgroup version) that contains all
of the text that needs to be translated, edited, proofread, or just gawked at
outside the Déjà Vu environment. The External View document can
easily be reimported into Déjà Vu once the translation / editing /
proofreading / gawking is done. In my humble opinion, this feature is one of
the most generous of any tool since it allows for working with any number of
third-party tools within a production chain.
One more tip: HTML is a very easy language. In fact,
it is so easy that it's a snap to create a filter that accounts the very
limited number of formatting and coding rules that HTML provides. Of course, this
is one of the reasons that HTML is almost always among the first couple of formats
that new tools support. The problem is that an HTML file is never only HTML.
There is almost always another kind of coding present, whether JavaScript,
XML, or some other kind of non-HTML entity. And that is where all of the preconfigured
HTML filters typically fall short and import and/or don't protect text that
should be left alone or protected. The solution that the makers of Déjà Vu have found for this is not particularly
user-friendly but, if used right, nonetheless powerful. You
can create a little text file, name it HTMHide.txt,and save it in the same directory
that your project file is located in. You can then enter regular expressions
into that file (the help system gives you a lot of support for this) which
exactly govern what kind of text should be translated and what not. Again, this
gets pretty technical, but for a project with many hundreds or thousands of
HTML files it's well worth the effort.
Oh, and one more thing: Déjà Vu crashes
sometimes. But you know what? You never lose information because everything is
saved automatically every time you enter it.
|
The Last Word on the Tool Kit
|
|
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
Last week these
readers added a link:
www.wordbonds.es
http://blog.alfredo.tv
© 2009 International Writers' Group | |