| 1. Zappity-Do! | |
This past week a new version of the wonder tool CodeZapper was released. Though I have mentioned it before, here is a quick reminder of what it is. Developer David Turner says this in the readme file:
CodeZapper is a set of Word macros (programs written in VBA to automate operations in applications) designed to "clean up" Word files before being imported into a translation environment program such as Déjà Vu X, memoQ, SDL Trados Studio, TagEditor, Swordfish, OmegaT, etc.
Word documents are often strewn with junk or "rogue" tags (so-called "smart tags," language tags, track changes tags, soft hyphenations, scaling and spacing changes, redundant bookmarks, etc.). This tagged information shows up in the DVX or memoQ grid as spurious {1}codes{2} around, or even in the mid{3}dle of, words, making sentences difficult to read and translate and generally negating many of the productivity benefits of the program.
OCR'd files [files processed with optical character recognition software] or files converted from PDF are even worse.
CodeZapper tries to remove as many of these tags as possible while retaining formatting and layout. It also contains a number of other macros which may be useful before and after importing files (temporarily transferring bulky images (photos, etc.) out of a file, to speed up import, and then back in the right place after translation, moving footnotes to a table at the end of the document and back after translation, for example).
The tool, which runs in Word 2000 and above, consists of more than two dozen macros sporting very "macro-ish" names (how about "HdeMnB" or "MvFtnts"), but once you get the hang of it you'll learn to appreciate them if you have the "rogue code" problem.
Here is a sampling of the features (and for any of these alterations there is always the opposite one that brings the file back into its original state once the translation is done):
- Removal of rogue codes on the active document or all documents in a folder
- Moving of footnotes to the end of the document
- Moving of text out of text boxes into a table
- Moving images out into a temporary file
- Extra removal of codes for converted files from PDF
- Hiding of (automatic and manual) bullets and automatic paragraph numbers
- Setting of one language ID throughout the document
- Saving of .doc to .rtf and vice versa
- Counting of number of textboxes in a document and the total number of words in text boxes
Until recently David gave away his tool away for free, but it's now a paid tool ($20). But I gotta tell ya, that's certainly worth the price if you are among the many who struggle with wicked codes.
|
| ADVERTISEMENT |
Translation Office 3000 -- Deliver Every Job on Time! Easy management of your freelance translation business. Automates your business routine. Helps you focus on translation, not administration! 25% discount, for Tool Kit readers only: http://www.translation3000.com/aitpn/160_22_3_40.htm
|
2. InDesign in the Cloud (Premium Edition)
| |
This is how it goes sometimes: Just a few years ago most of us had never heard about "InDesign." Today it's everywhere. It's come to the point that when a new translation environment tool is released, the first few formats that are supported typically include HTML, XML and InDesign, often before Microsoft Office!
And while it's much easier to work in and with InDesign files now with a more stable XML-based exchange format (.idml rather than .inx), it's still a pain -- at least if you as a translator want to have any role in how the document actually looks once you're done translating, or if you as a language service provider need to convert between the various formats that are necessary to have your translators work in the InDesign files in a format that is convenient for you and your translators. (And let's not even start talking about the multitudes of different InDesign versions that presently exist.)
So this is the challenge: create a way to avoid having translators purchase the extremely expensive (and hard-to-use) InDesign; let them see what they actually do in their translation and what it looks like; plus make it as painless as possible for the project manager to manage, prepare, and post-process the files. Enter web-to-print (web2print) technology, based on the high-powered InDesign Server. This technology allows you to view and edit InDesign documents in a web browser without having to have any additional product. Since the release of the InDesign Server, a multitude of products have sprung up that are using it (see a list of those right here) and, not surprisingly, some are geared toward the translation industry.
One of these products is one2edit, a German product with a new sales office here in the US. Across and SDL (Trados) are among its official partners, each offering an easy integration into the workflow, but as we will see in a second, that integration can be achieved with virtually any TEnT that supports XML or XLIFF.
But let's first look at one2edit on its own merits.
When service providers receive an InDesign file, they can upload it to one2edit and immediately view the file in the browser. (one2edit itself is Mac-based, but I was assured that the only potential problem with compatibility between the Windows and Mac worlds might be fonts.) The next step now consists of going through the file and deciding which parts are translatable and which parts are not (this is done with a simple drag-and-drop feature) and what kinds of rights are assigned to each step in the process (for instance, can the translator resize text boxes/stories or delete them and so on). As a next step, the content of an internal translation memory is applied to see whether there are any matches for any of the translatables.
After that, translators, editors, DTP professionals, and whoever else is needed are assigned (their data is stored in the system if it is set up that way) and are given access to the file on the server. Here the translator can now translate whatever has not been translated by the TM, then pass it on to the editor, and so on and so forth until the file is translated, the TM is updated, the file gets layouted in the browser interface and then sent to the printer as a PDF or native InDesign file.
Sounds good, but not great. What about a more sophisticated way of translating the content? Yes, there is a translation memory involved, but do I really have to work in a browser to translate the remaining content? If we're just talking about a few labels here and there it might be fine, but what if it's a very text-heavy document?
And this is where one2edit's additional level of sophistication comes into play. While the above workflow is indeed one of the possibilities, another one is much more likely in a professional translation environment.
Loading the file, determining what is to be translated, and assigning roles is the same in this second workflow. But the next step now is that a TMX file from an external tool can be imported as well as an SRX file, a segmentation rules exchange file that most TEnTs support (with the notable exception of Trados). The SRX file will force a certain segmentation of the translatable content in the InDesign file to match the TEnT segmentation and thus the translation memory, so there is a greater likelihood for matches.
The next (optional) step is the leverage step -- the application of the TM content to the file -- and after that a translation file is exported. The translation file can have one of three formats. It is either a bilingual Excel file (typically not a good idea because you would lose all formatting information), a bilingual XLIFF file, or a one2edit-specific XML-based file with the strange extension .o2et. In my tests, working with the .o2et was actually easier than with the XLIFF file. Trados (with the help of an .ini file that the folks from one2edit supply) and Across readily support this XML format, but I ran a test with memoQ and Déjà Vu and was able to seamlessly process the file without much preparation. And the same will be true for any tool that supports XML.
Once this .o2et (or XLIFF) file is translated it can be brought back into one2edit by either the translator or the project manager. All formatting is maintained and it can now be reviewed by the translator, the editor, or both (or none).
And once everything is done (and that includes the layout with possible last-minute changes to the text), the internal translation memory is updated with the very last stage of the project before it goes out to the printer (or wherever it goes). This is extremely helpful and eliminates one of the greatest pains in the translation process: making sure that the TM actually reflects the final project.
one2edit's internal TM can then be exported to TMX and imported into any TMX-supporting translation environment tool.
So, cuánto cuesta?
If you want to have your own server, the smallest edition is the Mini Edition where you get your own preinstalled Mac Mini Server for 800 Euro a month. But this solution serves only two concurrent users, so it's not really what most language providers would be looking for. Naturally there are larger solutions where you own the server, but these become truly pricey.
Not surprisingly, one2edit also offers a SaaS solution or software as a service. For 250 Euro a month you can rent a "Workspace" (haven't we heard that term elsewhere recently . . .) that can be used by 30 users, has 10 GB of storage, and allows you to create up to 2,500 pages per month (if you go beyond that amount you can purchase more with a per-page pricing model.)
Interesting? I sure think so. Because nothing is always just perfect, there is at least one large caveat: one2edit does not work with bi-directional languages, so no Arabic, Hebrew, Urdu, Persian, etc. The makers say that it's in the works, though.
|
3. The Joy of Using Text Editors (Premium Edition)
| |
This was the title of one of my columns in the ATA Chronicle awhile back. I agree that text editors -- programs that can be used only for text-based files -- are not really what you and I would identify as sources of joy (or let's put it like this: hopefully we can think of more immediate sources of joy than a software program), but working with text editors can be very, very satisfying.
I was reminded of this just last week when a well-known colleague contacted me because she was having a hard time importing a translation memory exchange (TMX) file into a translation environment tool. The TEnT actually gave her some specific error messages, but how do you modify a TMX file?
Well, TMX files are actually XML files. And XML files are actually text(-based) files. See, in the world of computing there are essentially two different kinds of files: "text-based files," also known as "flat files," and "binary files." Any file that is humanly readable (though not necessarily understandable) is a text-based file, and that would include .txt, .html, .xml, .rtf, .properties, and a huge variety of other files that, when opened in a text editor, display characters and numbers rather than crazy symbols. Binary files, on the other hand, are "compiled files" that can only be opened and edited in one or two specific applications (such as .doc in Word, .ppt in PowerPoint, .zip in a zipping program, or .exe as a standalone executable file). If you open a binary file in a text editor and save it, you have broken it -- forever and ever. And ever. (So it might be a good idea not to do it!)
So, to come back to the above-mentioned TMX file, I was able to open the file in a text editor, change the offending language designator (it said "French" rather than "FR-FR"), save and close it, and everything was hunky-dory.
Here is another case of something that just happened last week. Another well-known translator contacted me with this problem: Translation environment tools typically don't deal well with XML files that also contain HTML coding. I've written quite a bit about this in the past and have actually posted a summary of one strategy for dealing with these files on my website.
To use this strategy you'll need to be able to open the XML files in MS Word for pre-processing before you process them in a translation memory tool. Due to legal struggles suffered by Microsoft, it had to neuter the US versions of MS Word 2007 and 2010 so when it opens XML files, it strips all the tag data, which in turn makes the file completely unusable.
It turns out that there are two things Word looks for when opening the XML files: the extension (.xml) and the first line of the file (<?xml version="1.0" encoding="UTF-8"?>). If the extension is changed to, say, .txt and "xml" is taken out of that first line, you can use Word to open the XML file as a plain text file and preprocess it exactly the way I described in that above-mentioned document -- you'll just have to remember to add the "xml" back in and change the extension back to its original format.
And what does this have to do with text editors? You'll need to open the .xml file in a text editor so that you can take the "xml" out. And if you had to do that for many files individually? Then you could use the batch search-and-replace feature that most text editors offer! (See below.)
What else are text editors good for and which text editors should you use? Every version of Windows comes equipped with a text editor (Notepad, accessible under Start> Programs> Accessories; the Mac OSX equivalent is TextEdit) that can be used for a number of things (including the above-described alterations of the TMX or XML files). But if you would like to do something a little more advanced, you might want to look at a more powerful specimen of a text editor. Notepad++ might be the most obvious choice since it's free (and, as the name suggests, at least twice as good as the Windows Notepad . . . ) and is indeed powerful. There are many, many other text editors out there as well, some free and some charging a nominal fee. Two that I like are the Japanese Emeditor which, maybe not surprisingly, is particularly strong with Asian languages, and UltraEdit, an all-around great text editor. The one thing that you want to make sure when choosing a text editor is to see whether it can handle Unicode well (not all do: for instance, a previous favorite of mine, Multi-Edit, still does not).
Here are some of the things you can do with a text editor:
- Open large text-based files in a fraction of the time it takes in an Office program (Excel or Word).
- Search (and replace) in any number of text-based files all at once.
- Open HTML, XML, or other tagged and coded files. The text editor recognizes what is code and what is not and separates it visually by using different colors.
- Sort glossaries and delete duplicates, even across lines of different capitalization, all at once.
- Open binary files (but don't save them!) to find out what kind of file they are (if you don't know otherwise.) The first couple of characters often give a clue to their "true" identity: %PDF is a PDF file, BM is a Bitmap file, MZ is an EXE or DLL file, JFIF in the first line points to JPEG files, or PK to a ZIP file.
- View the source of websites to check out some clever coding or to see what kind of translation issues you are going to run into.
- Check for the validity of XML or HTML files (if you messed something up during translation).
- Compare files in an easy-to-see fashion to find out what kind of changes may have been introduced between different versions of the same file.
- Switch code pages (to and from Unicode or between "native" code pages).
And on and on. Here's what I've found to be true: If I can imagine some kind of logical operation within a text-based file, chances are that it can be done with a text editor.
|
| 4. This and That (Premium Edition) | |
Let's start with the errata section (it's actually just an erratum section.)
A couple of issues ago I wrote about the hybrid machine translation engine offered by PROMT that combines the two dominant MT technologies: statistical machine translation and rules-based machine translation (RbMT). I mentioned the results of tests that the PROMT developers ran and said
. . . when they compared the output of translations from their customized RbMT system and the hybrid system, [the tests] had a difference of 5 points on the BLEU scale (BLEU is the commonly used method to evaluate the quality of machine translation), still far away from most human output and still very much in need of post-editing
While the final conclusion is accurate, Olga from PROMT asked me to correct the "5 points" statement. She sent me some data reflecting PROMT's BLEU scores from English into Russian, German, and Spanish, and the difference between an already-customized rules-based system and the hybrid system is actually around 7. So there you go.
One thing that has bothered me since I started using Windows 7 is that Skype's default behavior is not to minimize itself into the system tray. Silly me! I simply didn't realize that there is an option under Tools> Options> Advanced> Keep Skype in the taskbar while I'm signed in that I needed to uncheck to have Skype go back to its good old ways. (I'm sure most of you are asking: What in the world is he talking about? But I know there's going to be one reader who is celebrating: YES! THANK YOU!)
I've written about web fonts before. These are fonts that free browsers from relying on the fonts users may or may not have installed on their computers. (Ever noticed that web sites are super boring when it comes to fonts? That's because their designers have to plan their layout in a font they know all users will have.) With online-based fonts, however, it does not matter which fonts users have on their computers; the websites always look as planned.
Google has just released its own web font project, which not only allows you to embed some of the beautiful fonts into your webpages but even enables you to download the fonts for your personal computer. (If you think new wallpaper brightens your living room, try changing the font in your main working environment to Vollkorn or Molengo!)
I've written about the free translation environment tool Similis before. This tool works particularly well with extracting term pairs to quickly build up a terminology database, and best of all, it's free. Unfortunately, extracting term pairs is really not a particularly intuitive process in Similis, so I was pleased that Tool Kit reader and DE>FR translator Jean-Marc Tapernoux created a document explaining everything from the installation process to how to extract the terminology lists from your documents. You can download it right here.
Lastly, I've been offering a special discount for ATA members to subscribe to the Premium edition of this newsletter. Now, in celebration of the upcoming FIT (Fédération Internationale des Traducteurs) conference in San Francisco in September, I'm also offering a discount for members of any FIT member association. (If you don't know whether you qualify, chances are pretty good if you are a member of an official translator's association. You can find a list of member associations right here.) Just let me know in the PayPal form when you order your subscription what association you're a member of.
|
| The Last Word on the Tool Kit | |
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
Here are two websites that added the Tool Kit link this week:
www.awiles.net
www.czechtranslation.com
© 2011 International Writers' Group
|
|