| 1. Data Is Good, Right? | |
Good question. Many novice users of translation environment tools would readily argue that some data is good, more data is better, and lots of data is best. I remember when I first started using TEnTs, I eagerly collected all kinds of translation memories that clients would send along for projects and imported them into "my" translation memory in the firm hope that this would eventually pay off and give me the leverage I was looking for. Sometimes that worked. More often it did not.
Today I would argue that the more data we have at our disposal, the more clearly we see the benefits of quality data versus "lots of data" . . . for translation memory purposes, that is. You see, if there were indeed so much repetition between projects and clients, good translation wouldn't be so difficult after all, and machine translation, whether of the generic Google-kind or the highly trained client-specific-kind, would have gained more ground than it already has. So while it's frustrating not to gain the leverage from external data that we hope for, we should also be thankful for it.
Still, one of the reasons why I insist on using the term "translation environment tool" rather than "translation memory tool" is that the simple translation memory feature that finds perfect and fuzzy matches is only one of the many relevant features we need to use in these tools. And quite frankly, it is not always the most helpful one (though very welcome when matches from qualified sources are found). If there were no other ways to get to data than to hope and pray that the tool finds the hoped-for match (and, again, hope and pray that the match is indeed a good match), we should indeed stick to very small and highly controlled translation memories. The nice thing is that TEnTs can do much more.
In the past few months I have written so much about subsegment matching that my lovely editor has started to complain: "Is there nothing else you can write about?" Well, yes, there is, but automatic subsegment matching that tools like Trados, memoQ, Lingotek, and Multitrans now provide has radically extended the usefulness of translation memories, including, and maybe especially, those that come from external sources.
Add to that the terminology extraction provided by default by some tools (such as Similis and Multitrans) and by others as add-on features (such as the SDL line of tools) or through third-party tools (such as some of the tools by Terminotix), the manual concordance feature that virtually every TEnT offers (i.e., the ability to manually search for terms or phrases within a translation memory while translating a project), and the self-correcting ability of translation memory matches through the insertion of terminology matches (which Déjà Vu and memoQ offer), and you get the idea that there is a lot more to translation memories than we originally thought. |
| ADVERTISEMENT |
Word Count? AnyCount!
Invoicing? Translation Office 3000!
Translation Management? Projetex!
Spring 2011 Promotion - 21% off - 7 days left
Click the link to get your copy today: goo.gl/oT5Ht
|
| 2. Where Is the Data? | |
There are essentially two different kinds of sources: those that have already been aligned in the form of existing translation memories of corpora and those that have not (i.e., translated products such as websites, software, or data sheets where you have access to both the original and the translated version). Let's start with . . .
|
2.1. Already Aligned Databases (Premium Content)
| |
There are huge corpora available that are often ready to be downloaded and used by anyone, even though often they were not created with the translator in mind. Many of these corpora were created for (or actually created by) academics dealing with natural language processing and, by extension, with machine translation. While that may not excite some of you, it's no reason to shun these assets.
To get a good overview of some of the larger corpora, go to OPUS. This is a very, very large resource of bilingual files in many language directions containing such varied materials as data from the European Medicines Agency, the European constitution, the European Parliament Proceedings, and various open-source localization and software documentation files. As mentioned before, the files are not especially made for translation memory -- most of them are in a text format - but they are nothing that could not be converted to a TM-compatible format or even the translation memory exchange format TMX (in fact, the files for the European Medicine Agency are in TMX). I wrote about this site awhile back, but it's great to see that is continuously expanding and updating. Some of the new entries that were added last year included a parallel corpus of the "Balkan languages," as well as some English-Greek and English-Chinese data.
Of course, some of these corpora are available from easily accessible stand-alone sites such as the EU translation corpus, which contains 22 European languages. Once downloaded, it allows you to create any of the 231 language pairs in TMX format that you deem helpful.
Or there is the famous Hansards corpus of the 36th Parliament of Canada, naturally between English and French and containing a total of 1.3 million pairs of aligned text chunks.
Are "political" corpora such as the Canadian or European ones really helpful for translation purposes? Honestly, I'm not sure. For instance, I do have all the English-German data of the EU corpus in a translation memory, and while I have never used it as a translation memory per se, I do regularly employ it as a reference database when I need information on the official title of a certain law in its translation between English and German. Could I also find that information through a web search? Most often just as fast. Still it feels good to have that data on my computer.
Of course, a lot of this and much other data is also available through the TAUS Data Associations website, a site that I've often mentioned in the past. It is possible to download TMX data from that site to be used as translation memories (you'll have to "donate" some of "your own" data in exchange -- and we all know that the definition of "your own" data is hard to come by), but I just realized the value of this resource solely as a source to conduct searches when I was working on a project these last few weeks that had a very close match of subject and client (in TDA lingo: Industry and Owner) to data in the TDA corpus. Using the combination of the ability to filter your search so that you have only relevant matches with IntelliWebSearch's ability to look for matches only by highlighting a term and hitting a keyboard shortcut, it essentially turned the TDA corpus into a large TM on which I could perform a concordance search (there is about a second of delay for TDA to filter and display the matches).
(If you are interested in a sample of this particular search query (here is an example), select Search Settings in IntelliWebSearch after you install it, copy the following lines to your clipboard, and then select Share> Import from Clipboard in IWS.
Label=CAD-CAM
Start=http://www.tausdata.org/index.php/language-search-engine?owner=40&q
Finish=&intelliwebsearch=1&pos=&lemma=no&industry=3&target_lang=de-de&source_lang=en-ug
Interword Separator=+
Naturally you will now have to adjust things so that they match your languages and other settings, plus you will have to assign a keyboard shortcut.)
By the way, there is a field in the display of TDA's search hits ("Computed translations") that can be slightly misleading when performing a filtered search. Here you can find already calculated data about the translation of the term/phrase in question. This is really helpful in many ways, but it's always data pertaining to the whole corpus and not only to your filtered query!
Another corpus that is somewhat similar to TDA is MyMemory, though it has a different purpose and certainly a different make-up.
Most of you have encountered MyMemory at some point, but for the one or two who have not, here's what it is: a colossal translation memory of hundreds of millions of segments that contains data from web alignments (app. 30% of the total data), corpora such as the EU corpus (app. 50%), and TMs that translators (including the parent company Translated) contributed.
You can access data in a search mask (or through a tool like the above-mentioned IWS), which allows you to enter a term, phrase, or complete segment and then returns translations in your target language. A machine translated version is first offered, followed then by TM matches. Those TM matches can be used for your translation, or you can choose to improve or delete them with very easy editing facilities.
It's also possible to upload an .html, .doc, .tmx (the recommended format), or .txt file and get a translation memory in return. According to your settings, this can contain only data from stuff that you have already contributed and/or machine translation data when no TM matches are found. If you choose to have the TM augmented with MT matches, they will be marked as such so you can assign a penalty to those in your TEnT or simply filter them out.
Besides making data available, the overall goal of the project is to use the data to feed Translated's own statistical machine translation engine as well as lower a threshold for acceptance of machine translation.
So, how helpful is this resource? Personally, I think not very much for the professional translator, even though I'm in awe about the scope of this project and the fact that it does seem to find a lot of "takers." There are two flaws that make this tool difficult to use productively. First, there is no filtering -- you can filter data by subject, but since there are rather vague rules on how to assign a subject to the data you can upload, the applied filter tends to be too unspecific. And second, while there is a peer-to-peer quality control (everyone can correct data), it's not data that I want to have in my translation memory, and so far it has also not served me particularly well as a reference tool. Still there are a number of tools that directly or indirectly allow you to access MyMemory's data automatically while you work on your project, including Trados, memoQ, and MultiTrans, and I would be quite eager to learn from users who are employing this feature how much benefit they derive from it. Let me know. (Since MultiTrans and Translation Workspace do the same thing with the TDA resource, I would also be interested in feedback there.)
One thing that MyMemory offers is links to the sites from which data originates, if the data was indeed aligned from web content. (By the way, you can also find the same feature on the English-into-German/Spanish/French/Portuguese search engine Linguee.) Why is this helpful? Because it gives you a great way of locating sites that could very well serve as a great resource. Which brings us to . . .
|
| ADVERTISEMENT | |
SDL Spring Offers now on!
Enjoy exclusive savings of over 30% when you buy SDL Trados Studio 2009 Freelance with SDL AutoSuggest Creator!
Are you an existing customer looking to upgrade to the latest version?
Upgrade online and receive 25% off for a limited time only.
To buy online or learn more about our products visit www.sdl.com/promo/toolkit4
|
| 2.2. Non-aligned Translated Data | |
Here I would outline two different strategies for dealing with this kind of data.
The first may be the most obvious: align the data. This works great for sites or documents that have exactly the kind of material you're looking for, by the client you are working for, and focused on the kind of product line you're translating. It would take me a few pages to write about the process for downloading the files and aligning them, a process that I describe in detail in my Tool Box book.
But here are the super basics: If it's just a handful of documents, go ahead and use the integrated alignment feature of your TEnT. If it is a whole or a large part of a website, you will want to use a web spider to download the sites (tools like HTTrack or Teleport) and a specialized alignment tool such as Terminotix's line of tools or NoBabel's AutoAligner to do the actual alignment work. But again, more about that in my book.
There is one trick that I don't mention in my book, though, which works amazingly well when aligning websites in languages where you really can't read the target language, but since it's a little technical and I'm likely to lose some of you, I put it in a separate article (see Crazy Alignment).
What I want to spend more time talking about here is becoming smart about URLs (web addresses). There are numerous parts of a URL that could identify the language of the webpage that it displays. And the cool thing is that if that is the case, chances are that the same webpage is also displayed in other languages (otherwise there is not much reason to note the language in the first place).
http://help.solidworks.com/2011/English/solidworks/sldworks/legacyhelp/sldworks/display/wireframe_view.htm
Clearly, it's not particularly hard to identify which part of the URL is the language identifier, and if you happen to know that that site is translated into the target language as well, you only need to change that identifier in the URL thus to reach the corresponding site
http://help.solidworks.com/2011/Chinese/solidworks/sldworks/legacyhelp/sldworks/display/wireframe_view.htm
or
http://help.solidworks.com/2011/German/solidworks/sldworks/legacyhelp/sldworks/display/wireframe_view.htm
Once you identify the source language site as a site with lots of potential terminology matches for your project, and you also notice that it has a search feature like it does in this case, you can set up a search macro for something like IntelliWebSearch:
Label=SolidWorks
Start=http://help.solidworks.com/Search.aspx?query=
Finish=&version=2011&lang=English&prod=SolidWorks
(Just follow the process detailed above of importing this into IWS.)
Here are some other examples:
http://windows.microsoft.com/en-US/windows/help/network-connection-problems-in-windows
This follows a language code made up of a reference for language and locale (you can find a reference right here), so that
http://windows.microsoft.com/fr-FR/windows/help/network-connection-problems-in-windows
http://windows.microsoft.com/pt-BR/windows/help/network-connection-problems-in-windows
or
http://windows.microsoft.com/ja-JP/windows/help/network-connection-problems-in-windows
are all valid URLs.
Here is another example that I have reported about before which presents a different sort of problem. Look at this great English and German parallel SAP glossary.
Since this is a glossary within an HTML frame, it's not quite as easy to get to specific entries. If you click on any of the actual English entries in the above page, the URL does not seem to change. However, some browsers allow you to view the frame as a standalone page. With Firefox, for instance, right-click in the frame and select: This Frame> Open Frame in New Tab (Safari: Open Frame in New Tab; Opera: Frame> Open in new tab; Internet Explorer and Chrome don't offer that option).
This might open
http://help.sap.com/saphelp_glossary/en/3b/57a67b78608045852d629395c6844b/content.htm
And sure enough, just by changing it to
http://help.sap.com/saphelp_glossary/de/3b/57a67b78608045852d629395c6844b/content.htm
we get to the translated page.
No matter what you do, however, always remember what happened to Jeromobot when he ran into some bad data.
|
| 3. Crazy Alignment (Premium Content) | |
In the last article I mentioned one "trick" to align websites or any other material in languages where you really can't read the target language. When I say "really" can't read the target language I mean languages like (in my case) Korean, Thai, or Indic languages where the (read: my) eye has very little to rely on. This "trick" assumes that there are various languages on that site, with at least one target language you can read that you will be able to use as a "control language."
First, download the site in three languages - a source and target you can read and one target that you can't read. Create two TMX files in your alignment tool, with the first being between the languages you know and the other with the same source but a target you don't know. Import these TMX files into a TEnT that can handle TMs with various sources. I used Déjà Vu, but many tools can handle this.
Next, delete all duplicates, i.e., translation units that are identical in source and target (in DV: Database> Find Duplicate Sentences); for the unknown language combination, also delete all source duplicates with different translations, remembering that you have no way of telling which is the "right" one (in DV: Database> Find Duplicate Sentences> Find Sets of Sentences with different translations then select Duplicates Only and delete all).
Next, sort the translation memory by source and make sure that you can see both of the targets listed as well. (In my case this is not possible with Déjà Vu, so I export the TM as a three-language TMX file and open it in something like Olifant.) What you should see now is a table-like list that looks like this:
Source segment 1 Known target A segment 1
Source segment 1 Unknown target B segment 1
Source segment 2 Known target A segment 2
Source segment 2 Unknown target B segment 2
and so on and so forth. You can now quickly browse through the file and a) locate where there are obvious misalignments between the source and target A and delete target A and target B translation units with that source segment, and b) locate all instances where the targets do not alternate, especially those where there is a target B but no target A, and delete those. This last step assumes that if there is no known target A that I can use to verify a certain likelihood of target B being correct, I delete a possible match rather than taking the risk of having an incorrect alignment.
Once you're done you can separate the language pairs again so that you end up with two TMX files: Source - Target 1 and Source - Target 2.
While it's likely that you deleted a certain amount of "good" alignments in the unknown language combination, you will end up with a surprisingly clean TM. I once actually paid a Korean translator to evaluate the accuracy of an English-Korean TM that I had created by using the above-described scenario, and the error rate was virtually identical to an aligned project with a known language combination (less than 1%).
|
4. New Password for the Tool Kit Archive
| |
As a subscriber to the Premium version of this newsletter you have access to an archive of Premium newsletters going back to May 2008.
You can access the archive right here. This month the user name is toolkit and the password is salmon.
New user names and passwords will be announced in future newsletters.
|
| The Last Word on the Tool Kit | |
If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
© 2011 International Writers' Group
|
|