ToolkitSmall

A computer newsletter for translation professionals

Issue 10-12-182
(the one hundred eighty-second edition)
Contents
1. I Have 10 Dreams! (Premium Edition)
2. More Books from the Best of Us for the Rest of Us
3. Perfect Match-Making
The Last Word on the Tool Kit
Between the Years

Here is a completely worthless question that's applicable for this "between-the-years" time: What do a puppy, a little rat, an elephant's trunk, a monkey tail, a duckling, a strudel, a maggot, a snail, a dog's head, and a cinnamon roll have in common?

Please don't inundate me with answers. I promise to publish the answer in the next newsletter. (Or you can ask someone from Bruce International, the originators of this riddle.)

Happy New Year, Amig@s


1. I Have 10 Dreams! (Premium Edition)

There have been plenty of Top Ten lists released by smart people in our industry during the past few weeks -- not surprising, considering that we are about to enter a new decade -- so I won't bore you with yet another one.

Instead, I thought it would be interesting to take a look at the translation technology landscape out there, kick back, and dream of things that I think are missing and would make our work a lot easier.

The difference from a typical "Top Ten" list? This is stuff that is probably not going to happen (though I would be truly happy to be proven wrong!)

Here is my dreams list, in no particular order:

1.      Open standards for server-based systems

2.      Full fruition of XLIFF

3.      Wide and customizable access to tool-external resources

4.      Feature parity of browser-based vs. hard-drive-based tools

5.      Integration of common word-processing features

6.      Morphological tool kits

7.      Rediscovery of terminology and integration into TM and MT workflows

8.      Completely integrated terminology extraction

9.      App stores for all kinds of tools

10.  Honest communication by tool vendors

Open standards for server-based systems

This is not exactly a new idea. For instance, in the last edition of the newsletter I wrote about how important the release of the Server API for Trados will be so that other tools can hook into its server-based workflows. But since most tools have some kind of server-based workflow, it's not just Trados we're talking about.

What exactly is a server-based workflow and why do we need to "hook" into it?

Well, we've been talking about translation exchange standards forever -- standards like TMX to exchange translation memories, TBX to exchange terminology databases, and XLIFF to exchange the actual translation files (and others, including SRX, GMX, and xml:tm). These standards are really important to facilitate the exchange of translation project between different (versions of) translation environment tools. Traditionally, the focus of these standards was the desktop so that projects could be exchanged from one system to another. With the advent of server-based projects, assets (translation memories and terminology databases) are stored online with real-time access to a translation professional, and suddenly things are no longer quite so clear-cut. While in some cases it is possible to download some or all of the data into one of these exchange formats, it defeats the purpose: the real-time collaboration between various translation professionals is therefore typically not granted. Ergo, the data is only accessible with the originating tool (which could be Trados, Déjà Vu, Across, memoQ, or any of the other tools that support this kind of workflow).

What is needed, therefore, is a standard that allows for any translation environment tool to hook into that workflow, i.e., query the underlying translation memories and terminology databases that remain on the server and write data to them as the project continues, enabling other folks who are working on the project to benefit from that data in real-time. Sound unlikely? Agreed. But it's no doubt possible if there is a real will on the side of tool makers. Is there? No, not really, and this is true for big and small vendors alike. Yes, they all support the above-mentioned translation data standards in some way or other and communicate this very publically and proudly. But with the future and present clearly going in the direction of server-based projects (at least for large projects), many vendors are only too happy to have found this new way of capturing their technology. You find that frustrating? I do, too. There is actually an email address that Alan Melby and I set up awhile back to allow you to voice your concerns so we can pass those on to the vendors. Feel free to let tool makers know through the address I listed, or talk to them directly. Are they gonna listen? That depends on how many of us are talking, I think.

Full fruition of XLIFF

XLIFF (XML Localisation Interchange File Format) is probably one of the best things that has happened to our industry in the last few years, and ironically it almost happened by accident. XLIFF really was only meant to support the translation process of software development formats, but it has long since morphed into an exchange format for all kinds of files (and in fact, many tools use XLIFF internally for all their supported file formats).

It's a really cool standard. Yeah, it's great that we can exchange TMs and termbases with standards, but how much greater is it to exchange the actual translation file between different tools? No more awkward converting of source files or need for filters to support file formats of other translation environment tools if everything went through XLIFF!

The problem with XLIFF is that it is extendable by definition. This means that you can create perfectly valid XLIFF documents that still cannot be read -- or read in their entirety -- by other tools. Now isn't that lovely! There is a really instructive presentation on this very topic that shows one quality assurance tool vendor's struggle when he needed to support the output of the exchange format XLIFF from various other tools. It tends to sound like an oxymoron when you have to create customized filters for an exchange format. The plea to tool vendors needs to be this: don't extend your XLIFF definitions. Use what XLIFF already offers.

(And yes, you might want to use the same email address as mentioned above.)

Wide and customizable access to tool-external resources

Some tools, especially Across and Fluency, already offer the ability to easily link to online and, in the case of Across, even hard-drive-based resources for third-party terminology resources. With tools like IntelliWebSearch (see the article below), it is also possible to do searches on terminology resources on virtually anything and everything from within any tool in the Windows environment. Still, it would be a helpful and easy add-on for all TEnTs to offer online dictionary and corpus searches as well as dictionary searches from within their environment. Why? Because not everyone is reading this newsletter to find great ways of getting to helpful resources (can you even imagine what life would be without this newsletter? ;-), and the more help we can provide to the uninitiated, the better.

I would not be surprised if we might actually see this in upcoming versions of tools!

Feature parity of browser-based vs. hard-drive-based tools

2010 saw at least two translation environment tools (XTM Cloud and Beetext) come out with browser-based translation interfaces that are almost as easy to work with as the interface of desktop tools. There are some things missing, though, including features such as a global search and replace. Clearly there will be more and more tools that will offer browser-based interfaces. The mantra for the developer of these tools is that less is better (i.e., don't over-clutter your interface) while at the same time offering all or most of the features of desktop-based tools. Some things work in favor of a browser-based interface, including spelling checking provided by the browser or an AutoSuggest feature that many browsers also offer by default. Still, this continues to be a tough nut to crack. And since virtually all tool vendors will be looking at an online-based translation process, they'd all be well-advised to look at the two tools mentioned above.

Integration of common word-processing features

This is an area that many tool developers have been working on, but it's also one that needs serious expansion. With MS Word all but out of the picture as the interface in which translation is being done (I know that many of you still use tools that allow for it, but even you will have to admit that there is a general trend away from MS Word), some features are missed by translation professionals in the tabular or otherwise structured interface of translation environment tools. Many tools now have decent spelling checkers, AutoText features, and some even AutoCorrect, and an increasing number also have visible formatting (WYSIWYG) for features such as bold and italic text, but that's not all that tools like MS Word or OpenOffice offer. What about grammar checks? What about track changes? What about more than just italics and bold as visible formatting? In many cases, tool developers have to deal with limitations on the editing areas that their particular environment can handle, but maybe it's time to look for more highly enabled environments.

Overall, I see a little improvement on this front, but it's likely that we'll see third-party tools (or add-ons -- see below) stepping into the fray.

Morphological tool kits

I have had many conversations with tool vendors about this area. Why don't the tools know morphological rules when it comes to recognition of termbase and TM entries or the automatic adjustment once matches are found? Some tools offer some functionality in this area in some languages (Across and Transit), but the majority of tools don't. And why? It's too expensive. Plus, any kind of development in this area will be language-specific -- so where do you stop? Which languages do you include?

Here is a suggestion: Why don't tool vendors develop tool kits to have users develop morphological language-specific rules that they can share with others or with their same language group who also will have some contributions to make? Too simple a solution to make it work? Let's see whether we can find some takers among the tool vendors. I certainly promise to give them all the coverage they need in my newsletter.

Rediscovery of terminology and integration into TM and MT workflows

How long have we/you been preached to about the importance of proper terminology work? Like, forever? It sure seems like it, and yet I think it would not be presumptuous to say that most of us are still not particularly prudent about it. And being prudent can mean a lot of things: Do we take terminology work seriously and really invest into it? How do we use terminology in our workflows? Do we have manual look-up procedures or do our tools automate look-up and utilize their findings? Do the different components of our tools (translation memory, grammar checking, morphological "knowledge," machine translation, etc.) "talk" to each other and work in concert with each other?

I think it's this last point where there can still be a lot done on the side of the tool vendors. Some tool vendors have offered features where termbases and translation memories "talk" to each other, so terms in a fuzzy TM match are switched on-the-fly if the old and new terms are "known" to the termbase (Déjà Vu does that, for instance). But there is clearly so much more that can be done. How about grammatical adjustments (singular vs. plural or gender-specific changes in articles, etc.)? And how about communication between machine translation matches and terminology databases?

Some kind of machine translation feature is present in virtually all translation environment tools. But it's by no means advanced. The feature usually provides for a suggestion from a machine translation engine when no translation memory match is found, but there is no automatic lookup or even replacement of terms that are found in the termbase. And it would be an easy thing to do (especially if the same kind of communication already happens between the translation memory and the termbase). Provided that we had good termbases (!), we might actually get decent hits even from machine translation engines like Google Translate or Bing Translator.

Now of course there are ways to integrate terminology into machine translation processes, but only by training processes and not on-the-fly. I foresee this feature as a potential game changer on two fronts: a) it could make machine translation a more reasonable additional feature for many of us, and b) many of us might actually start to take terminology work seriously.

Completely integrated terminology extraction

Terminology extraction has been around for a while. The idea behind this is that you can either take a set of documents (in source and target language) or a bilingual file (XLIFF, TMX, TTX) and have a program identify matches in those files, which in turn can serve as the basis of the glossary or terminology database for a given project. This is pretty ingenious, and yet most of the stand-alone programs that do that (for instance, MultiTerm Extract or SynchroTerm) have not exactly been flying off the shelves.

So why not include this feature into any file that is being processed by default, at least by creating a monolingual glossary that can then be leveraged against the TM or termbase? Tools like Déjà Vu and Text United (I'll write about that next year) can do it, and so can other vendors. Granted, it is a little more processing-heavy -- but not by much. And it once again would increase the use of terminology resources.

App stores for all kinds of tools

In the last newsletter I wrote about OpenExchange, SDL Trados' app store. Here developers can create and offer applications that can extend the usefulness of SDL Trados.

Other vendors: follow suit! Not only will you prevent SDL from building up its market dominance, but you will also help users of your tools. (Plus you might learn quite a bit about what users actually want and need.)

Honest communication by tool vendors

This is pretty self-explanatory: We have been hurt enough by false promises. This is certainly true for machine translation technology, but also for the technology used in translation environment tools.

The much-heralded "support for PDF" in many tools? Give me a break. No tool truly supports PDF files -- at least not in the way that other file types are supported. Complete openness because of XLIFF support? Well, I already talked about that a few paragraphs earlier.

Hint to the tool vendors: by definition, translators are smart people (even though we may not look like it!).

And will this happen next year or even in the next decade? I doubt it. But there is always room for hope.

ADVERTISEMENT

Tired of expensive tools that make you work harder for less?

 

Start saving time and money with Snowball!

 

Free 90-day trial: Lite (free), Freelance (€99), Pro (€199)

Download       Philosophy (Video)     Email

 

You translate. Snowball remembers.
2. More Books from the Best of Us for the Rest of Us

I had promised awhile back to write a series of reviews of books relevant to our industry. I reviewed Alex Eames' Business Success for Freelance Translators; now it's time for Judy and Dagmar Jenner's book, The Entrepreneurial Linguist: The Business-School Approach to Freelance Translation.

Many of you will be familiar with the Jenner twins through one of their blogs (Translation Times or Neue deutsche Rechtschreibung), their other publications and speaking engagements, and their active participation in many industry events, associations, and discussions.

The Jenner twins are different (no, not really from each other -- after all, that's the point of being a twin!) from many of us. For instance, they never work with language service providers but choose to work exclusively with end clients. Naturally this has strong implications for how they market themselves, and since they have language and business degrees, they not only do that very well but also decided to share it with the rest of us.

I really enjoyed reading the book and need to admit that I learned a lot.

The book is divided into various parts. It starts with a look at you as a translator (why you are a business as a language professional and how you should act as such), goes on in great detail on how to meet direct customers, how and how not to present yourself and your business to them, and how to negotiate prices with them (in short: don't), then talks about you as a translator again (activities in professional associations and balancing life and work), and finally summarizes everything in a helpful recap.

When I talk about books or software I often start with, "This XYZ is not for those among you who. . . ." Truthfully, however, I cannot think of many translation professionals for whom this book would not be helpful (maybe with the rare exception of those who already have so many clients who pay such high rates that there really is not much else to hope for -- but even for those the Jenners have advice: never stop marketing because you will lose some of your clients).

Here are some ideas that provided some aha moments for me:

  • Any business relationship that you as a translation professional have with someone is a business-to-business (B2B) relationship in which there are exactly two equals: you and your business partner. And it does not matter whether your business partner represents a small company with a handful of employees or Lionbridge or IBM. Do you really always carry yourself that way when negotiating deadlines and prices?
  • Client education is important, but only if it's done smoothly. Many of us become extremely dogmatic when "teaching" clients about our business. The reality is that clients will not appreciate our dogma and will not learn much at all.
  • Forget about résumés when dealing with direct customers, and offer a well-written company brochure instead (and, by the way, save on all kinds of things, but not on marketing materials).
  • And here is the coolest one: when a client asks you for too much in a conversation (pricing, deadline, etc.), just be quiet. Don't say anything. In most cases the client will reconsider and ask for more reasonable conditions (I think I'll try that one with my wife and kids as well!).

In the book, the Jenners spend a lot of time and give excellent advice on the need for publications (including detailed advice on how to write blogs and news releases), social interaction on the web ("the Web 2.0 is your friend" -- I really appreciated the section on Twitter and LinkedIn but did not quite understand the business case they made for Facebook), how to manage things like business expenses and other paperwork ("love to do your paperwork" -- man, I wish that could become true for me!), and the way we need to present ourselves and reciprocate favors that we ask for from others.

Both of the Jenners live in metropolitan areas (Dagmar lives in Vienna and Judy in Las Vegas), so much of what they describe in their marketing is with local contacts -- clearly something that is not applicable to us country bumpkins. However, we are also spoken of and blessed with advice on how to find contacts (travel to trade shows, build networks, etc.).

Neither Dagmar nor Judy would call herself a technology wizard, which becomes apparent in the few places where they talk about technology (there really is no need to have an extra application when saving something from Word or OpenOffice to PDF, and there is no "international version" of Dragon NaturallySpeaking), but that is not why you want to read this book.

The book is beautifully layouted and proofread (one of the Jenners mentioned somewhere that when the first copy came from the printer she immediately found a typo -- I think I found it, but it's literally the only one, which is no small feat for a 200-page book), and it is sold as a "real" printed book as well as a PDF download.

Today and tomorrow you still have time to buy the book and declare it as a business expense for 2010 (the Jenners would surely be proud of me for that savvy business advice), and I can think of no better way to start 2011 than by reading this helpful book.


ADVERTISEMENT

Only two weeks left to get WORDFAST CLASSIC and WORDFAST PRO together and SAVE €170. 

Beginning on January 1, new pricing will apply for Wordfast products so act now and save. 


Visit www.wordfast.com/pricing.html for more information.
3. Perfect Match-Making

Awhile back I mentioned that I started a conversation between Yan Yu, one of the people behind the development of the TDA (TAUS Data Association) online search engine, and Michael Farrell, the developer behind our ("our" as in "your and my") favorite search tool IntelliWebSearch. The purpose of their talks was to make it possible for IntelliWebSearch to search the amazing multi-lingual TDA corpus -- which meant that TDA had to make some changes to the way that its searches can implemented. As a little year-end present, I'm glad to announce that it can now be done.

The cool thing is that you can (of course) specify language combinations in your searches, but also things like industry, data owner, or content type.

This is how it's done. (If you've never used IntelliWebSearch, this will look confusing, but once you start using it this will make a lot of sense to you.)

In the Start field in the Search Settings for IntelliWebSearch, enter:

http://www.tausdata.org/index.php/language-search-engine?q=

In the Finish field, enter:

&intelliwebsearch=1&pos=&lemma=no&industry=3&owner=1&content_type=0&target_lang=fr-fr&source_lang=en-ug

Now you'll need to change the source and target languages (en-ug is a combination of British and US English, en-us is US English, en-uk is UK English), the industry (3 is "computer software" and the numbering is according to the listing on www.tausdata.org/index.php/language-search-engine), the owner (1 is "All"), the content type (0 is "All"), and you're done!

IWS

Enjoy. And I hope you'll have a great year!


The Last Word on the Tool Kit

If you would like to promote this newsletter by placing a link on your website, I will in turn mention your website in a future edition of the Tool Kit. Just paste the code you find here into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.

Here are two readers who recently loaded the code:

volkova.professorjournal.ru

© 2010 International Writers' Group