ATR Interpreting Telecommunications Research Labs


Basic Research on Speech Translation Technologies




Seiichi Yamamoto



The ATR Interpreting Telecommunications Research Laboratories completed its project involving basic research on speech translation technologies at the end of February 2000. NEC Corporation's trailblazing research and the basic research by ATR Interpreting Telephony Research Laboratories, which were founded in the 1980s, had demonstrated that speech translation could be realized, but the technologies developed in these research projects, the grammar-based speech translation technologies, were not able to properly translate naturally spoken language, such as simple daily conversations.

The ultimate goal of our research on speech translation technologies, built on the foundations of the previous studies mentioned above, was to establish the component technologies needed to translate natural spoken dialogue from daily life. In other words, our aim was to develop the basic technologies of a speech translation system that ordinary people without expert linguistic knowledge could use without special training. For this purpose, a huge corpus of data was collected from everyday dialogues, and a corpus-based approach was adopted for research on each component technology. During the first four years of the research, our efforts were focused on the component technologies necessary for recognizing, translating, and synthesizing natural spoken language. Then, the next three years were devoted no integrating these technologies into one system so that a comprehensive evaluation could be made. Since the research results for each component technology are described in detail by each department head in the following columns, this paper concentrates on the outcome of research on integrated speech translation technologies.

A bi-directional speech translation system was necessary to evaluate the preliminary technologies developed for translating spoken dialogues between two different languages and to grasp related technological problems. In 1998, this imperative led us to integrate the component technologies into a Japanese-English bi-directional speech translation system with a vocabulary of over 10,000 words for dialogues on travel reservations. We evaluated the system from various viewpoints, and the evaluations confirmed that the system was capable of real-time speech translation for travel bookings. The actual performance achieved an average subjective rating of 3.8 on a 1-to-5 scale, which indicated that the system was able to fulfill the task most of the time. By comparing the quality of the speech translation with the translation performance of many Japanese individuals at various levels of English, we found that the current machine translation was roughly equivalent in quality to the ability of a Japanese person who scores 500 to 600 points on the TOEIC (Test of English for International Communication). This figure is quite impressive considering the fact that the average TOEIC score of Japanese university students is about 570.

Fourteen years ago when ATR Interpreting Telephony Research Laboratories were established, an automatic translation telephone was still a dream. The current capability of our technology has reached the level of a Japanese person who has been studying English for nearly 10 years, although our system's application is limited to travel reservations.

When the technology developed by ATR Interpreting Telecommunications Research Laboratories for travel reservations was applied to the task of answering questions about area codes and so on, it demonstrated similar levels of translation performance. The results left no doubt that the speech translation technology developed by ATR ITL could be applied to other types of dialogues successfully.

Such application, however, demands the collection of an immense corpus of target dialogues. It is difficult to expand the scope of application to encompass every kind of dialogue by carrying out this time-consuming task for each domain. No efficient way has yet been developed to widen the range of use for the speech translation system, such as making effective use of newspaper text corpora that are abundantly available. This Is a major challenger for researchers in this field.

In summary, seven years of research at ATR Interpreting Telecommunications Research Laboratories have yielded such great progress in speech translation technology that the current system can basically fulfill a limited task and perform as well as a Japanese person who is able to score 500 to 600 points on the TOEIC.

A closer look, however, reveals that the system can outperform a Japanese with a TOEIC score of 500 or more only when it deals with relatively simple utterances of small entropy. Its performance is less impressive when more complicated speech is involved. It is good at solving basic problems thanks to its tremendous memory, but it cannot fare so well when handling more complicated issues. Further effort is needed to improve the system's performance in such cases.

Considering that the transmission of information from Japan to overseas will be increasingly more important, there is a great need for more sophisticated speech translation technology, such as simultaneous speech translation technology, besides speech translation of daily conversation with injected questions. The realization of this simultaneous translation technology, which can also be used for monologues composed of complicated utterances like news broadcasts and lectures, is still being awaited.

History of Research Activities