


ATR Interpreting Telecommunications Research Labs
Basic Research on Speech Translation Technologies
Seiichi Yamamoto
The ATR Interpreting Telecommunications Research Laboratories completed its project
involving basic research on speech translation technologies at the end of February
2000. NEC Corporation's trailblazing research and the basic research by ATR Interpreting
Telephony Research Laboratories, which were founded in the 1980s, had demonstrated
that speech translation could be realized, but the technologies developed in these
research projects, the grammar-based speech translation technologies, were not
able to properly translate naturally spoken language, such as simple daily conversations.
The ultimate goal of our research on speech translation technologies, built on
the foundations of the previous studies mentioned above, was to establish the
component technologies needed to translate natural spoken dialogue from daily
life. In other words, our aim was to develop the basic technologies of a speech
translation system that ordinary people without expert linguistic knowledge could
use without special training. For this purpose, a huge corpus of data was collected
from everyday dialogues, and a corpus-based approach was adopted for research
on each component technology. During the first four years of the research, our
efforts were focused on the component technologies necessary for recognizing,
translating, and synthesizing natural spoken language. Then, the next three years
were devoted no integrating these technologies into one system so that a comprehensive
evaluation could be made. Since the research results for each component technology
are described in detail by each department head in the following columns, this
paper concentrates on the outcome of research on integrated speech translation
technologies.
A bi-directional speech translation system was necessary to evaluate the preliminary
technologies developed for translating spoken dialogues between two different
languages and to grasp related technological problems. In 1998, this imperative
led us to integrate the component technologies into a Japanese-English bi-directional
speech translation system with a vocabulary of over 10,000 words for dialogues
on travel reservations. We evaluated the system from various viewpoints, and the
evaluations confirmed that the system was capable of real-time speech translation
for travel bookings. The actual performance achieved an average subjective rating
of 3.8 on a 1-to-5 scale, which indicated that the system was able to fulfill
the task most of the time. By comparing the quality of the speech translation
with the translation performance of many Japanese individuals at various levels
of English, we found that the current machine translation was roughly equivalent
in quality to the ability of a Japanese person who scores 500 to 600 points on
the TOEIC (Test of English for International Communication). This figure is quite
impressive considering the fact that the average TOEIC score of Japanese university
students is about 570.
Fourteen years ago when ATR Interpreting Telephony Research Laboratories were
established, an automatic translation telephone was still a dream. The current
capability of our technology has reached the level of a Japanese person who has
been studying English for nearly 10 years, although our system's application is
limited to travel reservations.
When the technology developed by ATR Interpreting Telecommunications Research
Laboratories for travel reservations was applied to the task of answering questions
about area codes and so on, it demonstrated similar levels of translation performance.
The results left no doubt that the speech translation technology developed by
ATR ITL could be applied to other types of dialogues successfully.
Such application, however, demands the collection of an immense corpus of target
dialogues. It is difficult to expand the scope of application to encompass every
kind of dialogue by carrying out this time-consuming task for each domain. No
efficient way has yet been developed to widen the range of use for the speech
translation system, such as making effective use of newspaper text corpora that
are abundantly available. This Is a major challenger for researchers in this field.
In summary, seven years of research at ATR Interpreting Telecommunications Research
Laboratories have yielded such great progress in speech translation technology
that the current system can basically fulfill a limited task and perform as well
as a Japanese person who is able to score 500 to 600 points on the TOEIC.
A closer look, however, reveals that the system can outperform a Japanese with
a TOEIC score of 500 or more only when it deals with relatively simple utterances
of small entropy. Its performance is less impressive when more complicated speech
is involved. It is good at solving basic problems thanks to its tremendous memory,
but it cannot fare so well when handling more complicated issues. Further effort
is needed to improve the system's performance in such cases.
Considering that the transmission of information from Japan to overseas will be
increasingly more important, there is a great need for more sophisticated speech
translation technology, such as simultaneous speech translation technology, besides
speech translation of daily conversation with injected questions. The realization
of this simultaneous translation technology, which can also be used for monologues
composed of complicated utterances like news broadcasts and lectures, is still
being awaited.
History of Research Activities

