


ATR Interpreting Telecommunications Research Labs
C-STAR International Joint Experiment toward Interpreting
Telecommunications
Akio Yokoo, Toshiyuki Takezawa, and Seiichi Yamamoto
Summary
ATR Interpreting Telecommunications Research Laboratories conducted an experiment
at Keihanna Plaza at Seika-cho, Kyoto, on Thursday July 22, 1999 to translate
spoken conversations in four languages including Japanese, English, German, and
Korean via connections with members of C-STAR (Consortium for Speech Translation
Advanced Research). In the experiment, conversations were exchanged among multiple
different languages while simulating a hotel reservation task. This paper gives
a research summary, experiment results, and suggests future work for developing
speech translation.
1. Targeting the Automatic Translation of Spoken Conversations
Communication among people using different languages is expected to broadly increase
in the 21st century. Accordingly, there will be demands for speech translation
technology capable of automatically translating conversations as a communication-aiding
technology.
The experiment described here was the second international joint experiment, following
the one conducted by ATR Interpreting Telephony Research Laboratories in January
1993, and was conducted to verify the technological progress made during the past
six years. In this speech translation experiment, conducted through connections
among multiple different languages, we hoped to achieve speech translation of
conversations in a natural manner similar to human behavior.
2. C-STAR Consortium for Speech Translation Advanced
Research
The progress made in basic research on speech translation technology by ATR Interpreting
Telephony research Laboratories, which was founded in 1986, has encouraged several
other research organizations to start studies in speech translation. The international
consortium C-STAR was founded in 1992 with ATR playing the lead role in advancing
research on speech translation. C-STAR was set up to promote the exchange of research
results among the participating research organizations, which, in addition to
ATR, included Carnegie Mellon University in Pittsburgh, USA, Siemens Co. in Munich,
Germany, and the University of Karlsruhe in Karlsruhe, Germany. These initial
members of C-STAR conducted an international joint experiment on speech translation
for the first time in the world in January 1993 in an attempt to publicize the
need for research activities in speech translation.
Following the foundation in 1993 of ATR Interpreting Telecommunications Research
Laboratories, the successor organization of the Interpreting Telephony research
Laboratories, the consortium was reorganized in 1994 with the four original research
organizations as its core members. The 2nd phase of C-STAR activities culminated
with a summary in 1999 that focused the consortium's research results to efficiently
advance cooperation in studies of real-time translation technologies for spontaneous
speech.
Research organizations that agreed to join the next international joint experiment
in 1999, enlarged to compile more research results, were the partner members.
The Electronics and Telecommunications research Institute (ETRI) in Taejon, Korea;
IRST in Trento, Italy; and CLIPS in Grenoble, France, later became partner members.
To efficiently advance research, C-STAR allocated to the research organizations
of each nation the responsibility for speech recognition and speech synthesis
in its own language and for translation of its own language into other languages.
C-STAR has operated as the hub of all these efforts toward communication among
multiple languages by compiling the research results achieved by the partner members.
3. Contiguration and Features of the Experimental System
ATR-MATRIX,1,2,3
ATR's speech translation system, can translate spontaneous Japanese speech into
English, German, and Korean. The system can recognize and translate spoken Japanese
even with interjections such as "er" ("anoh" in Japanese) and "let's see" ("ehto"
in Japanese), expressions characteristic of conversations such as "I have no choice"
("shoganai desu ne" in Japanese), or conversations with Japanese postpositional
particles dropped out. It can recognize and translate speech of about 13,000 vocabulary
words, which is sufficient for carrying out most conversations about hotel reservations
and travel-related tasks except for proper nouns such as the names of people.
The ATR-MATRIX system can be run on a personal computer or a workstation no process
translations Iin nearly real-time, and could even be done with a pocket-sized
or notebook-sized computer if the system configuration were changed.
Figure 1 shows the system configuration
used in a speech translation experiment with simultaneous connections among multiple
different languages. In the experiment, the speech translation systems at ATR,
Carnegie Mellon University, the University of Karlsruhe, and the Electronics and
Telecommunications Research Institute (ETRI) were simultaneously connected via
a 64-kbps ISDN system. The experiment went as follows:
Communications were made via a communication server dedicated to the experiment
to efficiently transmit data among the systems of the connected research organizations.
The speech translation system of each nation recognized speech in its own language,
translated the recognized speech to languages of other nations, and then transmitted
the translated results to the communication server in text form. The communication
server in text form. The communication server in turn sent the received text data
to the systems of the other nations, which retrieved only the information pertinent
to that nation and output the information translated in the nation's language
using synthesized voice.
Video conferencing systems displaying speakers and conference sites were also
used in the experiment by connecting the systems via a 384-kbps ISDN system.
4. Speech Translation Experiment
The overall experiment consisted of a speech translation component with multiple
different languages simultaneously connected and an experiment to translate spontaneous
speech.
The speech translation experiment showed that speech given in Japanese by a Japanese
was successfully translated into English, German, and Korean, while speech given
in English was successfully translated simultaneously into Japanese, German, and
Korean. The speech sample data in the experiment mainly included daily greeting
expressions such as "What time is it in America now?", "What time is it in Germany?",
and "We have fine weather in Japan now. How about in Korea?".
The photograph in figure 2 shows a scene
from the experiment. The experiment to translate spontaneous speech was conducted
between ATR and Carnegie Mellon University. We simulated a scene of a Japanese
visiting the USA to watch World Series baseball games and trying to reserve a
hotel room in New York through a travel agent. We had some speech recognition
errors during the experiment, but succeeded in confirming the translation of conversations
by repeating parts of the conversations and asking questions for confirmation.
We later conducted an experiment on speech translation by changing the configuration
of ATR's Japanese-English interactive speech translation system4
in two ways. First, we included the use of a wearable client computer as a necessary
part of the speech recognition and speech synthesis. Second, we transmitted data
via a wireless LAN to a desktop personal computer used as the speech translation
server. The conversations in the experiment were simulations of a scene of an
English-speaking person losing his way and asking a Japanese-speaking person for
directions. We succeeded in showing the viability of the translation with a portable
terminal.
We finally conducted an experiment on speech translation by using the system with
a front desk clerk at the Keihanna-Miyako Hotel and consequently proving that
the general public could use it.
5. Related Information and Future Work
About 100 people visited the experiment from public agencies, universities, and
various industries in addition to representatives of the mass media (TV and newspapers).
In the Q&A session, someone asked, "How long will it take before the system is
put to practical use?" We replied, "It is difficult to predict, as the time required
until practical use will vary depending on how it is used, but use of the system
as proven in the experiment is not too far away". Recent surveys5
have suggested that practical use will be achieved between the years 2010 and
2020. Another person asked, "Can we eventually discuss things freely with foreigners?"
to which we answered "Discussions on daily subjects will be possible, but with
such a corpus-based system such as that of ATR it will be difficult to discuss
specialized subjects or subjects involving extraordinary logic".
From now, we will proceed with investigations into the changes in human behavior
that occur during communications with people speaking different languages via
speech translation systems. We will then proceed with performance evaluations
and identify problems to be solved in running concrete applications of these systems.
Readers interested in the ATR-MATRIX speech translation system can access our
WWW site at the URL below for a summary description of the system: http://www.itl.atr.co.jp/matrix/
References

