ATR Interpreting Telecommunications Research Labs


C-STAR International Joint Experiment toward Interpreting Telecommunications




Akio Yokoo, Toshiyuki Takezawa, and Seiichi Yamamoto



Summary

ATR Interpreting Telecommunications Research Laboratories conducted an experiment at Keihanna Plaza at Seika-cho, Kyoto, on Thursday July 22, 1999 to translate spoken conversations in four languages including Japanese, English, German, and Korean via connections with members of C-STAR (Consortium for Speech Translation Advanced Research). In the experiment, conversations were exchanged among multiple different languages while simulating a hotel reservation task. This paper gives a research summary, experiment results, and suggests future work for developing speech translation.


1. Targeting the Automatic Translation of Spoken Conversations


Communication among people using different languages is expected to broadly increase in the 21st century. Accordingly, there will be demands for speech translation technology capable of automatically translating conversations as a communication-aiding technology.

The experiment described here was the second international joint experiment, following the one conducted by ATR Interpreting Telephony Research Laboratories in January 1993, and was conducted to verify the technological progress made during the past six years. In this speech translation experiment, conducted through connections among multiple different languages, we hoped to achieve speech translation of conversations in a natural manner similar to human behavior.


2. C-STAR Consortium for Speech Translation Advanced Research

The progress made in basic research on speech translation technology by ATR Interpreting Telephony research Laboratories, which was founded in 1986, has encouraged several other research organizations to start studies in speech translation. The international consortium C-STAR was founded in 1992 with ATR playing the lead role in advancing research on speech translation. C-STAR was set up to promote the exchange of research results among the participating research organizations, which, in addition to ATR, included Carnegie Mellon University in Pittsburgh, USA, Siemens Co. in Munich, Germany, and the University of Karlsruhe in Karlsruhe, Germany. These initial members of C-STAR conducted an international joint experiment on speech translation for the first time in the world in January 1993 in an attempt to publicize the need for research activities in speech translation.

Following the foundation in 1993 of ATR Interpreting Telecommunications Research Laboratories, the successor organization of the Interpreting Telephony research Laboratories, the consortium was reorganized in 1994 with the four original research organizations as its core members. The 2nd phase of C-STAR activities culminated with a summary in 1999 that focused the consortium's research results to efficiently advance cooperation in studies of real-time translation technologies for spontaneous speech.

Research organizations that agreed to join the next international joint experiment in 1999, enlarged to compile more research results, were the partner members. The Electronics and Telecommunications research Institute (ETRI) in Taejon, Korea; IRST in Trento, Italy; and CLIPS in Grenoble, France, later became partner members.

To efficiently advance research, C-STAR allocated to the research organizations of each nation the responsibility for speech recognition and speech synthesis in its own language and for translation of its own language into other languages. C-STAR has operated as the hub of all these efforts toward communication among multiple languages by compiling the research results achieved by the partner members.


3. Contiguration and Features of the Experimental System

ATR-MATRIX,1,2,3 ATR's speech translation system, can translate spontaneous Japanese speech into English, German, and Korean. The system can recognize and translate spoken Japanese even with interjections such as "er" ("anoh" in Japanese) and "let's see" ("ehto" in Japanese), expressions characteristic of conversations such as "I have no choice" ("shoganai desu ne" in Japanese), or conversations with Japanese postpositional particles dropped out. It can recognize and translate speech of about 13,000 vocabulary words, which is sufficient for carrying out most conversations about hotel reservations and travel-related tasks except for proper nouns such as the names of people.

The ATR-MATRIX system can be run on a personal computer or a workstation no process translations Iin nearly real-time, and could even be done with a pocket-sized or notebook-sized computer if the system configuration were changed.

Figure 1 shows the system configuration used in a speech translation experiment with simultaneous connections among multiple different languages. In the experiment, the speech translation systems at ATR, Carnegie Mellon University, the University of Karlsruhe, and the Electronics and Telecommunications Research Institute (ETRI) were simultaneously connected via a 64-kbps ISDN system. The experiment went as follows:
Communications were made via a communication server dedicated to the experiment to efficiently transmit data among the systems of the connected research organizations. The speech translation system of each nation recognized speech in its own language, translated the recognized speech to languages of other nations, and then transmitted the translated results to the communication server in text form. The communication server in text form. The communication server in turn sent the received text data to the systems of the other nations, which retrieved only the information pertinent to that nation and output the information translated in the nation's language using synthesized voice.

Video conferencing systems displaying speakers and conference sites were also used in the experiment by connecting the systems via a 384-kbps ISDN system.


4. Speech Translation Experiment


The overall experiment consisted of a speech translation component with multiple different languages simultaneously connected and an experiment to translate spontaneous speech.

The speech translation experiment showed that speech given in Japanese by a Japanese was successfully translated into English, German, and Korean, while speech given in English was successfully translated simultaneously into Japanese, German, and Korean. The speech sample data in the experiment mainly included daily greeting expressions such as "What time is it in America now?", "What time is it in Germany?", and "We have fine weather in Japan now. How about in Korea?".

The photograph in figure 2 shows a scene from the experiment. The experiment to translate spontaneous speech was conducted between ATR and Carnegie Mellon University. We simulated a scene of a Japanese visiting the USA to watch World Series baseball games and trying to reserve a hotel room in New York through a travel agent. We had some speech recognition errors during the experiment, but succeeded in confirming the translation of conversations by repeating parts of the conversations and asking questions for confirmation.

We later conducted an experiment on speech translation by changing the configuration of ATR's Japanese-English interactive speech translation system4 in two ways. First, we included the use of a wearable client computer as a necessary part of the speech recognition and speech synthesis. Second, we transmitted data via a wireless LAN to a desktop personal computer used as the speech translation server. The conversations in the experiment were simulations of a scene of an English-speaking person losing his way and asking a Japanese-speaking person for directions. We succeeded in showing the viability of the translation with a portable terminal.

We finally conducted an experiment on speech translation by using the system with a front desk clerk at the Keihanna-Miyako Hotel and consequently proving that the general public could use it.


5. Related Information and Future Work

About 100 people visited the experiment from public agencies, universities, and various industries in addition to representatives of the mass media (TV and newspapers). In the Q&A session, someone asked, "How long will it take before the system is put to practical use?" We replied, "It is difficult to predict, as the time required until practical use will vary depending on how it is used, but use of the system as proven in the experiment is not too far away". Recent surveys5 have suggested that practical use will be achieved between the years 2010 and 2020. Another person asked, "Can we eventually discuss things freely with foreigners?" to which we answered "Discussions on daily subjects will be possible, but with such a corpus-based system such as that of ATR it will be difficult to discuss specialized subjects or subjects involving extraordinary logic".

From now, we will proceed with investigations into the changes in human behavior that occur during communications with people speaking different languages via speech translation systems. We will then proceed with performance evaluations and identify problems to be solved in running concrete applications of these systems.

Readers interested in the ATR-MATRIX speech translation system can access our WWW site at the URL below for a summary description of the system: http://www.itl.atr.co.jp/matrix/


References