ATR Interpreting Telecommunications Research Labs


Research into Integration of Speech and Language Processing

Integration of Translation Systems for Spontaneous Speech



Akio Yokoo



We have built an experimental speech translation system (ATR-MATRIX) for both Japanese to English and English to Japanese, based on the elemental technologies of speech recognition, language translation, and speech synthesis. We have proven the viability of the system in PC applications running almost in real-time, with speech translation on a scale of 13,000 words. We are also moving ahead with research into "dialogue management technologies" required to handle speech in normal conversation, including methods of distinguishing the meaning of utterances based on the situation; for example, through "speech acts". Through conversation experiments using bi-directional translation systems, we have demonstrated the feasibility of applications for cooperative conversations, including confirmations and repeated utterances, on specified topics.


1. Integration of Speech Translation Systems


We have built a Japanese-to-English and English-to-Japanese speech translation system that combines three elemental technologies | speech recognition, language translation, and speech synthesis | and also incorporates utterance division, prosody information extraction, and dialogue management technologies described below. The system is comprised of the speech recognition, language translation, and speech synthesis subsystems, in addition to a graphical user interface subsystem and communication control subsystem, with a main controller that controls the entire system. Connections are made via each satellite controller that adjusts the interface between the above subsystems and the main controller. Each of the systems for Japanese to English and English to Japanese runs on a workstation or high-end PC and achieves almost real-time processing.

From the results of a preliminary conversation experiment, we gained an understanding that to promote smooth conversations through a speech translation system, it is important that the system respond quickly, be able to repeat utterances, and be able to allow the other party to interrupt utterances. The system configuration facilitates a control method that takes these functions into consideration.


2. Technologies to Transform Utterance Units into Language Translation Units

Spontaneous conversions sometimes contain utterances in which two or more phrases are connected; for example, "That's a little expensive. Don't you have a cheaper room?" Even in these cases, it is necessary to translate each phrase correctly. In the case of these types of boundary units, a pause greater than a certain length is sometimes inserted, but this is not always the case. Here, we have proposed a method of transforming utterance units based on pause information, and the N-gram of fine-grained part-of-speech subcategories. By incorporating into the speech translation system a method that combines a statistical model and a few heuristics, we have confirmed the effectiveness of the proposed approach.


3. Prosody Extraction Technology

In spontaneous conversation, an interrogative is sometimes expressed by raising the pitch of the voice at the end of a sentence, as in the Japanese sentence "Heya, aitemasu?" [without the Japanese sentence-final question particle "ka"]. Because the system can detect the change in tone of the voice (prosody) and make a judgment as to whether or not the sentence is an interrogative, this information can be passed on to the language translation subsystem to allow a translation as "Are rooms available?" rather than "Rooms are available".

4. Dialogue Management Technologies

We have been moving ahead with research using information derived by managing the situations in which conversations occur in speech recognition and language translation technologies.

For example, after an utterance inquiring about a price, there is a high probability that the utterance following that inquiry will be a reply. In this way, by utilizing the relationship between the previous utterance and the current utterance in terms of the content words or the sentence-final expressions, we have proposed a method in which speech recognition results are reorganized in order of priority, giving precedence to candidates with a high level of contextual consistency.

It is also necessary to differentiate the meanings of a given utterance based on the situation, such as the presence of speech acts. For example, the Japanese positive response "Hai" may have to be interpreted as "yes" (acceptance) or "I see" (acknowledgment). We created a speech act database as a method of automatically recognizing these sorts of speech acts, and used this to develop a method of learning the occurrence probability for preset speech acts.


5. Evaluation of the Speech Translation System

We conducted an evaluation of the Japanese-to-English and English-to-Japanese speech translation system that uses the configuration described above. The evaluation was made from two points of view: 1) to what extent can interpretation technology support communication between different languages? And 2) to what extent can each utterance be appropriately translated? We have shown the feasibility of applications for cooperative conversations, including confirmations and repeated utterances, on specified topics. Results of evaluations relating to a hotel reservation task showed a level of satisfaction among users scoring 3.8 on a scale of 5, indicating that while some measure of dissatisfaction remains, the goal of the task had been adequately achieved. Furthermore, on comparing the translation results with those of human beings with varying English abilities as measured by TOEIC scores, we found that the system's capabilities for the task were roughly equivalent to those of a human being with a TOEIC score between 500 and 600.

There are two suggested themes for future study: research into more advanced integrated technologies for speech recognition, language translation, and speech synthesis for use in utterance situations; and research into multilingual, bi-directional speech translation technologies for use with a variety of topics.