


ATR Interpreting Telecommunications Research Labs
Research into Integration of Speech and Language Processing
Integration of Translation Systems for Spontaneous Speech
Akio Yokoo
We have built an experimental speech translation system (ATR-MATRIX) for both
Japanese to English and English to Japanese, based on the elemental technologies
of speech recognition, language translation, and speech synthesis. We have proven
the viability of the system in PC applications running almost in real-time, with
speech translation on a scale of 13,000 words. We are also moving ahead with research
into "dialogue management technologies" required to handle speech in normal conversation,
including methods of distinguishing the meaning of utterances based on the situation;
for example, through "speech acts". Through conversation experiments using bi-directional
translation systems, we have demonstrated the feasibility of applications for
cooperative conversations, including confirmations and repeated utterances, on
specified topics.
1. Integration of Speech Translation Systems
We have built a Japanese-to-English and English-to-Japanese speech translation
system that combines three elemental technologies | speech recognition, language
translation, and speech synthesis | and also incorporates utterance division,
prosody information extraction, and dialogue management technologies described
below. The system is comprised of the speech recognition, language translation,
and speech synthesis subsystems, in addition to a graphical user interface subsystem
and communication control subsystem, with a main controller that controls the
entire system. Connections are made via each satellite controller that adjusts
the interface between the above subsystems and the main controller. Each of the
systems for Japanese to English and English to Japanese runs on a workstation
or high-end PC and achieves almost real-time processing.
From the results of a preliminary conversation experiment, we gained an understanding
that to promote smooth conversations through a speech translation system, it is
important that the system respond quickly, be able to repeat utterances, and be
able to allow the other party to interrupt utterances. The system configuration
facilitates a control method that takes these functions into consideration.
2. Technologies to Transform Utterance Units into Language
Translation Units
Spontaneous conversions sometimes contain utterances in which two or more phrases
are connected; for example, "That's a little expensive. Don't you have a cheaper
room?" Even in these cases, it is necessary to translate each phrase correctly.
In the case of these types of boundary units, a pause greater than a certain length
is sometimes inserted, but this is not always the case. Here, we have proposed
a method of transforming utterance units based on pause information, and the N-gram
of fine-grained part-of-speech subcategories. By incorporating into the speech
translation system a method that combines a statistical model and a few heuristics,
we have confirmed the effectiveness of the proposed approach.
3. Prosody Extraction Technology
In spontaneous conversation, an interrogative is sometimes expressed by raising
the pitch of the voice at the end of a sentence, as in the Japanese sentence "Heya,
aitemasu?" [without the Japanese sentence-final question particle "ka"]. Because
the system can detect the change in tone of the voice (prosody) and make a judgment
as to whether or not the sentence is an interrogative, this information can be
passed on to the language translation subsystem to allow a translation as "Are
rooms available?" rather than "Rooms are available".
4. Dialogue Management Technologies
We have been moving ahead with research using information derived by managing
the situations in which conversations occur in speech recognition and language
translation technologies.
For example, after an utterance inquiring about a price, there is a high probability
that the utterance following that inquiry will be a reply. In this way, by utilizing
the relationship between the previous utterance and the current utterance in terms
of the content words or the sentence-final expressions, we have proposed a method
in which speech recognition results are reorganized in order of priority, giving
precedence to candidates with a high level of contextual consistency.
It is also necessary to differentiate the meanings of a given utterance based
on the situation, such as the presence of speech acts. For example, the Japanese
positive response "Hai" may have to be interpreted as "yes" (acceptance) or "I
see" (acknowledgment). We created a speech act database as a method of automatically
recognizing these sorts of speech acts, and used this to develop a method of learning
the occurrence probability for preset speech acts.
5. Evaluation of the Speech Translation System
We conducted an evaluation of the Japanese-to-English and English-to-Japanese
speech translation system that uses the configuration described above. The evaluation
was made from two points of view: 1) to what extent can interpretation technology
support communication between different languages? And 2) to what extent can each
utterance be appropriately translated? We have shown the feasibility of applications
for cooperative conversations, including confirmations and repeated utterances,
on specified topics. Results of evaluations relating to a hotel reservation task
showed a level of satisfaction among users scoring 3.8 on a scale of 5, indicating
that while some measure of dissatisfaction remains, the goal of the task had been
adequately achieved. Furthermore, on comparing the translation results with those
of human beings with varying English abilities as measured by TOEIC scores, we
found that the system's capabilities for the task were roughly equivalent to those
of a human being with a TOEIC score between 500 and 600.
There are two suggested themes for future study: research into more advanced integrated
technologies for speech recognition, language translation, and speech synthesis
for use in utterance situations; and research into multilingual, bi-directional
speech translation technologies for use with a variety of topics.

