


Enabling Global Communication Using Spoken Language
Dr. Victor Zue
MIT Laboratory for Computer Science Cambridge, MA 02139 USA
In the Beginning
Since its founding in 1986, the Advanced Telecommunications Research Institute
International, or ATR, has devoted itself to the advancement in telecommunications
and information technology the will improve the efficacy for humans to easily
access and exchange information of all forms. From the onset, ATR has placed heavy
emphasis on enabling human-human communication, realizing that, as geopolitical
barriers continue to fall and national economies become increasingly intertwined,
there is a growing need for people of the world to communicate with one another.
Since speech is the most natural, efficient, flexible, and effortless means for
human communication, technologies must be developed to enable humans speaking
different languages to communicate verbally. From 1986 to 1992, the ATR Interpreting
Telephony Research Laboratories conducted research in this all-important and very
difficult research area, culminating in a successful demonstration of three-way
(English-German-Japanese) speech-to-speech translation, conducted jointly with
Carnegie Mellon University, University of Karlsruhe, and Siemens.
While the trans-global demonstration was historic, researchers are fully aware
of the myriad of technical obstacles that must be overcome if the vision of enabling
verbal communication by speakers of different languages, each in their own tongue,
is to be realized. As a result, the ATR Interpreting Telecommunications Research
Laboratories, or ATR-ITL, was established in 1993 to continue this quest. Over
the past seven years, ATR-ITL has continued to conduct research in speech-to-speech
translation with emphasis on improving the underlying recognition, synthesis,
and translation technologies as well as addressing system integration issues.
This research has led to the development of a new generation system, the ATR Multilingual
Automatic Translation System for Information Exchange, or ATR-MATRIX. Compared
to its predecessors, ATR-MATRIX shows improved recognition of spontaneous speech,
more natural sounding synthetic speech through sophisticated prosodic modeling,
and more robust translation of everyday expressions.
My association with ATR started in early 1987 when I, together with five staff
members and graduate students, traveled to Osaka and taught a one-week spectrogram
reading course at ATR. Over the years, I taught the course once again in Kyoto.
As a matter or course, I have made my annual pilgrimage to ATR with other MIT
researchers to receive and present research updates, view demonstrations, and
exchange ideas. Similarly, researchers from ATR have visited my group frequently,
some for extended periods of time. From my perspective, our relationship with
ATR in general and ITL in particular has been long-standing and fruitful. It is
from this unique perspective that I offer my personal observations and comments
on its research activities.
Looking into the Past
Figure 1 illustrates the major processes
involved in speech-to-speech translation. The speech recognition sub-system converts
the speech signal in the source language into hypothesized word sequences. The
machine translation sub-system converts the word sequences from the source language
to the target language based on syntactic and semantic analyses. Finally, the
speech synthesis module takes the text produced by the translation module and
generates the speech signal in the target language.
If I were to characterize Phase One of the ATR speech translation effort as ground
breaking, in that it demonstrated the feasibility of speech-to-speech translation,
then ATR-MATRIX is best characterized as being robust and realistic. During the
past seven years, researchers at ATR have devoted considerable effort to developing
the technologies that will enable their system to deal with realistic input and
produce natural-sounding output. For the speech recognition sub-system, this means
the ability to deal with extemporaneously generated speech, spoken by many speakers.
Spontaneous speech is difficult to recognize and understand because it often contains
hesitations and false starts, and it may be agrammatical. For translation, it
means the ability to detect and utilize the acoustic cues that convey the subtle
nuances in conversational speech, much of it prosodically encoded. For synthesis,
it means the ability to produce natural sounding speech, in multiple languages
and representing multiple speakers. From a system's perspective, it means the
ability to provide near real-time response, and to continuously evaluate its performance
on a large body of data. Many of the accomplishments of ATR-ITL over the past
seven years can be found in journal articles and proceedings of international
conferences. In the remainder of this section, I will highlight a few of what
I consider to be the particularly noteworthy achievements in speech recognition
and synthesis.
Like many state-of-the-art speech recognition systems worldwide, the recognition
sub-system of ATR-MATRIX is speaker independent in that it can accommodate an
unknown speaker. A unique aspect of ATR-MATRIX's speech recognition sub-system
is its ability to improve performance through rapid speaker adaptation beyond
the traditional use of gender-specific models. The speech recognizer accomplishes
this by first creating a multitude of acoustic models based on the training data.
Through features that identify the unique characteristics of the speaker, the
system rapidly selects and utilizes the acoustic models of a speaker (or a set
of speakers) that best match the acoustic properties of the incoming speech. Considerable
research has been carried out at ITL on hierarchical speaker clustering, in which
successively more specific speaker models are created and organized in a tree
structure during training. Efficient search procedures have been proposed, in
which the speaker tree is traversed, from top to bottom, in order to find the
most appropriate models for acoustic decoding, using a maximum likelihood criterion.
In addition, instantaneous speaker adaptation through maximum a posteriori estimation,
using techniques such as incremental Bayesian learning and transfer vector field
smoothing, have been explored to modify the acoustic models using a small amount
of the input data.
Another area in which researchers at ITL have made significant and innovative
contributions is in the development of context-dependent acoustic models for hidden
Markov modeling (HMM). Specifically, they proposed the maximum likelihood based,
successive state splitting (ML-SSS) method, in which an HMM state can be split
to permit robust acoustic modeling of allophones or diphthongs, depending on the
statistical properties of the output distribution. This technique has the property
of being able to model one factor (such as the exact phonetic context) independent
of other factors (such as speaker variance). While it was originally proposed
for modeling allophonic variations using a single Gaussian, the technique has
since been refined to include tied-mixture Gaussian representations, and it has
been successfully applied to the modeling of DNA and protein sequences using HMM
methods.
If I were to point to one single technical accomplishment of ATR-ITL that has
had the greatest impact on the research community, it would be speech synthesis.
For more than a decade, researchers at ATR have been pursuing a corpus-based,
concatenative approach to speech synthesis. The key idea behind their approach
is the generation of synthetic speech by directly concatenating non-uniform waveform
segments selected from a large inventory of pre-recorded and properly annotated
units subject to a distortion measure. Over the years, the annotation of the speech
corpus has become increasingly complex, and it now includes acoustic-phonetic,
prosodic, and syntactic information, or even paralinguistic factors such as emotion.
Their research, and the subsequent demonstration of the superior quality of the
output speech in multiple languages produced by the CHATR synthesizer, has caused
a paradigm shift worldwide. Today, the corpus-based approach based on non-uniform
unit selection is being pursued actively in Asia, Europe, and North America, and
the resulting systems are out-performing the traditional rule-based systems. In
my opinion, ATR's contribution to speech synthesis in the form of non-uniform
unit concatenation is comparable in its impact to IBM's contribution to speech
recognition in the form of hidden Markov modeling.
In addition to the many technical accomplishments, of which only a small subset
have been summarized in this section, a number of other accomplishments should
also be mentioned. ATR-ITL was one of the three founding members of the Consortium
for Speech Translation Advanced Research (C-STAR). Today, C-STAR membership includes
six partners and fourteen affiliates from nine countries, many of whom have active
collaborations with ATR-ITL.
ATR-ITL has always been an international institution that draws its strength from
active and sizeable participation of foreign scholars. In may many visits to ATR,
I have always been pleasantly surprised by the mix of nationalities on board the
shuttle bus at the Takanohara train station. In fact, two of the eight research
scientists in my group at MIT have spent extended periods of time at ITL. I cannot
think of another speech research institution around the world that is more global
in its constituencies than ATR. Many of these scholars contribute actively while
at ATR, then serve as ambassadors for information dissemination upon departure.
Today, ATR-ITL alumni can be found in all corners of the world including China,
France, Germany, India, the United States, and the U.K.
What Holds in the Future
The new millenium witnessed the formation of the ATR Spoken Language Translation
Laboratories to carry out further research in speech-to-speech translation. I
have no doubt that this new laboratory will continue the outstanding tradition
established by its predecessors. I will conclude by offering my personal comments
regarding some of the research challenges that need to be met.
According to the Ethnologue, there are more than 6,700 languages being spoken
by nearly 6 billion people in some 228 countries. To be sure, the distribution
is highly skewed. For example, each of the top 100 languages is spoken by more
than 7 million people, and the top 10 languages | Chinese (Mandarin), Spanish,
English, Bengali, Hindi, Portuguese, Russian, Japanese, and Chinese (Wu) | are
spoken by nearly 40 percent of the world's population. Nevertheless, it is clear
that the task of enabling global speech-to-speech communication is daunting when
measured by its sheer magnitude. At the moment, success has only been demonstrated
for a small set of the languages, and as such in an extremely limited set of domains.
To achieve the vision of natural and effortless global human communication by
speech, we must begin to address the problem of portability, i.e., the development
of the knowledge and infrastructure that will enable us to rapidly develop speech-to-speech
translation systems for a new language pair in a new domain.
The problem facing the speech translation community is not unique. Speech recognition,
speech synthesis, and machine translation researchers are all too aware of the
fact that present day technologies are highly language and task dependent. For
example, the development of speech recognition and language understanding technologies
require a large amount of annotated training data. Therefore, we must learn to
produce systems in a new language and domain given at most a small amount of language-
and domain-specific training data. Meeting this challenge will require some significant
efforts in different directions. For example, we must strive to cleanly separate
the algorithmic aspects of the system from the language- and domain-specific aspects.
We must also develop automatic or semi-automatic methods, grammars, semantic structures
for language understanding, and dialogue models required by a new language and
application domain. The issue of portability spans across different acoustic environments,
databases, knowledge domains, and languages. Real deployment of speech translation
technology cannot take place without adequately addressing this issue.
In some cases, fundamentally different approaches may need to be proposed and
considered. In speech recognition, for example, the prevailing approaches typically
assume that a word can be represented as a sequence of phonemes. While increasingly
complex units (e.g., tri-phones) are being employed to capture the context dependency
of phonemes, these approaches cannot readily exploit constraints that are known
to exist for a given language at different parts of the linguistic hierarchy.
Besides, by focusing on words as the lexical units, the recognizer will depend
heavily on the particular application domain. A solution to this problem may lie
in an approach that separates the domain-dependent aspects of the process from
those that are domain independent. This could be achieved by decomposing the recognition
problem into two distinct stages, the first being domain independent but language
dependent, and the second dependent on both. This first stage may consist of a
subword-based recognition kernel that will accept as input the speech signal and
produce as output a graph of alternative subword units, utilizing acoustic and
language models that are unique for a given language. Such a subword recognition
kernel offers several advantages. By focusing on the recognition of subword units,
the recognition kernel can be domain independent, as long as its acoustic and
language models are trained on sufficient data in a given language from multiple
domains. By removing the reliance on word level language models, one should be
able to assess more accurately the contribution of the acoustic models, which
will hopefully lead to faster improvement. By introducing multiple levels of subword
representations, we should be able to capture more complex and long distance constraints
in a parsimonious manner. Since subword units form a closed set for a given language,
their recognition will not encounter new "words". Therefore, the issues of out-of-vocabulary
items are minimized. In fact, it can potentially also deal with disfluencies in
spontaneous speech by recognizing partial words.
While many machine translation efforts worldwide strive to achieve broad coverage,
performance of the resulting systems has been far from satisfactory. The need
for information access triggered by the popularity of the World Wide Web will
undoubtedly provide added impetus for rapid improvement. In the United States,
for example, a government effort called Trans-lingual Information Detection, Extraction,
and Summarization, or TIDES, has just been initiated. However, the problem of
translating verbal human-human interaction is further compounded by the fact that
the system must be able to overcome recognition and understanding errors, manage
both sides of the dialogue, and capture the nuances of human communication. As
a result, speech translation systems are likely to be successful only in narrow
domains for the foreseeable future. If this is indeed the case, then an inter-lingua
based approach, which relies on a common, language-independent meaning representation,
may deserve further investigation, as it offers the flexibility of translating
into multiple target languages.
At present, creating a robust speech translation system can require a tremendous
amount of effort on the part of researchers. In order for this technology to ultimately
be successful, the process of porting existing technology to new domains and languages
must be made easier. In the related field of spoken dialogue research, for example,
different research groups worldwide have been attempting to make it easier for
non-experts to create new domains. Systems which modularize their dialogue manager
try to take advantage of the fact that a dialogue can often be broken down into
a smaller set of sub-dialogues (e.g., dates, addresses), in order to make it easier
to construct dialogues for a new domain. Similar research may be needed in this
area if we are to allow speech translation systems with complex dialogue strategies
to generalize to different languages and domains.
Congratulations on your illustrious past, ATR-ITL! We look forward to many more
years of productive research to emerge from your descendents in the future.

