Enabling Global Communication Using Spoken Language




Dr. Victor Zue
MIT Laboratory for Computer Science Cambridge, MA 02139 USA



In the Beginning

Since its founding in 1986, the Advanced Telecommunications Research Institute International, or ATR, has devoted itself to the advancement in telecommunications and information technology the will improve the efficacy for humans to easily access and exchange information of all forms. From the onset, ATR has placed heavy emphasis on enabling human-human communication, realizing that, as geopolitical barriers continue to fall and national economies become increasingly intertwined, there is a growing need for people of the world to communicate with one another. Since speech is the most natural, efficient, flexible, and effortless means for human communication, technologies must be developed to enable humans speaking different languages to communicate verbally. From 1986 to 1992, the ATR Interpreting Telephony Research Laboratories conducted research in this all-important and very difficult research area, culminating in a successful demonstration of three-way (English-German-Japanese) speech-to-speech translation, conducted jointly with Carnegie Mellon University, University of Karlsruhe, and Siemens.

While the trans-global demonstration was historic, researchers are fully aware of the myriad of technical obstacles that must be overcome if the vision of enabling verbal communication by speakers of different languages, each in their own tongue, is to be realized. As a result, the ATR Interpreting Telecommunications Research Laboratories, or ATR-ITL, was established in 1993 to continue this quest. Over the past seven years, ATR-ITL has continued to conduct research in speech-to-speech translation with emphasis on improving the underlying recognition, synthesis, and translation technologies as well as addressing system integration issues. This research has led to the development of a new generation system, the ATR Multilingual Automatic Translation System for Information Exchange, or ATR-MATRIX. Compared to its predecessors, ATR-MATRIX shows improved recognition of spontaneous speech, more natural sounding synthetic speech through sophisticated prosodic modeling, and more robust translation of everyday expressions.

My association with ATR started in early 1987 when I, together with five staff members and graduate students, traveled to Osaka and taught a one-week spectrogram reading course at ATR. Over the years, I taught the course once again in Kyoto. As a matter or course, I have made my annual pilgrimage to ATR with other MIT researchers to receive and present research updates, view demonstrations, and exchange ideas. Similarly, researchers from ATR have visited my group frequently, some for extended periods of time. From my perspective, our relationship with ATR in general and ITL in particular has been long-standing and fruitful. It is from this unique perspective that I offer my personal observations and comments on its research activities.


Looking into the Past


Figure 1 illustrates the major processes involved in speech-to-speech translation. The speech recognition sub-system converts the speech signal in the source language into hypothesized word sequences. The machine translation sub-system converts the word sequences from the source language to the target language based on syntactic and semantic analyses. Finally, the speech synthesis module takes the text produced by the translation module and generates the speech signal in the target language.

If I were to characterize Phase One of the ATR speech translation effort as ground breaking, in that it demonstrated the feasibility of speech-to-speech translation, then ATR-MATRIX is best characterized as being robust and realistic. During the past seven years, researchers at ATR have devoted considerable effort to developing the technologies that will enable their system to deal with realistic input and produce natural-sounding output. For the speech recognition sub-system, this means the ability to deal with extemporaneously generated speech, spoken by many speakers. Spontaneous speech is difficult to recognize and understand because it often contains hesitations and false starts, and it may be agrammatical. For translation, it means the ability to detect and utilize the acoustic cues that convey the subtle nuances in conversational speech, much of it prosodically encoded. For synthesis, it means the ability to produce natural sounding speech, in multiple languages and representing multiple speakers. From a system's perspective, it means the ability to provide near real-time response, and to continuously evaluate its performance on a large body of data. Many of the accomplishments of ATR-ITL over the past seven years can be found in journal articles and proceedings of international conferences. In the remainder of this section, I will highlight a few of what I consider to be the particularly noteworthy achievements in speech recognition and synthesis.

Like many state-of-the-art speech recognition systems worldwide, the recognition sub-system of ATR-MATRIX is speaker independent in that it can accommodate an unknown speaker. A unique aspect of ATR-MATRIX's speech recognition sub-system is its ability to improve performance through rapid speaker adaptation beyond the traditional use of gender-specific models. The speech recognizer accomplishes this by first creating a multitude of acoustic models based on the training data. Through features that identify the unique characteristics of the speaker, the system rapidly selects and utilizes the acoustic models of a speaker (or a set of speakers) that best match the acoustic properties of the incoming speech. Considerable research has been carried out at ITL on hierarchical speaker clustering, in which successively more specific speaker models are created and organized in a tree structure during training. Efficient search procedures have been proposed, in which the speaker tree is traversed, from top to bottom, in order to find the most appropriate models for acoustic decoding, using a maximum likelihood criterion. In addition, instantaneous speaker adaptation through maximum a posteriori estimation, using techniques such as incremental Bayesian learning and transfer vector field smoothing, have been explored to modify the acoustic models using a small amount of the input data.

Another area in which researchers at ITL have made significant and innovative contributions is in the development of context-dependent acoustic models for hidden Markov modeling (HMM). Specifically, they proposed the maximum likelihood based, successive state splitting (ML-SSS) method, in which an HMM state can be split to permit robust acoustic modeling of allophones or diphthongs, depending on the statistical properties of the output distribution. This technique has the property of being able to model one factor (such as the exact phonetic context) independent of other factors (such as speaker variance). While it was originally proposed for modeling allophonic variations using a single Gaussian, the technique has since been refined to include tied-mixture Gaussian representations, and it has been successfully applied to the modeling of DNA and protein sequences using HMM methods.

If I were to point to one single technical accomplishment of ATR-ITL that has had the greatest impact on the research community, it would be speech synthesis. For more than a decade, researchers at ATR have been pursuing a corpus-based, concatenative approach to speech synthesis. The key idea behind their approach is the generation of synthetic speech by directly concatenating non-uniform waveform segments selected from a large inventory of pre-recorded and properly annotated units subject to a distortion measure. Over the years, the annotation of the speech corpus has become increasingly complex, and it now includes acoustic-phonetic, prosodic, and syntactic information, or even paralinguistic factors such as emotion. Their research, and the subsequent demonstration of the superior quality of the output speech in multiple languages produced by the CHATR synthesizer, has caused a paradigm shift worldwide. Today, the corpus-based approach based on non-uniform unit selection is being pursued actively in Asia, Europe, and North America, and the resulting systems are out-performing the traditional rule-based systems. In my opinion, ATR's contribution to speech synthesis in the form of non-uniform unit concatenation is comparable in its impact to IBM's contribution to speech recognition in the form of hidden Markov modeling.

In addition to the many technical accomplishments, of which only a small subset have been summarized in this section, a number of other accomplishments should also be mentioned. ATR-ITL was one of the three founding members of the Consortium for Speech Translation Advanced Research (C-STAR). Today, C-STAR membership includes six partners and fourteen affiliates from nine countries, many of whom have active collaborations with ATR-ITL.

ATR-ITL has always been an international institution that draws its strength from active and sizeable participation of foreign scholars. In may many visits to ATR, I have always been pleasantly surprised by the mix of nationalities on board the shuttle bus at the Takanohara train station. In fact, two of the eight research scientists in my group at MIT have spent extended periods of time at ITL. I cannot think of another speech research institution around the world that is more global in its constituencies than ATR. Many of these scholars contribute actively while at ATR, then serve as ambassadors for information dissemination upon departure. Today, ATR-ITL alumni can be found in all corners of the world including China, France, Germany, India, the United States, and the U.K.


What Holds in the Future

The new millenium witnessed the formation of the ATR Spoken Language Translation Laboratories to carry out further research in speech-to-speech translation. I have no doubt that this new laboratory will continue the outstanding tradition established by its predecessors. I will conclude by offering my personal comments regarding some of the research challenges that need to be met.

According to the Ethnologue, there are more than 6,700 languages being spoken by nearly 6 billion people in some 228 countries. To be sure, the distribution is highly skewed. For example, each of the top 100 languages is spoken by more than 7 million people, and the top 10 languages | Chinese (Mandarin), Spanish, English, Bengali, Hindi, Portuguese, Russian, Japanese, and Chinese (Wu) | are spoken by nearly 40 percent of the world's population. Nevertheless, it is clear that the task of enabling global speech-to-speech communication is daunting when measured by its sheer magnitude. At the moment, success has only been demonstrated for a small set of the languages, and as such in an extremely limited set of domains. To achieve the vision of natural and effortless global human communication by speech, we must begin to address the problem of portability, i.e., the development of the knowledge and infrastructure that will enable us to rapidly develop speech-to-speech translation systems for a new language pair in a new domain.

The problem facing the speech translation community is not unique. Speech recognition, speech synthesis, and machine translation researchers are all too aware of the fact that present day technologies are highly language and task dependent. For example, the development of speech recognition and language understanding technologies require a large amount of annotated training data. Therefore, we must learn to produce systems in a new language and domain given at most a small amount of language- and domain-specific training data. Meeting this challenge will require some significant efforts in different directions. For example, we must strive to cleanly separate the algorithmic aspects of the system from the language- and domain-specific aspects. We must also develop automatic or semi-automatic methods, grammars, semantic structures for language understanding, and dialogue models required by a new language and application domain. The issue of portability spans across different acoustic environments, databases, knowledge domains, and languages. Real deployment of speech translation technology cannot take place without adequately addressing this issue.

In some cases, fundamentally different approaches may need to be proposed and considered. In speech recognition, for example, the prevailing approaches typically assume that a word can be represented as a sequence of phonemes. While increasingly complex units (e.g., tri-phones) are being employed to capture the context dependency of phonemes, these approaches cannot readily exploit constraints that are known to exist for a given language at different parts of the linguistic hierarchy. Besides, by focusing on words as the lexical units, the recognizer will depend heavily on the particular application domain. A solution to this problem may lie in an approach that separates the domain-dependent aspects of the process from those that are domain independent. This could be achieved by decomposing the recognition problem into two distinct stages, the first being domain independent but language dependent, and the second dependent on both. This first stage may consist of a subword-based recognition kernel that will accept as input the speech signal and produce as output a graph of alternative subword units, utilizing acoustic and language models that are unique for a given language. Such a subword recognition kernel offers several advantages. By focusing on the recognition of subword units, the recognition kernel can be domain independent, as long as its acoustic and language models are trained on sufficient data in a given language from multiple domains. By removing the reliance on word level language models, one should be able to assess more accurately the contribution of the acoustic models, which will hopefully lead to faster improvement. By introducing multiple levels of subword representations, we should be able to capture more complex and long distance constraints in a parsimonious manner. Since subword units form a closed set for a given language, their recognition will not encounter new "words". Therefore, the issues of out-of-vocabulary items are minimized. In fact, it can potentially also deal with disfluencies in spontaneous speech by recognizing partial words.

While many machine translation efforts worldwide strive to achieve broad coverage, performance of the resulting systems has been far from satisfactory. The need for information access triggered by the popularity of the World Wide Web will undoubtedly provide added impetus for rapid improvement. In the United States, for example, a government effort called Trans-lingual Information Detection, Extraction, and Summarization, or TIDES, has just been initiated. However, the problem of translating verbal human-human interaction is further compounded by the fact that the system must be able to overcome recognition and understanding errors, manage both sides of the dialogue, and capture the nuances of human communication. As a result, speech translation systems are likely to be successful only in narrow domains for the foreseeable future. If this is indeed the case, then an inter-lingua based approach, which relies on a common, language-independent meaning representation, may deserve further investigation, as it offers the flexibility of translating into multiple target languages.

At present, creating a robust speech translation system can require a tremendous amount of effort on the part of researchers. In order for this technology to ultimately be successful, the process of porting existing technology to new domains and languages must be made easier. In the related field of spoken dialogue research, for example, different research groups worldwide have been attempting to make it easier for non-experts to create new domains. Systems which modularize their dialogue manager try to take advantage of the fact that a dialogue can often be broken down into a smaller set of sub-dialogues (e.g., dates, addresses), in order to make it easier to construct dialogues for a new domain. Similar research may be needed in this area if we are to allow speech translation systems with complex dialogue strategies to generalize to different languages and domains.

Congratulations on your illustrious past, ATR-ITL! We look forward to many more years of productive research to emerge from your descendents in the future.