ATR Interpreting Telecommunications Research Labs


Prosody Interpretation and Speech Synthesis




Nick Campbell



Spoken language is not usually as carefully prepared as written text, and relies instead on prosody, or inflexion of the voice, to carry the meaning. Prosody allows the listener to interpret the speaker's intention in an appropriate way by indicating the meaning-groups, the focus of discourse, and the relations holding between the words and phrases in the speech. For the past seven years, Department 2 ATR-ITL has been researching the automatic processing of this linguistic and para-linguistic information under the PI and the PEGASUS projects.

The PI (or Prosody Interpretation) project featured the analysis, transmission, and generation of prosodic information in speech. It models the variations in tone-of-voice, in speaking rate, and in loudness, and maps between changes in the acoustic signal and changes in the intended meaning of groups of words in the speech. This work resulted in a set of algorithms for the extraction of speech-act, emphasis, and phrasing information from the spoken utterance input to the speech translation system, for the transfer of this information along with the translated result, and for its use as enriched input to the speech synthesiser to ensure that the translated utterance could be produced with the appropriate meaning and that ambiguities in the translated word-sequence could be minimized. However, the constrained style and short length of the utterance typical of the hotel-reservation task, which forms the common core of ITL research, make little use of such prosodic information and, apart from a simple statement/question detection module, the PI research has not yet been incorporated into the MATRIX translation system.

The Pegasus project focused on research into the basic algorithms for Prosody Extraction and generation, and Signal-processing for Unit Selection. It forms the basis of the CHATR speech synthesis system. Research was carried out into methods of extracting and modeling pitch and duration contours, on the realization of prosodic variation in speech waveforms by parametric methods including linear-prediction, cepstral decomposition, and PSOLA modification, as well as by the selection of units for direct concatenation from a large source speech database.

Like most of ITL's research, the PEGASUS project was largely corpus-based, involving statistical modeling of the meaningful variation in acoustic parameters by training of neural-networks, decision-tree classifiers, and Markov-chain models. Considerable effort was spent on the collection and annotation of large speech corpora for the analysis and prediction of speaking-style and speaker characteristics. Work with these corpora resulted in tools and techniques which were then directly applicable to externally-generated corpora, and use was also made of publically-distributed speech databases for different languages as part of the research.

ATR-ITL inherited the ATR Nu-Talk speech synthesiser from the predecessor Interpreting Telephony Research Laboratories, and much of the initial research concerned extending this system to speak English, and then to Korean, German, and Chinese-language versions. Although the text-analysis and prosody-processing is usually different for each new language, we were able to standardise the voice-creation modules, which are relatively language-independent, and thereby produced a generic system capable of being ported easily to new languages, speakers, and speaking styles. Since the input text is generated by the translation component, prosodic meaning can in theory be predicted directly from an interpretation of the input speech.

However, because the MATRIX speech translation is presently limited to a textual (word-based) representation of single translated utterances, subsequent research was devoted to designing dictionaries and pre-processing modules for predicting how this written text should be processed to produce speech in the various languages, and for making use of syntactic and punctuation cues for predicting the most appropriate phrasing and emphasis patterns from the written representation of the translated utterance. This interpretation requires considerable information about the discourse-history of the utterance, as well as world-knowledge, and may therefore be impossible for isolated single sentences. However, as a result of the PI project, we are able to generate a highly appropriate prosodic specification if the input text is also annotated with the appropriate phrasing and intentional information.

In the course of this research, several prosody-generation and prosody-labeling implementations were tested, including the well-known Fujisaki model of pitch-contour generation, and a mapping was developed between this and the new international standard ToBI prosodic stylization, with extensive application of the latter to Japanese and Korean speech corpora. This resulted in the training of prosody-prediction models directly from the ToBI annotations and enabled prediction of the appropriate ToBI tone and break-index sequences for a given text input, thus linking between abstract higher-level and phy6sical lower-level prosodic tiers for the generation of speech from text.

Taking the very large, multilingual and multi-speaker speech corpora as a starting point, the basic Nu-Talk speech synthesis system was extended to incorporate direct concatenation of raw waveform units selected from the corpora, without the use of prosody-modifying signal processing. This development resulted in extremely natural-sounding voice quality in the synthetic speech but required well-balanced and accurately-labeled corpora. Consequently, based on more than 100 speakers' data, effort was directed towards the automatic processing of corpora for concatenative synthesis. Algorithms were developed for balancing phonemic and prosodic content, for the alignment of labels, for the extraction of acoustic features, and for the removal of redundant segments in the resulting speech databases. This work resulted in the DATR database-processing toolkit which forms the counterpart to the CHATR speech synthesis engine.

Unit selection in It's theoretical aspects represents the intersection of linguistic and phonetic science, as it requires a precise abstract specification of the meaningful acoustic variability in speech for an efficient searchable index to be created. This mapping between feature-based representations of articulation and subtle distinctions of meaning required the development of perceptually-based physical measures of proximity between speech waveform segments. Several waveform transforms were tested, and the Bi-Spectrum, which incorporates phase information, was found to be the most efficient.

The resulting CHATR speech synthesis system represents a major paradigm-shift in synthesis technology, offering a very realistic and personalizable speech output. It was first licensed by AT&T and later incorporated in commercial products by Omron and NTT-Soft, among others. It Is currently being developed for use in a variety of task-specific applications, However, CHATR is by no means complete, and much work still remains to be carried out on improving its intonation through production of balanced speech corpora. No one yet knows the optimal specification for a task-independent corpus, but task-specific needs are finite and calculable.

As the practical applications of synthesis develop, we foresee the need to incorporate emotional speech in addition to prosodically-varied samples and research has been carried out in preparation for this. The features currently used to represent the various speech sounds are still not fully optimized, and the weights used to distinguish between them have yet to be finally tuned. Much research remains before this system can be considered mature, but it has shown a potential that justifies the effort required.

Because of the multilingual focus of this research, Department 2 at ITL has always been a multicultural lab, with visiting researchers including scientists and engineers, professors and students, whose first languages have included Japanese, Korean, French, German, English (of at least 5 varieties), Chinese, Russian, Indian, and Farsi. Each has left behind a valuable contribution to our project, and we would like to take this opportunity to thank them all, wherever they may be working now, for their part in our continuing work.