


ATR Interpreting Telecommunications Research Labs
Prosody Interpretation and Speech Synthesis
Nick Campbell
Spoken language is not usually as carefully prepared as written text, and relies
instead on prosody, or inflexion of the voice, to carry the meaning. Prosody allows
the listener to interpret the speaker's intention in an appropriate way by indicating
the meaning-groups, the focus of discourse, and the relations holding between
the words and phrases in the speech. For the past seven years, Department 2 ATR-ITL
has been researching the automatic processing of this linguistic and para-linguistic
information under the PI and the PEGASUS projects.
The PI (or Prosody Interpretation) project featured the analysis, transmission,
and generation of prosodic information in speech. It models the variations in
tone-of-voice, in speaking rate, and in loudness, and maps between changes in
the acoustic signal and changes in the intended meaning of groups of words in
the speech. This work resulted in a set of algorithms for the extraction of speech-act,
emphasis, and phrasing information from the spoken utterance input to the speech
translation system, for the transfer of this information along with the translated
result, and for its use as enriched input to the speech synthesiser to ensure
that the translated utterance could be produced with the appropriate meaning and
that ambiguities in the translated word-sequence could be minimized. However,
the constrained style and short length of the utterance typical of the hotel-reservation
task, which forms the common core of ITL research, make little use of such prosodic
information and, apart from a simple statement/question detection module, the
PI research has not yet been incorporated into the MATRIX translation system.
The Pegasus project focused on research into the basic algorithms for Prosody
Extraction and generation, and Signal-processing for Unit Selection. It forms
the basis of the CHATR speech synthesis system. Research was carried out into
methods of extracting and modeling pitch and duration contours, on the realization
of prosodic variation in speech waveforms by parametric methods including linear-prediction,
cepstral decomposition, and PSOLA modification, as well as by the selection of
units for direct concatenation from a large source speech database.
Like most of ITL's research, the PEGASUS project was largely corpus-based, involving
statistical modeling of the meaningful variation in acoustic parameters by training
of neural-networks, decision-tree classifiers, and Markov-chain models. Considerable
effort was spent on the collection and annotation of large speech corpora for
the analysis and prediction of speaking-style and speaker characteristics. Work
with these corpora resulted in tools and techniques which were then directly applicable
to externally-generated corpora, and use was also made of publically-distributed
speech databases for different languages as part of the research.
ATR-ITL inherited the ATR Nu-Talk speech synthesiser from the predecessor Interpreting
Telephony Research Laboratories, and much of the initial research concerned extending
this system to speak English, and then to Korean, German, and Chinese-language
versions. Although the text-analysis and prosody-processing is usually different
for each new language, we were able to standardise the voice-creation modules,
which are relatively language-independent, and thereby produced a generic system
capable of being ported easily to new languages, speakers, and speaking styles.
Since the input text is generated by the translation component, prosodic meaning
can in theory be predicted directly from an interpretation of the input speech.
However, because the MATRIX speech translation is presently limited to a textual
(word-based) representation of single translated utterances, subsequent research
was devoted to designing dictionaries and pre-processing modules for predicting
how this written text should be processed to produce speech in the various languages,
and for making use of syntactic and punctuation cues for predicting the most appropriate
phrasing and emphasis patterns from the written representation of the translated
utterance. This interpretation requires considerable information about the discourse-history
of the utterance, as well as world-knowledge, and may therefore be impossible
for isolated single sentences. However, as a result of the PI project, we are
able to generate a highly appropriate prosodic specification if the input text
is also annotated with the appropriate phrasing and intentional information.
In the course of this research, several prosody-generation and prosody-labeling
implementations were tested, including the well-known Fujisaki model of pitch-contour
generation, and a mapping was developed between this and the new international
standard ToBI prosodic stylization, with extensive application of the latter to
Japanese and Korean speech corpora. This resulted in the training of prosody-prediction
models directly from the ToBI annotations and enabled prediction of the appropriate
ToBI tone and break-index sequences for a given text input, thus linking between
abstract higher-level and phy6sical lower-level prosodic tiers for the generation
of speech from text.
Taking the very large, multilingual and multi-speaker speech corpora as a starting
point, the basic Nu-Talk speech synthesis system was extended to incorporate direct
concatenation of raw waveform units selected from the corpora, without the use
of prosody-modifying signal processing. This development resulted in extremely
natural-sounding voice quality in the synthetic speech but required well-balanced
and accurately-labeled corpora. Consequently, based on more than 100 speakers'
data, effort was directed towards the automatic processing of corpora for concatenative
synthesis. Algorithms were developed for balancing phonemic and prosodic content,
for the alignment of labels, for the extraction of acoustic features, and for
the removal of redundant segments in the resulting speech databases. This work
resulted in the DATR database-processing toolkit which forms the counterpart to
the CHATR speech synthesis engine.
Unit selection in It's theoretical aspects represents the intersection of linguistic
and phonetic science, as it requires a precise abstract specification of the meaningful
acoustic variability in speech for an efficient searchable index to be created.
This mapping between feature-based representations of articulation and subtle
distinctions of meaning required the development of perceptually-based physical
measures of proximity between speech waveform segments. Several waveform transforms
were tested, and the Bi-Spectrum, which incorporates phase information, was found
to be the most efficient.
The resulting CHATR speech synthesis system represents a major paradigm-shift
in synthesis technology, offering a very realistic and personalizable speech output.
It was first licensed by AT&T and later incorporated in commercial products by
Omron and NTT-Soft, among others. It Is currently being developed for use in a
variety of task-specific applications, However, CHATR is by no means complete,
and much work still remains to be carried out on improving its intonation through
production of balanced speech corpora. No one yet knows the optimal specification
for a task-independent corpus, but task-specific needs are finite and calculable.
As the practical applications of synthesis develop, we foresee the need to incorporate
emotional speech in addition to prosodically-varied samples and research has been
carried out in preparation for this. The features currently used to represent
the various speech sounds are still not fully optimized, and the weights used
to distinguish between them have yet to be finally tuned. Much research remains
before this system can be considered mature, but it has shown a potential that
justifies the effort required.
Because of the multilingual focus of this research, Department 2 at ITL has always
been a multicultural lab, with visiting researchers including scientists and engineers,
professors and students, whose first languages have included Japanese, Korean,
French, German, English (of at least 5 varieties), Chinese, Russian, Indian, and
Farsi. Each has left behind a valuable contribution to our project, and we would
like to take this opportunity to thank them all, wherever they may be working
now, for their part in our continuing work.

