ATR Human Information Processing Research Labs


Scientific Inquiry: Improving English Listening and Pronunciation Skills




Reiko Akahane-Yamada



To better understand how humans process spoken language, our research group is studying how language learners develop listening and pronunciation skills in a foreign language. For the past several years, we have been conducting extensive research on native Japanese speakers' acquisition of English sounds, using state-of-the-art technology. Our research findings not only advance our understanding of human speech processing, but they also help develop better technology for human-computer interaction and effective instructional tools for second-language education.


1. Introduction

There must be some way to resolve the complex that Japanese have with regard to the English language. At our laboratory, we are studying the ways language learners develop spe4ech perception and production abilities. Through a series of experiments, we have demonstrated that even adults can learn to properly perceive and produce speech segments in a second language. We found that this learning process, which was once thought to be unattainable for adults, can be facilitated through the use of state-of-the-art speech technologies, and through a consideration of human information processing mechanisms. Many important factors related to this learning process, such as the link between perception and production and the effects of age, have also been clarified. We believe that research in these areas will be helpful in resolving the "English complex of Japanese", and over the past two years, we have published three practical books under the title "Scientific Inquiry: Improving English Listening and Pronunciation Skills" (see figure 1).1,2,3 This paper will describe the research related to the learning methods taken up in these three books.


2. Learning Listening Skills

We conducted perception training experiments, with adult Japanese speakers (university students) as subjects, focusing on the identification ability of American English (AE)/r/ and /l/, which are particularly difficult for Japanese speakers. In this training, pairs of English words contrasting the /r/ and /l/ in various positions within words, such as "red" - "led", "poor" - "pool", etc., were used as training materials. We presented these training speech stimuli over headphones, had the subjects respond with one of the items in the pairs, and gave feedback as to the correct answers. We also used utterances produced by a number of talkers as the training stimuli, rather than using utterances from only one speaker. Subjects' identification ability of /r/ and /l/ improved significantly from the pre-training test to the post-training test, and the training effects generalized to untrained words and untrained talkers. We also found that the effects of the tr5aining were maintained even six months after the training was completed. The results obtained have overturned the common belief that "new sounds can only be learned during childhood".

Furthermore, despite the fact that we conducted only perception training, production ability (the level to which native speakers can correctly distinguish the speech segments uttered) by the subjects also improved from pre-test to post-test, suggesting that there is an obvious link between perception and production during the course of second language acquisition.


3. Learning Pronunciation Skills


We also conducted production training experiments to determine whether computer-based production training would be effective or not. Because it is difficult to monitor the movement of the tongue and other organs related to articulation of speech sounds, there has been a problem with methods of feeding back information to learners on whether their pronunciation was adequate or not. We devised several methods to resolve this problem. In the first stage of the learning process, we instructed the position and shape of the tongue for producing /r/ and /l/ sounds by using 3-D computer animation to create a talking head with a semitransparent image of the cheek. In this way, we displayed to the learners, in visible form, the movements of the tongue and other organs related to articulation (Figure 2). By doing so, learners became able to pronounce /r/ or /l/ alone (that is, a prolonged /r/ or /l/ sound) in a surprisingly short period of time.

In the next stage, we conducted production training of English word pairs contrasting in /r/ and /l/, by giving feedback to learners on whether or not they were able to pronounce the words adequately. We designed two types of feedback. As our second method, we trained subjects by providing feedback using a spectrographic display. Furthermore, as our third method, we used evaluation scores generated by a speech recognition system so that the learners could clearly understand the quality of their pronunciation in quantitative terms.

Before and after the above production training, a pre-test and post-test were conducted in which learners' utterances were recorded. These utterances were later evaluation by native AE speakers. It was found that the production ability of /r/ and /l/ sounds improved dramatically from pre-test post-test.

We also measured listening abilities before and after production training, and found that the production training improved perception ability in addition to production ability.


4. Learning and Age


All of the above experiments were conducted using university students as subjects. What would be the effects of this training, however, in the case of younger or older learners? To answer this question, we are currently exposing subjects in a wide range of age groups | from elementary school students to middle-aged and elderly persons | to exactly the same learning experiences as the above university students. The experiments are currently in progress, but at the present time, it has become clear that high school students have obtained an effect from training that is equivalent to that of the university students. It has also become clear that there is a significant effect from the training even in the case of learners in their 50s and 60s. It appears that in the context of language learning, the excuse "I'n too old" will soon become unacceptable.


5. Moving toward Distance Learning


In April 1998, we began an "open experiment concerning listening and speaking of English sounds using the Internet" (URL: http://bluebacks.hip.atr.co.jp/en/index.html) with two main goals: to investigate the possibilities of distance learning, and to gather data from a wide variety of individuals. At the present stage, we are still operating the system experimentally with Japanese persons as targets, but in the past 18 months, we have received over 3,000 items of thought-provoking experimental data. This demonstrates that the use of a network linking people in remote locations is effective as one form of language learning.


6. Relationship with the Native Language

We are often asked, "Why can't Japanese speak English well?" There are many different factors involved, including education methods. If we were to reply from the standpoint of speech science, however, we could say that the difference in linguistic structure between English and Japanese, our native language, is a major factor. In fact, we have found that Korean university students and Japanese university students use entirely different ways of distinguishing between /r/ and /l/. It is possible that we may find a more fundamental method of facilitating second language learning by investigating the interference effects of the native language.


7. Phoneme Learning and English Learning

We are also asked at times, "How does a study of phonemes (vowels and consonants) affect English communication abilities?" It goes without saying that an ability to distinguish phonemes is not enough to enable a person to speak English.

Nevertheless, consideration of human speech processing in general suggests to us the importance of a study of phonemes. The speech signals that enter the human ear are encoded, through an early stage of processing, as a series of phonemes, then as words and sentences, and later meaning is understood using background knowledge. English learning in the past has for the most part focused on the latter stage of this information processing; we could say that the learning of early stages has been neglected. But if learners are not able to properly encode phonemes, the initial process, there is a large possibility that an excessive burden is placed on the latter process. We have been able to obtain consistent results by focusing on the study of phonemes. We believe that in the future, it will be possible to dramatically increase English communication abilities by systematically linking the study of phonemes with the latter stages of information processing.


8. Conclusion

Our method of improving English listening and pronunciation skills is characterized by four elements:
(1) the method is based on knowledge about mechanisms underlying human information processing,
(2) the effects of training using this method have been proven through learning experiments,
(3) the method has effectively adopted advanced technologies such as speech recognition and multimedia, and
(4) it is highly efficient in its use of the information infrastructure.

The learning methods presented in the books mentioned in the introduction have not attained an increase in overall communication abilities in English, but nevertheless they have received excellent evaluation from the general public. After repeated printings of the books, the total number of publications for the three books has exceeded 80,000 copies. We would go so far as to say that this result might be an expression of a new place for "information" in the context of foreign language learning. As we approach the beginning of the 21st century, we are confident that the results of our research will lead to increased use of information technologies in language education, and to effective distance education as well.


Reference