This section reviews studies that measure successful repetition of targets, which are typically lexical or syntactic elements, although prosodic targets have also been defined in some of the studies (Ward and Litman 2007b, 2007a). Section 4.3.1 reviews two studies that have measured successful repetition ratios of lexical elements. Section 4.3.2 reviews two studies which, in addition to successful repetition, have also used linear regression in order to measure the effect of distance on the probability or frequency of repetition.
4.3.1 Successful repetition ratio
(Brennan 1996) studied lexical entrainment in recordings of spontaneous speech, as well as adoption of system terms (lexical convergence) by users of speech interfaces. The goal of the study was to investigate differences and similarities between these two processes, and implications of this for SDS, especially in relation to the vocabulary problem: the wealth of language is a problem for SDS, because a user may adopt several terms to describe the same concept. Lexical entrainment (or convergence) is a possible way of encouraging the user to use specific terms (by presenting this vocabulary to the user), thus shortening the list of candidate words that the ASR and ALU components have to process, which in turn would result in increased efficiency.
The spontaneous speech recordings were acquired using a task experiment, which involved two participants who could converse without visual contact and had to line up identical sets of picture cards in the same order. The purpose of these experiments was to further investigate previous theoretical predictions on lexical entrainment (Clark and Wilkes-Gibbs 1986; Brennan and Clark 1996). The latter explain lexical entrainment through “conceptual pacts”, or implicit “agreements” between interlocutors on terms that describe concepts in the discourse.
(Brennan 1996) conducted a series of wizard-of-oz experiments (database query), in which a (simulated) system employed two different correction strategies in order to encourage the user to adopt its terminology: “embedded” and “exposed”. Embedded corrections are repetitions of the query by the system with substitution of the user term with the system term, while exposed corrections are explicit clarification requests of the system that contain the system term (e.g. “did you mean /term/ ?). A speech-based, as well as a text-based interface were used.
The measurements comprised a ratio of successful adoption of the system term by users over the total amount of user turns. Also, the effect of delay (whether the user response came immediately after a correction or after several utterances) was investigated. In both text and speech cases, there was significantly more lexical convergence for the immediate condition (compared to delayed) and exposed corrections (compared to embedded). However, convergence was significant in all cases. (Brennan 1996) suggested that convergence only in the immediate condition would imply autonomous entrainment, while frequent convergence in the delayed condition would imply a more strategic process. The study concluded that (a) a system should output only terms that it can process as input, (b) should be consistent in its output and documentation, (c) repairs are essential, as shown from the higher convergence to exposed corrections, and (d) a system could adopt terms proposed by the user by adopting grounding strategies.
In (Fais 1996), lexical accommodation was investigated in view of human-machine interface design. Three experimental scenarios were conducted, in which a conversation was either (a) direct monolingual, between English-speaking subjects and conference “agents”, (b) bilingual, between English-speaking and Japanese-speaking agents, mediated by a human interpreter, and (c) mediated by a simulated machine translation system. The goal of the study was to study lexical accommodation in these three contexts in order to determine the effect of (1) desire for social approval, and (2) difficulty of communication, on the degree of accommodation. The measurement was a ratio of the number of (different) words spoken by both speakers over the overall number of (different) words in each dialogue. The direction of accommodation was assessed by defining that a speaker who uses a word previously spoken by the interlocutor is the one who accommodates. The results showed significant accommodation in all three scenarios. The highest accommodation occurred in the human-mediated scenario where, according to (Fais 1996), both social factors and communication efficiency are important. Higher accommodation was also found for the machine- mediated scenario, when compared to the direct dialogue scenario, despite the fact that social factors were irrelevant. In addition, accommodation was equal between interlocutors in the direct dialogue scenario, but in the other two the client accommodated to the agent. (Fais 1996) attributed
this finding to the fact that clients perceived interpreters (either human or machine) as having the dominant role. Thus clients accommodated to the lexical choices of the interpreters, in order to improve communication efficiency. The latter conclusion is in agreement with (Brennan 1996). (Fais 1996) also suggested that higher accommodation in SDS can be encouraged by use of an animated face or “persona”, replicating the human-interpreted setting (that shows the highest accommodation).
4.3.2 Linear regression of repetition over distance
(Reitter et al. 2006) explored priming of syntactic structures, in order to test various predictions of the Interactive Alignment Model (Pickering and Garrod 2004) that was outlined in section 3.5. Spontaneous (telephone) speech and task-oriented speech (a map-task, which was described in section 4.2.3) were used in order to test the effect of the situational constraints of the task on the degree of priming. The syntactic trees of all utterances in the corpus were converted to phrase rules and an algorithm search for repetition of these rules was conducted. Any sentence could be a valid candidate for a prime or target for priming (repetitions of entire phrases were excluded). Distance (expressed in number of turns or seconds) of priming was also taken into account. In addition, a distinction was made between comprehension-production (CP) priming, where one speaker produces the prime and the partner produces the target, and production-production (PP) priming, where both prime and target are produced by the same speaker.
Statistical analysis is performed by use of generalized linear mixed effects regression models (GLMM). This regression approach allows the calculation of coefficients of linear models, such as a model of the probability of repetition of a prime, based on discrete factors (such as type of corpus) or continuous explanatory variables (such as distance). The maximum distance used was 25 turns or 15 seconds. There were various outcomes from this study. The probability of priming was found to decay with distance in both corpora, and significant PP priming was found in both corpora. In the case of CP priming, higher confidence was found for the map-task corpus, when compared to spontaneous speech. This, according to (Reitter et al. 2006), validates the hypothesis of the Interactive Alignment Model that syntactic priming leads to semantic priming: when the cognitive workload is increased (task corpus), speakers reproduce each other's syntactic structures in order to align their situational models with less effort. In the unconstrained spontaneous speech, the cognitive workload is less, thus the speakers are less eager to adopt their partners' syntactic structures, but they still do so to a lesser extent mechanistically.
in relation to learning in tutorial sessions with a human tutor and an intelligent tutoring system (SDS). The features studied were lexical (word repetition) and prosodic (F0 and Intensity). The theoretical background of the studies was based on the Interactive Alignment Model (Pickering and Garrod 2004).
The measure of lexical convergence was the count of different word tokens repeated by the student in a window of up to 20 turns after the tutor's utterance (prime). For prosodic features, the minimum, maximum and average F0 and Intensity of the tutor's utterance were considered primes if their z-score normalized values where greater than one (an arbitrary threshold of one standard deviation). Again, the response of the student to the prosodic prime for a window of up to 30 turns (to capture variation in intensity) was measured. The effect of distance on the number of repetitions (either lexical or prosodic) of a prime in the speech of the tutor was measured as the slope of a line fitted by linear regression (least squares). The slopes typically have a negative value, which is an indication that prime repetition decays over time. The significance of this was assessed by calculating a p-value, as an indication of the probability of fitting that line if there was no effect of distance.
In order to assess the effect of convergence on learning, (Ward and Litman 2007a) used a corpus of students who completed two physics tests, one before and one after a tutoring session with a (human) tutor. Thus, the learning outcomes from the tutoring session were quantified by means of test-scores. An automatic feature selection algorithm (stepwise regression) was used to find which features, if any, affect the learning outcomes. The only factors that were identified by the algorithm were lexical repetition and response on mean intensity primes, for a window of 20 turns. The identified models were then tested on a different corpus, which contained dialogues between students and an automatic intelligent tutoring system. The latter was following the same tutoring session layout and procedure as in the sessions with a human tutor. The models remained significant in the test data, although there were some unexplained differences (e.g. change of sign in some coefficients). (Ward and Litman 2007b) concluded that there is evidence of a relationship between convergence and learning, despite contrasting differences of the models in the two corpora, which can possibly be explained by the differences in speech style between human-human and human- computer conversation.