LEY 44 DE 1990
9. BIBLIOGRAFÍA
I take it as axiomatic that effective and robust descriptions of any kind of
lexical item must be based on evidence, not intuition, and that corpora provide evidence of a suitable type and quality. At the heart of the study described in this book lies a database of several thousand FEIs, recording their features and characteristics as demonstrated in corpus data. This chapter briefly discusses the database, corpus, and some related computational issues. The following chapters set out the results of the research, and correlations to be drawn.
3.1 DATABASES OF FEIS
The concept of a database of FEIs is far from original. All dictionaries are essentially databases: there are also more focused databases, such as
machine-readable dictionaries of FEIs or those set up to study Russian FEIs, reported in Telija and Doroshenko ( 1992), or Czech idioms for lexicographical purposes, reported in Čermák ( 1994a). Everaert and Kuiper ( 1996) report the construction of a database of 14,000 English phrasal lexical items, and another of 10,000 Dutch items: the exact typologies of these items is unstated, but the numbers involved suggest that they include phrasal verbs as well as proverbs and conventions. None of these resources exactly parallels each other: each records and prioritizes only certain kinds of information.
3.1.1 The Set of FEIs
A total of 6776 FEIs were recorded in the database. I intended to include, as far as possible, a large proportion of the commonest FEIs in current British English, together with some commoner FEIs from American English. The database does not of course record the complete set of the FEIs of English,
which is uncharted, unquantified, and indeterminate. -44-
It was not feasible to assemble a set of FEIs for this study purely by empirical means, for example, by examining corpus data either manually or
automatically and retrieving all gestalts and only lexicalized or coded units. I therefore decided to use an existing published source as starting-point.
However, most such sources--general dictionaries, or specialist dictionaries of idioms--record and perpetuate items not necessarily found in current English. Of the dictionaries available at the time, only CCELD ( 1987) set out to analyse from first principles all (or virtually all) tokens in a corpus of current and
general English; the FEIs it included may be taken as a reasonable indication of which FEIs were actually in use in the 1980s. Furthermore, as one of the
editors of CCELD, I was in a privileged position and knew how and why FEIs in
CCELD were recorded in the way they were. I therefore built the database
around those items which CCELD included and identified as 'phrases' in the grammatical coding. In the event, about 10% of the phrases in CCELD were rejected as not conforming to the types of FEI under consideration, or as being insufficiently fixed or non-compositional; about 14% of database FEIs do not occur in CCELD.
Some of these additions were proverbs from a comparative study of proverbs in French and English that I undertook with Pierre Arnaud in 1991-2, and reported in Arnaud and Moon ( 1993). The English component of the study examined 240 proverbs in the light of evidence in OHPC, recording in a subdatabase the frequency, and the forms, clause and text positions, and genres in which each example of each proverb occurred. These 240 proverbs consisted of those proverbs best attested in an informant study previously undertaken by Arnaud. Other additions came from elsewhere: FEIs observed in OHPC but not treated in CCELD, and FEIs encountered in everyday interaction and reading. In this way, FEIS, or strings manifesting characteristics of FEIS, were included on the basis of at least two out of three possible pieces of evidence: their occurrence in OHPC; their inclusion in a corpus-based dictionary; and their familiarity to informants.
3.1.2 The Structure of the Database
'Database' is technically a misnomer as far as my research was concerned, since my database consists computationally of a series of structured text files which could be manipulated by means of
-45-
standard UNIX tools. 1 A detailed account of the database and report on the findings is given in Moon ( 1994b).
Different kinds of data relating to individual FEIs were recorded in up to 17 separate fields. Two fields were organizational, for example to handle miscellaneous comments concerning exploitation or peculiar distributions. Three fields related to form: the canonical or citation forms of FEIS, major
lexicogrammatical variations, and information concerning the realizations of any open slots. One field recorded typology, following the model set out in Section 1.4, and another recorded frequency as observed in OHPC: see Chapter 4. Three fields related to syntax: the clausal functions of FEIs
(according to a systemic model); passivization and other transformations and inflections; and significant collocations or colligations of FEIs. Syntactic
characteristics of FEIs are discussed in Chapter 5.
Four fields looked more closely at the semantics of four syntactic classes of FEIs. Where the FEI consisted of a predicator and complementation, the verbal process was recorded, following Halliday's model ( 1985: 101ff.; 1994:
106ff.). Since in many cases, especially with metaphors, the surface process described in the lexis does not accord with the deep process inherent in the actual meaning, I recorded both surface and deep processes. Similarly, where FEIs functioned adjectivally, adverbially, or as nominal groups, I recorded the kind of attribute, circumstantial, or entity they denoted, together with any mismatches between surface lexis and deep meaning. Mismatches are discussed in Section 7.6. Two fields contained information about discoursal functions and pragmatics: that is, the typical function of the FEI in discourse, and the contribution made to text ideationally, interpersonally, or
organizationally. The final field recorded attitude and evaluation, for example, positive or negative evaluation, ironic usage, or euphemistic or dysphemistic content. The data from these last three fields is discussed in Chapters 8, 9, and 10.
3.2 CORPUS AND TOOLS
Altenberg and Eeg-Olofsson discuss the need for corpus-based studies of FEIs ( 1990), commenting that there have been few to
____________________
1I am indebted to Ian G. Batten, who advised me on computational aspects of database design. I am also indebted to former colleagues at DEC/SRC on the Hector project, and in particular Mike Burrows, for their assistance with software.
-46-
date. Early work on FEIs was effectively based on the analysis of lists of known items, either observed in texts or in dictionaries ( Meier 1975; Norrick 1985). Collection of data was an erratic process and depended on the quantity and type of the texts encountered or the accuracy of the dictionaries consulted. As a result, some studies of FEIs in English are flawed or unbalanced because rare, obsolete, or even spurious FEIs are given equal status with common, current ones. For example, hand-collected sets of citations cannot give robust information concerning relative frequencies.
Inevitably, the development of corpus linguistics and increasing use of large corpora in lexicology and lexicography is changing all this. One of the most important and basic pieces of information to be derived from a corpus
concerns lexis: the frequencies and distributions of lemmas, and the forms and collocational patterns in which they occur. Profiles of the lexicon based on corpora can be used to prioritize: to distinguish the incontrovertibly significant
from the marginal (and gradations between). This has clear applications in pedagogy, artificial intelligence, contrastive linguistics, and other fields. Collocational studies of corpora shed light on lexical behaviour and pave the way for smarter models of the interaction between syntagm and paradigm. The linguistic phenomena attested in corpora can be used both to test existing abstract models and hypotheses concerning language, and to establish
empirically new models and hypotheses through description. The second approach is characteristic of collocational studies, but most studies of FEIs follow the first since they are founded on and characterized by a priori
assumptions. This is not necessarily bad: assumptions and hypotheses may require adjustment or modification, but they are not necessarily wrong. The literature of corpus linguistics shows decisively that there is a tension or conflict between received, introspectionderived beliefs about language and observed behaviour in corpora. One of the most significant results of corpus linguistics is the blurring of divisions and categories that were formerly thought discrete. This is reported, for example, by Sinclair ( 1986; 1991: 103), Halliday ( 1993), and, with particular reference to grammatical categories, by Aarts (cited in Aarts 1991: 45f.) and Sampson ( 1987: 219ff.): see Briscoe ( 1990) for comments on this last. In relation to FEIs, corpora show up clearly the fallacy of the notion of fixedness of form, and, I suggest, the notion that FEIs can be clearly distinguished from other kinds of linguistic item.
-47-