{"id":4003,"date":"2009-08-24T22:34:51","date_gmt":"2009-08-25T05:34:51","guid":{"rendered":"http:\/\/www.tesl-ej.org\/wordpress\/?page_id=4003"},"modified":"2009-12-20T06:56:19","modified_gmt":"2009-12-20T13:56:19","slug":"ej50r6","status":"publish","type":"page","link":"https:\/\/tesl-ej.org\/wordpress\/issues\/volume13\/ej50\/ej50r6\/","title":{"rendered":"Corpora for University Language Teachers"},"content":{"rendered":"<h3 class=\"datevol\">September 2009 &#8212; Volume 13, Number 2<\/h3>\n<table class=\"review\" align=center border=\"1\">\n<tr>\n<td colspan=4 align=center>\n<h2>Corpora for University Language Teachers<\/h2>\n<\/td>\n<\/tr>\n<tr>\n<td><b>Author:<\/b><\/td>\n<td colspan=2>Carol Taylor Torsello, Katherine Ackerley &amp; Erik Castello, Eds. (2008)<\/td>\n<td rowspan=4 align='center'>&nbsp;<br \/><img decoding=\"async\" class=\"tableimg\" src=\"http:\/\/tesl-ej.org\/ej50\/rpix\/corpora.jpg\" width=\"120\"><\/td>\n<\/tr>\n<tr>\n<td><b>Publisher:<\/b><\/td>\n<td colspan=2>Bern: Peter Lang<\/td>\n<\/tr>\n<tr>\n<th>Pages<\/th>\n<th align=\"center\">ISBN<\/th>\n<th align=\"center\">Price<\/th>\n<\/tr>\n<tr>\n<td>Pp. 309<\/td>\n<td align=center>978-3-03911-639-3 (paper)<\/td>\n<td>\n$80.95 U.S.<\/td>\n<\/tr>\n<\/table>\n<p>This collection of articles by researchers from nine Italian universities provides a useful overview of current approaches to using English corpora: searchable electronic collections of prose. The volume ended up, in effect, as a festschrift for Birmingham University linguist John Sinclair, who, before his death in 2007, was to have been a keynote speaker at the conference in Padua from which these papers were drawn. An introductory piece by Guy Aston not only recalls Sinclair and his influence\u2014\u201cJohn changed our view of the lexical item\u201d (p. 17)\u2014but also provides a good brief history of British English corpora, detailing in particular the relationship between the COBUILD reference books, the corpus they were based on, and the subsequent evolution of that corpus into the Bank of English; this intro chapter also traces the later emergence of a competitor, the British National Corpus (BNC).<\/p>\n<p>Although several recent collections have informatively discussed corpus work and language teaching (e.g., Sinclair, 2004), the current volume stands apart from other conference paper collections that simply report on a themed set of individual research projects. Since several of these papers were based on workshops, the book contains chapters that offer readers instructions on how to apply existing tools to their own corpus projects and language lessons. A later Aston essay, for example, reviews the new edition of the BNC, comparing its current texts to earlier versions and introducing the reader to XAIRA software for searching the XML tags used to code prose, thus allowing users to sort material by text variables such as genre, author, and date of composition. The first half of this article is an accessible introduction to the components of the BNC, whereas the second half assumes some experience with different query formulas. \u201cThe BNC,\u201d Aston observes, \u201cis a prolific resource\u2026 learners [and, I would add, teachers] need to be trained to use it\u2014to recognize and formulate problems, pose queries and interpret solutions\u201d (p. 235).<\/p>\n<p>While several chapters rely on results found in large general-language corpora like the BNC and the Bank of English, it is smaller, custom-made corpora that are discussed here most often. With much current ESL writing and vocabulary instruction emphasizing exposing students to specialized text types to help them gain mastery of the genres of their discipline, creating these Language for Special Purposes (LSP) corpora is well motivated. Some of the specialized corpora discussed in the book include the Padova Learner Debate Corpus (PLDC), which comprises computer forum posts by language learners engaged in debates (Dalziel &amp; Helm). A set of four other corpora (Ulrych &amp; Murphy) was gathered following the framework of mediated discourse analysis, (Scollon, 2001) to emphasize how monolingual texts as well as translated texts reveal editorial and social influences: (1) EuroParl, formal oral discourse from European parliamentary debates; (2) AbCoR, annual reports from multinational companies; (3) AMC, American movie transcripts and their dubbed Italian versions; and (4) EuroCom, essays, half of which were written by non-native English speakers working at the European Commission, the other half being versions of the same texts edited by native English speakers working as translators.<\/p>\n<p>Focusing on another LSP corpus, Tognini Bonelli analyzes terms specific to economics writing in a dataset from <em>The Economist<\/em>. And Taylor compares speech features of the artificial exchanges found in the genre film and television transcripts to the use and distribution of the same speech features in exchanges within the Bank of English. Pushing the definition of textual corpora beyond written and spoken forms, Baldry explores how concordancing can make use of multi-modal material, which can be indexed in ways that help students reinforce their text-based language learning. For example, such corpora can be sorted by images or themes, aligning film clips and the metatext that explicates them, or linking web videos with thematically connected vocabulary items.<\/p>\n<p>Focusing on the writing of language learners themselves, Castello created a corpus of\u00a0 25 essays from both American and British ESL proficiency exams. These learner essays were gathered to measure features of textual complexity. In other work examining writing in a non-native language, D\u2019Angelo created CADIS, the Corpus of Academic Discourse, to capture and compare the English of academic journal articles. That corpus allows the works to be sorted by both discipline as well as the native languages of the authors (English, Italian, or other first languages). Other chapters discuss not just the compiling of texts into a corpus, but using tags to annotate more specialized corpora: Prat Zagrebelsky discusses projects using tags to code common errors in language learners\u2019 college essays. In another tagging endeavor, not student-based, Brunetti discusses creating XML tags to show the inflectional and syntactic relations of each lexical item in a corpus of Old English poems, as well as in its Italian gloss.<\/p>\n<p>As with Brunetti\u2019s chapter, some of the essays cover projects relevant for language-related curriculums for native-speaker students as well as for English language learners, though most papers specifically focus on foreign language teaching and learning. For teachers planning to mine the results of this volume to model or help their students acquire individual English lexical items\u2014to see, for example, how learners\u2019 choices of modals compare to the edited usage of native speakers; which verbs most typically appear adjacent to the noun <em>survey<\/em>; or the different distributions of <em>fork out <\/em>vs. <em>pay<\/em>\u2014it is important to keep in mind that the book\u2019s contributors work mainly with British rather than North American varieties of English. American language practitioners who create or have created their own specialized corpus but seek a larger reference corpus of American phraseology should see Davies (2008), the Corpus of Contemporary American English (COCA), accessible on the web. However, as models of techniques for compiling a corpus based on specialized texts, and of tagging, concordancing, and searching for words that typically appear together in particular genres, these papers provide helpful guidelines for language teachers in any locale. These corpus creators successfully show how to bring to students\u2019 attention patterns of usage found in disciplines ranging from movie transcripts and criticism to economics and news reporting, as well as in more traditional classroom text types such as poetry and academic essays. While several pieces are geared towards the comparison tasks of translators, all the chapters should prove especially relevant for those L2 classroom projects and assignments that value capturing real life constructions over grammar book examples.<\/p>\n<p><strong>References<\/strong><\/p>\n<p>Davies, M. (2008- ). <em>The corpus of contemporary American English (COCA): 385 million words, 1990-present<\/em>. Available online at <a href=\"http:\/\/www.americancorpus.org\/\">http:\/\/www.americancorpus.org<\/a>.<\/p>\n<p>Sinclair, J. M. (Ed.). (2004). <em>How to use corpora in language teaching<\/em>. Amsterdam: John Benjamins.<\/p>\n<p>Scollon, R. (2001). <em>Mediated discourse: The nexus of practice<\/em>. London: Routledge.<\/p>\n<p><b>Laurel Smith Stvan<br \/>\nThe University of Texas at Arlington<br \/>\n&lt;stvan<img decoding=\"async\" class=\"atmark\" src=\"http:\/\/tesl-ej.org\/atmark.png\" border='0' width='12px'>uta.edu&gt;<\/b><\/p>\n<p>&copy; Copyright rests with authors. Please cite TESL-EJ appropriately.<\/p>\n<p><b>Editor&#8217;s Note:<\/b> The HTML version contains no page numbers. Please use the <a href=\"http:\/\/tesl-ej.org\/pdf\/ej50\/r6.pdf\">PDF version<\/a> of this article for citations. <\/p>\n","protected":false},"excerpt":{"rendered":"<p>September 2009 &#8212; Volume 13, Number 2 Corpora for University Language Teachers Author: Carol Taylor Torsello, Katherine Ackerley &amp; Erik Castello, Eds. (2008) &nbsp; Publisher: Bern: Peter Lang Pages ISBN Price Pp. 309 978-3-03911-639-3 (paper) $80.95 U.S. This collection of articles by researchers from nine Italian universities provides a useful overview of current approaches to [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":5192,"menu_order":16,"comment_status":"closed","ping_status":"closed","template":"","meta":{"_genesis_hide_title":false,"_genesis_hide_breadcrumbs":false,"_genesis_hide_singular_image":false,"_genesis_hide_footer_widgets":false,"_genesis_custom_body_class":"","_genesis_custom_post_class":"","_genesis_layout":"","footnotes":""},"class_list":["post-4003","page","type-page","status-publish","entry"],"featured_image_src":null,"featured_image_src_square":null,"_links":{"self":[{"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/pages\/4003","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/comments?post=4003"}],"version-history":[{"count":8,"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/pages\/4003\/revisions"}],"predecessor-version":[{"id":5581,"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/pages\/4003\/revisions\/5581"}],"up":[{"embeddable":true,"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/pages\/5192"}],"wp:attachment":[{"href":"https:\/\/tesl-ej.org\/wordpress\/wp-json\/wp\/v2\/media?parent=4003"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}