Showing posts with label corpora. Show all posts
Showing posts with label corpora. Show all posts

Thursday, 24 September 2015

Workshop: 'Computer-Based Literary Analysis', Bristol, 27 November 2015

This event is being organised by one of our graduate students, and there are still a few places left at time of writing. The event is free but registration is required.




Computer-based Literary Analysis

Professor Jan Christoph Meister,  University of Hamburg
 
Friday 27 November, 9.30 to 4.30
Biomedical Teaching Lab E2.1 PC Room, University of Bristol

This workshop is open to postgraduate students and staff working in the humanities, in particular in literary and language studies, including translation. The workshop is free.
 
It requires no previous experience of corpus linguistics software, and will be led by Professor Jan Christoph Meister (Hamburg), who heads the team which has developed the free, web-based software which will be used, known as CATMA (Computer Aided Textual Markup and Analysis).

CATMA is a practical and intuitive tool for literary scholars, students and other parties with an interest in text analysis and literary research. In contrast to most corpus linguistics software, the program combines standard features such as wordlists and collocation searches with the ability to mark up texts in a user-defined way, prior to analysing them quantitatively. It is also designed to facilitate collaborative work on literary texts and allows for the easy sharing of data and metadata. It can be used with individual texts or corpora of multiple texts, in a wide variety of languages.

The philosophy underlying its development is that computers can now be used to complement the traditional close reading of literary texts with quantitative analysis of various narrative, stylistic and linguistic features to gain a deeper understanding of texts and to develop and test various interpretations of them. Examples of the kinds of use to which CATMA can be put include:

          analysing various aspects of narrative form and structure such as events and actions
          analysing aspects of literary style of a text such as sentence length and lexical richness
          analysing language use
          searching texts for words or phrases and their collocates
          comparing translations with source texts

The workshop will provide a hands-on introduction to the software and its capabilities, and there will opportunities for you to discuss how you might use it in your own research.

CATMA can be found at: http://www.catma.de/ and also at http://www.catma.de/clea

Tea, coffee and a light buffet lunch will be provided. Participants are responsible for their own travel costs.

Booking a place:

The workshop will be limited to 20 people. If you wish to book a place, or have any questions, please contact roy.youdale[at]bristol.ac.uk.

Friday, 10 April 2015

Free webconference: Corpora and Tools in Translator Training, 14 April 2015

This just came round via the IATIS Training Committee and looks very useful. The event will run both onsite and online. There is no charge for participation.


IATIS Training Event: 

Corpora and Tools in Translator Training

IATIS Online/Onsite Training Event, 14 April 2015
Cologne University of Applied Sciences, 
Institute of Translation and Multilingual Communication

The use of corpora has now found its way into both the theoretical/descriptive and applied branches of translation studies. This online event is aimed at stimulating debate on current and future trends in corpus-based translator training and critically appraising current practices.
The training event will focus on strategies, usability and technology in the use of corpora and discuss the following aspects:
1) Corpus types which are particularly relevant to translator training.
2) Corpus use for learning to translate vs. learning corpus use for translating.
3) Corpus compilation and selection criteria.
4) Contextualisation of corpus texts and of the theoretical/methodological set-up underlying corpus- based research.
5) Corpus types, e.g. Do-it-yourself corpora, high-quality translation corpora, such as the Cologne Specialized Translation Corpus (CSTC).
6) The web as a source of corpus texts; the web as a macro-corpus.
7) Situating the use of corpora in translation competence models, e.g. PACTE or EMT.
8) Corpus tools, such as corpus analysis software and Web concordancers.
The event leaders (Prof Dr Silvia Bernardini, University of Bologna, Dr Ralph Krüger, Cologne University of Applied Sciences, Prof Dr Silvia Hansen-Schirra, University of Mainz) will present position papers (ca. 25 min each) covering a number of the issues suggested above and then invite discussion from participants.
More details including schedule, abstracts and login instructions at:

Friday, 10 September 2010

Corpus of Historical American English

This came round on a distribution list and I thought it might be handy for those of you who use corpora. (I see they have Spanish and Portuguese corpora too on the site).

We are pleased to announce the release of the 400 million word Corpus of Historical American English (1810-2009). The corpus has been funded by a generous grant from the US National Endowment for the Humanities (NEH), and it is freely available at http://corpus.byu.edu/coha/. COHA is the largest structured corpus of historical English, and it contains more than 100,000 texts from fiction, popular magazines, newspapers, and non-fiction books, with the same genre balance decade by decade from the 1810s-2000s.

COHA is also related to other large corpora that we have created or modified, including the 410 million word Corpus of Contemporary American English (COCA), the 100 million word TIME Magazine Corpus (1920s-2000s), the 100 million word British National Corpus (our architecture and interface), the 100 million word NEH-funded Corpus del Español (1200s-1900s), and the 45 million word NEH-funded Corpus do Português (1300s-1900s). For information on these corpora, see http://corpus.byu.edu.

COHA allows you to quickly and easily search the 400 million words of text from the 1810s-2000s to see how words, phrases and grammatical constructions have increased or decreased in frequency, how words have changed meaning over time, and how stylistic changes have taken place in the language. Users can see the overall (normalized) frequency by decade and year, as well as the frequency of each matching string, by decade.

The following are just a small sample of an unlimited number of queries, but they should give some idea of what the corpus can do.

* Lexical change: the rise and fall of words and phrases like the following:
 - (decrease since the 1800s): bosom, folly, grieved, bestow*, quaint, beauteous, fellow, sublime, lad, many a time, of no little, for (conj)
  - (an increase and then decrease): mustn't, naughty, boyish, agog, toddle, far-out, famed, wangle, swell (adj), lousy
  - (an increase to the present time): a lot of, unleash, sexual, calm down, screw up, freak out, mommy, skills, frustrating
  - (words reflecting historical and cultural shifts): emancipation, steamship, telegraph, flapper*, fascis*, teenage*, communis*, global warming

* Stylistic change (which gives the flavor of a different time period). Examples from the 1800s, which have decreased since then, are: [so ADJ as to V] (so good as to show me), [PRON be but] (they are but the last examples), [have quite V-ed] (until she had quite finished), [NOUN be that of] (her dress was that of a beggar), or [a most ADJ NOUN] (a most helpful child).

* Morphological change: which show how word roots, prefixes, and suffixes have been used over time, including comparisons between different periods, such as -heart- (1800s noble-hearted, 1900s heart-stopping), home- (1800s homebred, 1900s homeowner), or -able adjectives (1800s placable, 1900s predictable).

* Syntactic change (since the corpus is tagged and lemmatized), like [end up V-ing], [going to V], [V PRON into V-ing] (talked them into going), phrasal verbs with [up] (make up, show up), post-verbal negation with [need] (needn't mention), the 'get' passive (get hired), sentence-initial 'hopefully', and semi-modals like [need to] and [have to].

* Semantic change: how the meaning or usage of words have changed over time, by looking at changes in collocates (co-occurring words), like [sexual, gay, chip, engine, or web]. This can also signal cultural changes over time, such as nouns used with [woman] in the 1930s-50s compared to the 1960s-80s (fabrics, hips // liberation, abortion), or nouns used with [problem] in the 1810s-1920s compared to the 1920s-2000s (railway, trust // drugs, pollution).

 * Lexical change (again): users can also have the corpus generate a list of words that were used more in one period than another, even when they don't know what the specified words might be. For example, the corpus can generate lists of verbs in the 1970s-2000s compared to the 1930s-1960s (download, recycle // effectuate, redound), adjectives in the 1970s-2000s and the 1930s-1960s (online, affordable // leftist, communistic), or -ly adverbs in the 1900s and the 1800s (basically, reportedly // despondingly, sportively).

As can be seen, the corpus allows research on a wide range of phenomena in 400 million words of text from the last two centuries of American English. The corpus is freely available at http://corpus.byu.edu/coha/, and we invite you to use it for your research and teaching.