Gå till denna sida på svenska webben

Corpus-Based Methods

  • 7.5 credits

This course deals with corpus-based methods, that is, the large-scale study of written text, or spoken or signed utterances.

Contents: Data, methods and evidence in different linguistic traditions. Quantitative properties of language, frequencies, n-grams. Data collection for different types of corpora (including traditional sample corpora, monitor corpora and web corpora) and modalities (text, speech, signing). Representation of corpora in XML. Overview of computational linguistic methods for automatic segmentation and annotation of text, including tokenisation, part-of-speech tagging and syntactic analysis. Searching corpora using regular expressions. Analysis of corpora based on occurrences and co-occurrences. Relationship between corpus material and research questions. Ethics, copyright, licenses.

  • Course structure

    Teaching format

    The course is based on lectures and laborations.


    The course is examined through written exams and reports.

  • Contact

    Student affairs office, Departement of Linguistics

    Södra huset, C 378
    Visiting hours for students
    Tuesdays 9.00-10.00
    Wednesdays 13.00-15.00
    Thursdays 9.00-11.00 and 13.00-16.00

    +46 8 16 23 47

    Director of Studies, Second and Third Level

    Sofia Gustafsson-Capková
    Office: C254
    Phone: +46 8 16 34 88
    E-mail: ma@ling.su.se

Find more courses and programmes

Know what you want to study?

Find your study programme

What can I study?

Explore our subjects