tnWaC – Setswana corpus from the Web

The Setswana corpus (tnWaC), also known as the Tswana corpus, is a Setswana corpus made up of texts collected from the Internet. The corpus was prepared according to standards described in the document A Corpus Factory for Many Languages (Kilgarriff et al. at LREC 2010) and the corpus data was provided by Thapelo J. Otlogetswe (University of Botswana). The corpus includes a genre classification covering texts from various domains, such as prose, newspapers, spoken language, etc. The Setswana language is also known as Tswana.

This Setswana corpus has not been tagged or lemmatized yet. The texts are only tagged using shallow tagging which is based on regular expressions and frequency properties of tokens.

Tools to work with the Setswana corpus

A complete set of Sketch Engine tools is available for working with this Setswana corpus and generating:

tnWaC (2013)

  • version 2 has almost 11,5 million words

BARONI, Marco, et al. The WaCky wide web: a collection of very large linguistically processed web-crawled corporaLanguage resources and evaluation, 2009, 43.3: 209-226.

Corpus factory method

Adam Kilgarriff, Siva Reddy, Jan Pomikálek, and Avinesh PVS. A corpus factory for many languages. In LREC workshop on Web Services and Processing Pipelines, Malta, May 2010.

Search the Setswana corpus

Sketch Engine offers a range of tools to work with this Setswana corpus.

Other text corpora

Sketch Engine offers 800+ language corpora.

Use Sketch Engine in minutes

Generating collocations, frequency lists, examples in contexts, n-grams or extracting terms is easy with Sketch Engine. Use our Quick Start Guide to learn it in minutes.