tnWaC – Setswana corpus from the Web
The Setswana corpus (tnWaC), also known as the Tswana corpus, is a Setswana corpus made up of texts collected from the Internet. The corpus was prepared according to standards described in the document A Corpus Factory for Many Languages (Kilgarriff et al. at LREC 2010) and the corpus data was provided by Thapelo J. Otlogetswe (University of Botswana). The corpus includes a genre classification covering texts from various domains, such as prose, newspapers, spoken language, etc. The Setswana language is also known as Tswana.
This Setswana corpus has not been tagged or lemmatized yet. The texts are only tagged using shallow tagging which is based on regular expressions and frequency properties of tokens.
Tools to work with the Setswana corpus
A complete set of Sketch Engine tools is available for working with this Setswana corpus and generating:
- keywords – terminology extraction of one-word units
- word lists – lists of Setswana words organized by frequency
- n-grams – frequency lists of multi-word units
- concordance – examples in context
- text type analysis – statistics of metadata in the corpus
Changelog
tnWaC (2013)
- version 2 has almost 11,5 million words
Bibliography
BARONI, Marco, et al. The WaCky wide web: a collection of very large linguistically processed web-crawled corpora. Language resources and evaluation, 2009, 43.3: 209-226.
Corpus factory method
Adam Kilgarriff, Siva Reddy, Jan Pomikálek, and Avinesh PVS. A corpus factory for many languages. In LREC workshop on Web Services and Processing Pipelines, Malta, May 2010.
Search the Setswana corpus
Sketch Engine offers a range of tools to work with this Setswana corpus.
Use Sketch Engine in minutes
Generating collocations, frequency lists, examples in contexts, n-grams or extracting terms is easy with Sketch Engine. Use our Quick Start Guide to learn it in minutes.




