Case sensitive and insensitive corpus analysis
This blog post explains how to analyse corpora and take into account or ignore the difference between lowercase and uppercase. In other words, how to use Sketch Engine to: type wifi and find wifi, WIFI, WiFi and Wifi OR type WiFi and only find WiFi but not the other variants
Words, tags, lemmas, lemposes, lowercase
When using Sketch Engine, every now and then the user comes across the word attribute and its values: words, tags, lemmas, lempos, lowercase and some others depending on the corpus and language. This blog post explains how these positional attributes, to use the correct terminology, work in Sketch Engine and how the user can benefit […]
Build a corpus from the web
The web is a great source of readily available textual data but also a limitless warehouse of spam, machine-generated content and duplicated content unsuitable for linguistic analysis. This may generate some uncertainty about the quality of the language included in the corpora from the web. At Sketch Engine, we are very well aware of the […]
POS tags
This blog post defines what POS tags are, explains manual and automatic POS tagging and points readers to Sketch Engine where they can have their texts tagged automatically in many languages. What is a POS tag? A POS tag (or part-of-speech tag) is a special label assigned to each token (word) in a text corpus to […]