SpiderLing / SketchEngine web crawler

SpiderLing and SketchEngine are two user-agent identifiers used by our web crawler.
The crawler is operated by Lexical Computing CZ s.r.o., the company behind Sketch Engine. It collects text from publicly accessible web pages for building text corpora used for linguistic analysis.
This is not a generative-AI or search-engine crawler, it is a corpus-building crawler.

User agents

Our crawler uses one of these user-agent identifiers:
SpiderLing
SketchEngine

Crawling policy

  • we respect the crawling rules in robots.txt
  • we do not follow links marked with rel="nofollow"
  • within a single crawl project, we wait at least 20 seconds between requests to the same web domain, i.e. no more than 3 requests per minute per domain

How to block the crawler

To prevent our crawler from accessing your website, block both user agents in robots.txt:

User-agent: SpiderLing
Disallow: /

User-agent: SketchEngine
Disallow: /

Standard robots.txt rules can also be used to restrict access only to particular sections of your website.

Contact

The crawler is operated by:
Lexical Computing CZ s.r.o.
Botanická 554/68a
602 00 Brno
Czech Republic
For questions or problems related to crawling, contact:
support@sketchengine.eu