knesset: Knesset Committee Corpus
The Knesset Committee Corpus (knesset) is a Hebrew corpus made up of texts collected from the Knesset (Israeli parliament).
Part-of-speech tagset and lemmatization
The Knesset Committee Corpus was tagged and lemmatized with parsan.
Knesset Committee Corpus corpus sizes
| Number of words | 496+ million |
| Number of tokens | 573+ million |
| Number of sentences | 32+ million |
| Number of web pages | 47+ thousand |
Search the Knesset Committee Corpus
Sketch Engine offers a range of tools to work with this Hebrew corpus.
Tools to work with the Knesset Committee Corpus corpus from the web
A complete set of Sketch Engine tools is available to work with this Hebrew corpus to generate:
- keywords – terminology extraction of one-word units
- word lists – lists of Hebrew words, lemmas organized by frequency
- n-grams – frequency list of multi-word units
- concordance – examples in context
- text type analysis – statistics of metadata in the corpus
Changelog
Knesset Committee Corpus (knesset)
- published in Sketch Engine in August 2026
- data from Noam Ordan
Bibliography
Use Sketch Engine in minutes
Generate collocations, frequency lists, examples in contexts, n-grams or extract terms. Use our Quick Start Guide to learn it in minutes.




