A guide to text analysis within the tidy data framework, using the tidytext package and other tidy tools.
Directory
Search results
Published directory entries matching your search.
TextOptimizer is a browser extension for checking SEO signals, page data, rankings, or website metrics.
Non-profit electronic publishing service providing access to journals published in Thailand.
The MIT AI Risk Initiative produces authoritative data and frameworks to help you identify, prioritize, and manage the risks from AI.
Online archive of anarchism-related written works.
Online library of animal rights texts.
Search an online collection of books, documents, media and archival records.
Initiative to recreate the private book collection of Count Leopoldo Cicognara (1767–1834) in digital form.
The complete online archive of every issue of the global weekly newspaper.
Web service providing access to resources of national libraries across Europe.
Search an online collection of books, documents, media and archival records.
Initiative developed by the Internet Archive aiming to digitize 78 rpm singles from the period between 1880 and 1960.
ERIC record for a guide to finding specialized databases and other web content missed by search engines.
Search an online collection of books, documents, media and archival records.
The Key to Unlocking the Web's Secrets is an OSINT resource for online investigations, research links, search methods, and public data tools.
Digital colllection about Kurdish Writing and Kurdistan.
Website devoted to public domain Latin texts.
Search an online collection of books, documents, media and archival records.
ProPublica database of NYPD misconduct complaints and Civilian Complaint Review Board records.
Telegram-focused OSINT toolbox with resources for researching accounts, channels, and public data.
Archive of the writings of Abraham Lincoln.
OCCRP project page collecting reporting and documents from the Pegasus spyware investigation.
The Pika's OSINT ToolBox is an OSINT resource for online investigations, research links, search methods, and public data tools.
The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together.