Random samples from large datasets, for convenience.
-
bluelightai-dev/dclm-full-deduped-sample
Viewer • Updated • 4.92M • 133 -
bluelightai-dev/the-stack-dedup-sample
Viewer • Updated • 474k • 52 -
bluelightai-dev/common-corpus-sample-open-culture
Viewer • Updated • 462k • 57 -
bluelightai-dev/common-corpus-sample-open-government
Viewer • Updated • 373k • 74 • 1