RVL-CDIP document classifiers, with and without identification codes
Seven image classifiers, twelve linear SVM baselines, and two text models (BERT, RoBERTa) trained on RVL-CDIP (16 document categories), each in four versions that differ in the labels and in the images they were trained on, and two layout models (LayoutLM, LayoutLMv3) in the two corrected-label versions:
| folder | labels | training images |
|---|---|---|
labels-original_codes-kept |
original RVL-CDIP labels | de-identified pages |
labels-original_codes-removed |
original RVL-CDIP labels | de-identified pages with the identification codes whitened out |
labels-corrected_codes-kept |
corrected labels | de-identified pages |
labels-corrected_codes-removed |
corrected labels | de-identified pages with the identification codes whitened out |
Identification codes are the Bates numbers stamped on most pages; they are associated with the
category label and act as a shortcut (see stefan-hf/rvlcdip-id-codes).
The codes were located with stefan-hf/yolov8n-rvlcdip-idcodes.
All models were trained on a de-identified copy of the corpus, in which personal data had been
replaced with synthetic values, so none of them was trained on the original personal data.
The codes-removed models cannot use the codes, at a cost of 0.0β0.3 points in-domain on the
corrected labels. On RVL-CDIP-N, which was collected from other sources, the codes-kept models
are 1β6 points more accurate (second table below), so neither version is better on every measure.
Models and accuracy
Test accuracy (%) of the released checkpoint on the test split of its own training condition; in parentheses, the mean over three training seeds. The released checkpoint is always seed 42, the first run, chosen in advance and not by score. Models trained on the original labels are scored on the original test labels (39,999 pages) and models trained on the corrected labels on the corrected test labels (36,446 pages), so the two halves of the table are not directly comparable.
| model | folder | original, codes kept | original, codes removed | corrected, codes kept | corrected, codes removed | size |
|---|---|---|---|---|---|---|
| AlexNet | alexnet |
89.21 (89.30) | 88.41 (88.52) | 92.85 (92.82) | 92.73 (92.73) | 228 MB |
| GoogLeNet | googlenet |
89.34 (89.25) | 88.52 (88.50) | 93.19 (93.15) | 92.91 (92.97) | 23 MB |
| ResNet-50 | resnet50 |
91.01 (90.85) | 90.21 (90.20) | 94.26 (94.31) | 94.28 (94.19) | 94 MB |
| ResNeXt-50 (32x4d) | resnext50 |
91.48 (91.33) | 90.48 (90.63) | 94.64 (94.69) | 94.56 (94.52) | 92 MB |
| SqueezeNet 1.0 | squeezenet |
88.70 (88.66) | 87.74 (87.69) | 92.46 (92.59) | 92.50 (92.46) | 3 MB |
| VGG-16 | vgg16 |
91.18 (91.27) | 90.74 (90.81) | 94.61 (94.62) | 94.46 (94.42) | 537 MB |
| DiT-base | dit_base |
92.69 (92.76) | 92.20 (92.19) | 95.94 (95.92) | 95.74 (95.81) | 343 MB |
Corrected-label models on both test conditions and on RVL-CDIP-N (1,002 newly collected pages):
| model | kept β kept | kept β removed | removed β kept | removed β removed | RVL-CDIP-N, kept | RVL-CDIP-N, removed |
|---|---|---|---|---|---|---|
| AlexNet | 92.85 | 91.54 | 92.37 | 92.73 | 73.6 | 69.7 |
| GoogLeNet | 93.19 | 92.62 | 92.72 | 92.91 | 77.9 | 71.5 |
| ResNet-50 | 94.26 | 93.71 | 94.19 | 94.28 | 79.2 | 77.0 |
| ResNeXt-50 (32x4d) | 94.64 | 94.10 | 94.45 | 94.56 | 80.8 | 76.2 |
| SqueezeNet 1.0 | 92.46 | 91.57 | 92.32 | 92.50 | 77.0 | 75.0 |
| VGG-16 | 94.61 | 93.75 | 94.25 | 94.46 | 78.5 | 77.8 |
| DiT-base | 95.94 | 95.37 | 95.71 | 95.74 | 87.1 | 83.9 |
"kept β removed" is a model trained with codes and tested on pages without them.
Usage
rvlcdip_models.py in this repository builds the architecture, loads the weights, and applies the
training-time preprocessing (grayscale, bilinear resize to 224 Γ 224, replicate to three channels,
normalize).
from huggingface_hub import hf_hub_download
from PIL import Image
import importlib.util, torch
spec = importlib.util.spec_from_file_location(
"rvlcdip_models", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_models.py"))
lib = importlib.util.module_from_spec(spec); spec.loader.exec_module(lib)
model, cfg = lib.load("resnet50", "labels-corrected_codes-removed")
x = lib.preprocess(Image.open("page.png"), cfg).unsqueeze(0)
with torch.no_grad():
print(cfg["id2label"][str(model(x).argmax(1).item())])
The DiT-base folders are also in the transformers layout and load with
AutoModelForImageClassification.from_pretrained(..., subfolder="dit_base/labels-corrected_codes-removed").
Each folder's config.json (rvlcdip_config.json for DiT) records the label map, the
normalization, the selected epoch, and the accuracies above.
Linear SVM baselines
Twelve linear baselines trained on the same four configurations (one run each; liblinear is deterministic, so there are no seeds). The TF-IDF models read the OCR text of the page, not the image, and the fusion models read both. Test accuracy (%) on the test split of the model's own training condition:
| model | folder | input | original, codes kept | original, codes removed | corrected, codes kept | corrected, codes removed | size |
|---|---|---|---|---|---|---|---|
| TF-IDF + linear SVM | svm_tfidf |
OCR text | 88.95 | 88.78 | 93.40 | 93.35 | 4.1 MB |
| TF-IDF + handwriting tokens + linear SVM | svm_tfidf_hw |
OCR text | 89.77 | 89.72 | 94.65 | 94.64 | 4.1 MB |
| Linear SVM on CLIP ViT-L/14-336 embeddings | svm_clip_vitl14_336 |
image | 87.47 | 87.44 | 93.77 | 93.59 | 0.1 MB |
| TF-IDF + linear SVM, Tesseract text | svm_tfidf_t411 |
OCR text (Tesseract) | 79.81 | 79.71 | 84.68 | 84.51 | 4.0 MB |
| TF-IDF + CLIP embeddings, one linear SVM | svm_fusion |
OCR text and image | 92.35 | 92.27 | 96.76 | 96.70 | 4.1 MB |
| TF-IDF + handwriting tokens + CLIP embeddings, one linear SVM | svm_fusion_hw |
OCR text and image | 92.43 | 92.36 | 96.86 | 96.78 | 4.1 MB |
| Linear SVM on SigLIP so400m embeddings | svm_siglip_so400m |
image | 90.70 | 90.32 | 95.59 | 95.61 | 0.1 MB |
| TF-IDF + SigLIP embeddings, one linear SVM | svm_fusion_siglip_so400m |
OCR text and image | 93.09 | 92.78 | 97.12 | 97.07 | 4.1 MB |
| TF-IDF + handwriting tokens + SigLIP embeddings, one linear SVM | svm_fusion_hw_siglip_so400m |
OCR text and image | 93.12 | 92.76 | 97.14 | 97.14 | 4.1 MB |
| Linear SVM on ImageNet ResNet-50 features | svm_resnet50_imagenet |
image | 76.88 | 76.39 | 82.32 | 82.04 | 0.1 MB |
| TF-IDF + ImageNet ResNet-50 features, one linear SVM | svm_fusion_resnet50_imagenet |
OCR text and image | 92.11 | 91.94 | 96.15 | 96.18 | 4.2 MB |
| TF-IDF + handwriting tokens + ImageNet ResNet-50 features, one linear SVM | svm_fusion_hw_resnet50_imagenet |
OCR text and image | 92.22 | 92.08 | 96.29 | 96.28 | 4.2 MB |
- TF-IDF. Word uni- and bigrams, 50,000 features, sublinear term frequency, one-vs-rest
LinearSVC, withCchosen on validation accuracy. - Handwriting tokens. The same model with two pseudo-tokens in front of the text: the number of words on the page in logarithmic bins, and the share of words that Textract marks as handwriting.
- CLIP embeddings. The projected image embedding of the grayscale page from the frozen
openai/clip-vit-large-patch14-336, L2-normalized and standardized, then the same SVM. - SigLIP embeddings. The same recipe with the image embedding of the frozen
google/siglip-so400m-patch14-384(1,152 dimensions) in place of CLIP's. - ImageNet ResNet-50 features. The same recipe with the 2,048-dimensional pooled feature of
torchvision's ResNet-50 with its ImageNet weights (
IMAGENET1K_V1), an encoder that was never trained on documents; the grayscale page is resized to 224 Γ 224 without cropping. This is not the fine-tunedresnet50classifier of the first table. - Tesseract text. The TF-IDF model trained and tested on Tesseract 4.1.1 (
--psm 3) text in place of Amazon Textract text. All other text models here expect Textract text. - Fusion. One SVM on the TF-IDF features and the image embedding of the same page,
concatenated. The standardized embedding block is divided by the square root of its dimension
(β768 for CLIP, β1152 for SigLIP, β2048 for ResNet-50), so that its rows have about the same norm as the TF-IDF rows,
and multiplied by a weight
alpha;alphaandCare chosen together on validation accuracy and recorded inconfig.json. Thesvm_fusion_hwmodels add the handwriting tokens to the text. With the corrected labels, the CLIP fusion models reach 95.1β95.6% on RVL-CDIP-N and the SigLIP fusion models 97.0β97.3%, against 85.7β86.1% forsvm_tfidf, 92.1% forsvm_tfidf_hw, 88.3β88.7% forsvm_clip_vitl14_336, and 95.6β95.8% forsvm_siglip_so400m. The ResNet-50 features alone give 60.5β62.1% there, and 91.9β93.6% in the fusion models.
The weights are stored as safetensors and the vocabulary as JSON; rvlcdip_svms.py rebuilds the
models with scikit-learn and no pickled objects.
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location(
"rvlcdip_svms", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_svms.py"))
svms = importlib.util.module_from_spec(spec); spec.loader.exec_module(svms)
text_svm = svms.load_tfidf_svm("svm_tfidf", "labels-corrected_codes-removed")
print(text_svm.predict(["Dear Mr. Smith, thank you for your letter of March 3 ..."]))
image_svm = svms.load_clip_svm("labels-corrected_codes-removed") # downloads CLIP
# image_svm.predict([PIL.Image.open("page.png")])
fusion_svm = svms.load_fusion_svm("svm_fusion", "labels-corrected_codes-removed") # downloads CLIP
# fusion_svm.predict(["text of the page ..."], [PIL.Image.open("page.png")])
# the SigLIP and ResNet-50 versions load the same way and download their encoder
# (the ResNet-50 versions need torchvision)
# svms.load_image_svm("svm_siglip_so400m", "labels-corrected_codes-removed")
# svms.load_fusion_svm("svm_fusion_siglip_so400m", "labels-corrected_codes-removed")
# svms.load_image_svm("svm_resnet50_imagenet", "labels-corrected_codes-removed")
# svms.load_fusion_svm("svm_fusion_resnet50_imagenet", "labels-corrected_codes-removed")
For svm_tfidf_hw and the three svm_fusion_hw* models, put svms.hw_prefix(n_words, n_handwritten_words) (or
svms.hw_prefix_from_textract(response_text)) in front of each page's text. svm_tfidf_t411
loads with load_tfidf_svm like the other text models. The decision values are uncalibrated
one-vs-rest margins.
Text and layout models
Four fine-tuned transformers that read the OCR output of the page (Amazon Textract): BERT and
RoBERTa read the text, LayoutLM reads the words together with their bounding boxes, and LayoutLMv3
reads the words, their boxes, and the page image. BERT and RoBERTa are included in all four
versions, LayoutLM and LayoutLMv3 in the two corrected-label versions. For these models codes-removed means that the words of the identification
codes were deleted from the OCR output the model was trained on; for LayoutLMv3 the codes were
also whitened out of the page images. Test accuracy (%) of the released seed-42 checkpoints
trained on the corrected labels, on the corrected test labels (36,446 pages) under both test
conditions, and on RVL-CDIP-N:
| model | folder | input | kept β kept | kept β removed | removed β kept | removed β removed | RVL-CDIP-N, kept | RVL-CDIP-N, removed | size |
|---|---|---|---|---|---|---|---|---|---|
| LayoutLM (v1, base) | layoutlmv1 |
OCR words and boxes | 96.66 | 96.40 | 96.58 | 96.55 | 96.6 | 96.7 | 451 MB |
| LayoutLMv3 (base) | layoutlmv3 |
page image, OCR words and boxes | 97.51 | 97.38 | 97.43 | 97.49 | 97.2 | 96.3 | 507 MB |
| BERT-base (uncased) | bert_base |
OCR text | 95.61 | 95.24 | 95.35 | 95.31 | 91.8 | 91.4 | 439 MB |
| RoBERTa-base | roberta_base |
OCR text | 96.02 | 95.68 | 95.85 | 95.77 | 92.3 | 92.9 | 502 MB |
Over three training seeds, LayoutLM averages 96.67 (kept β kept) and 96.57 (removed β removed), and LayoutLMv3 averages 97.47 and 97.42.
The same for the BERT and RoBERTa checkpoints trained on the original labels, on the original test labels (39,999 pages):
| model | folder | kept β kept | kept β removed | removed β kept | removed β removed | RVL-CDIP-N, kept | RVL-CDIP-N, removed |
|---|---|---|---|---|---|---|---|
| BERT-base (uncased) | bert_base |
92.34 | 90.86 | 91.58 | 91.45 | 86.1 | 87.1 |
| RoBERTa-base | roberta_base |
92.55 | 90.90 | 91.70 | 91.52 | 88.6 | 87.2 |
Over three training seeds, BERT averages 92.33 (kept β kept) and 91.40 (removed β removed) with the original labels, and RoBERTa 92.48 and 91.58.
Each folder is in the transformers layout, with the tokenizer (for LayoutLMv3, the processor)
next to the weights, and loads with from_pretrained(..., subfolder=...). The example reads a page from
stefan-hf/rvlcdip-redact, whose
textract_* configurations contain the text, words, and boxes these models were trained on:
import torch
from datasets import load_dataset
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo = "stefan-hf/rvlcdip-classifiers"
page = next(iter(load_dataset("stefan-hf/rvlcdip-redact", "textract_redact_noid", split="test", streaming=True)))
# BERT or RoBERTa: the plain text of the page, first 512 tokens
folder = "roberta_base/labels-corrected_codes-removed"
tok = AutoTokenizer.from_pretrained(repo, subfolder=folder)
model = AutoModelForSequenceClassification.from_pretrained(repo, subfolder=folder).eval()
text = page["text"] + "\n" if page["text"] else ""
with torch.no_grad():
logits = model(**tok(text, truncation=True, max_length=512, return_tensors="pt")).logits
print(model.config.id2label[logits.argmax(1).item()], "| label:", page["label_fixed"])
# LayoutLM: the words and their boxes (0-1000), first 512 tokens
folder = "layoutlmv1/labels-corrected_codes-removed"
tok = AutoTokenizer.from_pretrained(repo, subfolder=folder)
model = AutoModelForSequenceClassification.from_pretrained(repo, subfolder=folder).eval()
words, boxes = page["tokens"] or [""], page["bboxes"] or [[0, 0, 0, 0]]
enc = tok(words, is_split_into_words=True, truncation=True, max_length=512, return_tensors="pt")
bbox = torch.tensor([[[0, 0, 0, 0] if w is None else boxes[w] for w in enc.word_ids(0)]])
with torch.no_grad():
logits = model(**enc, bbox=bbox).logits
print(model.config.id2label[logits.argmax(1).item()], "| label:", page["label_fixed"])
LayoutLMv3 also takes the page image, which is in the images_* configuration of the same
version; the rows of the two configurations are in the same order. The page is converted to
grayscale as in training, and the processor resizes it to 224 Γ 224:
import torch
from datasets import load_dataset
from transformers import AutoModelForSequenceClassification, AutoProcessor
repo, folder = "stefan-hf/rvlcdip-classifiers", "layoutlmv3/labels-corrected_codes-removed"
page = next(iter(load_dataset("stefan-hf/rvlcdip-redact", "textract_redact_noid", split="test", streaming=True)))
scan = next(iter(load_dataset("stefan-hf/rvlcdip-redact", "images_redact_noid", split="test", streaming=True)))
assert page["id"] == scan["id"]
processor = AutoProcessor.from_pretrained(repo, subfolder=folder) # apply_ocr is off: words and boxes are given
model = AutoModelForSequenceClassification.from_pretrained(repo, subfolder=folder).eval()
words, boxes = page["tokens"] or [""], page["bboxes"] or [[0, 0, 0, 0]]
image = scan["image"].convert("L").convert("RGB")
enc = processor(image, words, boxes=boxes, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
logits = model(**enc).logits
print(model.config.id2label[logits.argmax(1).item()], "| label:", page["label_fixed"])
The text files used in training end with a newline on most pages, which the RoBERTa tokenizer
turns into a token, so the example appends one; BERT's tokenizer ignores it. Each folder's
rvlcdip_config.json records the base model, the selected epoch, and the accuracies above.
Training
- Data. RVL-CDIP train split (319,999 pages; 296,605 with the corrected labels), model selected by validation accuracy after every epoch.
- CNNs. ImageNet-pretrained torchvision models, SGD (learning rate 0.001, momentum 0.9, polynomial decay), batch 64, 30 epochs, cross-entropy.
- DiT-base.
microsoft/dit-base, AdamW (learning rate 3e-5, cosine schedule with warm-up), batch 32, 10 epochs. - LayoutLM, BERT, RoBERTa.
microsoft/layoutlm-base-uncased,google-bert/bert-base-uncased, andFacebookAI/roberta-base, AdamW (learning rate 5e-5, linear warm-up over the first 10% of steps, then linear decay), gradient clipping at 1.0, batch 16, 5 epochs, full precision, the first 512 tokens of the page. All of these runs selected the last epoch. - LayoutLMv3.
microsoft/layoutlmv3-base, Adam (learning rate 1e-5, no schedule), batch 64, mixed precision, 10 epochs, the first 512 tokens of the page and the page at 224 Γ 224. The released checkpoints are from epochs 6 (codes-kept) and 5 (codes-removed) of 10. - Input. For the image models, grayscale page resized to 224 Γ 224. At this size the identification codes are two or three pixels tall.
Each exported file was checked by reloading it through rvlcdip_models.py and comparing its
predictions on 512 test pages with those of the original checkpoint (agreement 99.8β100%).
The SVM exports were checked the same way through rvlcdip_svms.py (identical predictions). For
the fusion models that check used the image embeddings stored at training time; computed again
from the page images, the predictions agreed on 63 or 64 of 64 pages per model with CLIP and on
61 to 64 of 64 pages with SigLIP and ResNet-50. The text and layout
exports were reloaded with from_pretrained and reproduced the logits of the original checkpoints
on 512 test pages.
Limitations
- Trained and evaluated on tobacco-litigation scans only; accuracy on other document sources is lower (see the RVL-CDIP-N column).
- The
codes-keptmodels can rely on the identification codes and lose accuracy when the codes are absent. - RVL-CDIP has substantial overlap between its train and test splits, which inflates all test accuracies here.
- The text and layout models were trained on Amazon Textract output and were not tested with other OCR engines; the TF-IDF model loses about 9 points when trained and tested on Tesseract text.
- GoogLeNet must be built with
transform_input=Trueand without auxiliary heads, as the loader does.
License
The weights, configuration files, and loader code in this repository are released under the MIT license.