How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="BDRC/tibetan-ocr")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM

processor = AutoProcessor.from_pretrained("BDRC/tibetan-ocr")
model = AutoModelForMultimodalLM.from_pretrained("BDRC/tibetan-ocr", device_map="auto")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
inputs = processor.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

Tibetan OCR โ€” Yigdzin 1

A vision-first OCR model for Tibetan pecha โ€” modern publications, woodblock prints, and manuscripts (uchen and u-med), including pecha layout and orthographic shorthands. Built by the Buddhist Digital Resource Center (BDRC) and trained on the bec_mixed_elie_v8 mix of curated real + synthetic pages, with the evaluation-benchmark pages held out of training.

TL;DR

On a 1,070-page hand-transcribed benchmark (production serving: vLLM + sequential regime + DRY guard + temperature-retry), median page CER is ~1.4% and clean-page (non-catastrophic) mean CER is ~4.0%, with 0 hard repetition loops. The benchmark pages are held out of training. It is a specialized OCR model, not a general chat VLM.

Model details

Architecture PaddleOCR-VL-1.6 (vision encoder + depth-upscaled decoder)
Parameters 810M (413M vision ยท ~360M 26-layer decoder ยท ~4M embeddings/head)
Vision share ~51โ€“59% of parameters ("big eyes, small mouth" โ€” vision-first)
Tokenizer Tibetan unicode-stack tokenizer, vocab 3,560 (pruned + BoCorpus-warmed)
Decoder growth grow26: SOLAR-style depth-upscale 18Lโ†’26L, then annealed
Position regime sequential image-token M-RoPE (1-D); max_pixels 1280 budget
Languages Tibetan (bo); Tibetan script
License Apache-2.0 (derivative of PaddleOCR-VL-1.6; upstream NOTICE retained)
Funder Khyentse Foundation ("The BDRC Etext Corpus")

Intended use

  • OCR of Tibetan pecha pages: uchen and u-med (dbu med) scripts across modern print, woodblock, and manuscript sources.
  • Batch OCR at corpus scale (BDRC is applying it to ~25M scanned pages).

Out of scope: general vision-language chat / VQA; non-Tibetan scripts; layout analysis of illustrations, tables, or diagrams; line/region detection (this model transcribes page or line crops it is given).

How to use

Prompt (used at training and inference), rendered through the repo's chat_template.jinja:

Extract all Tibetan text. Preserve line breaks.

This checkpoint uses the native PaddleOCR-VL architecture (model_type: paddleocr_vl, PaddleOCRVLForConditionalGeneration) โ€” no trust_remote_code needed, but it requires transformers โ‰ฅ 5.15 (native paddleocr_vl support) or vLLM โ‰ฅ 0.26. Decode greedy (temperature=0).

pip install "transformers>=5.15" torch torchvision accelerate pillow
# torchvision is required by the PaddleOCR-VL image processor.

โš  Position regime โ€” required for correct output. This model was trained in the sequential image-token M-RoPE regime (1-D positions), not the default "grid" regime. If the processor emits a non-zero mm_token_type_ids, structured pages loop badly (some pages regress from ~0.10 to ~0.62 CER). Serve it sequential:

  • HF: zero the mask before the forward โ€” inputs["mm_token_type_ids"] = torch.zeros_like(inputs["input_ids"]).
  • vLLM: install the vllm_paddleocr_seqpos plugin and set OCR_VLLM_IMAGE_TOKEN_POSITIONS=sequential.

HF transformers (reference / accuracy baseline)

import torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText

model_id = "BDRC/tibetan-ocr"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, dtype="bfloat16", device_map="cuda")

image = Image.open("page.jpg").convert("RGB")
messages = [{"role": "user", "content": [
    {"type": "image"},
    {"type": "text", "text": "Extract all Tibetan text. Preserve line breaks."}]}]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to("cuda")
inputs["mm_token_type_ids"] = torch.zeros_like(inputs["input_ids"])  # sequential regime (required)
out = model.generate(**inputs, do_sample=False, max_new_tokens=4096)
print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])

To add the DRY anti-loop guard under HF too (optional; dry_logits_processor.py ships in this repo):

from transformers import LogitsProcessorList
from dry_logits_processor import make_hf_dry_processor

prompt_len = inputs["input_ids"].shape[1]
out = model.generate(**inputs, do_sample=False, max_new_tokens=4096,
    logits_processor=LogitsProcessorList([make_hf_dry_processor(
        prompt_len=prompt_len, multiplier=0.8, base=1.75, allowed_length=12)]))

vLLM (production)

Serve with vLLM โ‰ฅ 0.26. Two things matter for production quality:

1. Sequential image-token positions (required). Stock vLLM serves PaddleOCR-VL in the grid regime and never reads mm_token_type_ids; this model needs the sequential regime. A tiny vLLM plugin ships in this repo under vllm_paddleocr_seqpos/:

pip install "git+https://huggingface.co/BDRC/tibetan-ocr#subdirectory=vllm_paddleocr_seqpos"
export OCR_VLLM_IMAGE_TOKEN_POSITIONS=sequential   # 'grid' / unset = no-op

Without it, structured (book/list) pages skip or merge lines and CER regresses badly.

2. DRY anti-loop guard + temperature retry (recommended). Greedy decoding runs away into repetition on a small fraction of pages. The surgical fix is the DRY ("Don't Repeat Yourself") penalty: it ties an n-gram ban on corpus CER, zeroes hard loops, and โ€” unlike repetition_penalty / no_repeat_ngram_size โ€” barely touches clean pages or legitimate repetitive scripture (mantras, litanies). It ships here as the self-contained dry_logits_processor.py (torch-only):

from vllm import LLM, SamplingParams
from dry_logits_processor import DRYLogitsProcessor   # ships in this repo

llm = LLM(model="BDRC/tibetan-ocr", logits_processors=[DRYLogitsProcessor])  # + seqpos env above
params = SamplingParams(temperature=0, max_tokens=4096, extra_args={
    "dry_multiplier": 0.8, "dry_base": 1.75, "dry_allowed_length": 12,  # production config, no breakers
})

For the pages that still loop under greedy+DRY, re-decode at temperature and keep the cleanest sample. The strong signal is how often DRY fired: the measured knee is โ‰ฅ 100 fires โ€” re-decode only those pages at temperature=0.4, n=3 (DRY still on) and pick the sample with the lowest leftover repetition. On this benchmark that is ~2% of pages and clears the residual hard loops. dry_logits_processor.py writes per-request fire counts (pass dry_stats_id / dry_stats_path in extra_args, read them back with load_dry_stats_dir); deploy/fast_inference/bench.py in the training repo is the reference implementation of the full greedy โ†’ fire-count gate โ†’ temperature-retry loop.

Do not reach for aggressive repetition_penalty / no_repeat_ngram_size: they "correct away" legitimate repeated scripture. DRY with the config above is the tested, surgical choice.

Training data

Curated mix (bec_mixed_elie_v8), assembled from BDRC alignment collections + BoCorpus-rendered synthetic pages, with an easyโ†’hard curriculum, u-med over-weighting, and repetition-aware filtering. The evaluation-benchmark pages are held out of training (images_exclude_from_train.csv). Released component datasets:

Evaluation

Evaluated on the BDRC hand-transcribed benchmark 20260315 (1,070 pages) in the production configuration (vLLM + vllm_paddleocr_seqpos sequential regime + DRY guard mult=0.8 base=1.75 allowed_length=12 + temperature-retry temp=0.4 n=3 on pages with โ‰ฅ100 DRY fires). Metrics from the benchmark's compute_cer.py (botok normalization, whitespace stripping, tsheg folding, placeholder removal); pages with CER > 50% are counted as catastrophic and reported separately.

Metric (1,070 pages) Yigdzin 1
Median page CER 0.0142
Mean page CER (all) 0.0628
Catastrophic (>50%) share 3.3% (35)
Clean mean CER (non-catastrophic) 0.0404
Clean mean SER 0.0764
Hard repetition loops 0

Full method + the multi-system leaderboard: BDRC Tibetan OCR benchmark and the leaderboard.

Limitations & recommendations

  • Residual repetition loops on a small fraction of pages. Production uses the DRY guard + temperature retry (see How to use), which clears them (0 hard loops on the benchmark) while leaving legitimate repetitive scripture (mantras) intact; keep any anti-loop mitigation this gentle.
  • Line-break joins on some dense (6-line) woodblock pecha.
  • Struggles on illustrations, non-Tibetan text, and illegible/damaged pages.

Citation

A comprehensive article is forthcoming (Springer Language Resources and Evaluation). Until then, please cite the model and BDRC:

@misc{bdrc_tibetan_ocr_2026,
  title  = {Tibetan OCR (Yigdzin 1)},
  author = {Roux, Elie and Werner, Eric},
  year   = {2026},
  howpublished = {Buddhist Digital Resource Center, Hugging Face},
  note   = {https://huggingface.co/BDRC/tibetan-ocr}
}

Acknowledgements

Built on PaddleOCR-VL (Apache-2.0). Funded by the Khyentse Foundation. Thanks to the BDRC, Dharmaduta, and collaborating transcription teams.

Downloads last month
262,604
Safetensors
Model size
0.8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for BDRC/tibetan-ocr

Finetuned
(13)
this model
Quantizations
1 model

Space using BDRC/tibetan-ocr 1

Collection including BDRC/tibetan-ocr