image2vienna 0.1.0: Vienna Classification edition 10 (EN) via intfloat/multilingual-e5-base
Both models above are listed as base models so that they show in the model tree (the Hub only has "merge" as a relation for two models); no weights are merged. The index in this repository derives from the first, the embedder. The second is the vision model, Florence-2 base (230M), whose detailed caption of the image is the query; the package runs it through transformers on CPU or GPU and the browser demo runs its ONNX twin.
Maps a trade mark image to ranked codes of the Vienna Classification (the figurative
elements of marks, WIPO). The vision model writes a caption naming the objects,
letters and colours of the image; that caption is embedded with intfloat/multilingual-e5-base
and scored against the classification entries, each embedded once from its full
path (category > division > section), walking the hierarchy to answer at the level
asked. No training.
This repository is a custom Inference Endpoints handler (handler.py) for the
second stage. Deploy it as an Inference Endpoint and send the description:
{"inputs": "A crescent moon with three stars above it. The stars are yellow, the moon blue.",
"parameters": {"level": "section", "top_k": 5}}
inputs is a string or a list of strings (one result list each). An image
({"inputs": {"image": "<base64>"}}) is accepted only when the endpoint can reach a
vision model named in IMAGE2VIENNA_DESCRIBER; standard endpoints cannot, so
describe locally or use the package. Parameters:
| parameter | values | default | meaning |
|---|---|---|---|
level |
category, division, section, auto | section | where in the hierarchy to answer; auto descends while the evidence supports it |
top_k |
integer | 10 | results per input (distinct branches, so possibly fewer) |
chunking |
whole, mean, max | whole | the description as one query, or by sentence (mean, or best sentence per entry) |
principal_only |
true, false | false | leave the auxiliary (A) sections out |
exclude_codes |
list of codes | [] | leave whole subtrees out, e.g. ["29"] for the colours |
gap |
float | none | drop results more than this below the best score |
auto_margin |
float | 0.02 | auto level: keep descending while a child is within this of its parent |
normalize |
true, false | true | collapse whitespace and lower-case shouting text |
The response is a list of {code, padded, level, auxiliary, score, similarity, title, path}: code in WIPO's form (1.1.4), padded in EUIPO's (01.01.04), auxiliary
true for an "A" section, path the full category-to-section text.
Use it locally
pip install "image2vienna[st] @ git+https://github.com/Jaimenms/image2vienna"
i2vienna download jaimenms/image2vienna-en # or --revision v0.1.0 to pin this edition
i2vienna classify logo.png --level section --top-k 5 # Florence-2 downloads on first use (460 MB), any laptop
echo "a lion's head above two crossed swords" | i2vienna classify - --level auto # a description skips the vision stage
from image2vienna import ViennaClassifier
clf = ViennaClassifier("10")
for m in clf.classify("logo.png", level="section", top_k=5):
print(m.pretty, m.auxiliary, round(m.score, 3), m.text)
print(clf.last_description)
clf.classify_text("three stars above a crescent moon", level="division")
Any vision model served by Ollama can replace Florence-2 (--describer ollama:qwen2.5vl:7b --prompt inventory); the evaluation measured Qwen2.5-VL 7B
that way and found it close to Florence-2 at thirty times the size.
How well it works
On 300 EUIPO figurative marks with the examiners' codes (Large Labelled Logo Dataset, editions 5 to 8; a fifth of the codes are EUIPO extensions no WIPO edition contains), the whole caption as the query:
| level | Florence-2 base hit@1 / @3 / @10 | Qwen2.5-VL 7B (reference) hit@1 / @3 / @10 | frequency baseline hit@1 / @10 |
|---|---|---|---|
| category | 37.0% / 69.7% / 74.3% | 44.3% / 68.7% / 72.7% | 32.0% / 85.3% |
| division | 31.0% / 53.3% / 62.7% | 33.0% / 50.7% / 60.3% | 26.0% / 61.3% |
| section | 11.1% / 20.5% / 30.6% | 9.8% / 20.2% / 33.0% | 13.1% / 39.7% |
It beats a frequency prior at rank 1 above the section and finds the right division
for most pictorial elements (animals, heraldry, human beings, celestial bodies, where
the prior finds none); the prior wins at depth because EUIPO's most frequent codes
are conventions (letters in a special font, quadrilaterals, colours). Every
configuration and the per-category table: PERFORMANCE.md in the GitHub repository.
Versions
Every publication of this repository is tagged v<package version> (v0.1.0,
v0.1.0-2, ...), so i2vienna download jaimenms/image2vienna-en --revision <tag> and the Hub's
revision parameter pin an edition. The classification edition is in the file names
and in image2vienna.json; the package code vendored under image2vienna/ is the
one the handler runs.
Contents: scheme/ (titles, notes and hierarchy), index/ (vectors, parquet; the
_notes file is the index whose texts append the explanatory notes), image2vienna/
(the package, vendored), image2vienna.json (which index to serve), handler.py.
Data: Vienna Classification by WIPO (https://www.wipo.int/web/classification-vienna/).
Embedder: intfloat/multilingual-e5-base from the Hub. Method, evaluation protocol and every
measurement: https://github.com/Jaimenms/image2vienna. A browser demo that needs no
server runs at https://huggingface.co/spaces/jaimenms/image2vienna.