image2vienna 0.1.0: Vienna Classification edition 10 (EN) via intfloat/multilingual-e5-base

Both models above are listed as base models so that they show in the model tree (the Hub only has "merge" as a relation for two models); no weights are merged. The index in this repository derives from the first, the embedder. The second is the vision model, Florence-2 base (230M), whose detailed caption of the image is the query; the package runs it through transformers on CPU or GPU and the browser demo runs its ONNX twin.

Maps a trade mark image to ranked codes of the Vienna Classification (the figurative elements of marks, WIPO). The vision model writes a caption naming the objects, letters and colours of the image; that caption is embedded with intfloat/multilingual-e5-base and scored against the classification entries, each embedded once from its full path (category > division > section), walking the hierarchy to answer at the level asked. No training.

This repository is a custom Inference Endpoints handler (handler.py) for the second stage. Deploy it as an Inference Endpoint and send the description:

{"inputs": "A crescent moon with three stars above it. The stars are yellow, the moon blue.",
 "parameters": {"level": "section", "top_k": 5}}

inputs is a string or a list of strings (one result list each). An image ({"inputs": {"image": "<base64>"}}) is accepted only when the endpoint can reach a vision model named in IMAGE2VIENNA_DESCRIBER; standard endpoints cannot, so describe locally or use the package. Parameters:

parameter values default meaning
level category, division, section, auto section where in the hierarchy to answer; auto descends while the evidence supports it
top_k integer 10 results per input (distinct branches, so possibly fewer)
chunking whole, mean, max whole the description as one query, or by sentence (mean, or best sentence per entry)
principal_only true, false false leave the auxiliary (A) sections out
exclude_codes list of codes [] leave whole subtrees out, e.g. ["29"] for the colours
gap float none drop results more than this below the best score
auto_margin float 0.02 auto level: keep descending while a child is within this of its parent
normalize true, false true collapse whitespace and lower-case shouting text

The response is a list of {code, padded, level, auxiliary, score, similarity, title, path}: code in WIPO's form (1.1.4), padded in EUIPO's (01.01.04), auxiliary true for an "A" section, path the full category-to-section text.

Use it locally

pip install "image2vienna[st] @ git+https://github.com/Jaimenms/image2vienna"
i2vienna download jaimenms/image2vienna-en                            # or --revision v0.1.0 to pin this edition
i2vienna classify logo.png --level section --top-k 5   # Florence-2 downloads on first use (460 MB), any laptop
echo "a lion's head above two crossed swords" | i2vienna classify - --level auto   # a description skips the vision stage
from image2vienna import ViennaClassifier
clf = ViennaClassifier("10")
for m in clf.classify("logo.png", level="section", top_k=5):
    print(m.pretty, m.auxiliary, round(m.score, 3), m.text)
print(clf.last_description)
clf.classify_text("three stars above a crescent moon", level="division")

Any vision model served by Ollama can replace Florence-2 (--describer ollama:qwen2.5vl:7b --prompt inventory); the evaluation measured Qwen2.5-VL 7B that way and found it close to Florence-2 at thirty times the size.

How well it works

On 300 EUIPO figurative marks with the examiners' codes (Large Labelled Logo Dataset, editions 5 to 8; a fifth of the codes are EUIPO extensions no WIPO edition contains), the whole caption as the query:

level Florence-2 base hit@1 / @3 / @10 Qwen2.5-VL 7B (reference) hit@1 / @3 / @10 frequency baseline hit@1 / @10
category 37.0% / 69.7% / 74.3% 44.3% / 68.7% / 72.7% 32.0% / 85.3%
division 31.0% / 53.3% / 62.7% 33.0% / 50.7% / 60.3% 26.0% / 61.3%
section 11.1% / 20.5% / 30.6% 9.8% / 20.2% / 33.0% 13.1% / 39.7%

It beats a frequency prior at rank 1 above the section and finds the right division for most pictorial elements (animals, heraldry, human beings, celestial bodies, where the prior finds none); the prior wins at depth because EUIPO's most frequent codes are conventions (letters in a special font, quadrilaterals, colours). Every configuration and the per-category table: PERFORMANCE.md in the GitHub repository.

Versions

Every publication of this repository is tagged v<package version> (v0.1.0, v0.1.0-2, ...), so i2vienna download jaimenms/image2vienna-en --revision <tag> and the Hub's revision parameter pin an edition. The classification edition is in the file names and in image2vienna.json; the package code vendored under image2vienna/ is the one the handler runs.

Contents: scheme/ (titles, notes and hierarchy), index/ (vectors, parquet; the _notes file is the index whose texts append the explanatory notes), image2vienna/ (the package, vendored), image2vienna.json (which index to serve), handler.py.

Data: Vienna Classification by WIPO (https://www.wipo.int/web/classification-vienna/). Embedder: intfloat/multilingual-e5-base from the Hub. Method, evaluation protocol and every measurement: https://github.com/Jaimenms/image2vienna. A browser demo that needs no server runs at https://huggingface.co/spaces/jaimenms/image2vienna.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jaimenms/image2vienna-en