Thai, Japanese, Korea, and Vietnamese in SmolLM3?

by wannaphong - opened Aug 5

Aug 5

Hello! I was read the pretraining config at https://huggingface.co/datasets/HuggingFaceTB/smollm3-configs. I found you was trained Thai, Japanese, Korea, and Vietnamese languages in the config but I see your instruction model isn't support their languages.

Is it not working with those languages (too small) or you don't have instruction dataset to try?

loubnabnl

Hugging Face Smol Models Research org Aug 6

•

edited Aug 6

We did not generate instruct data for those languages, so the instruct model doesn't support them (they also weren't upsampled in pretraining when compared to the official languages). But we did include them to open the possibility for continual pretraining.

wannaphong

Aug 6

•

edited Aug 6

We did not generate instruct data for those languages, so the instruct model doesn't support them (they also weren't upsampled in pretraining when compared to the official languages). But we did include them to open the possibility for continual pretraining.

Thank you! I was try to trained base model with Thai instruction dataset. The output is repetitions while generating (but it can still give good output). so I think I may continual pretraining the model if I have the resource :(.

wannaphong

Aug 6

@loubnabnl Could you check if Thai FineWeb2 in the config was trained on the full dataset or not? If it trained the full Thai subset in FineWeb2, I think I can use a few resources like the full Thai Wikipedia to do CPT model.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment