benchmarks

#50
by rfcoder0 - opened

i was training this to try to make it comparable with an 8b model like my rfcoder0/qwen3-4b-custom-Sile. i wanted to do the same thing with phi but it seems like the stock benchmarks are not correct arc 10 shot was way lower. Is there any others that have this issue.

Sign up or log in to comment