Running Agents 435 Reward Bench Leaderboard π 435 Explore and compare model scores on RewardBench benchmarks
Running Featured 129 Open-LLM performances are plateauing, letβs make the leaderboard steep again π 129 Explore and compare advanced language models on a new leaderboard
google-bert/bert-base-multilingual-cased Fill-Mask β’ 0.2B β’ Updated Feb 19, 2024 β’ 1.93M β’ β’ 603
Sleeping Agents 4 Qwen Arabic Semantic Suite β‘ 4 Process and analyze Arabic texts for similarity, classification, and more