Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Yifan Peng's picture

Yifan Peng

pyf98
10 12 25
ReplicaTom's profile picture mp1704's profile picture almjanx's profile picture
·
https://pyf98.github.io
  • pyf98
  • yifan-peng

AI & ML interests

Multimodal LLMs, Speech-to-Speech, Speech Recognition

Organizations

ESPnet's profile picture Blog-explorers's profile picture YODAS Sharing inc's profile picture Nvidia Data&Tools team's profile picture Carnegie Mellon University's profile picture

pyf98 's collections 1

Open Whisper-style Speech Models (OWSM)
Fully open Whisper-style speech foundation models developed by CMU WAVLab: https://www.wavlab.org/activities/2024/owsm/
  • Running on Zero
    Agents
    11

    OWSM v4

    🌍
    11

    The same demo as espnet/owsm-v4, kept at this older address

  • Runtime error
    Agents
    Featured
    55

    OWSM Demo

    🔊
    55

  • espnet/yodas_owsmv4

    Viewer • Updated Sep 1, 2025 • 4 • 2.35k • 18
  • espnet/owsm_ctc_v4_1B

    Automatic Speech Recognition • Updated about 5 hours ago • 11.1k • 14
Open Whisper-style Speech Models (OWSM)
Fully open Whisper-style speech foundation models developed by CMU WAVLab: https://www.wavlab.org/activities/2024/owsm/
  • Running on Zero
    Agents
    11

    OWSM v4

    🌍
    11

    The same demo as espnet/owsm-v4, kept at this older address

  • Runtime error
    Agents
    Featured
    55

    OWSM Demo

    🔊
    55

  • espnet/yodas_owsmv4

    Viewer • Updated Sep 1, 2025 • 4 • 2.35k • 18
  • espnet/owsm_ctc_v4_1B

    Automatic Speech Recognition • Updated about 5 hours ago • 11.1k • 14
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs