You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This repo contains models for generating hate speech and NLI adversarial examples. The base architecture is the GPT-2 causal language model. Hate speech models are trained on the DynaHate dataset, while NLI models are trained on AdversarialNLI. Further details can be found in this paper.

Models are intended for testing/improving robustness of neural classifiers only.

Downloads last month: -; Downloads are not tracked for this model. How to track

Inference Providers NEW

This model is not currently available via any of the supported third-party Inference Providers, and HF Inference API was unable to determine this model's library.