Refusal Classifier Model
A simple NLP model for classifying text responses as either refusal or valid.
Works on 90% of test cases.
π§ What It Does
This model predicts whether a given response from an LLM is:
- refusal β e.g., "I'm sorry, but I can't help with that."
- valid β e.g., "'I can't remember,' she said quietly."
Useful for detecting AI refusals and filtering them during generation.
π Training Details
- Built using scikit-learn with:
TfidfVectorizer(1β2 n-grams, lowercase)LogisticRegression
- Trained on a small dataset of synthetic labeled examples (1175/1517 refusal/valid pairs)
- Dataset stored in
dataset.jsonlwith format:{"text": "The capital of France is Paris.", "label": "valid"}
π How to Use
- Install requirements
pip install scikit-learn joblib
- Load the model
import joblib
joblib.load("refusal_classifier.pkl")
- Make a prediction
text = "I'm sorry, but I can't help with that."
label = model.predict([text])[0]
print(label) # "refusal"
π License
Licensed under CC BY 4.0: you may use this commercially or privately, but must include attribution in code or documentation.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support