Refusal Classifier Model

A simple NLP model for classifying text responses as either refusal or valid. Works on 90% of test cases.

🧠 What It Does

This model predicts whether a given response from an LLM is:

  • refusal – e.g., "I'm sorry, but I can't help with that."
  • valid – e.g., "'I can't remember,' she said quietly."

Useful for detecting AI refusals and filtering them during generation.

πŸ“š Training Details

  • Built using scikit-learn with:
    • TfidfVectorizer (1–2 n-grams, lowercase)
    • LogisticRegression
  • Trained on a small dataset of synthetic labeled examples (1175/1517 refusal/valid pairs)
  • Dataset stored in dataset.jsonl with format:
    {"text": "The capital of France is Paris.", "label": "valid"}
    

πŸš€ How to Use

  1. Install requirements
pip install scikit-learn joblib
  1. Load the model
import joblib
joblib.load("refusal_classifier.pkl")
  1. Make a prediction
text = "I'm sorry, but I can't help with that."
label = model.predict([text])[0]
print(label)  # "refusal"

πŸ“„ License

Licensed under CC BY 4.0: you may use this commercially or privately, but must include attribution in code or documentation.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support