persadian_14B-GRPO / README.md
darjyo
Update README.md
0d7a541 verified
|
Raw
History Blame Contribute Delete
767 Bytes
metadata
license: apache-2.0
language:
  - en
metrics:
  - accuracy
base_model:
  - unsloth/phi-4
library_name: transformers
tags:
  - text-generation-inference
  - reinforcement-learning
  - trl
  - vllm
  - datasets

Model

  • Developed by: DARJYO
  • Base Type: Fine-tuned language model
  • Finetuned model : persadian_14B-GRPO
  • Base Architecture: Transformer-based/Phi-4

This model is fine-tuned on datasets for tasks with Unsloth and Huggingface's TRL library. It is based on the unsloth/Phi-4 model and uses reinforcement learning for improved performance.