LLaMA 3 Knowledge-Editing GRPO Locality Checkpoint
This repository contains the Hugging Face actor checkpoint from the first checkpoint of the 5-more-epochs locality-only continuation experiment.
- Source Modal path:
/runs/checkpoints/ke-grpo-exact/llama3_phase2_hardlocality260_localityonly_epochs4to8_from_epoch3_h100x2_rollout8_genbs4/global_step_65/actor/huggingface - Training dataset:
official_m200_seed13_after_first2000_mined73_hardlocality260_locality_only - Checkpoint:
global_step_65 - Training objective: GRPO exact-match reward on hard locality/neighborhood prompts
- Base edited model lineage: LLaMA 3 + AlphaEdit/GRPO knowledge-editing experiments
Comparable N=200 Phase 3 evaluation for this checkpoint:
| Metric | Value |
|---|---|
| Rewrite prob / argmax | 99.50 / 97.00 |
| Paraphrase prob / argmax | 93.25 / 67.00 |
| Neighborhood prob / argmax | 82.40 / 22.60 |
| Gen entropy | 6.284 |
| Gen ref score | 0.314 |
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for ztiganj/llama3-ke-grpo-hardlocality-localityonly-step65
Base model
meta-llama/Meta-Llama-3-8B-Instruct