You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

LLaMA 3 Knowledge-Editing GRPO Locality Checkpoint

This repository contains the Hugging Face actor checkpoint from the first checkpoint of the 5-more-epochs locality-only continuation experiment.

  • Source Modal path: /runs/checkpoints/ke-grpo-exact/llama3_phase2_hardlocality260_localityonly_epochs4to8_from_epoch3_h100x2_rollout8_genbs4/global_step_65/actor/huggingface
  • Training dataset: official_m200_seed13_after_first2000_mined73_hardlocality260_locality_only
  • Checkpoint: global_step_65
  • Training objective: GRPO exact-match reward on hard locality/neighborhood prompts
  • Base edited model lineage: LLaMA 3 + AlphaEdit/GRPO knowledge-editing experiments

Comparable N=200 Phase 3 evaluation for this checkpoint:

Metric Value
Rewrite prob / argmax 99.50 / 97.00
Paraphrase prob / argmax 93.25 / 67.00
Neighborhood prob / argmax 82.40 / 22.60
Gen entropy 6.284
Gen ref score 0.314
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ztiganj/llama3-ke-grpo-hardlocality-localityonly-step65

Finetuned
(1148)
this model