ealharbi's picture
Upload folder using huggingface_hub
35deade verified
|
Raw History Blame Contribute Delete
1.67 kB
metadata
license: apache-2.0
base_model: mistralai/Mistral-7B-Instruct-v0.3
library_name: peft
pipeline_tag: text-generation
language:
  - en
tags:
  - crystallography
  - structural-biology
  - lora
  - peft

Mistral 7B — crystallography LoRA adapter

LoRA adapter for mistralai/Mistral-7B-Instruct-v0.3, fine-tuned on confirmed question-and-answer pairs from the CCP4BB, CCP-EM, phenixbb and COOT mailing lists.

Training

Method LoRA, r=64, α=128, dropout=0.05
Learning rate 2.675e-05
Max length 2048 tokens
Epochs 2, early stopping
Seed 1

Evaluation

BERTScore F1 against held-out confirmed answers (433 questions, roberta-large, score(prediction, reference)): 0.8367 → 0.8536.

BERTScore measures overlap with the reference wording, not factual correctness; the paper reports three LLM judges alongside it.

Use

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.3")
model = AutoModelForCausalLM.from_pretrained(
    "mistralai/Mistral-7B-Instruct-v0.3", device_map="auto", dtype="bfloat16")
model = PeftModel.from_pretrained(model, "ealharbi/crystallography-cryoem-mistral-7b-lora")

The system prompt used in training was:

You are a helpful crystallography assistant. Be concise and precise.

The training pairs are single-turn, so ask one self-contained question at a time; multi-turn prompts fall outside the fine-tuning format.

Caveat

Answers are generated and may be wrong. Verify anything consequential against the program documentation and the primary literature.