Mistral Large 123B โ€” crystallography LoRA adapter

LoRA adapter for mistralai/Mistral-Large-Instruct-2407, fine-tuned on confirmed question-and-answer pairs from the CCP4BB, CCP-EM, phenixbb and COOT mailing lists.

Released under the Mistral Research Licence: research and non-commercial use only, inherited from the base model.

Training

Method LoRA, r=64, ฮฑ=128, dropout=0.05
Learning rate 2.675e-05
Max length 2048 tokens
Epochs 2, early stopping
Seed 1

Evaluation

BERTScore F1 against held-out confirmed answers (433 questions, roberta-large, score(prediction, reference)): 0.8320 โ†’ 0.8570.

BERTScore measures overlap with the reference wording, not factual correctness; the paper reports three LLM judges alongside it.

Use

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("mistralai/Mistral-Large-Instruct-2407")
model = AutoModelForCausalLM.from_pretrained(
    "mistralai/Mistral-Large-Instruct-2407", device_map="auto", dtype="bfloat16")
model = PeftModel.from_pretrained(model, "ealharbi/crystallography-cryoem-mistral-large-123b-lora")

The system prompt used in training was:

You are a helpful crystallography assistant. Be concise and precise.

The training pairs are single-turn, so ask one self-contained question at a time; multi-turn prompts fall outside the fine-tuning format.

Caveat

Answers are generated and may be wrong. Verify anything consequential against the program documentation and the primary literature.

Downloads last month
34
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ealharbi/crystallography-cryoem-mistral-large-123b-lora

Adapter
(2)
this model

Collection including ealharbi/crystallography-cryoem-mistral-large-123b-lora