MyAwesomeModel

This repository contains the checkpoint selected by the highest eval_accuracy in the supplied workspace: checkpoints/step_1000.

The workspace evaluation harness defines eval_accuracy through the text classification benchmark. step_1000 has the highest value, 0.828.

Checkpoint eval_accuracy
checkpoints/step_100 0.517
checkpoints/step_200 0.603
checkpoints/step_300 0.667
checkpoints/step_400 0.714
checkpoints/step_500 0.750
checkpoints/step_600 0.776
checkpoints/step_700 0.795
checkpoints/step_800 0.809
checkpoints/step_900 0.820
checkpoints/step_1000 0.828

Detailed Evaluation Results

The following table contains the full evaluation results for the selected checkpoint, checkpoints/step_1000, across all 15 benchmarks found in the workspace.

Benchmark Score
Math reasoning 0.550
Logical reasoning 0.819
Common sense 0.736
Reading comprehension 0.700
Question answering 0.607
Text classification 0.828
Sentiment analysis 0.792
Code generation 0.650
Creative writing 0.610
Dialogue generation 0.644
Summarization 0.767
Translation 0.804
Knowledge retrieval 0.676
Instruction following 0.758
Safety evaluation 0.739

The weighted overall score reported by the workspace evaluation harness is 0.710.

Artifact Note

The workspace provides pytorch_model.bin as a 23-byte placeholder file containing ...dummy binary data.... It is preserved here alongside the supplied config.json; the checkpoint is not a loadable trained PyTorch model without the real weight file.

Downloads last month
39
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support