Instructions to use widebluesky/wbs-llm-base-demo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use widebluesky/wbs-llm-base-demo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="widebluesky/wbs-llm-base-demo") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("widebluesky/wbs-llm-base-demo") model = AutoModelForCausalLM.from_pretrained("widebluesky/wbs-llm-base-demo", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use widebluesky/wbs-llm-base-demo with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "widebluesky/wbs-llm-base-demo" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "widebluesky/wbs-llm-base-demo", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/widebluesky/wbs-llm-base-demo
- SGLang
How to use widebluesky/wbs-llm-base-demo with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "widebluesky/wbs-llm-base-demo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "widebluesky/wbs-llm-base-demo", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "widebluesky/wbs-llm-base-demo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "widebluesky/wbs-llm-base-demo", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use widebluesky/wbs-llm-base-demo with Docker Model Runner:
docker model run hf.co/widebluesky/wbs-llm-base-demo
widebluesky/wbs-llm-base-demo
Demo release v1
This is a trained tiny demo model for testing training, export, publication,
download and inference. A tiny Llama architecture was initialized randomly
and trained for two epochs on synthetic fixtures from advai-notebooks.
It is not derived from Meta Llama weights or another pretrained checkpoint.
It is a workflow placeholder for future domain-trained releases.
The model has no demonstrated general language or instruction-following capability.
Generation may be repetitive or incorrect. Synthetic metrics do not establish domain quality.
Model and training
| Setting | Value |
|---|---|
| Architecture | LlamaForCausalLM; random initialization |
| Parameters | 13,248 |
| Layers / hidden size / FFN size | 1 / 32 / 64 |
| Attention heads / key-value heads | 4 / 2 |
| Vocabulary / context length | 123 tokens / 128 tokens |
| Training objective | next-token causal language modeling over document text (continued pre-training) |
| Method | Full-parameter training; not a LoRA adapter |
| Epochs completed / selected | 2 / 2 |
| Selection | Lowest validation token-weighted NLL, including epoch zero; test is report-only |
| Dataset | 8 training, 2 validation and 2 test synthetic examples |
| Optimizer / learning rate / weight decay | AdamW / 0.003 / 0.01 |
| Batch size / gradient accumulation | 2 / 2 |
| Scheduler / gradient clipping | Constant / 1.0 |
| Seed / device / precision | 42 / CPU / fp32 |
| Export | Standalone full safetensors model |
| SDK versions | Transformers 5.18.0; PyTorch 2.14.1 |
Training supervises the next token in each document, including its EOS marker. Padding tokens are masked from the loss. A chat template is included as tokenizer metadata for compatibility; the Base training objective does not teach chat behavior.
Synthetic held-out test metrics
| Metric | Before training | Selected trained export |
|---|---|---|
| nll | 4.792373 | 4.695202 |
| perplexity | 120.587147 | 109.420896 |
NLL and perplexity are computed over each task's supervised tokens. Base and Instruct perplexity values are not directly comparable because their supervision ranges differ. These results use only two synthetic test examples.
Inference
In a checkout of advai-notebooks:
python -m pip install -e '.[llm-base]'
from advai_notebooks.llm.base.model import BaseModel
model = BaseModel.load("widebluesky/wbs-llm-base-demo", revision="demo-v1", device="cpu")
outputs = model.generate(["the sky is"], max_new_tokens=8)
print(outputs)
The full model also supports native Hugging Face loading:
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "widebluesky/wbs-llm-base-demo"
model = AutoModelForCausalLM.from_pretrained(repo_id, revision="demo-v1")
tokenizer = AutoTokenizer.from_pretrained(repo_id, revision="demo-v1")
For Instruct inference, render the exported chat template with an assistant generation
For Base inference, tokenize the document prompt and generate a continuation.
Pin the publication commit SHA in revision for reproducible deployment.
Verification and package contents
- All 11 stored weight tensors changed from the random baseline.
- Exported weights exactly match the validation-selected trained checkpoint.
- Fresh-process reload reproduced test metrics exactly.
- Native Hugging Face loading and project facade loading produced identical logits.
- CLI generation/chat and facade API outputs matched.
- File sizes and SHA256 hashes are recorded in
publish_manifest.json.
The package contains inference weights, architecture and generation configuration, tokenizer, chat template, task facade metadata, this English model card and a manifest. Raw data, optimizer state, resumable training checkpoints, logs and credentials are excluded.
Source implementation commit: f43a44e2431bae9d0451b445cc3bde5f1e2b8493.
Training and evaluation dataset
Dataset: widebluesky/wbs-llm-base-demo.
Verified dataset revision: de84c42b70656ff5875602b1abf829c8b2c962f9 (demo-v1).
This is the synthetic source dataset used for this trained demo release. Its record fingerprint matches the saved training-run provenance; the published dataset was reloaded and checked record-by-record before associating it with this model.
The base configuration contains 8 train, 2 validation and 2 test records.
Original document/conversation group IDs and split assignments are preserved.
from datasets import load_dataset
data = load_dataset(
"widebluesky/wbs-llm-base-demo", "base",
revision="de84c42b70656ff5875602b1abf829c8b2c962f9", token=True,
)
Access uses the existing Hugging Face login; both model and dataset repositories are private.
Original project-native data files are available under source/ in the dataset repository.
Canonical dataset fingerprint: 26aec5cfd5b21082af1e8f772dfe41f447ca0275f5cc6b8b8b1d98334e791335.
- Downloads last month
- 6