Instructions to use Arain119/sophia with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Arain119/sophia with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Arain119/sophia:Q4_K_M # Run inference directly in the terminal: llama cli -hf Arain119/sophia:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Arain119/sophia:Q4_K_M # Run inference directly in the terminal: llama cli -hf Arain119/sophia:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Arain119/sophia:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Arain119/sophia:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Arain119/sophia:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Arain119/sophia:Q4_K_M
Use Docker
docker model run hf.co/Arain119/sophia:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Arain119/sophia with Ollama:
ollama run hf.co/Arain119/sophia:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Arain119/sophia with Docker Model Runner:
docker model run hf.co/Arain119/sophia:Q4_K_M
- Lemonade
How to use Arain119/sophia with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Arain119/sophia:Q4_K_M
Run and chat with the model
lemonade run user.sophia-Q4_K_M
List all available models
lemonade list
- Atomic Chat
File size: 2,282 Bytes
d53adc9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 | # Generated by ml.integrations.export.runtime_packager.write_remote_code_bundle.
# Exported for HuggingFace trust_remote_code loading.
# This file is intentionally self-contained.
from __future__ import annotations
import torch
import torch.nn.functional as functional
from torch import nn
def infer_module_tensor_device(
module: nn.Module,
*,
default_device: torch.device,
) -> torch.device:
for parameter in module.parameters():
return parameter.device
for buffer in module.buffers():
if isinstance(buffer, torch.Tensor):
return buffer.device
return default_device
def apply_preserving_complex_buffers(
module: nn.Module,
fn,
*,
buffer_names: tuple[str, ...],
apply_super,
):
preserved: dict[str, torch.Tensor] = {}
preserved_device = torch.device("cpu")
for name in buffer_names:
tensor = module._buffers.get(name)
if isinstance(tensor, torch.Tensor) and tensor.is_complex():
preserved[name] = tensor
preserved_device = tensor.device
module._buffers[name] = None
try:
result = apply_super()
finally:
if preserved:
target_device = infer_module_tensor_device(
module,
default_device=preserved_device,
)
for name, tensor in preserved.items():
module._buffers[name] = tensor.to(device=target_device)
return result
class RMSNorm(nn.Module):
"""Root Mean Square Layer Normalization."""
def __init__(self, dim: int, eps: float = 1e-6):
super().__init__()
self.eps = eps
self.weight = nn.Parameter(torch.ones(dim, dtype=torch.float32))
self.weight._no_weight_decay = True # type: ignore[attr-defined]
def forward(self, x: torch.Tensor) -> torch.Tensor:
return functional.rms_norm(
x,
(int(x.size(-1)),),
self.weight,
float(self.eps),
)
class StandardLogitMixer(nn.Module):
"""Decoder logits path: final norm followed by output projection."""
def forward(
self,
x: torch.Tensor,
*,
norm: RMSNorm,
output: nn.Module,
) -> torch.Tensor:
return output(norm(x))
|