CODE: https://github.com/frank-morales2020/AST/blob/main/AAI_DEEPSEEK_TOPO.ipynb
Usage
# ============================================================================
# DEEPSEEK-V2-LITE FP8 — INFERENCE TEST
# ============================================================================
from transformers import AutoModelForCausalLM, AutoTokenizer, PreTrainedModel
from unittest.mock import patch
import torch, transformers, logging
transformers.logging.set_verbosity_error()
logging.getLogger("transformers").setLevel(logging.ERROR)
MODEL_ID = 'frankmorales2020/deepseek-v2-lite-fp8-topo2026'
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
_orig_init = PreTrainedModel._initialize_weights
def _safe_init(self, module):
try: _orig_init(self, module)
except NotImplementedError as e:
if 'Float8_e4m3fn' in str(e) or 'normal_kernel_cpu' in str(e):
module._is_hf_initialized = True
else: raise
with patch.object(PreTrainedModel, '_initialize_weights', _safe_init):
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID, trust_remote_code=True, torch_dtype=torch.bfloat16,
device_map='cuda:0', low_cpu_mem_usage=True,
)
model.eval()
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
tokenizer.pad_token_id = tokenizer.eos_token_id
TEST_PROMPTS = [
'The capital of France is',
'In machine learning, catastrophic forgetting refers to',
'DeepSeek-V2-Lite is a mixture-of-experts model that',
]
print('='*65)
print(f'DEEPSEEK-V2-LITE FP8 INFERENCE TEST | {MODEL_ID}')
print('='*65)
for prompt in TEST_PROMPTS:
inp = tokenizer(prompt, return_tensors='pt').to(device)
with torch.no_grad():
out = model.generate(
**inp, max_new_tokens=40, do_sample=False,
use_cache=False, repetition_penalty=1.3,
pad_token_id=tokenizer.eos_token_id,
)
response = tokenizer.decode(out[0][inp['input_ids'].shape[1]:], skip_special_tokens=True)
print(f'\nPrompt : {prompt}')
print(f'Output : {response}')
print('\n' + '='*65)
OUTPUT:
Compressing model: 3463it [00:17, 195.25it/s]
Loading checkpoint shards: 100% 5/5 [00:02<00:00, 2.46it/s]=================================================================
DEEPSEEK-V2-LITE FP8 INFERENCE TEST | frankmorales2020/deepseek-v2-lite-fp8-topo2026
=================================================================
Prompt : The capital of France is
Output : Paris.
Paris, the city that never sleeps! The French call it “La Ville-Lumière” (the City Of Light). It’s a place where you can find
Prompt : In machine learning, catastrophic forgetting refers to
Output : the phenomenon where a model trained on multiple tasks or datasets forgets previously learned information when it is presented with new data. This can occur because each task requires different parameters and weights for optimal performance; if these
Prompt : DeepSeek-V2-Lite is a mixture-of-experts model that
Output : can be used to generate text. It was trained on the Deepseek dataset, which contains over 10 billion tokens of English language data from various sources such as books and news articles.
The
=================================================================
# Run once, then RESTART RUNTIME, then run this cell:
# !pip install -q "compressed-tensors>=0.15.0"
import warnings, logging, os
warnings.filterwarnings("ignore")
logging.getLogger("transformers").setLevel(logging.ERROR)
os.environ["TRANSFORMERS_VERBOSITY"] = "error"
os.environ["TRANSFORMERS_NO_ADVISORY_WARNINGS"] = "1"
try:
from transformers.utils import logging as _hf_log
_hf_log.set_verbosity_error()
except Exception:
pass
import transformers.utils.import_utils as _iu
import transformers.utils as _tu
for _m in (_iu, _tu):
if not hasattr(_m, "is_torch_fx_available"):
_m.is_torch_fx_available = lambda: False
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
import transformers.modeling_utils as _mu
# FP8 fix: skip re-init of already-loaded FP8 weights.
_mu.PreTrainedModel._initialize_weights = lambda self, *a, **k: None
_mu.PreTrainedModel.initialize_weights = lambda self, *a, **k: None
REPO = "frankmorales2020/deepseek-v2-lite-fp8-topo2026"
PROMPTS = ["What is the capital of Japan?",
"What is the Spanish word for water?",
"What is Einstein's mass-energy equivalence?"]
try:
tok = AutoTokenizer.from_pretrained(REPO, trust_remote_code=True)
except Exception:
tok = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V2-Lite", trust_remote_code=True)
if tok.pad_token is None:
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained(
REPO, dtype=torch.float16, device_map="auto",
trust_remote_code=True, attn_implementation="eager").eval()
nan = torch.isnan(model.get_input_embeddings().weight.float()).any().item()
print("Embedding has NaN?", nan)
if not nan:
for p in PROMPTS:
ins = tok(p, return_tensors="pt").to(model.device)
out = model.generate(**ins, max_new_tokens=40, do_sample=False,
repetition_penalty=1.2, no_repeat_ngram_size=3,
use_cache=False,
pad_token_id=tok.pad_token_id)
txt = tok.decode(out[0], skip_special_tokens=True)
print(f"\nQ: {p}\nA: {txt[len(p):].strip() if txt.startswith(p) else txt.strip()}")
OUTPUT-EXPECTED
Compressing model: 100%|██████████| 3463/3463 [00:01<00:00, 1977.29it/s]
Loading weights: 100% 12217/12217 [00:01<00:00, 8114.97it/s]Embedding has NaN? False
Decompressing model: 100%|██████████| 3463/3463 [00:00<00:00, 13103.20it/s]
Q: What is the capital of Japan?
A: The answer to this question may seem obvious, but it’s actually a bit more complicated than you might think. The country has two capitals: Tokyo and Kyoto. Both cities are important in Japanese
Q: What is the Spanish word for water?
A: The English translation of “agua” in a sentence. Agua means Water, and it’s pronounced like this: ah-gwah (like how you say ‘gwa’
Q: What is Einstein's mass-energy equivalence?
A: Einstein’s Mass Energy Equivalence states that the energy of a body at rest (E) equals its mass times the speed of light squared. This equation can be written as: E = mc
- Downloads last month
- 14
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for frankmorales2020/deepseek-v2-lite-fp8-topo2026
Unable to build the model tree, the base model loops to the model itself. Learn more.