Text Generation
Transformers
Safetensors
PyTorch
nemotron_h
nvidia
conversational
custom_code
Eval Results

doesn't do kv caching on transformers

#14
by adaface-neurips - opened

If use_cache=True, I got a warning:
NemotronH requires an initialized NemotronHHybridDynamicCache to return a cache. None was provided, so no cache will be returned.
The speed is around 1-2 tokens/s on H200. So it seems kv cache is not enabled due to this warning.
I tried to monkey patch the code (e.g. using past_key_values as cache_params), but the code always runs into various errors, so I gave up.
Any suggestions? Can I only do fast inference using vllm?
Thanks.

NVIDIA org

Hi @adaface-neurips

Thank you for the message. We're aware that the current HF implementation has an issue with KV/MambaCache and are trying to address it.

Please note that HF is primarily for prototyping and other inference engines (vLLM, TRT-LLM, SGLang, Llama.cpp) optimized for Nemotron 3 Nano are available and support KV Cache. Please check Quick Start and try other inference engines meanwhile.

Thank you @suhara for your reply!
I understand that other inference engines are recommended. If I wish to finetune nemotron (e.g. with RL on math datasets), what framework would you recommend to use for training? Thank you very much again.

NVIDIA org

Hi @adaface-neurips
NemoRL supports fine-tuning nemotron-3-nano. Please see doc here
Also, the nemotron modeling code has been upstreamed to transformers lib. Please upgrade transformers to v5.3.0+. You don't need to add trust_remote_code=True when loading the model

Hi @adaface-neurips
NemoRL supports fine-tuning nemotron-3-nano. Please see doc here
Also, the nemotron modeling code has been upstreamed to transformers lib. Please upgrade transformers to v5.3.0+. You don't need to add trust_remote_code=True when loading the model

I'm using Transformer 5.4.0, but the following issue still persists. Could you please let me know what the reason might be?
WARNING:transformers_modules._1.modeling_nemotron_h:NemotronH requires an initialized `NemotronHHybridDynamicCache` to return a cache. None was provided, so no cache will be returned.

Hi @adaface-neurips
NemoRL supports fine-tuning nemotron-3-nano. Please see doc here
Also, the nemotron modeling code has been upstreamed to transformers lib. Please upgrade transformers to v5.3.0+. You don't need to add trust_remote_code=True when loading the model

I'm using Transformer 5.4.0, but the following issue still persists. Could you please let me know what the reason might be?
WARNING:transformers_modules._1.modeling_nemotron_h:NemotronH requires an initialized `NemotronHHybridDynamicCache` to return a cache. None was provided, so no cache will be returned.

fix?

Sign up or log in to comment