Text Generation
Transformers
Safetensors
PyTorch
nemotron_h
nvidia
conversational
custom_code
Eval Results

Benchmark: fine-tuned Nemotron 3 Nano vs. nine frontier LLMs on B2B propensity ranking

#70
by chrisanz19 - opened

Benchmark: fine-tuned Nemotron vs. nine frontier LLMs on B2B propensity ranking

We benchmarked a fine-tuned Nemotron 3 Nano (with Nemotron 3.5 Lightning as a bounded evidence channel) against nine flagship frontier language models on a blinded 18-business, 36,704-prospect B2B propensity task. Fine-tuned Nemotron led all four metrics against every frontier model in every prompting condition, on point estimates.

Paper: https://doi.org/10.5281/zenodo.22076381
Code, prompts, and paired bootstrap: https://github.com/NextLM/b2b-propensity-benchmark

Sample-size caveats and all confidence intervals in the paper.

Sign up or log in to comment