--- datasets: - HuggingFaceTB/smollm-corpus language: - en library_name: transformers license: apache-2.0 tags: - pruna-ai - safetensors --- # Model Card for AINovice2005/SmolLM-360M-smashed This model was created using the [pruna](https://github.com/PrunaAI/pruna) library. Pruna is a model optimization framework built for developers, enabling you to deliver more efficient models with minimal implementation overhead. ## Usage First things first, you need to install the pruna library: ```bash pip install pruna ``` You can [use the transformers library to load the model](https://huggingface.co/AINovice2005/SmolLM-360M-smashed?library=transformers) but this might not include all optimizations by default. To ensure that all optimizations are applied, use the pruna library to load the model using the following code: ```python from pruna import PrunaModel loaded_model = PrunaModel.from_pretrained( "AINovice2005/SmolLM-360M-smashed" ) # we can then run inference using the methods supported by the base model ``` For inference, you can use the inference methods of the original model like shown in [the original model card](https://huggingface.co/HuggingFaceTB/SmolLM-360M?library=transformers). Alternatively, you can visit [the Pruna documentation](https://docs.pruna.ai/en/stable/) for more information. ## Smash Configuration The compression configuration of the model is stored in the `smash_config.json` file, which describes the optimization methods that were applied to the model. ```bash { "batcher": null, "cacher": null, "compiler": "torch_compile", "factorizer": null, "kernel": null, "pruner": null, "quantizer": "hqq", "hqq_backend": "torchao_int4", "hqq_compute_dtype": "torch.bfloat16", "hqq_force_hf_implementation": false, "hqq_group_size": 128, "hqq_use_torchao_kernels": true, "hqq_weight_bits": 4, "torch_compile_backend": "inductor", "torch_compile_dynamic": false, "torch_compile_fullgraph": true, "torch_compile_make_portable": false, "torch_compile_max_kv_cache_size": 800, "torch_compile_mode": "default", "torch_compile_seqlen_manual_cuda_graph": 400, "torch_compile_target": "module_list", "batch_size": 1, "device": "cuda:0", "device_map": null, "save_fns": [ "hqq", "save_before_apply" ], "load_fns": [ "hqq" ], "reapply_after_load": { "factorizer": null, "pruner": null, "quantizer": null, "kernel": null, "cacher": null, "compiler": "torch_compile", "batcher": null } } ``` ## 🌍 Join the Pruna AI community! [![Twitter](https://img.shields.io/twitter/follow/PrunaAI?style=social)](https://twitter.com/PrunaAI) [![GitHub](https://img.shields.io/github/followers/PrunaAI?label=Follow%20%40PrunaAI&style=social)](https://github.com/PrunaAI) [![LinkedIn](https://img.shields.io/badge/LinkedIn-Connect-blue)](https://www.linkedin.com/company/93832878/admin/feed/posts/?feedType=following) [![Discord](https://img.shields.io/badge/Discord-Join%20Us-blue?style=social&logo=discord)](https://discord.gg/JFQmtFKCjd) [![Reddit](https://img.shields.io/reddit/subreddit-subscribers/PrunaAI?style=social)](https://www.reddit.com/r/PrunaAI/)