#train-novi
@NoviAIBot Hello, i would like to want you to train a 1M parameter model. Training data: https://huggingface.co/datasets/codelion/fineweb-edu-100M Name: Yesvi-1M
I’m unable to start the training job right now because the training script (TRAIN_SCRIPT) hasn’t been configured for this environment. Without a configured training script, the train_model tool cannot execute.
If you’d like to proceed with training, you’ll need to have an administrator set up the appropriate training script (or provide one) so that the bot can launch the job. Once that’s in place, feel free to let me know and I’ll kick off the training for a 1 M‑parameter model named Yesvi‑1M using the dataset codelion/fineweb-edu-100M.
Is there anything else I can help with in the meantime (e.g., opening a discussion, checking the repo, etc.)?
@NoviAIBot Train a 1M parameter model named Yesvi-1M using the dataset codelion/fineweb-edu-100M. For this test, train on only 100,000 tokens. Use your requested parameter count to automatically configure the model architecture, and publish the finished model to GGUFGuy/Yesvi-1M.
I attempted to start the training job for GGUFGuy/Yesvi-1M with the requested configuration (1 M parameter model, trained on 100 k tokens from codelion/fineweb-edu-100M). The training script was able to:
- Load the GPT‑2 tokenizer
- Infer a small model architecture that matches the ~1 M parameter target (n_embd=20, n_layer=1, n_head=1, n_inner=40, …)
- Begin loading the dataset in streaming mode
However, the script failed during the setup of the Hugging Face TrainingArguments with the following error:
TypeError: TrainingArguments.__init__() got an unexpected keyword argument 'overwrite_output_dir'
This indicates that the training script is passing an argument (overwrite_output_dir) that the current version of the transformers library does not recognize. Until the script is updated to use a supported argument (or the library version is adjusted), the train_model tool cannot complete the job.
Next steps you can take
- Ask an administrator to inspect and fix the training script (
train_model.py) – specifically the call toTrainingArguments– so it uses valid parameters for the installedtransformersversion. - If you have an alternative training script you’d like to use, you could provide it (or point to a repository) and we could try again once the script is in place.
- If you’d like to proceed with a different model size, dataset, or token limit, let me know and I’ll re‑attempt training once the script issue is resolved.
Please let me know how you’d like to move forward!
@NoviAIBot Train a 1M parameter model named Yesvi-1M using the dataset codelion/fineweb-edu-100M. For this test, train on only 100,000 tokens. Use your requested parameter count to automatically configure the model architecture, and publish the finished model to GGUFGuy/Yesvi-1M.
I understand you’d like to train a 1 M‑parameter model called Yesvi‑1M on the first 100 k tokens of codelion/fineweb-edu-100M and publish the result to GGUFGuy/Yesvi-1M.
Unfortunately, the training script that powers the train_model tool is currently failing with:
TypeError: TrainingArguments.__init__() got an unexpected keyword argument 'overwrite_output_dir'
This means the script is passing an argument (overwrite_output_dir) that the installed version of the transformers library does not recognize. Until the script is updated to use a supported argument (or the library version is adjusted), the train_model tool cannot complete the job.
What can be done next?
- Ask an administrator to inspect and fix the training script (
train_model.py)—specifically the call toTrainingArguments—so it uses valid parameters for the currenttransformersversion. - If you have an alternative training script you’d like to use, you can share it (or point to a repository) and we can try again once the script is in place.
- Once the script is corrected, just let me know and I’ll launch the training for GGUFGuy/Yesvi-1M with the requested settings (1 M parameters, 100 k tokens, automatic architecture inference).
Is there anything else I can help with while we wait for the script to be fixed (e.g., opening a discussion, checking the repo, etc.)?
I tried launching the training again with the updated configuration you provided, but the script is still encountering an issue when it sets up the Hugging Face TrainingArguments:
TypeError: TrainingArguments.__init__() got an unexpected keyword argument 'warmup_ratio'
This means the version of the transformers library available in this environment does not recognize the warmup_ratio argument (it was introduced in a later version). Until the training script is adjusted to use only arguments supported by the installed transformers version (for example, replacing warmup_ratio with warmup_steps or removing it altogether), the train_model tool cannot complete the job.
What you can do next:
Verify that the fixed script is actually in place.
If you’ve updatedtrain_model.pyto remove or replace unsupported arguments likewarmup_ratio, make sure the updated file is the one the bot is using (e.g., push the changes to the repository or otherwise ensure the environment picks up the new version).Provide an alternative training script (or a link to a repository) that uses only supported
transformersarguments for the current version, and I’ll attempt to run it.Specify a different
transformersversion if you have the ability to change the environment, but note that the bot cannot install packages on its own.
Once the script is compatible with the installed transformers library, just let me know and I’ll relaunch the training for GGUFGuy/Yesvi-1M with:
- 1 M‑parameter model (auto‑inferred architecture)
- Dataset:
codelion/fineweb-edu-100M - Training tokens: 100 000
- Output repo:
GGUFGuy/Yesvi-1M
Is there anything else I can help with while we wait for the script to be confirmed as fixed? (e.g., opening a discussion, checking the repo, etc.)
I tried to launch the training job for GGUFGuy/Yesvi‑1M with the settings you requested (1 M‑parameter model, 100 k tokens from codelion/fineweb-edu-100M, auto‑inferred architecture).
The script got past the tokenizer and model‑architecture selection steps, but it failed when creating the Hugging Face Trainer:
TypeError: Trainer.__init__() got an unexpected keyword argument 'tokenizer'
This means the current version of the transformers/trainer library in this environment does not accept a tokenizer argument when constructing a Trainer. Until the training script is adjusted to use only the arguments supported by the installed library (for example, by removing the unsupported tokenizer kwarg or replacing it with the correct one), the train_model tool cannot complete the job.
What can be done next?
Verify the fix – Make sure the updated
train_model.py(where theTrainercall lives) is actually the file the bot is using. Push any changes to the repository or otherwise ensure the environment picks up the new version.Replace unsupported kwargs – Look for lines similar to
trainer = Trainer(..., tokenizer=tokenizer, ...)and either delete the
tokenizerargument or pass it in the way the currenttransformersversion expects (often the tokenizer is handled via thedata_collatoror is unnecessary).Provide an alternative script – If you have a different training script that works with the installed library version, share it (or a link to a repo) and I’ll run it instead.
Clarify the destination repo – You asked to publish the model “under your account.” Publishing to a namespace you don’t own (e.g.,
NoviAIBot/…) isn’t permitted unless you own that namespace or are an organization admin. If you’d like the model hosted in your own namespace, just confirm the repo ID (e.g.,GGUFGuy/Yesvi-1M) and I’ll attempt the training again once the script is compatible.
Let me know once the script has been adjusted (or if you’d like to specify a different target repo), and I’ll relaunch the training immediately. If there’s anything else I can help with in the meantime—opening a discussion, checking the repo, etc.—just ask!
@NoviAIBot I fixed the tokenizer thing. Now train! Also btw ur username is "NoviAIBot" so you could publish there
I gave the training another shot with the settings you asked for (1 M‑parameter model, 100 k tokens from codelion/fineweb-edu-100M, targeting GGUFGuy/Yesvi-1M). The script got past the tokenizer and model‑architecture selection, but it hit the same obstacle as before:
ValueError: The train_dataset does not implement __len__, so the number of training steps cannot be inferred; max_steps has to be specified.
The training script is trying to instantiate a 🤗 Transformers Trainer with a streaming dataset that doesn’t report its length. Because the Trainer can’t infer how many steps to run, it requires you to explicitly set max_steps (or provide a dataset with a known __len__).
What needs to happen next
Update the training script so that, when using a streaming dataset, it either:
- calculates an appropriate
max_stepsfrom the requested token count, batch size, and sequence length, or - falls back to a non‑streaming version of the dataset for this short test run.
- calculates an appropriate
Make sure the updated
train_model.pyis the one the bot is using (push the changes to the repository or otherwise refresh the environment).Once the script can handle the streaming dataset without throwing the
__len__error, just let me know and I’ll launch the training again with the same configuration:- Model: ~1 M parameters (auto‑inferred architecture)
- Dataset:
codelion/fineweb-edu-100M - Training tokens: 100 000
- Output repo:
GGUFGuy/Yesvi-1M
If you’d like, I can also help you open a discussion, check the repo, or try any other troubleshooting steps while we wait for the script fix. Just say the word!
Nice!
Also we have a chat it at chat.smilyai.org