diff --git "a/logs/train_host0_latest.log" "b/logs/train_host0_latest.log" --- "a/logs/train_host0_latest.log" +++ "b/logs/train_host0_latest.log" @@ -1,3 +1,4 @@ +s are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -9,9 +10,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28100 | loss 5.6982 | text 2.1542 audio 5.2674 | grad 1.394 | 1.46s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.01 G=0.58 | nonfinite=0 +step 29100 | loss 5.7216 | text 2.1422 audio 5.2931 | grad 2.007 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -37,8 +37,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28125 | loss 5.7156 | text 2.1776 audio 5.2801 | grad 2.318 | 1.46s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.01 | nonfinite=0 +step 29125 | loss 5.7354 | text 2.1949 audio 5.2965 | grad 2.488 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.14 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -64,8 +64,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28150 | loss 5.7549 | text 2.1779 audio 5.3194 | grad 1.565 | 1.48s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 +step 29150 | loss 5.7076 | text 2.1771 audio 5.2722 | grad 2.926 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.01 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -91,8 +91,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28175 | loss 5.7338 | text 2.1718 audio 5.2995 | grad 2.218 | 1.50s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 +step 29175 | loss 5.7156 | text 2.1626 audio 5.2831 | grad 3.717 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.00 G=1.40 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -118,8 +118,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28200 | loss 5.7280 | text 2.1694 audio 5.2942 | grad 2.254 | 1.46s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.00 G=1.01 | nonfinite=0 +step 29200 | loss 5.6874 | text 2.1905 audio 5.2493 | grad 2.453 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -145,8 +145,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28225 | loss 5.7242 | text 2.1906 audio 5.2860 | grad 2.414 | 1.50s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.08 | nonfinite=0 +step 29225 | loss 5.7260 | text 2.1649 audio 5.2930 | grad 2.932 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.15 | spike L=-0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -172,21 +172,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28250 | loss 5.7200 | text 2.1909 audio 5.2818 | grad 2.138 | 1.47s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.95 | nonfinite=0 - running validation at step 28250... - val/composite=3.0358 (text=0.6628 audio=4.6179) val/loss=4.7504 cb0_acc=38.7% text_acc=90.3% - [val] per-codebook acc: cb0=38.7% cb1=19.9% cb2=16.9% cb3=10.9% cb4=8.2% cb5=6.7% cb6=5.8% cb7=5.6% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ifv9mdee/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b2edc-6cc7baaa6a4193ca555df9a9;f823947f-5a79-4045-878d-a19abf63f6c1) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 29250 | loss 5.7181 | text 2.1439 audio 5.2893 | grad 2.164 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.79 | nonfinite=0 + running validation at step 29250... + val/composite=3.0334 (text=0.6558 audio=4.6184) val/loss=4.7496 cb0_acc=38.7% text_acc=90.3% + [val] per-codebook acc: cb0=38.7% cb1=19.9% cb2=16.8% cb3=10.8% cb4=8.1% cb5=6.7% cb6=5.7% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0325) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -212,8 +203,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28275 | loss 5.7108 | text 2.1422 audio 5.2824 | grad 3.679 | 21.38s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.65 | nonfinite=0 +step 29275 | loss 5.7088 | text 2.1322 audio 5.2823 | grad 3.921 | 2.35s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.09 | spike L=-0.00 G=1.46 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -239,8 +230,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28300 | loss 5.7558 | text 2.1648 audio 5.3229 | grad 1.571 | 1.50s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.66 | nonfinite=0 +step 29300 | loss 5.7549 | text 2.1839 audio 5.3182 | grad 2.127 | 1.59s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -266,8 +257,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28325 | loss 5.7543 | text 2.1999 audio 5.3143 | grad 2.306 | 1.49s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 +step 29325 | loss 5.6933 | text 2.1503 audio 5.2633 | grad 1.508 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.55 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -293,8 +284,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28350 | loss 5.7402 | text 2.1853 audio 5.3031 | grad 1.754 | 1.57s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.76 | nonfinite=0 +step 29350 | loss 5.7047 | text 2.1504 audio 5.2747 | grad 2.847 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.15 | spike L=-0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -320,8 +311,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28375 | loss 5.7149 | text 2.1223 audio 5.2904 | grad 2.230 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.99 | nonfinite=0 +step 29375 | loss 5.7100 | text 2.1591 audio 5.2782 | grad 2.673 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -347,8 +338,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28400 | loss 5.7458 | text 2.1778 audio 5.3102 | grad 1.663 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.74 | nonfinite=0 +step 29400 | loss 5.7273 | text 2.1360 audio 5.3002 | grad 1.735 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.66 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -374,8 +365,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28425 | loss 5.7266 | text 2.1802 audio 5.2905 | grad 1.730 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.79 | nonfinite=0 +step 29425 | loss 5.7523 | text 2.1795 audio 5.3164 | grad 1.612 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=+0.01 G=0.63 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -401,8 +392,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28450 | loss 5.7010 | text 2.1936 audio 5.2623 | grad 1.702 | 1.54s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.01 G=0.80 | nonfinite=0 +step 29450 | loss 5.7343 | text 2.1852 audio 5.2972 | grad 2.594 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -428,8 +419,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28475 | loss 5.7597 | text 2.1612 audio 5.3275 | grad 3.507 | 1.53s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.07 | spike L=+0.00 G=1.67 | nonfinite=0 +step 29475 | loss 5.7163 | text 2.1439 audio 5.2875 | grad 2.112 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.85 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -455,12 +446,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28500 | loss 5.7497 | text 2.1814 audio 5.3134 | grad 1.431 | 1.55s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.64 | nonfinite=0 - running validation at step 28500... - val/composite=3.0382 (text=0.6602 audio=4.6235) val/loss=4.7556 cb0_acc=38.7% text_acc=90.3% - [val] per-codebook acc: cb0=38.7% cb1=19.7% cb2=16.8% cb3=11.0% cb4=8.2% cb5=6.8% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0358) +step 29500 | loss 5.7033 | text 2.1426 audio 5.2748 | grad 4.748 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.14 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=-0.00 G=1.95 | nonfinite=0 + running validation at step 29500... + val/composite=3.0330 (text=0.6584 audio=4.6160) val/loss=4.7477 cb0_acc=38.7% text_acc=90.4% + [val] per-codebook acc: cb0=38.7% cb1=20.0% cb2=16.9% cb3=10.9% cb4=8.2% cb5=6.8% cb6=5.7% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0325) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -486,8 +477,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28525 | loss 5.7250 | text 2.1899 audio 5.2871 | grad 1.655 | 2.39s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 +step 29525 | loss 5.7381 | text 2.1626 audio 5.3055 | grad 2.429 | 2.33s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -513,8 +504,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28550 | loss 5.7238 | text 2.1682 audio 5.2902 | grad 2.687 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.28 | nonfinite=0 +step 29550 | loss 5.6856 | text 2.1265 audio 5.2603 | grad 2.026 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -540,8 +531,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28575 | loss 5.7343 | text 2.1947 audio 5.2953 | grad 1.846 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.85 | nonfinite=0 +step 29575 | loss 5.7211 | text 2.1651 audio 5.2881 | grad 2.022 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -567,8 +558,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28600 | loss 5.7462 | text 2.1888 audio 5.3085 | grad 1.650 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.77 | nonfinite=0 +step 29600 | loss 5.7355 | text 2.1309 audio 5.3093 | grad 2.207 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -594,8 +585,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28625 | loss 5.7438 | text 2.1725 audio 5.3093 | grad 4.212 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.18 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=+0.00 G=2.02 | nonfinite=0 +step 29625 | loss 5.7147 | text 2.1485 audio 5.2850 | grad 1.976 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -621,8 +612,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28650 | loss 5.7516 | text 2.1744 audio 5.3167 | grad 3.602 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.14 | spike L=+0.00 G=1.57 | nonfinite=0 +step 29650 | loss 5.7366 | text 2.1921 audio 5.2981 | grad 1.785 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -648,8 +639,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28675 | loss 5.7130 | text 2.2028 audio 5.2724 | grad 1.650 | 1.49s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.68 | nonfinite=0 +step 29675 | loss 5.7341 | text 2.1653 audio 5.3010 | grad 1.758 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -675,8 +666,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28700 | loss 5.7224 | text 2.1803 audio 5.2864 | grad 1.498 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.64 | nonfinite=0 +step 29700 | loss 5.7398 | text 2.1773 audio 5.3043 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -702,8 +693,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28725 | loss 5.7378 | text 2.1832 audio 5.3011 | grad 3.430 | 1.52s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.51 | nonfinite=0 +step 29725 | loss 5.7252 | text 2.1935 audio 5.2865 | grad 1.577 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -729,12 +720,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28750 | loss 5.7436 | text 2.1957 audio 5.3045 | grad 3.054 | 1.51s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.28 | nonfinite=0 - running validation at step 28750... - val/composite=3.0401 (text=0.6651 audio=4.6234) val/loss=4.7565 cb0_acc=38.6% text_acc=90.2% - [val] per-codebook acc: cb0=38.6% cb1=19.7% cb2=16.7% cb3=10.9% cb4=8.3% cb5=6.7% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0358) +step 29750 | loss 5.7185 | text 2.1669 audio 5.2851 | grad 1.776 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 + running validation at step 29750... + val/composite=3.0292 (text=0.6607 audio=4.6083) val/loss=4.7404 cb0_acc=38.8% text_acc=90.4% + [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=10.9% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_it1ltfi7/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b3dc6-4c6482853ef24b3136885705;752b698a-0123-4e78-8974-6c9f1cb2aedf) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -760,8 +760,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28775 | loss 5.6989 | text 2.1380 audio 5.2713 | grad 4.402 | 2.45s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.12 | spike L=-0.01 G=1.80 | nonfinite=0 +step 29775 | loss 5.7285 | text 2.1674 audio 5.2950 | grad 2.373 | 21.34s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.10 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -787,8 +787,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28800 | loss 5.7532 | text 2.1850 audio 5.3162 | grad 4.631 | 1.50s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.14 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.09 | spike L=+0.00 G=1.75 | nonfinite=0 +step 29800 | loss 5.7073 | text 2.1300 audio 5.2813 | grad 2.486 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -814,8 +814,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28825 | loss 5.7355 | text 2.1855 audio 5.2984 | grad 1.890 | 1.54s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.66 | nonfinite=0 +step 29825 | loss 5.7351 | text 2.1504 audio 5.3050 | grad 2.191 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -841,8 +841,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28850 | loss 5.7498 | text 2.1720 audio 5.3154 | grad 2.545 | 1.54s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.09 | spike L=+0.00 G=0.93 | nonfinite=0 +step 29850 | loss 5.7432 | text 2.1591 audio 5.3114 | grad 2.085 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -868,8 +868,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28875 | loss 5.7074 | text 2.1567 audio 5.2761 | grad 2.836 | 1.51s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.00 G=1.04 | nonfinite=0 +step 29875 | loss 5.7321 | text 2.1600 audio 5.3001 | grad 2.919 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=+0.00 G=1.33 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -895,8 +895,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28900 | loss 5.7284 | text 2.1573 audio 5.2969 | grad 2.563 | 1.56s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.13 | spike L=-0.00 G=0.94 | nonfinite=0 +step 29900 | loss 5.7085 | text 2.1478 audio 5.2790 | grad 2.019 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -922,8 +922,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28925 | loss 5.7304 | text 2.1568 audio 5.2990 | grad 2.996 | 1.53s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.00 G=1.10 | nonfinite=0 +step 29925 | loss 5.7406 | text 2.1655 audio 5.3075 | grad 2.963 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -949,8 +949,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28950 | loss 5.7531 | text 2.1547 audio 5.3221 | grad 2.764 | 1.54s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 +step 29950 | loss 5.7224 | text 2.1685 audio 5.2887 | grad 2.859 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.14 | spike L=-0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -976,8 +976,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 28975 | loss 5.7483 | text 2.1758 audio 5.3131 | grad 2.025 | 1.55s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.74 | nonfinite=0 +step 29975 | loss 5.7182 | text 2.1755 audio 5.2831 | grad 1.804 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1003,16 +1003,43 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29000 | loss 5.7355 | text 2.1813 audio 5.2993 | grad 3.833 | 1.55s/step | peak -1.0G | host_rss 44.8G - [diag] gradnorm depth=0.17 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=+0.00 G=1.43 | nonfinite=0 - running validation at step 29000... - val/composite=3.0325 (text=0.6579 audio=4.6156) val/loss=4.7472 cb0_acc=38.7% text_acc=90.3% - [val] per-codebook acc: cb0=38.7% cb1=20.0% cb2=16.8% cb3=10.9% cb4=8.2% cb5=6.8% cb6=5.7% cb7=5.6% +step 30000 | loss 5.7549 | text 2.1721 audio 5.3205 | grad 2.832 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.01 G=1.22 | nonfinite=0 + [audio-demo] ar_cb0_acc=6.0% (35s) + running validation at step 30000... + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB  + + ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB  + + ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB  + + ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s + ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB + pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_030000 + val/composite=3.0203 (text=0.6417 audio=4.6060) val/loss=4.7344 cb0_acc=38.8% text_acc=90.5% + [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=10.9% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_whjqxje2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_vpw1dggr/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b3571-3f2f796174b88dcd164801eb;a3100d00-eac8-45e9-a0ef-60d5e33f8af9) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4150-7051ba0e2cb825374650c539;7f57bd5b-8fd3-43c1-b4ed-23eecab768d2) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -1020,7 +1047,7 @@ Make sure your token has the correct permissions. * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_99qadlou/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_029000 +[ckpt] uploading /tmp/ckpt_lczsg7r8/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_030000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1033,12 +1060,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_029000 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_030000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b3751-7f0c2d307f6398ad3a032efe;aa24e2fe-5d50-475d-af1e-602543b26c37) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4332-2102d4f8046862615ec783b4;bbf2f91d-62d4-4747-a05d-8b2bb6abbaad) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -1054,9 +1082,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30025 | loss 5.7367 | text 2.1700 audio 5.3027 | grad 2.055 | 41.20s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29025 | loss 5.7472 | text 2.1997 audio 5.3072 | grad 1.957 | 39.76s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1081,9 +1109,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30050 | loss 5.7188 | text 2.1281 audio 5.2931 | grad 1.639 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29050 | loss 5.7129 | text 2.1611 audio 5.2806 | grad 1.938 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1108,9 +1136,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30075 | loss 5.7561 | text 2.1976 audio 5.3166 | grad 1.470 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29075 | loss 5.7754 | text 2.1543 audio 5.3446 | grad 3.378 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=+0.01 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1135,9 +1163,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30100 | loss 5.7109 | text 2.1508 audio 5.2807 | grad 2.789 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29100 | loss 5.7216 | text 2.1422 audio 5.2931 | grad 2.007 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1162,9 +1190,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30125 | loss 5.7309 | text 2.1609 audio 5.2987 | grad 2.277 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29125 | loss 5.7354 | text 2.1949 audio 5.2965 | grad 2.488 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.14 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1189,9 +1217,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30150 | loss 5.7423 | text 2.1902 audio 5.3042 | grad 1.819 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29150 | loss 5.7076 | text 2.1771 audio 5.2722 | grad 2.926 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.01 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1216,9 +1244,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30175 | loss 5.7186 | text 2.1692 audio 5.2848 | grad 2.105 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29175 | loss 5.7156 | text 2.1626 audio 5.2831 | grad 3.717 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.00 G=1.40 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1243,9 +1271,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30200 | loss 5.7341 | text 2.1761 audio 5.2988 | grad 1.584 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29200 | loss 5.6874 | text 2.1905 audio 5.2493 | grad 2.453 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1270,9 +1298,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + [ALERT] grad spike 4.1x (grad 8.69 vs EMA 2.13) step 30225 +step 30225 | loss 5.7041 | text 2.1497 audio 5.2742 | grad 8.692 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.07 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.06 | spike L=-0.00 G=4.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29225 | loss 5.7260 | text 2.1649 audio 5.2930 | grad 2.932 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.15 | spike L=-0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1297,13 +1326,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30250 | loss 5.7108 | text 2.1812 audio 5.2746 | grad 1.567 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.56 | nonfinite=0 + running validation at step 30250... + val/composite=3.0263 (text=0.6452 audio=4.6137) val/loss=4.7427 cb0_acc=38.8% text_acc=90.7% + [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29250 | loss 5.7181 | text 2.1439 audio 5.2893 | grad 2.164 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.79 | nonfinite=0 - running validation at step 29250... - val/composite=3.0334 (text=0.6558 audio=4.6184) val/loss=4.7496 cb0_acc=38.7% text_acc=90.3% - [val] per-codebook acc: cb0=38.7% cb1=19.9% cb2=16.8% cb3=10.8% cb4=8.1% cb5=6.7% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0325) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1328,9 +1357,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30275 | loss 5.7370 | text 2.1381 audio 5.3094 | grad 1.647 | 2.35s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29275 | loss 5.7088 | text 2.1322 audio 5.2823 | grad 3.921 | 2.35s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.09 | spike L=-0.00 G=1.46 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1355,9 +1384,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30300 | loss 5.7740 | text 2.1916 audio 5.3357 | grad 3.100 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=+0.01 G=1.21 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29300 | loss 5.7549 | text 2.1839 audio 5.3182 | grad 2.127 | 1.59s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1382,9 +1411,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30325 | loss 5.6988 | text 2.1558 audio 5.2676 | grad 2.002 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29325 | loss 5.6933 | text 2.1503 audio 5.2633 | grad 1.508 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.55 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1409,9 +1438,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30350 | loss 5.7251 | text 2.1526 audio 5.2946 | grad 2.659 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.08 | spike L=-0.00 G=1.04 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29350 | loss 5.7047 | text 2.1504 audio 5.2747 | grad 2.847 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.15 | spike L=-0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1436,9 +1465,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30375 | loss 5.7275 | text 2.1723 audio 5.2930 | grad 1.326 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=-0.00 G=0.52 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29375 | loss 5.7100 | text 2.1591 audio 5.2782 | grad 2.673 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1463,9 +1492,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30400 | loss 5.7535 | text 2.1742 audio 5.3186 | grad 1.857 | 1.44s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29400 | loss 5.7273 | text 2.1360 audio 5.3002 | grad 1.735 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.66 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1490,9 +1519,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30425 | loss 5.6897 | text 2.1361 audio 5.2625 | grad 1.655 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29425 | loss 5.7523 | text 2.1795 audio 5.3164 | grad 1.612 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=+0.01 G=0.63 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1517,9 +1546,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30450 | loss 5.7256 | text 2.1539 audio 5.2948 | grad 3.586 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.55 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29450 | loss 5.7343 | text 2.1852 audio 5.2972 | grad 2.594 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1544,9 +1573,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30475 | loss 5.7147 | text 2.1494 audio 5.2848 | grad 2.515 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29475 | loss 5.7163 | text 2.1439 audio 5.2875 | grad 2.112 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.85 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1571,13 +1600,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30500 | loss 5.7465 | text 2.1749 audio 5.3115 | grad 1.691 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.69 | nonfinite=0 + running validation at step 30500... + val/composite=3.0241 (text=0.6463 audio=4.6093) val/loss=4.7386 cb0_acc=38.8% text_acc=90.6% + [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29500 | loss 5.7033 | text 2.1426 audio 5.2748 | grad 4.748 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.14 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=-0.00 G=1.95 | nonfinite=0 - running validation at step 29500... - val/composite=3.0330 (text=0.6584 audio=4.6160) val/loss=4.7477 cb0_acc=38.7% text_acc=90.4% - [val] per-codebook acc: cb0=38.7% cb1=20.0% cb2=16.9% cb3=10.9% cb4=8.2% cb5=6.8% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0325) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1602,9 +1631,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30525 | loss 5.7592 | text 2.1927 audio 5.3207 | grad 1.734 | 2.37s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29525 | loss 5.7381 | text 2.1626 audio 5.3055 | grad 2.429 | 2.33s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1629,9 +1658,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30550 | loss 5.7289 | text 2.1768 audio 5.2936 | grad 1.688 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29550 | loss 5.6856 | text 2.1265 audio 5.2603 | grad 2.026 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1656,9 +1685,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30575 | loss 5.7234 | text 2.1289 audio 5.2976 | grad 2.456 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29575 | loss 5.7211 | text 2.1651 audio 5.2881 | grad 2.022 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1683,9 +1712,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30600 | loss 5.6997 | text 2.1438 audio 5.2709 | grad 1.741 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29600 | loss 5.7355 | text 2.1309 audio 5.3093 | grad 2.207 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1710,9 +1739,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30625 | loss 5.7534 | text 2.1707 audio 5.3192 | grad 1.956 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29625 | loss 5.7147 | text 2.1485 audio 5.2850 | grad 1.976 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1737,9 +1766,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30650 | loss 5.7253 | text 2.1289 audio 5.2995 | grad 2.165 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29650 | loss 5.7366 | text 2.1921 audio 5.2981 | grad 1.785 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1764,9 +1793,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29675 | loss 5.7341 | text 2.1653 audio 5.3010 | grad 1.758 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.74 | nonfinite=0 +step 30675 | loss 5.7269 | text 2.1563 audio 5.2956 | grad 3.860 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1792,8 +1820,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29700 | loss 5.7398 | text 2.1773 audio 5.3043 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 +step 30700 | loss 5.7398 | text 2.1505 audio 5.3097 | grad 2.235 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1819,8 +1847,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29725 | loss 5.7252 | text 2.1935 audio 5.2865 | grad 1.577 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.70 | nonfinite=0 +step 30725 | loss 5.7489 | text 2.1535 audio 5.3182 | grad 2.031 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1846,16 +1874,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29750 | loss 5.7185 | text 2.1669 audio 5.2851 | grad 1.776 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 - running validation at step 29750... - val/composite=3.0292 (text=0.6607 audio=4.6083) val/loss=4.7404 cb0_acc=38.8% text_acc=90.4% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=10.9% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +step 30750 | loss 5.6953 | text 2.1131 audio 5.2727 | grad 3.574 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.17 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.55 | nonfinite=0 + running validation at step 30750... + val/composite=3.0163 (text=0.6371 audio=4.6024) val/loss=4.7299 cb0_acc=38.8% text_acc=90.9% + [val] per-codebook acc: cb0=38.8% cb1=19.9% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_it1ltfi7/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_v28gzt94/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b3dc6-4c6482853ef24b3136885705;752b698a-0123-4e78-8974-6c9f1cb2aedf) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b49a8-49c8b3a512d01dcf62ba2de6;c1d11542-5c8f-4019-8577-05bf7d3f521f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -1886,8 +1914,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29775 | loss 5.7285 | text 2.1674 audio 5.2950 | grad 2.373 | 21.34s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.10 | nonfinite=0 +step 30775 | loss 5.7336 | text 2.1703 audio 5.2995 | grad 2.588 | 21.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1913,8 +1941,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29800 | loss 5.7073 | text 2.1300 audio 5.2813 | grad 2.486 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.14 | nonfinite=0 +step 30800 | loss 5.7168 | text 2.1540 audio 5.2860 | grad 2.014 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1940,8 +1968,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29825 | loss 5.7351 | text 2.1504 audio 5.3050 | grad 2.191 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.99 | nonfinite=0 +step 30825 | loss 5.7184 | text 2.1635 audio 5.2857 | grad 1.822 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1967,8 +1995,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29850 | loss 5.7432 | text 2.1591 audio 5.3114 | grad 2.085 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.94 | nonfinite=0 +step 30850 | loss 5.7180 | text 2.1629 audio 5.2854 | grad 1.949 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1994,8 +2022,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29875 | loss 5.7321 | text 2.1600 audio 5.3001 | grad 2.919 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=+0.00 G=1.33 | nonfinite=0 +step 30875 | loss 5.7182 | text 2.1146 audio 5.2953 | grad 1.590 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2021,8 +2049,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29900 | loss 5.7085 | text 2.1478 audio 5.2790 | grad 2.019 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 +step 30900 | loss 5.6918 | text 2.1653 audio 5.2588 | grad 3.649 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.63 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2048,8 +2076,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29925 | loss 5.7406 | text 2.1655 audio 5.3075 | grad 2.963 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.32 | nonfinite=0 +step 30925 | loss 5.7255 | text 2.1531 audio 5.2949 | grad 1.466 | 1.43s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2075,8 +2103,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29950 | loss 5.7224 | text 2.1685 audio 5.2887 | grad 2.859 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.14 | spike L=-0.00 G=1.23 | nonfinite=0 +step 30950 | loss 5.7009 | text 2.1495 audio 5.2710 | grad 2.898 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2102,8 +2130,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29975 | loss 5.7182 | text 2.1755 audio 5.2831 | grad 1.804 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 +step 30975 | loss 5.7335 | text 2.1454 audio 5.3044 | grad 2.994 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2129,51 +2157,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30000 | loss 5.7549 | text 2.1721 audio 5.3205 | grad 2.832 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.01 G=1.22 | nonfinite=0 - [audio-demo] ar_cb0_acc=6.0% (35s) - running validation at step 30000... - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB  - - ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB  - - ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB  - - ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s - ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB - pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_030000 - val/composite=3.0203 (text=0.6417 audio=4.6060) val/loss=4.7344 cb0_acc=38.8% text_acc=90.5% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=10.9% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_vpw1dggr/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4150-7051ba0e2cb825374650c539;7f57bd5b-8fd3-43c1-b4ed-23eecab768d2) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 31000 | loss 5.7318 | text 2.1491 audio 5.3019 | grad 1.703 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.71 | nonfinite=0 + running validation at step 31000... + val/composite=3.0191 (text=0.6409 audio=4.6045) val/loss=4.7327 cb0_acc=39.0% text_acc=90.8% + [val] per-codebook acc: cb0=39.0% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0163) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_lczsg7r8/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_030000 +[ckpt] uploading /tmp/ckpt_y6mkwqap/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2185,18 +2177,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_030000 -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4332-2102d4f8046862615ec783b4;bbf2f91d-62d4-4747-a05d-8b2bb6abbaad) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4d0b-4f6ce7c1015680b77c18485b;4de181e9-748f-43df-8a2c-956408fcae48) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2208,10 +2198,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30025 | loss 5.7367 | text 2.1700 audio 5.3027 | grad 2.055 | 41.20s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31025 | loss 5.7013 | text 2.1689 audio 5.2675 | grad 2.218 | 20.63s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2235,10 +2225,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30050 | loss 5.7188 | text 2.1281 audio 5.2931 | grad 1.639 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31050 | loss 5.6785 | text 2.1305 audio 5.2524 | grad 2.106 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2262,10 +2252,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30075 | loss 5.7561 | text 2.1976 audio 5.3166 | grad 1.470 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31075 | loss 5.7067 | text 2.1458 audio 5.2776 | grad 2.192 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2289,10 +2279,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30100 | loss 5.7109 | text 2.1508 audio 5.2807 | grad 2.789 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31100 | loss 5.7458 | text 2.1828 audio 5.3092 | grad 2.135 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.06 projection=0.06 text_embed=0.14 | spike L=+0.01 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2316,10 +2306,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30125 | loss 5.7309 | text 2.1609 audio 5.2987 | grad 2.277 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31125 | loss 5.7439 | text 2.2002 audio 5.3039 | grad 2.645 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2343,10 +2333,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30150 | loss 5.7423 | text 2.1902 audio 5.3042 | grad 1.819 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31150 | loss 5.7190 | text 2.1559 audio 5.2878 | grad 1.900 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2370,10 +2360,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30175 | loss 5.7186 | text 2.1692 audio 5.2848 | grad 2.105 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31175 | loss 5.7093 | text 2.1537 audio 5.2786 | grad 2.073 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2397,10 +2387,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30200 | loss 5.7341 | text 2.1761 audio 5.2988 | grad 1.584 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31200 | loss 5.7151 | text 2.1503 audio 5.2850 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2424,11 +2414,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 4.1x (grad 8.69 vs EMA 2.13) step 30225 -step 30225 | loss 5.7041 | text 2.1497 audio 5.2742 | grad 8.692 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.07 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.06 | spike L=-0.00 G=4.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31225 | loss 5.7159 | text 2.1445 audio 5.2870 | grad 1.819 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2452,14 +2441,14 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30250 | loss 5.7108 | text 2.1812 audio 5.2746 | grad 1.567 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.56 | nonfinite=0 - running validation at step 30250... - val/composite=3.0263 (text=0.6452 audio=4.6137) val/loss=4.7427 cb0_acc=38.8% text_acc=90.7% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31250 | loss 5.7228 | text 2.1591 audio 5.2910 | grad 2.032 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 + running validation at step 31250... + val/composite=3.0194 (text=0.6437 audio=4.6031) val/loss=4.7319 cb0_acc=38.7% text_acc=90.7% + [val] per-codebook acc: cb0=38.7% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0163) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2483,10 +2472,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30275 | loss 5.7370 | text 2.1381 audio 5.3094 | grad 1.647 | 2.35s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31275 | loss 5.7330 | text 2.1806 audio 5.2969 | grad 1.797 | 2.38s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2510,10 +2499,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30300 | loss 5.7740 | text 2.1916 audio 5.3357 | grad 3.100 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=+0.01 G=1.21 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31300 | loss 5.7208 | text 2.1705 audio 5.2867 | grad 2.195 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2537,10 +2526,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30325 | loss 5.6988 | text 2.1558 audio 5.2676 | grad 2.002 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31325 | loss 5.7276 | text 2.1491 audio 5.2978 | grad 2.566 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.18 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2564,10 +2553,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30350 | loss 5.7251 | text 2.1526 audio 5.2946 | grad 2.659 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.08 | spike L=-0.00 G=1.04 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31350 | loss 5.7069 | text 2.1178 audio 5.2834 | grad 1.952 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2591,10 +2580,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30375 | loss 5.7275 | text 2.1723 audio 5.2930 | grad 1.326 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=-0.00 G=0.52 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31375 | loss 5.7308 | text 2.1203 audio 5.3067 | grad 1.813 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2618,10 +2607,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30400 | loss 5.7535 | text 2.1742 audio 5.3186 | grad 1.857 | 1.44s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31400 | loss 5.7433 | text 2.1584 audio 5.3116 | grad 2.645 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2645,10 +2634,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30425 | loss 5.6897 | text 2.1361 audio 5.2625 | grad 1.655 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31425 | loss 5.7413 | text 2.1406 audio 5.3132 | grad 1.860 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2672,10 +2661,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30450 | loss 5.7256 | text 2.1539 audio 5.2948 | grad 3.586 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.55 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31450 | loss 5.7094 | text 2.1542 audio 5.2786 | grad 1.536 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2699,10 +2688,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30475 | loss 5.7147 | text 2.1494 audio 5.2848 | grad 2.515 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31475 | loss 5.7086 | text 2.1526 audio 5.2781 | grad 3.518 | 1.59s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2726,14 +2715,23 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30500 | loss 5.7465 | text 2.1749 audio 5.3115 | grad 1.691 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.69 | nonfinite=0 - running validation at step 30500... - val/composite=3.0241 (text=0.6463 audio=4.6093) val/loss=4.7386 cb0_acc=38.8% text_acc=90.6% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31500 | loss 5.7234 | text 2.1505 audio 5.2933 | grad 2.082 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.93 | nonfinite=0 + running validation at step 31500... + val/composite=3.0129 (text=0.6302 audio=4.6013) val/loss=4.7273 cb0_acc=38.9% text_acc=90.9% + [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_97c5nxlf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b51f5-564c334001068ae16e82ce0f;aafdec3d-7141-4994-a73d-192be2632797) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2757,10 +2755,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30525 | loss 5.7592 | text 2.1927 audio 5.3207 | grad 1.734 | 2.37s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31525 | loss 5.7418 | text 2.1563 audio 5.3106 | grad 2.122 | 21.38s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2784,10 +2782,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30550 | loss 5.7289 | text 2.1768 audio 5.2936 | grad 1.688 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31550 | loss 5.7494 | text 2.1284 audio 5.3237 | grad 2.364 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2811,10 +2809,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30575 | loss 5.7234 | text 2.1289 audio 5.2976 | grad 2.456 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31575 | loss 5.7309 | text 2.1538 audio 5.3001 | grad 1.760 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2838,10 +2836,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30600 | loss 5.6997 | text 2.1438 audio 5.2709 | grad 1.741 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31600 | loss 5.7379 | text 2.1287 audio 5.3122 | grad 3.343 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.53 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2865,10 +2863,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30625 | loss 5.7534 | text 2.1707 audio 5.3192 | grad 1.956 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31625 | loss 5.6686 | text 2.1377 audio 5.2410 | grad 2.997 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.30 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2892,10 +2890,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30650 | loss 5.7253 | text 2.1289 audio 5.2995 | grad 2.165 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31650 | loss 5.6969 | text 2.1637 audio 5.2641 | grad 1.750 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2919,10 +2917,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30675 | loss 5.7269 | text 2.1563 audio 5.2956 | grad 3.860 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31675 | loss 5.7165 | text 2.1522 audio 5.2861 | grad 2.261 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2946,10 +2944,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30700 | loss 5.7398 | text 2.1505 audio 5.3097 | grad 2.235 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31700 | loss 5.6886 | text 2.1403 audio 5.2606 | grad 2.322 | 1.57s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2973,10 +2971,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30725 | loss 5.7489 | text 2.1535 audio 5.3182 | grad 2.031 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31725 | loss 5.7274 | text 2.1259 audio 5.3022 | grad 1.478 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3000,23 +2998,14 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30750 | loss 5.6953 | text 2.1131 audio 5.2727 | grad 3.574 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.17 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.55 | nonfinite=0 - running validation at step 30750... - val/composite=3.0163 (text=0.6371 audio=4.6024) val/loss=4.7299 cb0_acc=38.8% text_acc=90.9% - [val] per-codebook acc: cb0=38.8% cb1=19.9% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_v28gzt94/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b49a8-49c8b3a512d01dcf62ba2de6;c1d11542-5c8f-4019-8577-05bf7d3f521f) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31750 | loss 5.6918 | text 2.1276 audio 5.2663 | grad 2.754 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.24 | nonfinite=0 + running validation at step 31750... + val/composite=3.0149 (text=0.6321 audio=4.6034) val/loss=4.7299 cb0_acc=39.0% text_acc=90.9% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0129) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3040,10 +3029,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30775 | loss 5.7336 | text 2.1703 audio 5.2995 | grad 2.588 | 21.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31775 | loss 5.7190 | text 2.1568 audio 5.2876 | grad 1.540 | 2.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3067,10 +3056,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30800 | loss 5.7168 | text 2.1540 audio 5.2860 | grad 2.014 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31800 | loss 5.7086 | text 2.1397 audio 5.2807 | grad 1.333 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3094,10 +3083,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30825 | loss 5.7184 | text 2.1635 audio 5.2857 | grad 1.822 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31825 | loss 5.7059 | text 2.1289 audio 5.2801 | grad 1.661 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3121,10 +3110,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30850 | loss 5.7180 | text 2.1629 audio 5.2854 | grad 1.949 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31850 | loss 5.7276 | text 2.1494 audio 5.2978 | grad 1.604 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=+0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3148,10 +3137,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30875 | loss 5.7182 | text 2.1146 audio 5.2953 | grad 1.590 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + [ALERT] grad spike 3.4x (grad 6.92 vs EMA 2.02) step 31875 +step 31875 | loss 5.7635 | text 2.1532 audio 5.3329 | grad 6.916 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.10 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.08 | spike L=+0.01 G=3.42 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3175,10 +3165,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30900 | loss 5.6918 | text 2.1653 audio 5.2588 | grad 3.649 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.63 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31900 | loss 5.7007 | text 2.1354 audio 5.2736 | grad 1.531 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.10 text_embed=0.12 | spike L=-0.00 G=0.61 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3202,10 +3192,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30925 | loss 5.7255 | text 2.1531 audio 5.2949 | grad 1.466 | 1.43s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31925 | loss 5.7284 | text 2.1511 audio 5.2981 | grad 2.452 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3229,10 +3219,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30950 | loss 5.7009 | text 2.1495 audio 5.2710 | grad 2.898 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31950 | loss 5.7026 | text 2.1339 audio 5.2759 | grad 2.128 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3256,10 +3246,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30975 | loss 5.7335 | text 2.1454 audio 5.3044 | grad 2.994 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31975 | loss 5.7325 | text 2.1323 audio 5.3060 | grad 3.383 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.00 G=1.42 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3283,15 +3273,17 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31000 | loss 5.7318 | text 2.1491 audio 5.3019 | grad 1.703 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.71 | nonfinite=0 - running validation at step 31000... - val/composite=3.0191 (text=0.6409 audio=4.6045) val/loss=4.7327 cb0_acc=39.0% text_acc=90.8% - [val] per-codebook acc: cb0=39.0% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0163) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32000 | loss 5.7196 | text 2.1309 audio 5.2935 | grad 2.305 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.93 | nonfinite=0 + running validation at step 32000... + val/composite=3.0141 (text=0.6317 audio=4.6024) val/loss=4.7287 cb0_acc=38.9% text_acc=90.9% + [val] per-codebook acc: cb0=38.9% cb1=20.0% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0129) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_y6mkwqap/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 +[ckpt] uploading /tmp/ckpt_y3j7s14l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3303,11 +3295,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4d0b-4f6ce7c1015680b77c18485b;4de181e9-748f-43df-8a2c-956408fcae48) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5701-676c84df289d499c023dd0e8;593c1b3a-6ef8-4474-8a06-3d170a535244) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -3326,8 +3318,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31025 | loss 5.7013 | text 2.1689 audio 5.2675 | grad 2.218 | 20.63s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.95 | nonfinite=0 +step 32025 | loss 5.7313 | text 2.1524 audio 5.3008 | grad 1.656 | 20.67s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3353,8 +3345,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31050 | loss 5.6785 | text 2.1305 audio 5.2524 | grad 2.106 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 +step 32050 | loss 5.7048 | text 2.1554 audio 5.2737 | grad 2.361 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3380,9 +3372,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31075 | loss 5.7067 | text 2.1458 audio 5.2776 | grad 2.192 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.95 | nonfinite=0 -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32075 | loss 5.7039 | text 2.0979 audio 5.2843 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3407,9 +3398,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31100 | loss 5.7458 | text 2.1828 audio 5.3092 | grad 2.135 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.06 projection=0.06 text_embed=0.14 | spike L=+0.01 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3434,9 +3425,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31125 | loss 5.7439 | text 2.2002 audio 5.3039 | grad 2.645 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3461,9 +3452,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31150 | loss 5.7190 | text 2.1559 audio 5.2878 | grad 1.900 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3488,9 +3479,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31175 | loss 5.7093 | text 2.1537 audio 5.2786 | grad 2.073 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3515,9 +3506,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31200 | loss 5.7151 | text 2.1503 audio 5.2850 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3542,9 +3533,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31225 | loss 5.7159 | text 2.1445 audio 5.2870 | grad 1.819 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3569,13 +3560,22 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31250 | loss 5.7228 | text 2.1591 audio 5.2910 | grad 2.032 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 - running validation at step 31250... - val/composite=3.0194 (text=0.6437 audio=4.6031) val/loss=4.7319 cb0_acc=38.7% text_acc=90.7% - [val] per-codebook acc: cb0=38.7% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0163) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 + running validation at step 32250... + val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3600,9 +3600,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31275 | loss 5.7330 | text 2.1806 audio 5.2969 | grad 1.797 | 2.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3627,9 +3627,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31300 | loss 5.7208 | text 2.1705 audio 5.2867 | grad 2.195 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3654,9 +3654,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31325 | loss 5.7276 | text 2.1491 audio 5.2978 | grad 2.566 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.18 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3681,9 +3681,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31350 | loss 5.7069 | text 2.1178 audio 5.2834 | grad 1.952 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3708,9 +3708,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31375 | loss 5.7308 | text 2.1203 audio 5.3067 | grad 1.813 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3735,9 +3735,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31400 | loss 5.7433 | text 2.1584 audio 5.3116 | grad 2.645 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3762,9 +3762,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31425 | loss 5.7413 | text 2.1406 audio 5.3132 | grad 1.860 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3789,9 +3789,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31450 | loss 5.7094 | text 2.1542 audio 5.2786 | grad 1.536 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3816,9 +3816,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31475 | loss 5.7086 | text 2.1526 audio 5.2781 | grad 3.518 | 1.59s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3843,16 +3843,17 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31500 | loss 5.7234 | text 2.1505 audio 5.2933 | grad 2.082 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 31500... - val/composite=3.0129 (text=0.6302 audio=4.6013) val/loss=4.7273 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 + running validation at step 32500... + val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_97c5nxlf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b51f5-564c334001068ae16e82ce0f;aafdec3d-7141-4994-a73d-192be2632797) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -3883,8 +3884,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31525 | loss 5.7418 | text 2.1563 audio 5.3106 | grad 2.122 | 21.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 +step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3910,8 +3911,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31550 | loss 5.7494 | text 2.1284 audio 5.3237 | grad 2.364 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.06 | nonfinite=0 +step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3937,8 +3938,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31575 | loss 5.7309 | text 2.1538 audio 5.3001 | grad 1.760 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 +step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3964,8 +3965,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31600 | loss 5.7379 | text 2.1287 audio 5.3122 | grad 3.343 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.53 | nonfinite=0 +step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3991,8 +3992,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31625 | loss 5.6686 | text 2.1377 audio 5.2410 | grad 2.997 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.30 | nonfinite=0 +step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4018,8 +4019,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31650 | loss 5.6969 | text 2.1637 audio 5.2641 | grad 1.750 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 +step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4045,8 +4046,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31675 | loss 5.7165 | text 2.1522 audio 5.2861 | grad 2.261 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 + [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 +step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4072,8 +4074,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31700 | loss 5.6886 | text 2.1403 audio 5.2606 | grad 2.322 | 1.57s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.01 | nonfinite=0 +step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4099,8 +4101,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31725 | loss 5.7274 | text 2.1259 audio 5.3022 | grad 1.478 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.64 | nonfinite=0 +step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4126,12 +4128,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31750 | loss 5.6918 | text 2.1276 audio 5.2663 | grad 2.754 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.24 | nonfinite=0 - running validation at step 31750... - val/composite=3.0149 (text=0.6321 audio=4.6034) val/loss=4.7299 cb0_acc=39.0% text_acc=90.9% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0129) +step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 + running validation at step 32750... + val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4157,8 +4168,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31775 | loss 5.7190 | text 2.1568 audio 5.2876 | grad 1.540 | 2.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 +step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4184,8 +4195,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31800 | loss 5.7086 | text 2.1397 audio 5.2807 | grad 1.333 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.60 | nonfinite=0 +step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4211,8 +4222,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31825 | loss 5.7059 | text 2.1289 audio 5.2801 | grad 1.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4238,8 +4249,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31850 | loss 5.7276 | text 2.1494 audio 5.2978 | grad 1.604 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=+0.00 G=0.77 | nonfinite=0 +step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4265,9 +4276,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 3.4x (grad 6.92 vs EMA 2.02) step 31875 -step 31875 | loss 5.7635 | text 2.1532 audio 5.3329 | grad 6.916 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.10 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.08 | spike L=+0.01 G=3.42 | nonfinite=0 +step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4293,8 +4303,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31900 | loss 5.7007 | text 2.1354 audio 5.2736 | grad 1.531 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.10 text_embed=0.12 | spike L=-0.00 G=0.61 | nonfinite=0 +step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4320,8 +4330,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31925 | loss 5.7284 | text 2.1511 audio 5.2981 | grad 2.452 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.02 | nonfinite=0 +step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4347,8 +4357,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31950 | loss 5.7026 | text 2.1339 audio 5.2759 | grad 2.128 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=0.88 | nonfinite=0 +step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4374,8 +4384,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31975 | loss 5.7325 | text 2.1323 audio 5.3060 | grad 3.383 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.00 G=1.42 | nonfinite=0 +step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4401,15 +4411,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32000 | loss 5.7196 | text 2.1309 audio 5.2935 | grad 2.305 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 32000... - val/composite=3.0141 (text=0.6317 audio=4.6024) val/loss=4.7287 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.0% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0129) +step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 + running validation at step 33000... + val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% + [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_y3j7s14l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 +[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4420,17 +4430,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5701-676c84df289d499c023dd0e8;593c1b3a-6ef8-4474-8a06-3d170a535244) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4444,9 +4453,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32025 | loss 5.7313 | text 2.1524 audio 5.3008 | grad 1.656 | 20.67s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4471,9 +4480,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32050 | loss 5.7048 | text 2.1554 audio 5.2737 | grad 2.361 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4498,9 +4507,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32075 | loss 5.7039 | text 2.0979 audio 5.2843 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4525,9 +4534,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4552,9 +4561,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4579,9 +4588,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4606,9 +4615,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4633,9 +4642,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4660,9 +4669,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4687,16 +4696,17 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 - running validation at step 32250... - val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 + running validation at step 33250... + val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -4727,8 +4737,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4754,8 +4764,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 +step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4781,8 +4791,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 +step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4808,8 +4818,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 +step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4835,8 +4845,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4862,8 +4872,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 +step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4889,8 +4899,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 +step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4916,8 +4926,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 +step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4943,8 +4953,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 +step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4970,16 +4980,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 - running validation at step 32500... - val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 + running validation at step 33500... + val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -5010,8 +5020,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5037,8 +5047,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 +step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5064,8 +5074,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 +step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5091,8 +5101,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 +step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5118,8 +5128,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 +step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5145,8 +5155,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 +step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5172,9 +5182,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 -step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 +step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5200,8 +5209,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 +step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5227,8 +5236,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 +step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5254,21 +5263,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 - running validation at step 32750... - val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 + running validation at step 33750... + val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5294,8 +5294,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 +step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5321,8 +5321,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 +step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5348,8 +5348,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 +step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5375,8 +5375,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 +step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5402,8 +5402,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 +step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5429,8 +5429,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 +step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5456,8 +5456,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 +step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5483,8 +5483,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 +step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5510,8 +5510,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5537,15 +5537,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 - running validation at step 33000... - val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) +step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 + running validation at step 34000... + val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% + [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 +[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5556,22 +5565,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5580,9 +5588,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5607,9 +5615,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5634,9 +5642,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5661,9 +5669,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5688,9 +5696,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5715,9 +5723,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5742,9 +5750,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5769,9 +5777,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5796,9 +5804,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5823,22 +5831,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 - running validation at step 33250... - val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 + running validation at step 34250... + val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5863,9 +5862,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5890,9 +5889,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5917,9 +5916,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5944,9 +5943,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5971,9 +5970,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5998,9 +5997,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6025,9 +6024,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6052,9 +6051,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6079,9 +6078,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6106,22 +6105,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 - running validation at step 33500... - val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 + running validation at step 34500... + val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6146,9 +6136,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6173,9 +6163,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6200,9 +6190,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6227,9 +6217,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6254,9 +6244,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6281,9 +6271,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6308,9 +6298,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6335,9 +6325,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6362,9 +6352,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6389,13 +6379,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 - running validation at step 33750... - val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 + running validation at step 34750... + val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% + [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6420,9 +6410,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6447,9 +6437,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6474,9 +6464,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6501,9 +6491,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6528,9 +6518,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6555,9 +6545,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6582,9 +6572,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6609,9 +6599,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6636,9 +6626,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6663,16 +6653,44 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 - running validation at step 34000... - val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% - [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 + [audio-demo] ar_cb0_acc=12.0% (35s) + running validation at step 35000... + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB + pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 + val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -6680,7 +6698,7 @@ Make sure your token has the correct permissions. * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6691,19 +6709,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7c65-1c38aecc5ca4acdb5008e0b4;48e1227e-d30a-4783-b035-8b02985e5e1d) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6715,11 +6730,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35025 | loss 5.7000 | text 2.1331 audio 5.2734 | grad 2.588 | 41.18s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6742,11 +6757,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35050 | loss 5.7063 | text 2.1069 audio 5.2850 | grad 2.479 | 1.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6769,11 +6784,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + [ALERT] grad spike 3.0x (grad 6.76 vs EMA 2.25) step 35075 +step 35075 | loss 5.7092 | text 2.1201 audio 5.2851 | grad 6.761 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.11 lora=0.98 model_audio_embed=0.06 projection=0.02 text_embed=0.12 | spike L=+0.00 G=3.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6796,11 +6812,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35100 | loss 5.7286 | text 2.1354 audio 5.3015 | grad 2.020 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6823,11 +6839,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35125 | loss 5.6800 | text 2.1305 audio 5.2539 | grad 1.847 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6850,11 +6866,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35150 | loss 5.7182 | text 2.1559 audio 5.2870 | grad 1.641 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6877,11 +6893,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35175 | loss 5.6631 | text 2.1155 audio 5.2400 | grad 2.109 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6904,11 +6920,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35200 | loss 5.7188 | text 2.1798 audio 5.2828 | grad 2.145 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6931,11 +6947,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35225 | loss 5.7170 | text 2.1077 audio 5.2955 | grad 1.866 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6958,15 +6974,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 - running validation at step 34250... - val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35250 | loss 5.6647 | text 2.0979 audio 5.2451 | grad 1.703 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.73 | nonfinite=0 + running validation at step 35250... + val/composite=2.9985 (text=0.6053 audio=4.5939) val/loss=4.7150 cb0_acc=39.1% text_acc=91.6% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9973) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6989,11 +7005,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35275 | loss 5.7344 | text 2.1597 audio 5.3025 | grad 4.716 | 2.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=+0.01 G=2.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7016,11 +7032,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35300 | loss 5.6869 | text 2.1172 audio 5.2634 | grad 1.765 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7043,11 +7059,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35325 | loss 5.7376 | text 2.1312 audio 5.3113 | grad 1.650 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7070,11 +7086,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35350 | loss 5.6878 | text 2.1130 audio 5.2652 | grad 1.950 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7097,11 +7113,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35375 | loss 5.6788 | text 2.1222 audio 5.2544 | grad 1.724 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7124,11 +7140,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35400 | loss 5.6788 | text 2.1020 audio 5.2584 | grad 1.961 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7151,11 +7167,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35425 | loss 5.6969 | text 2.0995 audio 5.2770 | grad 1.766 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7178,11 +7194,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35450 | loss 5.6848 | text 2.0913 audio 5.2665 | grad 1.521 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7205,11 +7221,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35475 | loss 5.7112 | text 2.1194 audio 5.2873 | grad 2.124 | 1.60s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7232,12 +7248,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 - running validation at step 34500... - val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35500 | loss 5.7002 | text 2.0979 audio 5.2806 | grad 2.371 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 + running validation at step 35500... + val/composite=2.9923 (text=0.5962 audio=4.5896) val/loss=4.7089 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_3crsay5y/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b816c-2712d01710663cc45f59dc84;4de1714a-3953-4c64-bba2-5dfc7dfb76bb) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7263,8 +7291,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 +step 35525 | loss 5.7157 | text 2.1252 audio 5.2907 | grad 1.712 | 21.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7290,8 +7318,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 +step 35550 | loss 5.7160 | text 2.1306 audio 5.2898 | grad 2.455 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7317,8 +7345,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 +step 35575 | loss 5.7021 | text 2.1011 audio 5.2819 | grad 2.139 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7344,8 +7372,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 +step 35600 | loss 5.6976 | text 2.1119 audio 5.2752 | grad 1.692 | 1.59s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7371,8 +7399,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 +step 35625 | loss 5.7269 | text 2.1576 audio 5.2954 | grad 1.948 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7398,8 +7426,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 +step 35650 | loss 5.7038 | text 2.0968 audio 5.2844 | grad 2.009 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7425,8 +7453,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 +step 35675 | loss 5.6964 | text 2.1176 audio 5.2729 | grad 1.781 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7452,8 +7480,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 +step 35700 | loss 5.7023 | text 2.1459 audio 5.2731 | grad 1.470 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7479,8 +7507,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 +step 35725 | loss 5.6936 | text 2.1079 audio 5.2720 | grad 1.577 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7506,12 +7534,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 - running validation at step 34750... - val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% - [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) +step 35750 | loss 5.7018 | text 2.1473 audio 5.2724 | grad 3.195 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.64 | nonfinite=0 + running validation at step 35750... + val/composite=2.9894 (text=0.5988 audio=4.5832) val/loss=4.7030 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_q4bfdk_b/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b84eb-23e8713c6407ba4a163897ad;da59c5d6-9f19-4363-b138-8063d8495773) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7537,8 +7574,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 +step 35775 | loss 5.6928 | text 2.1394 audio 5.2649 | grad 2.077 | 21.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7564,8 +7601,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 +step 35800 | loss 5.7009 | text 2.1499 audio 5.2709 | grad 3.615 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7591,8 +7628,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 +step 35825 | loss 5.7072 | text 2.1152 audio 5.2842 | grad 1.769 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7618,8 +7655,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 +step 35850 | loss 5.7080 | text 2.1325 audio 5.2815 | grad 1.590 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7645,8 +7682,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 +step 35875 | loss 5.7191 | text 2.1272 audio 5.2937 | grad 2.377 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7672,8 +7709,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 +step 35900 | loss 5.6961 | text 2.0964 audio 5.2768 | grad 1.901 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7699,8 +7736,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 +step 35925 | loss 5.6746 | text 2.1152 audio 5.2516 | grad 1.636 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7726,8 +7763,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 +step 35950 | loss 5.6555 | text 2.1133 audio 5.2328 | grad 1.572 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7753,8 +7790,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 +step 35975 | loss 5.6890 | text 2.1208 audio 5.2648 | grad 2.115 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7780,49 +7817,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 - [audio-demo] ar_cb0_acc=12.0% (35s) - running validation at step 35000... - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB - pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 - val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% - [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 36000 | loss 5.7153 | text 2.1097 audio 5.2934 | grad 2.235 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.10 | nonfinite=0 + running validation at step 36000... + val/composite=2.9897 (text=0.5974 audio=4.5846) val/loss=4.7041 cb0_acc=39.2% text_acc=91.8% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9894) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 +[ckpt] uploading /tmp/ckpt_skck_i9h/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3