diff --git "a/logs/train_host0_latest.log" "b/logs/train_host0_latest.log" --- "a/logs/train_host0_latest.log" +++ "b/logs/train_host0_latest.log" @@ -1,4 +1,4 @@ -ch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +puts since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -12,8 +12,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31100 | loss 5.7458 | text 2.1828 audio 5.3092 | grad 2.135 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.06 projection=0.06 text_embed=0.14 | spike L=+0.01 G=0.93 | nonfinite=0 +step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -39,8 +39,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31125 | loss 5.7439 | text 2.2002 audio 5.3039 | grad 2.645 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.16 | nonfinite=0 +step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -66,8 +66,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31150 | loss 5.7190 | text 2.1559 audio 5.2878 | grad 1.900 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.82 | nonfinite=0 +step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -93,8 +93,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31175 | loss 5.7093 | text 2.1537 audio 5.2786 | grad 2.073 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.91 | nonfinite=0 +step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -120,8 +120,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31200 | loss 5.7151 | text 2.1503 audio 5.2850 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.18 | nonfinite=0 +step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -147,8 +147,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31225 | loss 5.7159 | text 2.1445 audio 5.2870 | grad 1.819 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.79 | nonfinite=0 +step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -174,12 +174,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31250 | loss 5.7228 | text 2.1591 audio 5.2910 | grad 2.032 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 - running validation at step 31250... - val/composite=3.0194 (text=0.6437 audio=4.6031) val/loss=4.7319 cb0_acc=38.7% text_acc=90.7% - [val] per-codebook acc: cb0=38.7% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0163) +step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 + running validation at step 32250... + val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -205,8 +214,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31275 | loss 5.7330 | text 2.1806 audio 5.2969 | grad 1.797 | 2.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.81 | nonfinite=0 +step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -232,8 +241,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31300 | loss 5.7208 | text 2.1705 audio 5.2867 | grad 2.195 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 +step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -259,8 +268,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31325 | loss 5.7276 | text 2.1491 audio 5.2978 | grad 2.566 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.18 | spike L=+0.00 G=1.17 | nonfinite=0 +step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -286,8 +295,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31350 | loss 5.7069 | text 2.1178 audio 5.2834 | grad 1.952 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=-0.00 G=0.88 | nonfinite=0 +step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -313,8 +322,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31375 | loss 5.7308 | text 2.1203 audio 5.3067 | grad 1.813 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 +step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -340,8 +349,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31400 | loss 5.7433 | text 2.1584 audio 5.3116 | grad 2.645 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.23 | nonfinite=0 +step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -367,8 +376,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31425 | loss 5.7413 | text 2.1406 audio 5.3132 | grad 1.860 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.84 | nonfinite=0 +step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -394,8 +403,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31450 | loss 5.7094 | text 2.1542 audio 5.2786 | grad 1.536 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.71 | nonfinite=0 +step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -421,8 +430,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31475 | loss 5.7086 | text 2.1526 audio 5.2781 | grad 3.518 | 1.59s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.67 | nonfinite=0 +step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -448,16 +457,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31500 | loss 5.7234 | text 2.1505 audio 5.2933 | grad 2.082 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 31500... - val/composite=3.0129 (text=0.6302 audio=4.6013) val/loss=4.7273 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 + running validation at step 32500... + val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_97c5nxlf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b51f5-564c334001068ae16e82ce0f;aafdec3d-7141-4994-a73d-192be2632797) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -488,8 +497,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31525 | loss 5.7418 | text 2.1563 audio 5.3106 | grad 2.122 | 21.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 +step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -515,8 +524,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31550 | loss 5.7494 | text 2.1284 audio 5.3237 | grad 2.364 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.06 | nonfinite=0 +step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -542,8 +551,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31575 | loss 5.7309 | text 2.1538 audio 5.3001 | grad 1.760 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 +step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -569,8 +578,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31600 | loss 5.7379 | text 2.1287 audio 5.3122 | grad 3.343 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.53 | nonfinite=0 +step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -596,8 +605,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31625 | loss 5.6686 | text 2.1377 audio 5.2410 | grad 2.997 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.30 | nonfinite=0 +step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -623,8 +632,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31650 | loss 5.6969 | text 2.1637 audio 5.2641 | grad 1.750 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 +step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -650,8 +659,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31675 | loss 5.7165 | text 2.1522 audio 5.2861 | grad 2.261 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 + [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 +step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -677,8 +687,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31700 | loss 5.6886 | text 2.1403 audio 5.2606 | grad 2.322 | 1.57s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.01 | nonfinite=0 +step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -704,8 +714,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31725 | loss 5.7274 | text 2.1259 audio 5.3022 | grad 1.478 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.64 | nonfinite=0 +step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -731,12 +741,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31750 | loss 5.6918 | text 2.1276 audio 5.2663 | grad 2.754 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.24 | nonfinite=0 - running validation at step 31750... - val/composite=3.0149 (text=0.6321 audio=4.6034) val/loss=4.7299 cb0_acc=39.0% text_acc=90.9% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0129) +step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 + running validation at step 32750... + val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -762,8 +781,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31775 | loss 5.7190 | text 2.1568 audio 5.2876 | grad 1.540 | 2.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 +step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -789,8 +808,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31800 | loss 5.7086 | text 2.1397 audio 5.2807 | grad 1.333 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.60 | nonfinite=0 +step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -816,8 +835,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31825 | loss 5.7059 | text 2.1289 audio 5.2801 | grad 1.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -843,8 +862,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31850 | loss 5.7276 | text 2.1494 audio 5.2978 | grad 1.604 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=+0.00 G=0.77 | nonfinite=0 +step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -870,9 +889,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 3.4x (grad 6.92 vs EMA 2.02) step 31875 -step 31875 | loss 5.7635 | text 2.1532 audio 5.3329 | grad 6.916 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.10 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.08 | spike L=+0.01 G=3.42 | nonfinite=0 +step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -898,8 +916,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31900 | loss 5.7007 | text 2.1354 audio 5.2736 | grad 1.531 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.10 text_embed=0.12 | spike L=-0.00 G=0.61 | nonfinite=0 +step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -925,8 +943,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31925 | loss 5.7284 | text 2.1511 audio 5.2981 | grad 2.452 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.02 | nonfinite=0 +step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -952,8 +970,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31950 | loss 5.7026 | text 2.1339 audio 5.2759 | grad 2.128 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=0.88 | nonfinite=0 +step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -979,8 +997,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31975 | loss 5.7325 | text 2.1323 audio 5.3060 | grad 3.383 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.00 G=1.42 | nonfinite=0 +step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1006,15 +1024,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32000 | loss 5.7196 | text 2.1309 audio 5.2935 | grad 2.305 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 32000... - val/composite=3.0141 (text=0.6317 audio=4.6024) val/loss=4.7287 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.0% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0129) +step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 + running validation at step 33000... + val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% + [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_y3j7s14l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 +[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1025,17 +1043,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5701-676c84df289d499c023dd0e8;593c1b3a-6ef8-4474-8a06-3d170a535244) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1049,8 +1066,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32025 | loss 5.7313 | text 2.1524 audio 5.3008 | grad 1.656 | 20.67s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.67 | nonfinite=0 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1076,8 +1094,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32050 | loss 5.7048 | text 2.1554 audio 5.2737 | grad 2.361 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 +step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1103,8 +1121,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32075 | loss 5.7039 | text 2.0979 audio 5.2843 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.77 | nonfinite=0 +step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1130,8 +1148,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 +step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1157,8 +1175,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 +step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1184,8 +1202,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 +step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1211,8 +1229,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 +step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1238,8 +1256,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 +step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1265,8 +1283,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 +step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1292,16 +1310,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 - running validation at step 32250... - val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 + running validation at step 33250... + val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -1332,8 +1350,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1359,8 +1377,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 +step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1386,8 +1404,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 +step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1413,8 +1431,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 +step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1440,8 +1458,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1467,8 +1485,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 +step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1494,8 +1512,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 +step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1521,8 +1539,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 +step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1548,8 +1566,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 +step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1575,16 +1593,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 - running validation at step 32500... - val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 + running validation at step 33500... + val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -1615,8 +1633,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1642,8 +1660,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 +step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1669,8 +1687,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 +step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1696,8 +1714,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 +step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1723,8 +1741,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 +step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1750,8 +1768,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 +step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1777,9 +1795,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 -step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 +step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1805,8 +1822,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 +step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1832,8 +1849,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 +step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1859,21 +1876,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 - running validation at step 32750... - val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 + running validation at step 33750... + val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1899,8 +1907,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 +step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1926,8 +1934,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 +step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1953,8 +1961,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 +step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1980,8 +1988,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 +step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2007,8 +2015,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 +step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2034,8 +2042,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 +step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2061,8 +2069,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 +step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2088,8 +2096,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 +step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2115,8 +2123,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2142,15 +2150,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 - running validation at step 33000... - val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) +step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 + running validation at step 34000... + val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% + [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 +[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2161,16 +2178,19 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2182,11 +2202,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2209,11 +2229,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2236,11 +2256,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2263,11 +2283,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2290,11 +2310,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2317,11 +2337,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2344,11 +2364,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2371,11 +2391,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2398,11 +2418,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2425,24 +2445,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 + running validation at step 34250... + val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 - running validation at step 33250... - val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2465,11 +2476,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2492,11 +2503,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2519,11 +2530,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2546,11 +2557,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2573,11 +2584,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2600,11 +2611,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2627,11 +2638,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2654,11 +2665,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2681,11 +2692,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2708,24 +2719,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 + running validation at step 34500... + val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 - running validation at step 33500... - val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2748,11 +2750,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2775,11 +2777,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2802,11 +2804,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2829,11 +2831,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2856,11 +2858,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2883,11 +2885,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2910,11 +2912,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2937,11 +2939,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2964,11 +2966,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2991,15 +2993,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 + running validation at step 34750... + val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% + [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 - running validation at step 33750... - val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3022,11 +3024,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3049,11 +3051,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3076,11 +3078,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3103,11 +3105,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3130,11 +3132,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 +step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3160,8 +3159,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 +step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3187,8 +3186,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 +step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3214,8 +3213,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 +step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3241,8 +3240,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 +step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3268,16 +3267,43 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 - running validation at step 34000... - val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% - [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 + [audio-demo] ar_cb0_acc=12.0% (35s) + running validation at step 35000... + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB + pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 + val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -3285,7 +3311,7 @@ Make sure your token has the correct permissions. * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3296,19 +3322,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7c65-1c38aecc5ca4acdb5008e0b4;48e1227e-d30a-4783-b035-8b02985e5e1d) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3320,11 +3343,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35025 | loss 5.7000 | text 2.1331 audio 5.2734 | grad 2.588 | 41.18s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3347,11 +3370,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35050 | loss 5.7063 | text 2.1069 audio 5.2850 | grad 2.479 | 1.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3374,11 +3397,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + [ALERT] grad spike 3.0x (grad 6.76 vs EMA 2.25) step 35075 +step 35075 | loss 5.7092 | text 2.1201 audio 5.2851 | grad 6.761 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.11 lora=0.98 model_audio_embed=0.06 projection=0.02 text_embed=0.12 | spike L=+0.00 G=3.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3401,11 +3425,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35100 | loss 5.7286 | text 2.1354 audio 5.3015 | grad 2.020 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3428,11 +3452,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35125 | loss 5.6800 | text 2.1305 audio 5.2539 | grad 1.847 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3455,11 +3479,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35150 | loss 5.7182 | text 2.1559 audio 5.2870 | grad 1.641 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3482,11 +3506,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35175 | loss 5.6631 | text 2.1155 audio 5.2400 | grad 2.109 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3509,11 +3533,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35200 | loss 5.7188 | text 2.1798 audio 5.2828 | grad 2.145 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3536,11 +3560,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35225 | loss 5.7170 | text 2.1077 audio 5.2955 | grad 1.866 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3563,15 +3587,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 - running validation at step 34250... - val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35250 | loss 5.6647 | text 2.0979 audio 5.2451 | grad 1.703 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.73 | nonfinite=0 + running validation at step 35250... + val/composite=2.9985 (text=0.6053 audio=4.5939) val/loss=4.7150 cb0_acc=39.1% text_acc=91.6% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9973) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3594,11 +3618,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35275 | loss 5.7344 | text 2.1597 audio 5.3025 | grad 4.716 | 2.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=+0.01 G=2.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3621,11 +3645,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35300 | loss 5.6869 | text 2.1172 audio 5.2634 | grad 1.765 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3648,11 +3672,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35325 | loss 5.7376 | text 2.1312 audio 5.3113 | grad 1.650 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3675,11 +3699,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35350 | loss 5.6878 | text 2.1130 audio 5.2652 | grad 1.950 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3702,11 +3726,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35375 | loss 5.6788 | text 2.1222 audio 5.2544 | grad 1.724 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3729,11 +3753,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35400 | loss 5.6788 | text 2.1020 audio 5.2584 | grad 1.961 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3756,11 +3780,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35425 | loss 5.6969 | text 2.0995 audio 5.2770 | grad 1.766 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3783,11 +3807,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35450 | loss 5.6848 | text 2.0913 audio 5.2665 | grad 1.521 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3810,11 +3834,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35475 | loss 5.7112 | text 2.1194 audio 5.2873 | grad 2.124 | 1.60s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3837,15 +3861,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 - running validation at step 34500... - val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35500 | loss 5.7002 | text 2.0979 audio 5.2806 | grad 2.371 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 + running validation at step 35500... + val/composite=2.9923 (text=0.5962 audio=4.5896) val/loss=4.7089 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_3crsay5y/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b816c-2712d01710663cc45f59dc84;4de1714a-3953-4c64-bba2-5dfc7dfb76bb) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3868,11 +3901,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35525 | loss 5.7157 | text 2.1252 audio 5.2907 | grad 1.712 | 21.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3895,11 +3928,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35550 | loss 5.7160 | text 2.1306 audio 5.2898 | grad 2.455 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3922,11 +3955,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35575 | loss 5.7021 | text 2.1011 audio 5.2819 | grad 2.139 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3949,11 +3982,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35600 | loss 5.6976 | text 2.1119 audio 5.2752 | grad 1.692 | 1.59s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3976,11 +4009,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35625 | loss 5.7269 | text 2.1576 audio 5.2954 | grad 1.948 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4003,11 +4036,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35650 | loss 5.7038 | text 2.0968 audio 5.2844 | grad 2.009 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4030,11 +4063,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35675 | loss 5.6964 | text 2.1176 audio 5.2729 | grad 1.781 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4057,11 +4090,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35700 | loss 5.7023 | text 2.1459 audio 5.2731 | grad 1.470 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4084,11 +4117,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35725 | loss 5.6936 | text 2.1079 audio 5.2720 | grad 1.577 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4111,15 +4144,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 - running validation at step 34750... - val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% - [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35750 | loss 5.7018 | text 2.1473 audio 5.2724 | grad 3.195 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.64 | nonfinite=0 + running validation at step 35750... + val/composite=2.9894 (text=0.5988 audio=4.5832) val/loss=4.7030 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_q4bfdk_b/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b84eb-23e8713c6407ba4a163897ad;da59c5d6-9f19-4363-b138-8063d8495773) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4142,11 +4184,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35775 | loss 5.6928 | text 2.1394 audio 5.2649 | grad 2.077 | 21.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4169,11 +4211,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35800 | loss 5.7009 | text 2.1499 audio 5.2709 | grad 3.615 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4196,11 +4238,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35825 | loss 5.7072 | text 2.1152 audio 5.2842 | grad 1.769 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4223,11 +4265,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35850 | loss 5.7080 | text 2.1325 audio 5.2815 | grad 1.590 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4250,11 +4292,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35875 | loss 5.7191 | text 2.1272 audio 5.2937 | grad 2.377 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4277,11 +4319,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35900 | loss 5.6961 | text 2.0964 audio 5.2768 | grad 1.901 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4304,11 +4346,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35925 | loss 5.6746 | text 2.1152 audio 5.2516 | grad 1.636 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4331,11 +4373,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35950 | loss 5.6555 | text 2.1133 audio 5.2328 | grad 1.572 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4358,11 +4400,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35975 | loss 5.6890 | text 2.1208 audio 5.2648 | grad 2.115 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4385,51 +4427,18 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 - [audio-demo] ar_cb0_acc=12.0% (35s) - running validation at step 35000... - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB - pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 - val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% - [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36000 | loss 5.7153 | text 2.1097 audio 5.2934 | grad 2.235 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.10 | nonfinite=0 + running validation at step 36000... + val/composite=2.9897 (text=0.5974 audio=4.5846) val/loss=4.7041 cb0_acc=39.2% text_acc=91.8% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9894) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 +[ckpt] uploading /tmp/ckpt_skck_i9h/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4440,36 +4449,32 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7c65-1c38aecc5ca4acdb5008e0b4;48e1227e-d30a-4783-b035-8b02985e5e1d) +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8865-1ec56561651c1a56354c727e;d2e4e91e-8960-46ef-b922-28020e2e7473) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35025 | loss 5.7000 | text 2.1331 audio 5.2734 | grad 2.588 | 41.18s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36025 | loss 5.6831 | text 2.1034 audio 5.2624 | grad 1.684 | 20.86s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.06 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4491,12 +4496,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35050 | loss 5.7063 | text 2.1069 audio 5.2850 | grad 2.479 | 1.48s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36050 | loss 5.7093 | text 2.0998 audio 5.2893 | grad 1.919 | 1.59s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4518,13 +4523,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 3.0x (grad 6.76 vs EMA 2.25) step 35075 -step 35075 | loss 5.7092 | text 2.1201 audio 5.2851 | grad 6.761 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.11 lora=0.98 model_audio_embed=0.06 projection=0.02 text_embed=0.12 | spike L=+0.00 G=3.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36075 | loss 5.6986 | text 2.1149 audio 5.2756 | grad 2.041 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4546,12 +4550,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35100 | loss 5.7286 | text 2.1354 audio 5.3015 | grad 2.020 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36100 | loss 5.7108 | text 2.1152 audio 5.2878 | grad 1.744 | 1.50s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4573,12 +4577,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35125 | loss 5.6800 | text 2.1305 audio 5.2539 | grad 1.847 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36125 | loss 5.6697 | text 2.1321 audio 5.2433 | grad 2.215 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.01 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4600,12 +4604,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35150 | loss 5.7182 | text 2.1559 audio 5.2870 | grad 1.641 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36150 | loss 5.6836 | text 2.1104 audio 5.2616 | grad 1.771 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4627,12 +4631,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35175 | loss 5.6631 | text 2.1155 audio 5.2400 | grad 2.109 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36175 | loss 5.7177 | text 2.1067 audio 5.2964 | grad 1.322 | 1.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4654,12 +4658,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35200 | loss 5.7188 | text 2.1798 audio 5.2828 | grad 2.145 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36200 | loss 5.6609 | text 2.1179 audio 5.2373 | grad 2.159 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.01 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4681,12 +4685,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35225 | loss 5.7170 | text 2.1077 audio 5.2955 | grad 1.866 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36225 | loss 5.7207 | text 2.1407 audio 5.2925 | grad 2.000 | 1.50s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4708,16 +4712,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35250 | loss 5.6647 | text 2.0979 audio 5.2451 | grad 1.703 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.73 | nonfinite=0 - running validation at step 35250... - val/composite=2.9985 (text=0.6053 audio=4.5939) val/loss=4.7150 cb0_acc=39.1% text_acc=91.6% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9973) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36250 | loss 5.6867 | text 2.0964 audio 5.2674 | grad 2.338 | 1.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.20 | nonfinite=0 + running validation at step 36250... + val/composite=2.9903 (text=0.5982 audio=4.5851) val/loss=4.7047 cb0_acc=39.2% text_acc=91.8% + [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9894) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4739,12 +4743,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35275 | loss 5.7344 | text 2.1597 audio 5.3025 | grad 4.716 | 2.47s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=+0.01 G=2.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36275 | loss 5.7501 | text 2.1488 audio 5.3204 | grad 1.751 | 2.30s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4766,12 +4770,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35300 | loss 5.6869 | text 2.1172 audio 5.2634 | grad 1.765 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36300 | loss 5.6584 | text 2.0785 audio 5.2427 | grad 1.793 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4793,12 +4797,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35325 | loss 5.7376 | text 2.1312 audio 5.3113 | grad 1.650 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36325 | loss 5.7092 | text 2.1386 audio 5.2815 | grad 1.547 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4820,12 +4824,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35350 | loss 5.6878 | text 2.1130 audio 5.2652 | grad 1.950 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36350 | loss 5.7118 | text 2.1361 audio 5.2846 | grad 2.371 | 1.60s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.24 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4847,12 +4851,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35375 | loss 5.6788 | text 2.1222 audio 5.2544 | grad 1.724 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36375 | loss 5.7245 | text 2.1202 audio 5.3005 | grad 1.915 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4874,12 +4878,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35400 | loss 5.6788 | text 2.1020 audio 5.2584 | grad 1.961 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36400 | loss 5.7019 | text 2.1270 audio 5.2765 | grad 2.564 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.15 | spike L=+0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4901,12 +4905,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35425 | loss 5.6969 | text 2.0995 audio 5.2770 | grad 1.766 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36425 | loss 5.7086 | text 2.1082 audio 5.2870 | grad 1.945 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4928,12 +4932,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35450 | loss 5.6848 | text 2.0913 audio 5.2665 | grad 1.521 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36450 | loss 5.6757 | text 2.1256 audio 5.2505 | grad 1.849 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4955,12 +4959,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35475 | loss 5.7112 | text 2.1194 audio 5.2873 | grad 2.124 | 1.60s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36475 | loss 5.6761 | text 2.1255 audio 5.2510 | grad 1.958 | 1.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4982,16 +4986,20 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35500 | loss 5.7002 | text 2.0979 audio 5.2806 | grad 2.371 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 - running validation at step 35500... - val/composite=2.9923 (text=0.5962 audio=4.5896) val/loss=4.7089 cb0_acc=39.1% text_acc=91.8% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36500 | loss 5.6908 | text 2.0905 audio 5.2727 | grad 1.988 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=1.00 | nonfinite=0 + running validation at step 36500... + val/composite=2.9857 (text=0.5922 audio=4.5813) val/loss=4.6998 cb0_acc=39.4% text_acc=91.8% + [val] per-codebook acc: cb0=39.4% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_3crsay5y/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_pnu0pz72/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b816c-2712d01710663cc45f59dc84;4de1714a-3953-4c64-bba2-5dfc7dfb76bb) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8d52-3c934e2072b2c29c7b9ed4a4;c8972a1a-b255-4d8a-a62a-f52dd68711e3) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -5022,8 +5030,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35525 | loss 5.7157 | text 2.1252 audio 5.2907 | grad 1.712 | 21.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.80 | nonfinite=0 +step 36525 | loss 5.6839 | text 2.1440 audio 5.2551 | grad 2.686 | 21.43s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5049,8 +5057,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35550 | loss 5.7160 | text 2.1306 audio 5.2898 | grad 2.455 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 +step 36550 | loss 5.6939 | text 2.1047 audio 5.2730 | grad 2.494 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.21 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5076,8 +5084,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35575 | loss 5.7021 | text 2.1011 audio 5.2819 | grad 2.139 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 +step 36575 | loss 5.6677 | text 2.0906 audio 5.2496 | grad 1.994 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5103,8 +5111,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35600 | loss 5.6976 | text 2.1119 audio 5.2752 | grad 1.692 | 1.59s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36600 | loss 5.7210 | text 2.1266 audio 5.2957 | grad 2.306 | 1.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.01 G=1.10 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5130,8 +5138,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35625 | loss 5.7269 | text 2.1576 audio 5.2954 | grad 1.948 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.93 | nonfinite=0 +step 36625 | loss 5.7019 | text 2.1429 audio 5.2733 | grad 1.521 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5157,8 +5165,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35650 | loss 5.7038 | text 2.0968 audio 5.2844 | grad 2.009 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.97 | nonfinite=0 +step 36650 | loss 5.7157 | text 2.1187 audio 5.2920 | grad 1.489 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.46 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5184,8 +5192,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35675 | loss 5.6964 | text 2.1176 audio 5.2729 | grad 1.781 | 1.56s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.86 | nonfinite=0 +step 36675 | loss 5.6895 | text 2.1200 audio 5.2655 | grad 1.530 | 1.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.14 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5211,8 +5219,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35700 | loss 5.7023 | text 2.1459 audio 5.2731 | grad 1.470 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.72 | nonfinite=0 +step 36700 | loss 5.6871 | text 2.1240 audio 5.2623 | grad 1.735 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5238,8 +5246,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35725 | loss 5.6936 | text 2.1079 audio 5.2720 | grad 1.577 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36725 | loss 5.6945 | text 2.1135 audio 5.2718 | grad 2.142 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5265,16 +5273,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35750 | loss 5.7018 | text 2.1473 audio 5.2724 | grad 3.195 | 1.56s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.64 | nonfinite=0 - running validation at step 35750... - val/composite=2.9894 (text=0.5988 audio=4.5832) val/loss=4.7030 cb0_acc=39.1% text_acc=91.8% - [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +step 36750 | loss 5.6832 | text 2.1276 audio 5.2577 | grad 1.583 | 1.55s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 + running validation at step 36750... + val/composite=2.9854 (text=0.5932 audio=4.5801) val/loss=4.6988 cb0_acc=39.2% text_acc=92.0% + [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.8% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_q4bfdk_b/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_jwbvue5g/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b84eb-23e8713c6407ba4a163897ad;da59c5d6-9f19-4363-b138-8063d8495773) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b90c5-6a65d78d5cb61e8b6beabdcb;44000a67-8fdf-4e95-91a1-63030a9bba23) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -5305,8 +5313,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35775 | loss 5.6928 | text 2.1394 audio 5.2649 | grad 2.077 | 21.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.00 | nonfinite=0 +step 36775 | loss 5.7233 | text 2.1374 audio 5.2958 | grad 2.439 | 21.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5332,8 +5340,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35800 | loss 5.7009 | text 2.1499 audio 5.2709 | grad 3.615 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.75 | nonfinite=0 +step 36800 | loss 5.7377 | text 2.1306 audio 5.3116 | grad 2.060 | 1.45s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.01 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5359,8 +5367,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35825 | loss 5.7072 | text 2.1152 audio 5.2842 | grad 1.769 | 1.54s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.79 | nonfinite=0 +step 36825 | loss 5.7254 | text 2.1173 audio 5.3019 | grad 2.148 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5386,8 +5394,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35850 | loss 5.7080 | text 2.1325 audio 5.2815 | grad 1.590 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.73 | nonfinite=0 +step 36850 | loss 5.6654 | text 2.0803 audio 5.2493 | grad 1.928 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.01 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5413,8 +5421,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35875 | loss 5.7191 | text 2.1272 audio 5.2937 | grad 2.377 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 +step 36875 | loss 5.7116 | text 2.1150 audio 5.2885 | grad 1.605 | 1.45s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5440,8 +5448,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35900 | loss 5.6961 | text 2.0964 audio 5.2768 | grad 1.901 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 +step 36900 | loss 5.6732 | text 2.0938 audio 5.2545 | grad 1.952 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5467,8 +5475,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35925 | loss 5.6746 | text 2.1152 audio 5.2516 | grad 1.636 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 +step 36925 | loss 5.6833 | text 2.0830 audio 5.2667 | grad 2.282 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5494,8 +5502,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35950 | loss 5.6555 | text 2.1133 audio 5.2328 | grad 1.572 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.76 | nonfinite=0 +step 36950 | loss 5.6657 | text 2.0712 audio 5.2514 | grad 2.144 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.01 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5521,8 +5529,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35975 | loss 5.6890 | text 2.1208 audio 5.2648 | grad 2.115 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.05 | nonfinite=0 +step 36975 | loss 5.6745 | text 2.0822 audio 5.2581 | grad 2.598 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.30 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5548,15 +5556,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36000 | loss 5.7153 | text 2.1097 audio 5.2934 | grad 2.235 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.10 | nonfinite=0 - running validation at step 36000... - val/composite=2.9897 (text=0.5974 audio=4.5846) val/loss=4.7041 cb0_acc=39.2% text_acc=91.8% - [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9894) +step 37000 | loss 5.6766 | text 2.0915 audio 5.2583 | grad 2.042 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 + running validation at step 37000... + val/composite=2.9927 (text=0.5950 audio=4.5912) val/loss=4.7101 cb0_acc=39.3% text_acc=91.9% + [val] per-codebook acc: cb0=39.3% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9854) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_skck_i9h/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +[ckpt] uploading /tmp/ckpt_w5h8lisx/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5569,13 +5577,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8865-1ec56561651c1a56354c727e;d2e4e91e-8960-46ef-b922-28020e2e7473) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b942b-46693a050c6a3b041fd98fa7;77e27da6-0027-4237-bdef-9d4d0a682e02) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -5591,10 +5597,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36025 | loss 5.6831 | text 2.1034 audio 5.2624 | grad 1.684 | 20.86s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.06 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37025 | loss 5.6793 | text 2.0733 audio 5.2647 | grad 1.558 | 20.62s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5618,10 +5624,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36050 | loss 5.7093 | text 2.0998 audio 5.2893 | grad 1.919 | 1.59s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37050 | loss 5.6463 | text 2.0525 audio 5.2358 | grad 1.873 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.22 | spike L=-0.01 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5645,10 +5651,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36075 | loss 5.6986 | text 2.1149 audio 5.2756 | grad 2.041 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37075 | loss 5.6827 | text 2.0917 audio 5.2643 | grad 1.554 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5672,10 +5678,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36100 | loss 5.7108 | text 2.1152 audio 5.2878 | grad 1.744 | 1.50s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37100 | loss 5.6575 | text 2.0491 audio 5.2476 | grad 1.608 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5699,10 +5705,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36125 | loss 5.6697 | text 2.1321 audio 5.2433 | grad 2.215 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.01 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37125 | loss 5.6627 | text 2.0852 audio 5.2457 | grad 1.778 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5726,10 +5732,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36150 | loss 5.6836 | text 2.1104 audio 5.2616 | grad 1.771 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37150 | loss 5.7070 | text 2.0835 audio 5.2903 | grad 1.519 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5753,10 +5759,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36175 | loss 5.7177 | text 2.1067 audio 5.2964 | grad 1.322 | 1.47s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37175 | loss 5.6957 | text 2.1050 audio 5.2747 | grad 4.542 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.10 | spike L=+0.00 G=2.44 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5780,10 +5786,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36200 | loss 5.6609 | text 2.1179 audio 5.2373 | grad 2.159 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.01 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37200 | loss 5.6781 | text 2.0752 audio 5.2630 | grad 1.735 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5807,10 +5813,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36225 | loss 5.7207 | text 2.1407 audio 5.2925 | grad 2.000 | 1.50s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37225 | loss 5.7024 | text 2.0994 audio 5.2825 | grad 1.957 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5834,14 +5840,23 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36250 | loss 5.6867 | text 2.0964 audio 5.2674 | grad 2.338 | 1.47s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.20 | nonfinite=0 - running validation at step 36250... - val/composite=2.9903 (text=0.5982 audio=4.5851) val/loss=4.7047 cb0_acc=39.2% text_acc=91.8% - [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9894) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37250 | loss 5.6916 | text 2.0812 audio 5.2753 | grad 2.125 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.02 | nonfinite=0 + running validation at step 37250... + val/composite=2.9837 (text=0.5807 audio=4.5857) val/loss=4.7018 cb0_acc=39.3% text_acc=92.1% + [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_7jqz98hv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b977d-18e768341c50c8c634b94256;1deef6d0-e9ee-4962-adfc-011792a7ff9d) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5865,10 +5880,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36275 | loss 5.7501 | text 2.1488 audio 5.3204 | grad 1.751 | 2.30s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37275 | loss 5.6859 | text 2.0602 audio 5.2739 | grad 3.134 | 21.36s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.50 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5892,10 +5907,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36300 | loss 5.6584 | text 2.0785 audio 5.2427 | grad 1.793 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37300 | loss 5.6772 | text 2.0848 audio 5.2602 | grad 2.241 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5919,10 +5934,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36325 | loss 5.7092 | text 2.1386 audio 5.2815 | grad 1.547 | 1.52s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37325 | loss 5.6471 | text 2.0604 audio 5.2350 | grad 1.747 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.01 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5946,10 +5961,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36350 | loss 5.7118 | text 2.1361 audio 5.2846 | grad 2.371 | 1.60s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.24 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37350 | loss 5.6828 | text 2.0960 audio 5.2636 | grad 1.923 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5973,10 +5988,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36375 | loss 5.7245 | text 2.1202 audio 5.3005 | grad 1.915 | 1.54s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37375 | loss 5.6923 | text 2.0948 audio 5.2733 | grad 2.252 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6000,10 +6015,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36400 | loss 5.7019 | text 2.1270 audio 5.2765 | grad 2.564 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.15 | spike L=+0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37400 | loss 5.6938 | text 2.0868 audio 5.2765 | grad 1.727 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6027,10 +6042,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36425 | loss 5.7086 | text 2.1082 audio 5.2870 | grad 1.945 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37425 | loss 5.6795 | text 2.0630 audio 5.2669 | grad 2.381 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6054,8 +6069,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36450 | loss 5.6757 | text 2.1256 audio 5.2505 | grad 1.849 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.92 | nonfinite=0 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37450 | loss 5.6937 | text 2.0768 audio 5.2784 | grad 1.354 | 1.53s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6081,8 +6098,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36475 | loss 5.6761 | text 2.1255 audio 5.2510 | grad 1.958 | 1.48s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 +step 37475 | loss 5.6899 | text 2.0712 audio 5.2757 | grad 3.236 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6108,16 +6125,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36500 | loss 5.6908 | text 2.0905 audio 5.2727 | grad 1.988 | 1.54s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=1.00 | nonfinite=0 - running validation at step 36500... - val/composite=2.9857 (text=0.5922 audio=4.5813) val/loss=4.6998 cb0_acc=39.4% text_acc=91.8% - [val] per-codebook acc: cb0=39.4% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +step 37500 | loss 5.6730 | text 2.0959 audio 5.2538 | grad 1.822 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=-0.00 G=0.84 | nonfinite=0 + running validation at step 37500... + val/composite=2.9835 (text=0.5934 audio=4.5769) val/loss=4.6956 cb0_acc=39.4% text_acc=92.0% + [val] per-codebook acc: cb0=39.4% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_pnu0pz72/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_4vbm_v2l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8d52-3c934e2072b2c29c7b9ed4a4;c8972a1a-b255-4d8a-a62a-f52dd68711e3) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b9ae5-6c9b90af7b04c5cd3db15ea0;96a0f28b-4ee4-4786-b616-80bc413a96be) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -6148,8 +6165,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36525 | loss 5.6839 | text 2.1440 audio 5.2551 | grad 2.686 | 21.43s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.35 | nonfinite=0 +step 37525 | loss 5.6611 | text 2.0692 audio 5.2472 | grad 1.530 | 21.54s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6175,8 +6192,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36550 | loss 5.6939 | text 2.1047 audio 5.2730 | grad 2.494 | 1.52s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.21 | nonfinite=0 +step 37550 | loss 5.6846 | text 2.0928 audio 5.2660 | grad 1.602 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6202,8 +6219,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36575 | loss 5.6677 | text 2.0906 audio 5.2496 | grad 1.994 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.95 | nonfinite=0 +step 37575 | loss 5.6758 | text 2.1009 audio 5.2556 | grad 2.979 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.16 | spike L=-0.00 G=1.47 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6229,8 +6246,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36600 | loss 5.7210 | text 2.1266 audio 5.2957 | grad 2.306 | 1.58s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.01 G=1.10 | nonfinite=0 +step 37600 | loss 5.6719 | text 2.0873 audio 5.2544 | grad 2.455 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6256,8 +6273,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36625 | loss 5.7019 | text 2.1429 audio 5.2733 | grad 1.521 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 +step 37625 | loss 5.7134 | text 2.0897 audio 5.2954 | grad 1.739 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6283,8 +6300,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36650 | loss 5.7157 | text 2.1187 audio 5.2920 | grad 1.489 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.46 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.73 | nonfinite=0 +step 37650 | loss 5.6801 | text 2.0723 audio 5.2656 | grad 2.018 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6310,8 +6327,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36675 | loss 5.6895 | text 2.1200 audio 5.2655 | grad 1.530 | 1.58s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.14 | spike L=-0.00 G=0.77 | nonfinite=0 +step 37675 | loss 5.6702 | text 2.0706 audio 5.2561 | grad 1.751 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6337,8 +6354,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36700 | loss 5.6871 | text 2.1240 audio 5.2623 | grad 1.735 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 +step 37700 | loss 5.6739 | text 2.0701 audio 5.2598 | grad 2.293 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6364,8 +6381,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36725 | loss 5.6945 | text 2.1135 audio 5.2718 | grad 2.142 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 +step 37725 | loss 5.6521 | text 2.0569 audio 5.2408 | grad 1.874 | 1.53s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.88 model_audio_embed=0.06 projection=0.09 text_embed=0.16 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6391,16 +6408,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36750 | loss 5.6832 | text 2.1276 audio 5.2577 | grad 1.583 | 1.55s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 - running validation at step 36750... - val/composite=2.9854 (text=0.5932 audio=4.5801) val/loss=4.6988 cb0_acc=39.2% text_acc=92.0% - [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.8% +step 37750 | loss 5.6452 | text 2.0799 audio 5.2292 | grad 3.034 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.47 | nonfinite=0 + running validation at step 37750... + val/composite=2.9778 (text=0.5780 audio=4.5777) val/loss=4.6933 cb0_acc=39.3% text_acc=92.1% + [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.8% cb6=5.9% cb7=5.8% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_jwbvue5g/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_inpnozqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b90c5-6a65d78d5cb61e8b6beabdcb;44000a67-8fdf-4e95-91a1-63030a9bba23) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b9e54-4bab431c0d0325ff4098a2ea;192d1300-4784-4a97-88ec-81abd1901e75) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -6431,8 +6448,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36775 | loss 5.7233 | text 2.1374 audio 5.2958 | grad 2.439 | 21.48s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.28 | nonfinite=0 +step 37775 | loss 5.6960 | text 2.1047 audio 5.2751 | grad 2.769 | 21.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6458,8 +6475,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36800 | loss 5.7377 | text 2.1306 audio 5.3116 | grad 2.060 | 1.45s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.01 G=1.05 | nonfinite=0 +step 37800 | loss 5.6842 | text 2.0801 audio 5.2682 | grad 1.344 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=+0.00 G=0.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6485,8 +6502,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36825 | loss 5.7254 | text 2.1173 audio 5.3019 | grad 2.148 | 1.52s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.09 | nonfinite=0 +step 37825 | loss 5.6790 | text 2.1186 audio 5.2553 | grad 2.178 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6512,8 +6529,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36850 | loss 5.6654 | text 2.0803 audio 5.2493 | grad 1.928 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.01 G=0.97 | nonfinite=0 +step 37850 | loss 5.6950 | text 2.1125 audio 5.2725 | grad 2.040 | 1.53s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6539,8 +6556,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36875 | loss 5.7116 | text 2.1150 audio 5.2885 | grad 1.605 | 1.45s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 +step 37875 | loss 5.6836 | text 2.0703 audio 5.2695 | grad 3.089 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.45 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6566,8 +6583,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36900 | loss 5.6732 | text 2.0938 audio 5.2545 | grad 1.952 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=1.00 | nonfinite=0 +step 37900 | loss 5.6898 | text 2.0709 audio 5.2756 | grad 1.446 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.48 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6593,8 +6610,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36925 | loss 5.6833 | text 2.0830 audio 5.2667 | grad 2.282 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.17 | nonfinite=0 +step 37925 | loss 5.6369 | text 2.0557 audio 5.2258 | grad 2.639 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6620,8 +6637,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36950 | loss 5.6657 | text 2.0712 audio 5.2514 | grad 2.144 | 1.46s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.01 G=1.08 | nonfinite=0 +step 37950 | loss 5.6719 | text 2.0937 audio 5.2532 | grad 4.150 | 1.45s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=-0.00 G=1.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6647,8 +6664,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36975 | loss 5.6745 | text 2.0822 audio 5.2581 | grad 2.598 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.30 | nonfinite=0 +step 37975 | loss 5.6779 | text 2.0709 audio 5.2637 | grad 1.744 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6674,15 +6691,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37000 | loss 5.6766 | text 2.0915 audio 5.2583 | grad 2.042 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 - running validation at step 37000... - val/composite=2.9927 (text=0.5950 audio=4.5912) val/loss=4.7101 cb0_acc=39.3% text_acc=91.9% - [val] per-codebook acc: cb0=39.3% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9854) +step 38000 | loss 5.6804 | text 2.0998 audio 5.2604 | grad 2.820 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.21 | nonfinite=0 + running validation at step 38000... + val/composite=2.9791 (text=0.5835 audio=4.5762) val/loss=4.6928 cb0_acc=39.3% text_acc=92.0% + [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9778) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_w5h8lisx/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 +[ckpt] uploading /tmp/ckpt_t6hfiptu/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_038000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6695,11 +6712,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_038000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b942b-46693a050c6a3b041fd98fa7;77e27da6-0027-4237-bdef-9d4d0a682e02) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5ba1bc-2d6afe0407d6411c322f0963;1f345238-9f49-4c05-b723-f995d4483d18) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -6717,9 +6733,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37025 | loss 5.6793 | text 2.0733 audio 5.2647 | grad 1.558 | 20.62s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38025 | loss 5.6918 | text 2.0742 audio 5.2769 | grad 2.330 | 20.68s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6744,9 +6760,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37050 | loss 5.6463 | text 2.0525 audio 5.2358 | grad 1.873 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.22 | spike L=-0.01 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38050 | loss 5.6601 | text 2.0805 audio 5.2440 | grad 1.500 | 1.54s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.63 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6771,9 +6787,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37075 | loss 5.6827 | text 2.0917 audio 5.2643 | grad 1.554 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38075 | loss 5.6890 | text 2.1015 audio 5.2687 | grad 1.704 | 1.51s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6798,9 +6814,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37100 | loss 5.6575 | text 2.0491 audio 5.2476 | grad 1.608 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38100 | loss 5.6609 | text 2.0811 audio 5.2446 | grad 2.264 | 1.45s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6825,9 +6841,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37125 | loss 5.6627 | text 2.0852 audio 5.2457 | grad 1.778 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38125 | loss 5.6824 | text 2.0940 audio 5.2636 | grad 1.465 | 1.49s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.66 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6852,9 +6868,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37150 | loss 5.7070 | text 2.0835 audio 5.2903 | grad 1.519 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38150 | loss 5.6563 | text 2.0778 audio 5.2407 | grad 1.632 | 1.42s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6879,9 +6895,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37175 | loss 5.6957 | text 2.1050 audio 5.2747 | grad 4.542 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.10 | spike L=+0.00 G=2.44 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38175 | loss 5.6619 | text 2.0882 audio 5.2443 | grad 1.952 | 1.48s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6906,9 +6922,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37200 | loss 5.6781 | text 2.0752 audio 5.2630 | grad 1.735 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38200 | loss 5.6659 | text 2.1104 audio 5.2438 | grad 2.356 | 1.49s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.00 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6933,9 +6949,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37225 | loss 5.7024 | text 2.0994 audio 5.2825 | grad 1.957 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38225 | loss 5.6877 | text 2.0908 audio 5.2695 | grad 1.914 | 1.43s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6960,22 +6976,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37250 | loss 5.6916 | text 2.0812 audio 5.2753 | grad 2.125 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.02 | nonfinite=0 - running validation at step 37250... - val/composite=2.9837 (text=0.5807 audio=4.5857) val/loss=4.7018 cb0_acc=39.3% text_acc=92.1% - [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_7jqz98hv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b977d-18e768341c50c8c634b94256;1deef6d0-e9ee-4962-adfc-011792a7ff9d) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38250 | loss 5.6801 | text 2.0828 audio 5.2636 | grad 1.628 | 1.51s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.78 | nonfinite=0 + running validation at step 38250... + val/composite=2.9816 (text=0.5869 audio=4.5781) val/loss=4.6955 cb0_acc=39.4% text_acc=92.1% + [val] per-codebook acc: cb0=39.4% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9778) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7000,9 +7007,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37275 | loss 5.6859 | text 2.0602 audio 5.2739 | grad 3.134 | 21.36s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.50 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38275 | loss 5.6750 | text 2.0800 audio 5.2590 | grad 3.931 | 2.23s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.17 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=+0.00 G=1.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7027,9 +7034,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37300 | loss 5.6772 | text 2.0848 audio 5.2602 | grad 2.241 | 1.46s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38300 | loss 5.6922 | text 2.0851 audio 5.2752 | grad 2.501 | 1.50s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.00 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7054,9 +7061,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37325 | loss 5.6471 | text 2.0604 audio 5.2350 | grad 1.747 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.01 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38325 | loss 5.6549 | text 2.0744 audio 5.2400 | grad 2.303 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7081,9 +7088,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37350 | loss 5.6828 | text 2.0960 audio 5.2636 | grad 1.923 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38350 | loss 5.6523 | text 2.0899 audio 5.2343 | grad 1.799 | 1.44s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7108,9 +7115,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37375 | loss 5.6923 | text 2.0948 audio 5.2733 | grad 2.252 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38375 | loss 5.6499 | text 2.0640 audio 5.2370 | grad 1.905 | 1.53s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7135,9 +7142,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37400 | loss 5.6938 | text 2.0868 audio 5.2765 | grad 1.727 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38400 | loss 5.6520 | text 2.0630 audio 5.2394 | grad 2.365 | 1.49s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7162,9 +7169,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37425 | loss 5.6795 | text 2.0630 audio 5.2669 | grad 2.381 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38425 | loss 5.6757 | text 2.0639 audio 5.2629 | grad 1.445 | 1.50s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.46 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7189,9 +7196,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37450 | loss 5.6937 | text 2.0768 audio 5.2784 | grad 1.354 | 1.53s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38450 | loss 5.6607 | text 2.0936 audio 5.2420 | grad 1.638 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7216,9 +7223,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37475 | loss 5.6899 | text 2.0712 audio 5.2757 | grad 3.236 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38475 | loss 5.6472 | text 2.0780 audio 5.2316 | grad 1.912 | 1.44s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7243,22 +7250,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37500 | loss 5.6730 | text 2.0959 audio 5.2538 | grad 1.822 | 1.46s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=-0.00 G=0.84 | nonfinite=0 - running validation at step 37500... - val/composite=2.9835 (text=0.5934 audio=4.5769) val/loss=4.6956 cb0_acc=39.4% text_acc=92.0% - [val] per-codebook acc: cb0=39.4% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_4vbm_v2l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b9ae5-6c9b90af7b04c5cd3db15ea0;96a0f28b-4ee4-4786-b616-80bc413a96be) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38500 | loss 5.6695 | text 2.0789 audio 5.2537 | grad 1.799 | 1.53s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.87 | nonfinite=0 + running validation at step 38500... + val/composite=2.9827 (text=0.5876 audio=4.5795) val/loss=4.6970 cb0_acc=39.3% text_acc=92.0% + [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9778) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7283,9 +7281,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37525 | loss 5.6611 | text 2.0692 audio 5.2472 | grad 1.530 | 21.54s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38525 | loss 5.6421 | text 2.0767 audio 5.2268 | grad 1.934 | 2.26s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7310,9 +7308,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37550 | loss 5.6846 | text 2.0928 audio 5.2660 | grad 1.602 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38550 | loss 5.6749 | text 2.0708 audio 5.2607 | grad 1.501 | 1.47s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.46 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.17 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7337,9 +7335,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37575 | loss 5.6758 | text 2.1009 audio 5.2556 | grad 2.979 | 1.46s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.16 | spike L=-0.00 G=1.47 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38575 | loss 5.6646 | text 2.0925 audio 5.2461 | grad 1.519 | 1.55s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7364,9 +7362,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37600 | loss 5.6719 | text 2.0873 audio 5.2544 | grad 2.455 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38600 | loss 5.6306 | text 2.0843 audio 5.2137 | grad 1.995 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.01 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7391,9 +7389,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37625 | loss 5.7134 | text 2.0897 audio 5.2954 | grad 1.739 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38625 | loss 5.6835 | text 2.1067 audio 5.2622 | grad 1.654 | 1.51s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.53 lora=0.84 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.85 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7418,9 +7416,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37650 | loss 5.6801 | text 2.0723 audio 5.2656 | grad 2.018 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38650 | loss 5.6796 | text 2.0794 audio 5.2637 | grad 1.375 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7445,9 +7443,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37675 | loss 5.6702 | text 2.0706 audio 5.2561 | grad 1.751 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38675 | loss 5.6819 | text 2.0777 audio 5.2664 | grad 1.549 | 1.54s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7472,9 +7470,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37700 | loss 5.6739 | text 2.0701 audio 5.2598 | grad 2.293 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38700 | loss 5.6835 | text 2.0812 audio 5.2673 | grad 2.113 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7499,9 +7497,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37725 | loss 5.6521 | text 2.0569 audio 5.2408 | grad 1.874 | 1.53s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.43 lora=0.88 model_audio_embed=0.06 projection=0.09 text_embed=0.16 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38725 | loss 5.6632 | text 2.0486 audio 5.2535 | grad 1.734 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7526,22 +7524,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37750 | loss 5.6452 | text 2.0799 audio 5.2292 | grad 3.034 | 1.50s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.47 | nonfinite=0 - running validation at step 37750... - val/composite=2.9778 (text=0.5780 audio=4.5777) val/loss=4.6933 cb0_acc=39.3% text_acc=92.1% - [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.8% cb6=5.9% cb7=5.8% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_inpnozqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b9e54-4bab431c0d0325ff4098a2ea;192d1300-4784-4a97-88ec-81abd1901e75) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38750 | loss 5.6952 | text 2.0686 audio 5.2815 | grad 2.250 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.22 | nonfinite=0 + running validation at step 38750... + val/composite=2.9841 (text=0.5856 audio=4.5831) val/loss=4.7002 cb0_acc=39.3% text_acc=92.1% + [val] per-codebook acc: cb0=39.3% cb1=20.4% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 6/10 (best 2.9778) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7566,9 +7555,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37775 | loss 5.6960 | text 2.1047 audio 5.2751 | grad 2.769 | 21.58s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38775 | loss 5.6627 | text 2.0807 audio 5.2465 | grad 2.009 | 2.27s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7593,9 +7582,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37800 | loss 5.6842 | text 2.0801 audio 5.2682 | grad 1.344 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=+0.00 G=0.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38800 | loss 5.6479 | text 2.0619 audio 5.2355 | grad 3.865 | 1.44s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.03 projection=0.04 text_embed=0.08 | spike L=-0.00 G=2.04 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7620,9 +7609,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37825 | loss 5.6790 | text 2.1186 audio 5.2553 | grad 2.178 | 1.46s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38825 | loss 5.6851 | text 2.0675 audio 5.2716 | grad 2.043 | 1.53s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7647,9 +7636,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37850 | loss 5.6950 | text 2.1125 audio 5.2725 | grad 2.040 | 1.53s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38850 | loss 5.6735 | text 2.0862 audio 5.2563 | grad 1.782 | 1.48s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.85 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7674,9 +7663,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37875 | loss 5.6836 | text 2.0703 audio 5.2695 | grad 3.089 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.45 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38875 | loss 5.6620 | text 2.0786 audio 5.2462 | grad 3.562 | 1.50s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.00 G=1.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7701,9 +7690,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37900 | loss 5.6898 | text 2.0709 audio 5.2756 | grad 1.446 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.48 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38900 | loss 5.6669 | text 2.0859 audio 5.2497 | grad 2.829 | 1.48s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.14 | spike L=-0.00 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7728,9 +7717,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37925 | loss 5.6369 | text 2.0557 audio 5.2258 | grad 2.639 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38925 | loss 5.6480 | text 2.0809 audio 5.2318 | grad 2.846 | 1.50s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.24 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7755,9 +7744,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37950 | loss 5.6719 | text 2.0937 audio 5.2532 | grad 4.150 | 1.45s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=-0.00 G=1.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38950 | loss 5.7113 | text 2.0936 audio 5.2925 | grad 1.387 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.04 projection=0.09 text_embed=0.11 | spike L=+0.01 G=0.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7782,9 +7771,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37975 | loss 5.6779 | text 2.0709 audio 5.2637 | grad 1.744 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 38975 | loss 5.6685 | text 2.0859 audio 5.2513 | grad 5.426 | 1.49s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.14 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.10 | spike L=-0.00 G=2.43 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7809,13 +7798,23 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 38000 | loss 5.6804 | text 2.0998 audio 5.2604 | grad 2.820 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.21 | nonfinite=0 - running validation at step 38000... - val/composite=2.9791 (text=0.5835 audio=4.5762) val/loss=4.6928 cb0_acc=39.3% text_acc=92.0% - [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9778) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 39000 | loss 5.6494 | text 2.1020 audio 5.2290 | grad 1.849 | 1.52s/step | peak -1.0G | host_rss 47.5G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.72 | nonfinite=0 + running validation at step 39000... + val/composite=2.9716 (text=0.5735 audio=4.5702) val/loss=4.6849 cb0_acc=39.4% text_acc=92.3% + [val] per-codebook acc: cb0=39.4% cb1=20.3% cb2=17.2% cb3=11.2% cb4=8.5% cb5=6.9% cb6=5.9% cb7=5.8% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_t6hfiptu/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_038000 +[ckpt] uploading /tmp/ckpt_8ifz816z/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5ba9ad-24e7f7063076401166167b14;e383a446-c4af-4718-b4ae-f0642afeff45) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_byd68hym/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_039000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3