diff --git "a/logs/train_host0_latest.log" "b/logs/train_host0_latest.log" --- "a/logs/train_host0_latest.log" +++ "b/logs/train_host0_latest.log" @@ -1,4 +1,4 @@ -1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +ch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -12,14 +12,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30100 | loss 5.7109 | text 2.1508 audio 5.2807 | grad 2.789 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.28 | nonfinite=0 +step 31100 | loss 5.7458 | text 2.1828 audio 5.3092 | grad 2.135 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.06 projection=0.06 text_embed=0.14 | spike L=+0.01 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -45,8 +39,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30125 | loss 5.7309 | text 2.1609 audio 5.2987 | grad 2.277 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.01 | nonfinite=0 +step 31125 | loss 5.7439 | text 2.2002 audio 5.3039 | grad 2.645 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -72,8 +66,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30150 | loss 5.7423 | text 2.1902 audio 5.3042 | grad 1.819 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 +step 31150 | loss 5.7190 | text 2.1559 audio 5.2878 | grad 1.900 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -99,8 +93,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30175 | loss 5.7186 | text 2.1692 audio 5.2848 | grad 2.105 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.95 | nonfinite=0 +step 31175 | loss 5.7093 | text 2.1537 audio 5.2786 | grad 2.073 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -126,8 +120,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30200 | loss 5.7341 | text 2.1761 audio 5.2988 | grad 1.584 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 +step 31200 | loss 5.7151 | text 2.1503 audio 5.2850 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -153,9 +147,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 4.1x (grad 8.69 vs EMA 2.13) step 30225 -step 30225 | loss 5.7041 | text 2.1497 audio 5.2742 | grad 8.692 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.07 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.06 | spike L=-0.00 G=4.07 | nonfinite=0 +step 31225 | loss 5.7159 | text 2.1445 audio 5.2870 | grad 1.819 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -181,12 +174,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30250 | loss 5.7108 | text 2.1812 audio 5.2746 | grad 1.567 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.56 | nonfinite=0 - running validation at step 30250... - val/composite=3.0263 (text=0.6452 audio=4.6137) val/loss=4.7427 cb0_acc=38.8% text_acc=90.7% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0203) +step 31250 | loss 5.7228 | text 2.1591 audio 5.2910 | grad 2.032 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 + running validation at step 31250... + val/composite=3.0194 (text=0.6437 audio=4.6031) val/loss=4.7319 cb0_acc=38.7% text_acc=90.7% + [val] per-codebook acc: cb0=38.7% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0163) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -212,8 +205,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30275 | loss 5.7370 | text 2.1381 audio 5.3094 | grad 1.647 | 2.35s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 +step 31275 | loss 5.7330 | text 2.1806 audio 5.2969 | grad 1.797 | 2.38s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -239,8 +232,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30300 | loss 5.7740 | text 2.1916 audio 5.3357 | grad 3.100 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=+0.01 G=1.21 | nonfinite=0 +step 31300 | loss 5.7208 | text 2.1705 audio 5.2867 | grad 2.195 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -266,8 +259,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30325 | loss 5.6988 | text 2.1558 audio 5.2676 | grad 2.002 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.76 | nonfinite=0 +step 31325 | loss 5.7276 | text 2.1491 audio 5.2978 | grad 2.566 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.18 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -293,8 +286,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30350 | loss 5.7251 | text 2.1526 audio 5.2946 | grad 2.659 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.08 | spike L=-0.00 G=1.04 | nonfinite=0 +step 31350 | loss 5.7069 | text 2.1178 audio 5.2834 | grad 1.952 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -320,8 +313,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30375 | loss 5.7275 | text 2.1723 audio 5.2930 | grad 1.326 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=-0.00 G=0.52 | nonfinite=0 +step 31375 | loss 5.7308 | text 2.1203 audio 5.3067 | grad 1.813 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -347,8 +340,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30400 | loss 5.7535 | text 2.1742 audio 5.3186 | grad 1.857 | 1.44s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.76 | nonfinite=0 +step 31400 | loss 5.7433 | text 2.1584 audio 5.3116 | grad 2.645 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -374,8 +367,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30425 | loss 5.6897 | text 2.1361 audio 5.2625 | grad 1.655 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.69 | nonfinite=0 +step 31425 | loss 5.7413 | text 2.1406 audio 5.3132 | grad 1.860 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -401,8 +394,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30450 | loss 5.7256 | text 2.1539 audio 5.2948 | grad 3.586 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.55 | nonfinite=0 +step 31450 | loss 5.7094 | text 2.1542 audio 5.2786 | grad 1.536 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -428,8 +421,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30475 | loss 5.7147 | text 2.1494 audio 5.2848 | grad 2.515 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.00 G=1.03 | nonfinite=0 +step 31475 | loss 5.7086 | text 2.1526 audio 5.2781 | grad 3.518 | 1.59s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -455,12 +448,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30500 | loss 5.7465 | text 2.1749 audio 5.3115 | grad 1.691 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.69 | nonfinite=0 - running validation at step 30500... - val/composite=3.0241 (text=0.6463 audio=4.6093) val/loss=4.7386 cb0_acc=38.8% text_acc=90.6% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0203) +step 31500 | loss 5.7234 | text 2.1505 audio 5.2933 | grad 2.082 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.93 | nonfinite=0 + running validation at step 31500... + val/composite=3.0129 (text=0.6302 audio=4.6013) val/loss=4.7273 cb0_acc=38.9% text_acc=90.9% + [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_97c5nxlf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b51f5-564c334001068ae16e82ce0f;aafdec3d-7141-4994-a73d-192be2632797) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -486,8 +488,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30525 | loss 5.7592 | text 2.1927 audio 5.3207 | grad 1.734 | 2.37s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.73 | nonfinite=0 +step 31525 | loss 5.7418 | text 2.1563 audio 5.3106 | grad 2.122 | 21.38s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -513,8 +515,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30550 | loss 5.7289 | text 2.1768 audio 5.2936 | grad 1.688 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.73 | nonfinite=0 +step 31550 | loss 5.7494 | text 2.1284 audio 5.3237 | grad 2.364 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -540,8 +542,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30575 | loss 5.7234 | text 2.1289 audio 5.2976 | grad 2.456 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.09 | nonfinite=0 +step 31575 | loss 5.7309 | text 2.1538 audio 5.3001 | grad 1.760 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -567,8 +569,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30600 | loss 5.6997 | text 2.1438 audio 5.2709 | grad 1.741 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 +step 31600 | loss 5.7379 | text 2.1287 audio 5.3122 | grad 3.343 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.53 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -594,8 +596,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30625 | loss 5.7534 | text 2.1707 audio 5.3192 | grad 1.956 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 +step 31625 | loss 5.6686 | text 2.1377 audio 5.2410 | grad 2.997 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.30 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -621,8 +623,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30650 | loss 5.7253 | text 2.1289 audio 5.2995 | grad 2.165 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.99 | nonfinite=0 +step 31650 | loss 5.6969 | text 2.1637 audio 5.2641 | grad 1.750 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -648,8 +650,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30675 | loss 5.7269 | text 2.1563 audio 5.2956 | grad 3.860 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.77 | nonfinite=0 +step 31675 | loss 5.7165 | text 2.1522 audio 5.2861 | grad 2.261 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -675,8 +677,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30700 | loss 5.7398 | text 2.1505 audio 5.3097 | grad 2.235 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.95 | nonfinite=0 +step 31700 | loss 5.6886 | text 2.1403 audio 5.2606 | grad 2.322 | 1.57s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -702,8 +704,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30725 | loss 5.7489 | text 2.1535 audio 5.3182 | grad 2.031 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.00 G=0.87 | nonfinite=0 +step 31725 | loss 5.7274 | text 2.1259 audio 5.3022 | grad 1.478 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -729,21 +731,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30750 | loss 5.6953 | text 2.1131 audio 5.2727 | grad 3.574 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.17 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.55 | nonfinite=0 - running validation at step 30750... - val/composite=3.0163 (text=0.6371 audio=4.6024) val/loss=4.7299 cb0_acc=38.8% text_acc=90.9% - [val] per-codebook acc: cb0=38.8% cb1=19.9% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_v28gzt94/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b49a8-49c8b3a512d01dcf62ba2de6;c1d11542-5c8f-4019-8577-05bf7d3f521f) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 31750 | loss 5.6918 | text 2.1276 audio 5.2663 | grad 2.754 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.24 | nonfinite=0 + running validation at step 31750... + val/composite=3.0149 (text=0.6321 audio=4.6034) val/loss=4.7299 cb0_acc=39.0% text_acc=90.9% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0129) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -769,8 +762,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30775 | loss 5.7336 | text 2.1703 audio 5.2995 | grad 2.588 | 21.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.06 | nonfinite=0 +step 31775 | loss 5.7190 | text 2.1568 audio 5.2876 | grad 1.540 | 2.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -796,8 +789,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30800 | loss 5.7168 | text 2.1540 audio 5.2860 | grad 2.014 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.82 | nonfinite=0 +step 31800 | loss 5.7086 | text 2.1397 audio 5.2807 | grad 1.333 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -823,8 +816,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30825 | loss 5.7184 | text 2.1635 audio 5.2857 | grad 1.822 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 +step 31825 | loss 5.7059 | text 2.1289 audio 5.2801 | grad 1.661 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -850,8 +843,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30850 | loss 5.7180 | text 2.1629 audio 5.2854 | grad 1.949 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.83 | nonfinite=0 +step 31850 | loss 5.7276 | text 2.1494 audio 5.2978 | grad 1.604 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=+0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -877,8 +870,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30875 | loss 5.7182 | text 2.1146 audio 5.2953 | grad 1.590 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.69 | nonfinite=0 + [ALERT] grad spike 3.4x (grad 6.92 vs EMA 2.02) step 31875 +step 31875 | loss 5.7635 | text 2.1532 audio 5.3329 | grad 6.916 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.10 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.08 | spike L=+0.01 G=3.42 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -904,8 +898,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30900 | loss 5.6918 | text 2.1653 audio 5.2588 | grad 3.649 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.63 | nonfinite=0 +step 31900 | loss 5.7007 | text 2.1354 audio 5.2736 | grad 1.531 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.10 text_embed=0.12 | spike L=-0.00 G=0.61 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -931,8 +925,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30925 | loss 5.7255 | text 2.1531 audio 5.2949 | grad 1.466 | 1.43s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 +step 31925 | loss 5.7284 | text 2.1511 audio 5.2981 | grad 2.452 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -958,8 +952,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30950 | loss 5.7009 | text 2.1495 audio 5.2710 | grad 2.898 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.27 | nonfinite=0 +step 31950 | loss 5.7026 | text 2.1339 audio 5.2759 | grad 2.128 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -985,8 +979,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30975 | loss 5.7335 | text 2.1454 audio 5.3044 | grad 2.994 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.27 | nonfinite=0 +step 31975 | loss 5.7325 | text 2.1323 audio 5.3060 | grad 3.383 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.00 G=1.42 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1012,15 +1006,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31000 | loss 5.7318 | text 2.1491 audio 5.3019 | grad 1.703 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.71 | nonfinite=0 - running validation at step 31000... - val/composite=3.0191 (text=0.6409 audio=4.6045) val/loss=4.7327 cb0_acc=39.0% text_acc=90.8% - [val] per-codebook acc: cb0=39.0% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0163) +step 32000 | loss 5.7196 | text 2.1309 audio 5.2935 | grad 2.305 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.93 | nonfinite=0 + running validation at step 32000... + val/composite=3.0141 (text=0.6317 audio=4.6024) val/loss=4.7287 cb0_acc=38.9% text_acc=90.9% + [val] per-codebook acc: cb0=38.9% cb1=20.0% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0129) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_y6mkwqap/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 +[ckpt] uploading /tmp/ckpt_y3j7s14l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1032,11 +1026,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4d0b-4f6ce7c1015680b77c18485b;4de181e9-748f-43df-8a2c-956408fcae48) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5701-676c84df289d499c023dd0e8;593c1b3a-6ef8-4474-8a06-3d170a535244) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -1055,8 +1049,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31025 | loss 5.7013 | text 2.1689 audio 5.2675 | grad 2.218 | 20.63s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.95 | nonfinite=0 +step 32025 | loss 5.7313 | text 2.1524 audio 5.3008 | grad 1.656 | 20.67s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1082,8 +1076,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31050 | loss 5.6785 | text 2.1305 audio 5.2524 | grad 2.106 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 +step 32050 | loss 5.7048 | text 2.1554 audio 5.2737 | grad 2.361 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1109,8 +1103,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31075 | loss 5.7067 | text 2.1458 audio 5.2776 | grad 2.192 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.95 | nonfinite=0 +step 32075 | loss 5.7039 | text 2.0979 audio 5.2843 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1136,8 +1130,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31100 | loss 5.7458 | text 2.1828 audio 5.3092 | grad 2.135 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.06 projection=0.06 text_embed=0.14 | spike L=+0.01 G=0.93 | nonfinite=0 +step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1163,8 +1157,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31125 | loss 5.7439 | text 2.2002 audio 5.3039 | grad 2.645 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.16 | nonfinite=0 +step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1190,8 +1184,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31150 | loss 5.7190 | text 2.1559 audio 5.2878 | grad 1.900 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.82 | nonfinite=0 +step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1217,8 +1211,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31175 | loss 5.7093 | text 2.1537 audio 5.2786 | grad 2.073 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.91 | nonfinite=0 +step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1244,8 +1238,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31200 | loss 5.7151 | text 2.1503 audio 5.2850 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.18 | nonfinite=0 +step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1271,8 +1265,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31225 | loss 5.7159 | text 2.1445 audio 5.2870 | grad 1.819 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.79 | nonfinite=0 +step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1298,12 +1292,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31250 | loss 5.7228 | text 2.1591 audio 5.2910 | grad 2.032 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 - running validation at step 31250... - val/composite=3.0194 (text=0.6437 audio=4.6031) val/loss=4.7319 cb0_acc=38.7% text_acc=90.7% - [val] per-codebook acc: cb0=38.7% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0163) +step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 + running validation at step 32250... + val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1329,8 +1332,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31275 | loss 5.7330 | text 2.1806 audio 5.2969 | grad 1.797 | 2.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.81 | nonfinite=0 +step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1356,8 +1359,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31300 | loss 5.7208 | text 2.1705 audio 5.2867 | grad 2.195 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 +step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1383,8 +1386,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31325 | loss 5.7276 | text 2.1491 audio 5.2978 | grad 2.566 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.18 | spike L=+0.00 G=1.17 | nonfinite=0 +step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1410,8 +1413,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31350 | loss 5.7069 | text 2.1178 audio 5.2834 | grad 1.952 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=-0.00 G=0.88 | nonfinite=0 +step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1437,8 +1440,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31375 | loss 5.7308 | text 2.1203 audio 5.3067 | grad 1.813 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 +step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1464,8 +1467,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31400 | loss 5.7433 | text 2.1584 audio 5.3116 | grad 2.645 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.23 | nonfinite=0 +step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1491,8 +1494,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31425 | loss 5.7413 | text 2.1406 audio 5.3132 | grad 1.860 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.84 | nonfinite=0 +step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1518,8 +1521,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31450 | loss 5.7094 | text 2.1542 audio 5.2786 | grad 1.536 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.71 | nonfinite=0 +step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1545,8 +1548,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31475 | loss 5.7086 | text 2.1526 audio 5.2781 | grad 3.518 | 1.59s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.67 | nonfinite=0 +step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1572,16 +1575,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31500 | loss 5.7234 | text 2.1505 audio 5.2933 | grad 2.082 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 31500... - val/composite=3.0129 (text=0.6302 audio=4.6013) val/loss=4.7273 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 + running validation at step 32500... + val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_97c5nxlf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b51f5-564c334001068ae16e82ce0f;aafdec3d-7141-4994-a73d-192be2632797) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -1612,8 +1615,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31525 | loss 5.7418 | text 2.1563 audio 5.3106 | grad 2.122 | 21.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 +step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1639,8 +1642,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31550 | loss 5.7494 | text 2.1284 audio 5.3237 | grad 2.364 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.06 | nonfinite=0 +step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1666,8 +1669,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31575 | loss 5.7309 | text 2.1538 audio 5.3001 | grad 1.760 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 +step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1693,8 +1696,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31600 | loss 5.7379 | text 2.1287 audio 5.3122 | grad 3.343 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.53 | nonfinite=0 +step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1720,8 +1723,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31625 | loss 5.6686 | text 2.1377 audio 5.2410 | grad 2.997 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.30 | nonfinite=0 +step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1747,8 +1750,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31650 | loss 5.6969 | text 2.1637 audio 5.2641 | grad 1.750 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 +step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1774,8 +1777,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31675 | loss 5.7165 | text 2.1522 audio 5.2861 | grad 2.261 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 + [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 +step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1801,8 +1805,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31700 | loss 5.6886 | text 2.1403 audio 5.2606 | grad 2.322 | 1.57s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.01 | nonfinite=0 +step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1828,8 +1832,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31725 | loss 5.7274 | text 2.1259 audio 5.3022 | grad 1.478 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.64 | nonfinite=0 +step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1855,12 +1859,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31750 | loss 5.6918 | text 2.1276 audio 5.2663 | grad 2.754 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.24 | nonfinite=0 - running validation at step 31750... - val/composite=3.0149 (text=0.6321 audio=4.6034) val/loss=4.7299 cb0_acc=39.0% text_acc=90.9% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0129) +step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 + running validation at step 32750... + val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1886,8 +1899,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31775 | loss 5.7190 | text 2.1568 audio 5.2876 | grad 1.540 | 2.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 +step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1913,8 +1926,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31800 | loss 5.7086 | text 2.1397 audio 5.2807 | grad 1.333 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.60 | nonfinite=0 +step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1940,8 +1953,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31825 | loss 5.7059 | text 2.1289 audio 5.2801 | grad 1.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1967,8 +1980,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31850 | loss 5.7276 | text 2.1494 audio 5.2978 | grad 1.604 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=+0.00 G=0.77 | nonfinite=0 +step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1994,9 +2007,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 3.4x (grad 6.92 vs EMA 2.02) step 31875 -step 31875 | loss 5.7635 | text 2.1532 audio 5.3329 | grad 6.916 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.10 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.08 | spike L=+0.01 G=3.42 | nonfinite=0 +step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2022,8 +2034,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31900 | loss 5.7007 | text 2.1354 audio 5.2736 | grad 1.531 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.10 text_embed=0.12 | spike L=-0.00 G=0.61 | nonfinite=0 +step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2049,8 +2061,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31925 | loss 5.7284 | text 2.1511 audio 5.2981 | grad 2.452 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.02 | nonfinite=0 +step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2076,8 +2088,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31950 | loss 5.7026 | text 2.1339 audio 5.2759 | grad 2.128 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=0.88 | nonfinite=0 +step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2103,8 +2115,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31975 | loss 5.7325 | text 2.1323 audio 5.3060 | grad 3.383 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.00 G=1.42 | nonfinite=0 +step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2130,15 +2142,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32000 | loss 5.7196 | text 2.1309 audio 5.2935 | grad 2.305 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 32000... - val/composite=3.0141 (text=0.6317 audio=4.6024) val/loss=4.7287 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.0% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0129) +step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 + running validation at step 33000... + val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% + [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_y3j7s14l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 +[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2149,17 +2161,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5701-676c84df289d499c023dd0e8;593c1b3a-6ef8-4474-8a06-3d170a535244) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2173,8 +2184,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32025 | loss 5.7313 | text 2.1524 audio 5.3008 | grad 1.656 | 20.67s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.67 | nonfinite=0 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2200,8 +2212,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32050 | loss 5.7048 | text 2.1554 audio 5.2737 | grad 2.361 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 +step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2227,8 +2239,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32075 | loss 5.7039 | text 2.0979 audio 5.2843 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.77 | nonfinite=0 +step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2254,8 +2266,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 +step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2281,8 +2293,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 +step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2308,8 +2320,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 +step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2335,8 +2347,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 +step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2362,8 +2374,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 +step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2389,8 +2401,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 +step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2416,16 +2428,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 - running validation at step 32250... - val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 + running validation at step 33250... + val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -2456,8 +2468,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2483,8 +2495,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 +step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2510,8 +2522,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 +step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2537,8 +2549,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 +step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2564,8 +2576,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2591,8 +2603,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 +step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2618,8 +2630,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 +step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2645,8 +2657,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 +step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2672,8 +2684,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 +step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2699,16 +2711,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 - running validation at step 32500... - val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 + running validation at step 33500... + val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -2739,8 +2751,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2766,8 +2778,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 +step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2793,8 +2805,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 +step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2820,8 +2832,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 +step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2847,8 +2859,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 +step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2874,8 +2886,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 +step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2901,9 +2913,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 -step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 +step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2929,8 +2940,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 +step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2956,8 +2967,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 +step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2983,21 +2994,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 - running validation at step 32750... - val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 + running validation at step 33750... + val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3023,8 +3025,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 +step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3050,8 +3052,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 +step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3077,8 +3079,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 +step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3104,8 +3106,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 +step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3131,8 +3133,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 +step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3158,8 +3160,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 +step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3185,8 +3187,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 +step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3212,8 +3214,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 +step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3239,8 +3241,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3266,15 +3268,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 - running validation at step 33000... - val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) +step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 + running validation at step 34000... + val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% + [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 +[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3285,21 +3296,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3309,8 +3320,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 +step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3336,8 +3347,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 +step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3363,8 +3374,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 +step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3390,8 +3401,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 +step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3417,8 +3428,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 +step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3444,8 +3455,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 +step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3471,8 +3482,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 +step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3498,8 +3509,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 +step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3525,8 +3536,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 +step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3552,21 +3563,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 - running validation at step 33250... - val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 + running validation at step 34250... + val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3592,8 +3594,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 +step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3619,8 +3621,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 +step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3646,8 +3648,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 +step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3673,8 +3675,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 +step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3700,8 +3702,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 +step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3727,8 +3729,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 +step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3754,8 +3756,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 +step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3781,8 +3783,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 +step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3808,8 +3810,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 +step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3835,21 +3837,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 - running validation at step 33500... - val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 + running validation at step 34500... + val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3875,8 +3868,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 +step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3902,8 +3895,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 +step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3929,8 +3922,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 +step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3956,8 +3949,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 +step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3983,8 +3976,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 +step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4010,8 +4003,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 +step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4037,8 +4030,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 +step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4064,8 +4057,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 +step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4091,8 +4084,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 +step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4118,12 +4111,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 - running validation at step 33750... - val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) +step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 + running validation at step 34750... + val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% + [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4149,8 +4142,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 +step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4176,8 +4169,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 +step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4203,8 +4196,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 +step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4230,8 +4223,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 +step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4257,8 +4250,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 +step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4284,8 +4277,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 +step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4311,8 +4304,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 +step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4338,8 +4331,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 +step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4365,8 +4358,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 +step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4392,16 +4385,43 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 - running validation at step 34000... - val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% - [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) +step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 + [audio-demo] ar_cb0_acc=12.0% (35s) + running validation at step 35000... + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB + pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 + val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -4409,7 +4429,7 @@ Make sure your token has the correct permissions. * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4420,19 +4440,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7c65-1c38aecc5ca4acdb5008e0b4;48e1227e-d30a-4783-b035-8b02985e5e1d) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4444,11 +4461,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35025 | loss 5.7000 | text 2.1331 audio 5.2734 | grad 2.588 | 41.18s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4471,11 +4488,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35050 | loss 5.7063 | text 2.1069 audio 5.2850 | grad 2.479 | 1.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4498,11 +4515,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + [ALERT] grad spike 3.0x (grad 6.76 vs EMA 2.25) step 35075 +step 35075 | loss 5.7092 | text 2.1201 audio 5.2851 | grad 6.761 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.11 lora=0.98 model_audio_embed=0.06 projection=0.02 text_embed=0.12 | spike L=+0.00 G=3.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4525,11 +4543,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35100 | loss 5.7286 | text 2.1354 audio 5.3015 | grad 2.020 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4552,11 +4570,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35125 | loss 5.6800 | text 2.1305 audio 5.2539 | grad 1.847 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4579,11 +4597,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35150 | loss 5.7182 | text 2.1559 audio 5.2870 | grad 1.641 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4606,11 +4624,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35175 | loss 5.6631 | text 2.1155 audio 5.2400 | grad 2.109 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4633,11 +4651,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35200 | loss 5.7188 | text 2.1798 audio 5.2828 | grad 2.145 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4660,11 +4678,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35225 | loss 5.7170 | text 2.1077 audio 5.2955 | grad 1.866 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4687,15 +4705,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 - running validation at step 34250... - val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35250 | loss 5.6647 | text 2.0979 audio 5.2451 | grad 1.703 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.73 | nonfinite=0 + running validation at step 35250... + val/composite=2.9985 (text=0.6053 audio=4.5939) val/loss=4.7150 cb0_acc=39.1% text_acc=91.6% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9973) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4718,11 +4736,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35275 | loss 5.7344 | text 2.1597 audio 5.3025 | grad 4.716 | 2.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=+0.01 G=2.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4745,11 +4763,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35300 | loss 5.6869 | text 2.1172 audio 5.2634 | grad 1.765 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4772,11 +4790,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35325 | loss 5.7376 | text 2.1312 audio 5.3113 | grad 1.650 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4799,11 +4817,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35350 | loss 5.6878 | text 2.1130 audio 5.2652 | grad 1.950 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4826,11 +4844,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35375 | loss 5.6788 | text 2.1222 audio 5.2544 | grad 1.724 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4853,11 +4871,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35400 | loss 5.6788 | text 2.1020 audio 5.2584 | grad 1.961 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4880,11 +4898,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35425 | loss 5.6969 | text 2.0995 audio 5.2770 | grad 1.766 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4907,11 +4925,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35450 | loss 5.6848 | text 2.0913 audio 5.2665 | grad 1.521 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4934,11 +4952,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35475 | loss 5.7112 | text 2.1194 audio 5.2873 | grad 2.124 | 1.60s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4961,15 +4979,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 - running validation at step 34500... - val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35500 | loss 5.7002 | text 2.0979 audio 5.2806 | grad 2.371 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 + running validation at step 35500... + val/composite=2.9923 (text=0.5962 audio=4.5896) val/loss=4.7089 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_3crsay5y/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b816c-2712d01710663cc45f59dc84;4de1714a-3953-4c64-bba2-5dfc7dfb76bb) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4992,11 +5019,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35525 | loss 5.7157 | text 2.1252 audio 5.2907 | grad 1.712 | 21.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5019,11 +5046,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35550 | loss 5.7160 | text 2.1306 audio 5.2898 | grad 2.455 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5046,11 +5073,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35575 | loss 5.7021 | text 2.1011 audio 5.2819 | grad 2.139 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5073,11 +5100,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35600 | loss 5.6976 | text 2.1119 audio 5.2752 | grad 1.692 | 1.59s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5100,11 +5127,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35625 | loss 5.7269 | text 2.1576 audio 5.2954 | grad 1.948 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.93 | nonfinite=0 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5127,10 +5155,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35650 | loss 5.7038 | text 2.0968 audio 5.2844 | grad 2.009 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5154,10 +5182,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35675 | loss 5.6964 | text 2.1176 audio 5.2729 | grad 1.781 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5181,10 +5209,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35700 | loss 5.7023 | text 2.1459 audio 5.2731 | grad 1.470 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5208,10 +5236,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35725 | loss 5.6936 | text 2.1079 audio 5.2720 | grad 1.577 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5235,14 +5263,23 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 - running validation at step 34750... - val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% - [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35750 | loss 5.7018 | text 2.1473 audio 5.2724 | grad 3.195 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.64 | nonfinite=0 + running validation at step 35750... + val/composite=2.9894 (text=0.5988 audio=4.5832) val/loss=4.7030 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_q4bfdk_b/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b84eb-23e8713c6407ba4a163897ad;da59c5d6-9f19-4363-b138-8063d8495773) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5266,10 +5303,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35775 | loss 5.6928 | text 2.1394 audio 5.2649 | grad 2.077 | 21.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5293,10 +5330,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35800 | loss 5.7009 | text 2.1499 audio 5.2709 | grad 3.615 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5320,10 +5357,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35825 | loss 5.7072 | text 2.1152 audio 5.2842 | grad 1.769 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5347,10 +5384,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35850 | loss 5.7080 | text 2.1325 audio 5.2815 | grad 1.590 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5374,10 +5411,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35875 | loss 5.7191 | text 2.1272 audio 5.2937 | grad 2.377 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5401,10 +5438,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35900 | loss 5.6961 | text 2.0964 audio 5.2768 | grad 1.901 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5428,10 +5465,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35925 | loss 5.6746 | text 2.1152 audio 5.2516 | grad 1.636 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5455,10 +5492,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35950 | loss 5.6555 | text 2.1133 audio 5.2328 | grad 1.572 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5482,10 +5519,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35975 | loss 5.6890 | text 2.1208 audio 5.2648 | grad 2.115 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5509,51 +5546,17 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 - [audio-demo] ar_cb0_acc=12.0% (35s) - running validation at step 35000... - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB - pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 - val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% - [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 36000 | loss 5.7153 | text 2.1097 audio 5.2934 | grad 2.235 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.10 | nonfinite=0 + running validation at step 36000... + val/composite=2.9897 (text=0.5974 audio=4.5846) val/loss=4.7041 cb0_acc=39.2% text_acc=91.8% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9894) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 +[ckpt] uploading /tmp/ckpt_skck_i9h/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5564,22 +5567,22 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7c65-1c38aecc5ca4acdb5008e0b4;48e1227e-d30a-4783-b035-8b02985e5e1d) +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8865-1ec56561651c1a56354c727e;d2e4e91e-8960-46ef-b922-28020e2e7473) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5588,8 +5591,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35025 | loss 5.7000 | text 2.1331 audio 5.2734 | grad 2.588 | 41.18s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.18 | nonfinite=0 +step 36025 | loss 5.6831 | text 2.1034 audio 5.2624 | grad 1.684 | 20.86s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.06 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5615,8 +5618,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35050 | loss 5.7063 | text 2.1069 audio 5.2850 | grad 2.479 | 1.48s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.11 | nonfinite=0 +step 36050 | loss 5.7093 | text 2.0998 audio 5.2893 | grad 1.919 | 1.59s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5642,9 +5645,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 3.0x (grad 6.76 vs EMA 2.25) step 35075 -step 35075 | loss 5.7092 | text 2.1201 audio 5.2851 | grad 6.761 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.11 lora=0.98 model_audio_embed=0.06 projection=0.02 text_embed=0.12 | spike L=+0.00 G=3.00 | nonfinite=0 +step 36075 | loss 5.6986 | text 2.1149 audio 5.2756 | grad 2.041 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5670,8 +5672,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35100 | loss 5.7286 | text 2.1354 audio 5.3015 | grad 2.020 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.75 | nonfinite=0 +step 36100 | loss 5.7108 | text 2.1152 audio 5.2878 | grad 1.744 | 1.50s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5697,8 +5699,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35125 | loss 5.6800 | text 2.1305 audio 5.2539 | grad 1.847 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.70 | nonfinite=0 +step 36125 | loss 5.6697 | text 2.1321 audio 5.2433 | grad 2.215 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.01 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5724,8 +5726,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35150 | loss 5.7182 | text 2.1559 audio 5.2870 | grad 1.641 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.64 | nonfinite=0 +step 36150 | loss 5.6836 | text 2.1104 audio 5.2616 | grad 1.771 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5751,8 +5753,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35175 | loss 5.6631 | text 2.1155 audio 5.2400 | grad 2.109 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.86 | nonfinite=0 +step 36175 | loss 5.7177 | text 2.1067 audio 5.2964 | grad 1.322 | 1.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5778,8 +5780,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35200 | loss 5.7188 | text 2.1798 audio 5.2828 | grad 2.145 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.88 | nonfinite=0 +step 36200 | loss 5.6609 | text 2.1179 audio 5.2373 | grad 2.159 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.01 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5805,8 +5807,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35225 | loss 5.7170 | text 2.1077 audio 5.2955 | grad 1.866 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.78 | nonfinite=0 +step 36225 | loss 5.7207 | text 2.1407 audio 5.2925 | grad 2.000 | 1.50s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5832,12 +5834,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35250 | loss 5.6647 | text 2.0979 audio 5.2451 | grad 1.703 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.73 | nonfinite=0 - running validation at step 35250... - val/composite=2.9985 (text=0.6053 audio=4.5939) val/loss=4.7150 cb0_acc=39.1% text_acc=91.6% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9973) +step 36250 | loss 5.6867 | text 2.0964 audio 5.2674 | grad 2.338 | 1.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.20 | nonfinite=0 + running validation at step 36250... + val/composite=2.9903 (text=0.5982 audio=4.5851) val/loss=4.7047 cb0_acc=39.2% text_acc=91.8% + [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9894) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5863,8 +5865,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35275 | loss 5.7344 | text 2.1597 audio 5.3025 | grad 4.716 | 2.47s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=+0.01 G=2.07 | nonfinite=0 +step 36275 | loss 5.7501 | text 2.1488 audio 5.3204 | grad 1.751 | 2.30s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5890,8 +5892,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35300 | loss 5.6869 | text 2.1172 audio 5.2634 | grad 1.765 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.70 | nonfinite=0 +step 36300 | loss 5.6584 | text 2.0785 audio 5.2427 | grad 1.793 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5917,8 +5919,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35325 | loss 5.7376 | text 2.1312 audio 5.3113 | grad 1.650 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 +step 36325 | loss 5.7092 | text 2.1386 audio 5.2815 | grad 1.547 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5944,8 +5946,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35350 | loss 5.6878 | text 2.1130 audio 5.2652 | grad 1.950 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 +step 36350 | loss 5.7118 | text 2.1361 audio 5.2846 | grad 2.371 | 1.60s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.24 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5971,8 +5973,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35375 | loss 5.6788 | text 2.1222 audio 5.2544 | grad 1.724 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 +step 36375 | loss 5.7245 | text 2.1202 audio 5.3005 | grad 1.915 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5998,8 +6000,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35400 | loss 5.6788 | text 2.1020 audio 5.2584 | grad 1.961 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 +step 36400 | loss 5.7019 | text 2.1270 audio 5.2765 | grad 2.564 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.15 | spike L=+0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6025,8 +6027,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35425 | loss 5.6969 | text 2.0995 audio 5.2770 | grad 1.766 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36425 | loss 5.7086 | text 2.1082 audio 5.2870 | grad 1.945 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6052,8 +6054,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35450 | loss 5.6848 | text 2.0913 audio 5.2665 | grad 1.521 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.69 | nonfinite=0 +step 36450 | loss 5.6757 | text 2.1256 audio 5.2505 | grad 1.849 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6079,8 +6081,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35475 | loss 5.7112 | text 2.1194 audio 5.2873 | grad 2.124 | 1.60s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 +step 36475 | loss 5.6761 | text 2.1255 audio 5.2510 | grad 1.958 | 1.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6106,16 +6108,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35500 | loss 5.7002 | text 2.0979 audio 5.2806 | grad 2.371 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 - running validation at step 35500... - val/composite=2.9923 (text=0.5962 audio=4.5896) val/loss=4.7089 cb0_acc=39.1% text_acc=91.8% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% +step 36500 | loss 5.6908 | text 2.0905 audio 5.2727 | grad 1.988 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=1.00 | nonfinite=0 + running validation at step 36500... + val/composite=2.9857 (text=0.5922 audio=4.5813) val/loss=4.6998 cb0_acc=39.4% text_acc=91.8% + [val] per-codebook acc: cb0=39.4% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_3crsay5y/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_pnu0pz72/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b816c-2712d01710663cc45f59dc84;4de1714a-3953-4c64-bba2-5dfc7dfb76bb) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8d52-3c934e2072b2c29c7b9ed4a4;c8972a1a-b255-4d8a-a62a-f52dd68711e3) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -6146,8 +6148,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35525 | loss 5.7157 | text 2.1252 audio 5.2907 | grad 1.712 | 21.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.80 | nonfinite=0 +step 36525 | loss 5.6839 | text 2.1440 audio 5.2551 | grad 2.686 | 21.43s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6173,8 +6175,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35550 | loss 5.7160 | text 2.1306 audio 5.2898 | grad 2.455 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 +step 36550 | loss 5.6939 | text 2.1047 audio 5.2730 | grad 2.494 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.21 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6200,8 +6202,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35575 | loss 5.7021 | text 2.1011 audio 5.2819 | grad 2.139 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 +step 36575 | loss 5.6677 | text 2.0906 audio 5.2496 | grad 1.994 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6227,8 +6229,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35600 | loss 5.6976 | text 2.1119 audio 5.2752 | grad 1.692 | 1.59s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36600 | loss 5.7210 | text 2.1266 audio 5.2957 | grad 2.306 | 1.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.01 G=1.10 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6254,8 +6256,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35625 | loss 5.7269 | text 2.1576 audio 5.2954 | grad 1.948 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.93 | nonfinite=0 +step 36625 | loss 5.7019 | text 2.1429 audio 5.2733 | grad 1.521 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6281,8 +6283,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35650 | loss 5.7038 | text 2.0968 audio 5.2844 | grad 2.009 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.97 | nonfinite=0 +step 36650 | loss 5.7157 | text 2.1187 audio 5.2920 | grad 1.489 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.46 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6308,8 +6310,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35675 | loss 5.6964 | text 2.1176 audio 5.2729 | grad 1.781 | 1.56s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.86 | nonfinite=0 +step 36675 | loss 5.6895 | text 2.1200 audio 5.2655 | grad 1.530 | 1.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.14 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6335,8 +6337,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35700 | loss 5.7023 | text 2.1459 audio 5.2731 | grad 1.470 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.72 | nonfinite=0 +step 36700 | loss 5.6871 | text 2.1240 audio 5.2623 | grad 1.735 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6362,8 +6364,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35725 | loss 5.6936 | text 2.1079 audio 5.2720 | grad 1.577 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36725 | loss 5.6945 | text 2.1135 audio 5.2718 | grad 2.142 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6389,16 +6391,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35750 | loss 5.7018 | text 2.1473 audio 5.2724 | grad 3.195 | 1.56s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.64 | nonfinite=0 - running validation at step 35750... - val/composite=2.9894 (text=0.5988 audio=4.5832) val/loss=4.7030 cb0_acc=39.1% text_acc=91.8% - [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +step 36750 | loss 5.6832 | text 2.1276 audio 5.2577 | grad 1.583 | 1.55s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 + running validation at step 36750... + val/composite=2.9854 (text=0.5932 audio=4.5801) val/loss=4.6988 cb0_acc=39.2% text_acc=92.0% + [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.8% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_q4bfdk_b/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_jwbvue5g/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b84eb-23e8713c6407ba4a163897ad;da59c5d6-9f19-4363-b138-8063d8495773) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b90c5-6a65d78d5cb61e8b6beabdcb;44000a67-8fdf-4e95-91a1-63030a9bba23) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -6429,8 +6431,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35775 | loss 5.6928 | text 2.1394 audio 5.2649 | grad 2.077 | 21.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.00 | nonfinite=0 +step 36775 | loss 5.7233 | text 2.1374 audio 5.2958 | grad 2.439 | 21.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6456,8 +6458,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35800 | loss 5.7009 | text 2.1499 audio 5.2709 | grad 3.615 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.75 | nonfinite=0 +step 36800 | loss 5.7377 | text 2.1306 audio 5.3116 | grad 2.060 | 1.45s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.01 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6483,8 +6485,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35825 | loss 5.7072 | text 2.1152 audio 5.2842 | grad 1.769 | 1.54s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.79 | nonfinite=0 +step 36825 | loss 5.7254 | text 2.1173 audio 5.3019 | grad 2.148 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6510,8 +6512,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35850 | loss 5.7080 | text 2.1325 audio 5.2815 | grad 1.590 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.73 | nonfinite=0 +step 36850 | loss 5.6654 | text 2.0803 audio 5.2493 | grad 1.928 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.01 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6537,8 +6539,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35875 | loss 5.7191 | text 2.1272 audio 5.2937 | grad 2.377 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 +step 36875 | loss 5.7116 | text 2.1150 audio 5.2885 | grad 1.605 | 1.45s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6564,8 +6566,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35900 | loss 5.6961 | text 2.0964 audio 5.2768 | grad 1.901 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 +step 36900 | loss 5.6732 | text 2.0938 audio 5.2545 | grad 1.952 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6591,8 +6593,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35925 | loss 5.6746 | text 2.1152 audio 5.2516 | grad 1.636 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 +step 36925 | loss 5.6833 | text 2.0830 audio 5.2667 | grad 2.282 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6618,8 +6620,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35950 | loss 5.6555 | text 2.1133 audio 5.2328 | grad 1.572 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.76 | nonfinite=0 +step 36950 | loss 5.6657 | text 2.0712 audio 5.2514 | grad 2.144 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.01 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6645,8 +6647,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35975 | loss 5.6890 | text 2.1208 audio 5.2648 | grad 2.115 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.05 | nonfinite=0 +step 36975 | loss 5.6745 | text 2.0822 audio 5.2581 | grad 2.598 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.30 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6672,15 +6674,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36000 | loss 5.7153 | text 2.1097 audio 5.2934 | grad 2.235 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.10 | nonfinite=0 - running validation at step 36000... - val/composite=2.9897 (text=0.5974 audio=4.5846) val/loss=4.7041 cb0_acc=39.2% text_acc=91.8% - [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9894) +step 37000 | loss 5.6766 | text 2.0915 audio 5.2583 | grad 2.042 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 + running validation at step 37000... + val/composite=2.9927 (text=0.5950 audio=4.5912) val/loss=4.7101 cb0_acc=39.3% text_acc=91.9% + [val] per-codebook acc: cb0=39.3% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9854) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_skck_i9h/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +[ckpt] uploading /tmp/ckpt_w5h8lisx/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6693,13 +6695,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8865-1ec56561651c1a56354c727e;d2e4e91e-8960-46ef-b922-28020e2e7473) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b942b-46693a050c6a3b041fd98fa7;77e27da6-0027-4237-bdef-9d4d0a682e02) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -6715,10 +6715,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36025 | loss 5.6831 | text 2.1034 audio 5.2624 | grad 1.684 | 20.86s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.06 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37025 | loss 5.6793 | text 2.0733 audio 5.2647 | grad 1.558 | 20.62s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6742,10 +6742,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36050 | loss 5.7093 | text 2.0998 audio 5.2893 | grad 1.919 | 1.59s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37050 | loss 5.6463 | text 2.0525 audio 5.2358 | grad 1.873 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.22 | spike L=-0.01 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6769,10 +6769,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36075 | loss 5.6986 | text 2.1149 audio 5.2756 | grad 2.041 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37075 | loss 5.6827 | text 2.0917 audio 5.2643 | grad 1.554 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6796,10 +6796,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36100 | loss 5.7108 | text 2.1152 audio 5.2878 | grad 1.744 | 1.50s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37100 | loss 5.6575 | text 2.0491 audio 5.2476 | grad 1.608 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6823,10 +6823,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36125 | loss 5.6697 | text 2.1321 audio 5.2433 | grad 2.215 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.01 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37125 | loss 5.6627 | text 2.0852 audio 5.2457 | grad 1.778 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6850,10 +6850,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36150 | loss 5.6836 | text 2.1104 audio 5.2616 | grad 1.771 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37150 | loss 5.7070 | text 2.0835 audio 5.2903 | grad 1.519 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6877,10 +6877,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36175 | loss 5.7177 | text 2.1067 audio 5.2964 | grad 1.322 | 1.47s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37175 | loss 5.6957 | text 2.1050 audio 5.2747 | grad 4.542 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.10 | spike L=+0.00 G=2.44 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6904,10 +6904,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36200 | loss 5.6609 | text 2.1179 audio 5.2373 | grad 2.159 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.01 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37200 | loss 5.6781 | text 2.0752 audio 5.2630 | grad 1.735 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6931,10 +6931,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36225 | loss 5.7207 | text 2.1407 audio 5.2925 | grad 2.000 | 1.50s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37225 | loss 5.7024 | text 2.0994 audio 5.2825 | grad 1.957 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6958,14 +6958,23 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36250 | loss 5.6867 | text 2.0964 audio 5.2674 | grad 2.338 | 1.47s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.20 | nonfinite=0 - running validation at step 36250... - val/composite=2.9903 (text=0.5982 audio=4.5851) val/loss=4.7047 cb0_acc=39.2% text_acc=91.8% - [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9894) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37250 | loss 5.6916 | text 2.0812 audio 5.2753 | grad 2.125 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.02 | nonfinite=0 + running validation at step 37250... + val/composite=2.9837 (text=0.5807 audio=4.5857) val/loss=4.7018 cb0_acc=39.3% text_acc=92.1% + [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_7jqz98hv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b977d-18e768341c50c8c634b94256;1deef6d0-e9ee-4962-adfc-011792a7ff9d) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6989,10 +6998,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36275 | loss 5.7501 | text 2.1488 audio 5.3204 | grad 1.751 | 2.30s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37275 | loss 5.6859 | text 2.0602 audio 5.2739 | grad 3.134 | 21.36s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.50 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7016,10 +7025,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36300 | loss 5.6584 | text 2.0785 audio 5.2427 | grad 1.793 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37300 | loss 5.6772 | text 2.0848 audio 5.2602 | grad 2.241 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7043,10 +7052,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36325 | loss 5.7092 | text 2.1386 audio 5.2815 | grad 1.547 | 1.52s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37325 | loss 5.6471 | text 2.0604 audio 5.2350 | grad 1.747 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.01 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7070,10 +7079,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36350 | loss 5.7118 | text 2.1361 audio 5.2846 | grad 2.371 | 1.60s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.24 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37350 | loss 5.6828 | text 2.0960 audio 5.2636 | grad 1.923 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7097,10 +7106,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36375 | loss 5.7245 | text 2.1202 audio 5.3005 | grad 1.915 | 1.54s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37375 | loss 5.6923 | text 2.0948 audio 5.2733 | grad 2.252 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7124,10 +7133,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36400 | loss 5.7019 | text 2.1270 audio 5.2765 | grad 2.564 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.15 | spike L=+0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37400 | loss 5.6938 | text 2.0868 audio 5.2765 | grad 1.727 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7151,10 +7160,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36425 | loss 5.7086 | text 2.1082 audio 5.2870 | grad 1.945 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37425 | loss 5.6795 | text 2.0630 audio 5.2669 | grad 2.381 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7178,10 +7187,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36450 | loss 5.6757 | text 2.1256 audio 5.2505 | grad 1.849 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37450 | loss 5.6937 | text 2.0768 audio 5.2784 | grad 1.354 | 1.53s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7205,10 +7214,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36475 | loss 5.6761 | text 2.1255 audio 5.2510 | grad 1.958 | 1.48s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37475 | loss 5.6899 | text 2.0712 audio 5.2757 | grad 3.236 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7232,16 +7241,18 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36500 | loss 5.6908 | text 2.0905 audio 5.2727 | grad 1.988 | 1.54s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=1.00 | nonfinite=0 - running validation at step 36500... - val/composite=2.9857 (text=0.5922 audio=4.5813) val/loss=4.6998 cb0_acc=39.4% text_acc=91.8% - [val] per-codebook acc: cb0=39.4% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 37500 | loss 5.6730 | text 2.0959 audio 5.2538 | grad 1.822 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=-0.00 G=0.84 | nonfinite=0 + running validation at step 37500... + val/composite=2.9835 (text=0.5934 audio=4.5769) val/loss=4.6956 cb0_acc=39.4% text_acc=92.0% + [val] per-codebook acc: cb0=39.4% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_pnu0pz72/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_4vbm_v2l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8d52-3c934e2072b2c29c7b9ed4a4;c8972a1a-b255-4d8a-a62a-f52dd68711e3) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b9ae5-6c9b90af7b04c5cd3db15ea0;96a0f28b-4ee4-4786-b616-80bc413a96be) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -7272,8 +7283,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36525 | loss 5.6839 | text 2.1440 audio 5.2551 | grad 2.686 | 21.43s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.35 | nonfinite=0 +step 37525 | loss 5.6611 | text 2.0692 audio 5.2472 | grad 1.530 | 21.54s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7299,8 +7310,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36550 | loss 5.6939 | text 2.1047 audio 5.2730 | grad 2.494 | 1.52s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.21 | nonfinite=0 +step 37550 | loss 5.6846 | text 2.0928 audio 5.2660 | grad 1.602 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7326,8 +7337,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36575 | loss 5.6677 | text 2.0906 audio 5.2496 | grad 1.994 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.95 | nonfinite=0 +step 37575 | loss 5.6758 | text 2.1009 audio 5.2556 | grad 2.979 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.16 | spike L=-0.00 G=1.47 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7353,8 +7364,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36600 | loss 5.7210 | text 2.1266 audio 5.2957 | grad 2.306 | 1.58s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.01 G=1.10 | nonfinite=0 +step 37600 | loss 5.6719 | text 2.0873 audio 5.2544 | grad 2.455 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7380,8 +7391,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36625 | loss 5.7019 | text 2.1429 audio 5.2733 | grad 1.521 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 +step 37625 | loss 5.7134 | text 2.0897 audio 5.2954 | grad 1.739 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7407,8 +7418,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36650 | loss 5.7157 | text 2.1187 audio 5.2920 | grad 1.489 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.46 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.73 | nonfinite=0 +step 37650 | loss 5.6801 | text 2.0723 audio 5.2656 | grad 2.018 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7434,8 +7445,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36675 | loss 5.6895 | text 2.1200 audio 5.2655 | grad 1.530 | 1.58s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.14 | spike L=-0.00 G=0.77 | nonfinite=0 +step 37675 | loss 5.6702 | text 2.0706 audio 5.2561 | grad 1.751 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7461,8 +7472,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36700 | loss 5.6871 | text 2.1240 audio 5.2623 | grad 1.735 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 +step 37700 | loss 5.6739 | text 2.0701 audio 5.2598 | grad 2.293 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7488,8 +7499,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36725 | loss 5.6945 | text 2.1135 audio 5.2718 | grad 2.142 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 +step 37725 | loss 5.6521 | text 2.0569 audio 5.2408 | grad 1.874 | 1.53s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.88 model_audio_embed=0.06 projection=0.09 text_embed=0.16 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7515,16 +7526,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36750 | loss 5.6832 | text 2.1276 audio 5.2577 | grad 1.583 | 1.55s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 - running validation at step 36750... - val/composite=2.9854 (text=0.5932 audio=4.5801) val/loss=4.6988 cb0_acc=39.2% text_acc=92.0% - [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.8% +step 37750 | loss 5.6452 | text 2.0799 audio 5.2292 | grad 3.034 | 1.50s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.47 | nonfinite=0 + running validation at step 37750... + val/composite=2.9778 (text=0.5780 audio=4.5777) val/loss=4.6933 cb0_acc=39.3% text_acc=92.1% + [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.8% cb6=5.9% cb7=5.8% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_jwbvue5g/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_inpnozqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b90c5-6a65d78d5cb61e8b6beabdcb;44000a67-8fdf-4e95-91a1-63030a9bba23) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b9e54-4bab431c0d0325ff4098a2ea;192d1300-4784-4a97-88ec-81abd1901e75) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -7555,8 +7566,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36775 | loss 5.7233 | text 2.1374 audio 5.2958 | grad 2.439 | 21.48s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.28 | nonfinite=0 +step 37775 | loss 5.6960 | text 2.1047 audio 5.2751 | grad 2.769 | 21.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7582,8 +7593,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36800 | loss 5.7377 | text 2.1306 audio 5.3116 | grad 2.060 | 1.45s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.01 G=1.05 | nonfinite=0 +step 37800 | loss 5.6842 | text 2.0801 audio 5.2682 | grad 1.344 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=+0.00 G=0.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7609,8 +7620,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36825 | loss 5.7254 | text 2.1173 audio 5.3019 | grad 2.148 | 1.52s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.09 | nonfinite=0 +step 37825 | loss 5.6790 | text 2.1186 audio 5.2553 | grad 2.178 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7636,8 +7647,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36850 | loss 5.6654 | text 2.0803 audio 5.2493 | grad 1.928 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.01 G=0.97 | nonfinite=0 +step 37850 | loss 5.6950 | text 2.1125 audio 5.2725 | grad 2.040 | 1.53s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7663,8 +7674,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36875 | loss 5.7116 | text 2.1150 audio 5.2885 | grad 1.605 | 1.45s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 +step 37875 | loss 5.6836 | text 2.0703 audio 5.2695 | grad 3.089 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.45 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7690,8 +7701,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36900 | loss 5.6732 | text 2.0938 audio 5.2545 | grad 1.952 | 1.52s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=1.00 | nonfinite=0 +step 37900 | loss 5.6898 | text 2.0709 audio 5.2756 | grad 1.446 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.48 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7717,8 +7728,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36925 | loss 5.6833 | text 2.0830 audio 5.2667 | grad 2.282 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.17 | nonfinite=0 +step 37925 | loss 5.6369 | text 2.0557 audio 5.2258 | grad 2.639 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7744,8 +7755,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36950 | loss 5.6657 | text 2.0712 audio 5.2514 | grad 2.144 | 1.46s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.01 G=1.08 | nonfinite=0 +step 37950 | loss 5.6719 | text 2.0937 audio 5.2532 | grad 4.150 | 1.45s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=-0.00 G=1.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7771,8 +7782,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36975 | loss 5.6745 | text 2.0822 audio 5.2581 | grad 2.598 | 1.49s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.30 | nonfinite=0 +step 37975 | loss 5.6779 | text 2.0709 audio 5.2637 | grad 1.744 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7798,13 +7809,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 37000 | loss 5.6766 | text 2.0915 audio 5.2583 | grad 2.042 | 1.47s/step | peak -1.0G | host_rss 47.0G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 - running validation at step 37000... - val/composite=2.9927 (text=0.5950 audio=4.5912) val/loss=4.7101 cb0_acc=39.3% text_acc=91.9% - [val] per-codebook acc: cb0=39.3% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9854) +step 38000 | loss 5.6804 | text 2.0998 audio 5.2604 | grad 2.820 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.21 | nonfinite=0 + running validation at step 38000... + val/composite=2.9791 (text=0.5835 audio=4.5762) val/loss=4.6928 cb0_acc=39.3% text_acc=92.0% + [val] per-codebook acc: cb0=39.3% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9778) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_w5h8lisx/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 +[ckpt] uploading /tmp/ckpt_t6hfiptu/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_038000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3