diff --git "a/logs/train_host0_latest.log" "b/logs/train_host0_latest.log" --- "a/logs/train_host0_latest.log" +++ "b/logs/train_host0_latest.log" @@ -1,4 +1,4 @@ -s are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -10,8 +10,6 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29100 | loss 5.7216 | text 2.1422 audio 5.2931 | grad 2.007 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -20,6 +18,14 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30100 | loss 5.7109 | text 2.1508 audio 5.2807 | grad 2.789 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.28 | nonfinite=0 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -37,8 +43,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29125 | loss 5.7354 | text 2.1949 audio 5.2965 | grad 2.488 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.14 | spike L=-0.00 G=0.94 | nonfinite=0 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 30125 | loss 5.7309 | text 2.1609 audio 5.2987 | grad 2.277 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -64,8 +72,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29150 | loss 5.7076 | text 2.1771 audio 5.2722 | grad 2.926 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.01 G=1.12 | nonfinite=0 +step 30150 | loss 5.7423 | text 2.1902 audio 5.3042 | grad 1.819 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -91,8 +99,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29175 | loss 5.7156 | text 2.1626 audio 5.2831 | grad 3.717 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.00 G=1.40 | nonfinite=0 +step 30175 | loss 5.7186 | text 2.1692 audio 5.2848 | grad 2.105 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -118,8 +126,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29200 | loss 5.6874 | text 2.1905 audio 5.2493 | grad 2.453 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.89 | nonfinite=0 +step 30200 | loss 5.7341 | text 2.1761 audio 5.2988 | grad 1.584 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -145,8 +153,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29225 | loss 5.7260 | text 2.1649 audio 5.2930 | grad 2.932 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.15 | spike L=-0.00 G=1.07 | nonfinite=0 + [ALERT] grad spike 4.1x (grad 8.69 vs EMA 2.13) step 30225 +step 30225 | loss 5.7041 | text 2.1497 audio 5.2742 | grad 8.692 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.07 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.06 | spike L=-0.00 G=4.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -172,12 +181,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29250 | loss 5.7181 | text 2.1439 audio 5.2893 | grad 2.164 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.79 | nonfinite=0 - running validation at step 29250... - val/composite=3.0334 (text=0.6558 audio=4.6184) val/loss=4.7496 cb0_acc=38.7% text_acc=90.3% - [val] per-codebook acc: cb0=38.7% cb1=19.9% cb2=16.8% cb3=10.8% cb4=8.1% cb5=6.7% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0325) +step 30250 | loss 5.7108 | text 2.1812 audio 5.2746 | grad 1.567 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.56 | nonfinite=0 + running validation at step 30250... + val/composite=3.0263 (text=0.6452 audio=4.6137) val/loss=4.7427 cb0_acc=38.8% text_acc=90.7% + [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -203,8 +212,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29275 | loss 5.7088 | text 2.1322 audio 5.2823 | grad 3.921 | 2.35s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.09 | spike L=-0.00 G=1.46 | nonfinite=0 +step 30275 | loss 5.7370 | text 2.1381 audio 5.3094 | grad 1.647 | 2.35s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -230,8 +239,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29300 | loss 5.7549 | text 2.1839 audio 5.3182 | grad 2.127 | 1.59s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.01 G=0.76 | nonfinite=0 +step 30300 | loss 5.7740 | text 2.1916 audio 5.3357 | grad 3.100 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=+0.01 G=1.21 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -257,8 +266,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29325 | loss 5.6933 | text 2.1503 audio 5.2633 | grad 1.508 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.55 | nonfinite=0 +step 30325 | loss 5.6988 | text 2.1558 audio 5.2676 | grad 2.002 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -284,8 +293,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29350 | loss 5.7047 | text 2.1504 audio 5.2747 | grad 2.847 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.15 | spike L=-0.00 G=1.09 | nonfinite=0 +step 30350 | loss 5.7251 | text 2.1526 audio 5.2946 | grad 2.659 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.08 | spike L=-0.00 G=1.04 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -311,8 +320,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29375 | loss 5.7100 | text 2.1591 audio 5.2782 | grad 2.673 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.01 | nonfinite=0 +step 30375 | loss 5.7275 | text 2.1723 audio 5.2930 | grad 1.326 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=-0.00 G=0.52 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -338,8 +347,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29400 | loss 5.7273 | text 2.1360 audio 5.3002 | grad 1.735 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.66 | nonfinite=0 +step 30400 | loss 5.7535 | text 2.1742 audio 5.3186 | grad 1.857 | 1.44s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -365,8 +374,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29425 | loss 5.7523 | text 2.1795 audio 5.3164 | grad 1.612 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=+0.01 G=0.63 | nonfinite=0 +step 30425 | loss 5.6897 | text 2.1361 audio 5.2625 | grad 1.655 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -392,8 +401,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29450 | loss 5.7343 | text 2.1852 audio 5.2972 | grad 2.594 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.05 | nonfinite=0 +step 30450 | loss 5.7256 | text 2.1539 audio 5.2948 | grad 3.586 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.55 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -419,8 +428,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29475 | loss 5.7163 | text 2.1439 audio 5.2875 | grad 2.112 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.85 | nonfinite=0 +step 30475 | loss 5.7147 | text 2.1494 audio 5.2848 | grad 2.515 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -446,12 +455,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29500 | loss 5.7033 | text 2.1426 audio 5.2748 | grad 4.748 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.14 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=-0.00 G=1.95 | nonfinite=0 - running validation at step 29500... - val/composite=3.0330 (text=0.6584 audio=4.6160) val/loss=4.7477 cb0_acc=38.7% text_acc=90.4% - [val] per-codebook acc: cb0=38.7% cb1=20.0% cb2=16.9% cb3=10.9% cb4=8.2% cb5=6.8% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0325) +step 30500 | loss 5.7465 | text 2.1749 audio 5.3115 | grad 1.691 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.69 | nonfinite=0 + running validation at step 30500... + val/composite=3.0241 (text=0.6463 audio=4.6093) val/loss=4.7386 cb0_acc=38.8% text_acc=90.6% + [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -477,8 +486,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29525 | loss 5.7381 | text 2.1626 audio 5.3055 | grad 2.429 | 2.33s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.91 | nonfinite=0 +step 30525 | loss 5.7592 | text 2.1927 audio 5.3207 | grad 1.734 | 2.37s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -504,8 +513,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29550 | loss 5.6856 | text 2.1265 audio 5.2603 | grad 2.026 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 +step 30550 | loss 5.7289 | text 2.1768 audio 5.2936 | grad 1.688 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -531,8 +540,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29575 | loss 5.7211 | text 2.1651 audio 5.2881 | grad 2.022 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 +step 30575 | loss 5.7234 | text 2.1289 audio 5.2976 | grad 2.456 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -558,8 +567,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29600 | loss 5.7355 | text 2.1309 audio 5.3093 | grad 2.207 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=+0.00 G=0.87 | nonfinite=0 +step 30600 | loss 5.6997 | text 2.1438 audio 5.2709 | grad 1.741 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -585,8 +594,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29625 | loss 5.7147 | text 2.1485 audio 5.2850 | grad 1.976 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 +step 30625 | loss 5.7534 | text 2.1707 audio 5.3192 | grad 1.956 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -612,8 +621,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29650 | loss 5.7366 | text 2.1921 audio 5.2981 | grad 1.785 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.73 | nonfinite=0 +step 30650 | loss 5.7253 | text 2.1289 audio 5.2995 | grad 2.165 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -639,8 +648,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29675 | loss 5.7341 | text 2.1653 audio 5.3010 | grad 1.758 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.74 | nonfinite=0 +step 30675 | loss 5.7269 | text 2.1563 audio 5.2956 | grad 3.860 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -666,8 +675,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29700 | loss 5.7398 | text 2.1773 audio 5.3043 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 +step 30700 | loss 5.7398 | text 2.1505 audio 5.3097 | grad 2.235 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -693,8 +702,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29725 | loss 5.7252 | text 2.1935 audio 5.2865 | grad 1.577 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.70 | nonfinite=0 +step 30725 | loss 5.7489 | text 2.1535 audio 5.3182 | grad 2.031 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -720,16 +729,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29750 | loss 5.7185 | text 2.1669 audio 5.2851 | grad 1.776 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 - running validation at step 29750... - val/composite=3.0292 (text=0.6607 audio=4.6083) val/loss=4.7404 cb0_acc=38.8% text_acc=90.4% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=10.9% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +step 30750 | loss 5.6953 | text 2.1131 audio 5.2727 | grad 3.574 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.17 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.55 | nonfinite=0 + running validation at step 30750... + val/composite=3.0163 (text=0.6371 audio=4.6024) val/loss=4.7299 cb0_acc=38.8% text_acc=90.9% + [val] per-codebook acc: cb0=38.8% cb1=19.9% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_it1ltfi7/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_v28gzt94/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b3dc6-4c6482853ef24b3136885705;752b698a-0123-4e78-8974-6c9f1cb2aedf) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b49a8-49c8b3a512d01dcf62ba2de6;c1d11542-5c8f-4019-8577-05bf7d3f521f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -760,8 +769,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29775 | loss 5.7285 | text 2.1674 audio 5.2950 | grad 2.373 | 21.34s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.10 | nonfinite=0 +step 30775 | loss 5.7336 | text 2.1703 audio 5.2995 | grad 2.588 | 21.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -787,8 +796,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29800 | loss 5.7073 | text 2.1300 audio 5.2813 | grad 2.486 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.14 | nonfinite=0 +step 30800 | loss 5.7168 | text 2.1540 audio 5.2860 | grad 2.014 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -814,8 +823,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29825 | loss 5.7351 | text 2.1504 audio 5.3050 | grad 2.191 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.99 | nonfinite=0 +step 30825 | loss 5.7184 | text 2.1635 audio 5.2857 | grad 1.822 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -841,8 +850,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29850 | loss 5.7432 | text 2.1591 audio 5.3114 | grad 2.085 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.94 | nonfinite=0 +step 30850 | loss 5.7180 | text 2.1629 audio 5.2854 | grad 1.949 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -868,8 +877,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29875 | loss 5.7321 | text 2.1600 audio 5.3001 | grad 2.919 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=+0.00 G=1.33 | nonfinite=0 +step 30875 | loss 5.7182 | text 2.1146 audio 5.2953 | grad 1.590 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -895,8 +904,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29900 | loss 5.7085 | text 2.1478 audio 5.2790 | grad 2.019 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 +step 30900 | loss 5.6918 | text 2.1653 audio 5.2588 | grad 3.649 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.63 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -922,8 +931,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29925 | loss 5.7406 | text 2.1655 audio 5.3075 | grad 2.963 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.32 | nonfinite=0 +step 30925 | loss 5.7255 | text 2.1531 audio 5.2949 | grad 1.466 | 1.43s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -949,8 +958,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29950 | loss 5.7224 | text 2.1685 audio 5.2887 | grad 2.859 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.14 | spike L=-0.00 G=1.23 | nonfinite=0 +step 30950 | loss 5.7009 | text 2.1495 audio 5.2710 | grad 2.898 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -976,8 +985,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 29975 | loss 5.7182 | text 2.1755 audio 5.2831 | grad 1.804 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 +step 30975 | loss 5.7335 | text 2.1454 audio 5.3044 | grad 2.994 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1003,51 +1012,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30000 | loss 5.7549 | text 2.1721 audio 5.3205 | grad 2.832 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.01 G=1.22 | nonfinite=0 - [audio-demo] ar_cb0_acc=6.0% (35s) - running validation at step 30000... - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB  - - ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...es/step_030000/source.wav: 100%|██████████| 230kB / 230kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB  - - ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...step_030000/target_gt.wav: 100%|██████████| 192kB / 192kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB  - - ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s - ...step_030000/generated.wav: 100%|██████████| 192kB / 192kB - pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_030000 - val/composite=3.0203 (text=0.6417 audio=4.6060) val/loss=4.7344 cb0_acc=38.8% text_acc=90.5% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=10.9% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_vpw1dggr/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4150-7051ba0e2cb825374650c539;7f57bd5b-8fd3-43c1-b4ed-23eecab768d2) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 31000 | loss 5.7318 | text 2.1491 audio 5.3019 | grad 1.703 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.71 | nonfinite=0 + running validation at step 31000... + val/composite=3.0191 (text=0.6409 audio=4.6045) val/loss=4.7327 cb0_acc=39.0% text_acc=90.8% + [val] per-codebook acc: cb0=39.0% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0163) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_lczsg7r8/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_030000 +[ckpt] uploading /tmp/ckpt_y6mkwqap/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1059,18 +1032,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_030000 -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4332-2102d4f8046862615ec783b4;bbf2f91d-62d4-4747-a05d-8b2bb6abbaad) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4d0b-4f6ce7c1015680b77c18485b;4de181e9-748f-43df-8a2c-956408fcae48) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1082,10 +1053,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30025 | loss 5.7367 | text 2.1700 audio 5.3027 | grad 2.055 | 41.20s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31025 | loss 5.7013 | text 2.1689 audio 5.2675 | grad 2.218 | 20.63s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1109,10 +1080,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30050 | loss 5.7188 | text 2.1281 audio 5.2931 | grad 1.639 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31050 | loss 5.6785 | text 2.1305 audio 5.2524 | grad 2.106 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1136,10 +1107,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30075 | loss 5.7561 | text 2.1976 audio 5.3166 | grad 1.470 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31075 | loss 5.7067 | text 2.1458 audio 5.2776 | grad 2.192 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1163,10 +1134,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30100 | loss 5.7109 | text 2.1508 audio 5.2807 | grad 2.789 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31100 | loss 5.7458 | text 2.1828 audio 5.3092 | grad 2.135 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.06 projection=0.06 text_embed=0.14 | spike L=+0.01 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1190,10 +1161,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30125 | loss 5.7309 | text 2.1609 audio 5.2987 | grad 2.277 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31125 | loss 5.7439 | text 2.2002 audio 5.3039 | grad 2.645 | 1.46s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1217,10 +1188,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30150 | loss 5.7423 | text 2.1902 audio 5.3042 | grad 1.819 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31150 | loss 5.7190 | text 2.1559 audio 5.2878 | grad 1.900 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1244,10 +1215,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30175 | loss 5.7186 | text 2.1692 audio 5.2848 | grad 2.105 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31175 | loss 5.7093 | text 2.1537 audio 5.2786 | grad 2.073 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1271,10 +1242,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30200 | loss 5.7341 | text 2.1761 audio 5.2988 | grad 1.584 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31200 | loss 5.7151 | text 2.1503 audio 5.2850 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1298,11 +1269,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 4.1x (grad 8.69 vs EMA 2.13) step 30225 -step 30225 | loss 5.7041 | text 2.1497 audio 5.2742 | grad 8.692 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.07 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.06 | spike L=-0.00 G=4.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31225 | loss 5.7159 | text 2.1445 audio 5.2870 | grad 1.819 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1326,14 +1296,14 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30250 | loss 5.7108 | text 2.1812 audio 5.2746 | grad 1.567 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.56 | nonfinite=0 - running validation at step 30250... - val/composite=3.0263 (text=0.6452 audio=4.6137) val/loss=4.7427 cb0_acc=38.8% text_acc=90.7% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31250 | loss 5.7228 | text 2.1591 audio 5.2910 | grad 2.032 | 1.47s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 + running validation at step 31250... + val/composite=3.0194 (text=0.6437 audio=4.6031) val/loss=4.7319 cb0_acc=38.7% text_acc=90.7% + [val] per-codebook acc: cb0=38.7% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0163) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1357,10 +1327,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30275 | loss 5.7370 | text 2.1381 audio 5.3094 | grad 1.647 | 2.35s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31275 | loss 5.7330 | text 2.1806 audio 5.2969 | grad 1.797 | 2.38s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1384,10 +1354,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30300 | loss 5.7740 | text 2.1916 audio 5.3357 | grad 3.100 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=+0.01 G=1.21 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31300 | loss 5.7208 | text 2.1705 audio 5.2867 | grad 2.195 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1411,10 +1381,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30325 | loss 5.6988 | text 2.1558 audio 5.2676 | grad 2.002 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31325 | loss 5.7276 | text 2.1491 audio 5.2978 | grad 2.566 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.18 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1438,10 +1408,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30350 | loss 5.7251 | text 2.1526 audio 5.2946 | grad 2.659 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.08 | spike L=-0.00 G=1.04 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31350 | loss 5.7069 | text 2.1178 audio 5.2834 | grad 1.952 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1465,10 +1435,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30375 | loss 5.7275 | text 2.1723 audio 5.2930 | grad 1.326 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=-0.00 G=0.52 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31375 | loss 5.7308 | text 2.1203 audio 5.3067 | grad 1.813 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1492,10 +1462,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30400 | loss 5.7535 | text 2.1742 audio 5.3186 | grad 1.857 | 1.44s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31400 | loss 5.7433 | text 2.1584 audio 5.3116 | grad 2.645 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1519,10 +1489,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30425 | loss 5.6897 | text 2.1361 audio 5.2625 | grad 1.655 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31425 | loss 5.7413 | text 2.1406 audio 5.3132 | grad 1.860 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1546,10 +1516,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30450 | loss 5.7256 | text 2.1539 audio 5.2948 | grad 3.586 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.55 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31450 | loss 5.7094 | text 2.1542 audio 5.2786 | grad 1.536 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1573,10 +1543,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30475 | loss 5.7147 | text 2.1494 audio 5.2848 | grad 2.515 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31475 | loss 5.7086 | text 2.1526 audio 5.2781 | grad 3.518 | 1.59s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1600,14 +1570,23 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30500 | loss 5.7465 | text 2.1749 audio 5.3115 | grad 1.691 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.69 | nonfinite=0 - running validation at step 30500... - val/composite=3.0241 (text=0.6463 audio=4.6093) val/loss=4.7386 cb0_acc=38.8% text_acc=90.6% - [val] per-codebook acc: cb0=38.8% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0203) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31500 | loss 5.7234 | text 2.1505 audio 5.2933 | grad 2.082 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.93 | nonfinite=0 + running validation at step 31500... + val/composite=3.0129 (text=0.6302 audio=4.6013) val/loss=4.7273 cb0_acc=38.9% text_acc=90.9% + [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_97c5nxlf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b51f5-564c334001068ae16e82ce0f;aafdec3d-7141-4994-a73d-192be2632797) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1631,10 +1610,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30525 | loss 5.7592 | text 2.1927 audio 5.3207 | grad 1.734 | 2.37s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31525 | loss 5.7418 | text 2.1563 audio 5.3106 | grad 2.122 | 21.38s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1658,10 +1637,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30550 | loss 5.7289 | text 2.1768 audio 5.2936 | grad 1.688 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31550 | loss 5.7494 | text 2.1284 audio 5.3237 | grad 2.364 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1685,10 +1664,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30575 | loss 5.7234 | text 2.1289 audio 5.2976 | grad 2.456 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31575 | loss 5.7309 | text 2.1538 audio 5.3001 | grad 1.760 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1712,10 +1691,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30600 | loss 5.6997 | text 2.1438 audio 5.2709 | grad 1.741 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31600 | loss 5.7379 | text 2.1287 audio 5.3122 | grad 3.343 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.53 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1739,10 +1718,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30625 | loss 5.7534 | text 2.1707 audio 5.3192 | grad 1.956 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31625 | loss 5.6686 | text 2.1377 audio 5.2410 | grad 2.997 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.30 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1766,10 +1745,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30650 | loss 5.7253 | text 2.1289 audio 5.2995 | grad 2.165 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31650 | loss 5.6969 | text 2.1637 audio 5.2641 | grad 1.750 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1793,10 +1772,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30675 | loss 5.7269 | text 2.1563 audio 5.2956 | grad 3.860 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.16 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31675 | loss 5.7165 | text 2.1522 audio 5.2861 | grad 2.261 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1820,10 +1799,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30700 | loss 5.7398 | text 2.1505 audio 5.3097 | grad 2.235 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31700 | loss 5.6886 | text 2.1403 audio 5.2606 | grad 2.322 | 1.57s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.01 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1847,10 +1826,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30725 | loss 5.7489 | text 2.1535 audio 5.3182 | grad 2.031 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31725 | loss 5.7274 | text 2.1259 audio 5.3022 | grad 1.478 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1874,23 +1853,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30750 | loss 5.6953 | text 2.1131 audio 5.2727 | grad 3.574 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.17 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.55 | nonfinite=0 - running validation at step 30750... - val/composite=3.0163 (text=0.6371 audio=4.6024) val/loss=4.7299 cb0_acc=38.8% text_acc=90.9% - [val] per-codebook acc: cb0=38.8% cb1=19.9% cb2=16.8% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_v28gzt94/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b49a8-49c8b3a512d01dcf62ba2de6;c1d11542-5c8f-4019-8577-05bf7d3f521f) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31750 | loss 5.6918 | text 2.1276 audio 5.2663 | grad 2.754 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.24 | nonfinite=0 + running validation at step 31750... + val/composite=3.0149 (text=0.6321 audio=4.6034) val/loss=4.7299 cb0_acc=39.0% text_acc=90.9% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0129) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1914,9 +1885,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30775 | loss 5.7336 | text 2.1703 audio 5.2995 | grad 2.588 | 21.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.06 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31775 | loss 5.7190 | text 2.1568 audio 5.2876 | grad 1.540 | 2.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1941,9 +1912,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30800 | loss 5.7168 | text 2.1540 audio 5.2860 | grad 2.014 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31800 | loss 5.7086 | text 2.1397 audio 5.2807 | grad 1.333 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1968,9 +1939,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30825 | loss 5.7184 | text 2.1635 audio 5.2857 | grad 1.822 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31825 | loss 5.7059 | text 2.1289 audio 5.2801 | grad 1.661 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -1995,9 +1966,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30850 | loss 5.7180 | text 2.1629 audio 5.2854 | grad 1.949 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31850 | loss 5.7276 | text 2.1494 audio 5.2978 | grad 1.604 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=+0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2022,9 +1993,10 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30875 | loss 5.7182 | text 2.1146 audio 5.2953 | grad 1.590 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + [ALERT] grad spike 3.4x (grad 6.92 vs EMA 2.02) step 31875 +step 31875 | loss 5.7635 | text 2.1532 audio 5.3329 | grad 6.916 | 1.58s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.10 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.08 | spike L=+0.01 G=3.42 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2049,9 +2021,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30900 | loss 5.6918 | text 2.1653 audio 5.2588 | grad 3.649 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.63 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31900 | loss 5.7007 | text 2.1354 audio 5.2736 | grad 1.531 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.10 text_embed=0.12 | spike L=-0.00 G=0.61 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2076,9 +2048,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30925 | loss 5.7255 | text 2.1531 audio 5.2949 | grad 1.466 | 1.43s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31925 | loss 5.7284 | text 2.1511 audio 5.2981 | grad 2.452 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2103,9 +2075,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30950 | loss 5.7009 | text 2.1495 audio 5.2710 | grad 2.898 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.05 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31950 | loss 5.7026 | text 2.1339 audio 5.2759 | grad 2.128 | 1.54s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2130,9 +2102,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 30975 | loss 5.7335 | text 2.1454 audio 5.3044 | grad 2.994 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.27 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 31975 | loss 5.7325 | text 2.1323 audio 5.3060 | grad 3.383 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.00 G=1.42 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2157,15 +2129,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31000 | loss 5.7318 | text 2.1491 audio 5.3019 | grad 1.703 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.71 | nonfinite=0 - running validation at step 31000... - val/composite=3.0191 (text=0.6409 audio=4.6045) val/loss=4.7327 cb0_acc=39.0% text_acc=90.8% - [val] per-codebook acc: cb0=39.0% cb1=20.0% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0163) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 32000 | loss 5.7196 | text 2.1309 audio 5.2935 | grad 2.305 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.93 | nonfinite=0 + running validation at step 32000... + val/composite=3.0141 (text=0.6317 audio=4.6024) val/loss=4.7287 cb0_acc=38.9% text_acc=90.9% + [val] per-codebook acc: cb0=38.9% cb1=20.0% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0129) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_y6mkwqap/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 +[ckpt] uploading /tmp/ckpt_y3j7s14l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2177,11 +2150,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_031000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b4d0b-4f6ce7c1015680b77c18485b;4de181e9-748f-43df-8a2c-956408fcae48) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5701-676c84df289d499c023dd0e8;593c1b3a-6ef8-4474-8a06-3d170a535244) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -2200,8 +2173,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31025 | loss 5.7013 | text 2.1689 audio 5.2675 | grad 2.218 | 20.63s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.95 | nonfinite=0 +step 32025 | loss 5.7313 | text 2.1524 audio 5.3008 | grad 1.656 | 20.67s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2227,8 +2200,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31050 | loss 5.6785 | text 2.1305 audio 5.2524 | grad 2.106 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 +step 32050 | loss 5.7048 | text 2.1554 audio 5.2737 | grad 2.361 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2254,8 +2227,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31075 | loss 5.7067 | text 2.1458 audio 5.2776 | grad 2.192 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.95 | nonfinite=0 +step 32075 | loss 5.7039 | text 2.0979 audio 5.2843 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2281,8 +2254,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31100 | loss 5.7458 | text 2.1828 audio 5.3092 | grad 2.135 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.06 projection=0.06 text_embed=0.14 | spike L=+0.01 G=0.93 | nonfinite=0 +step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2308,8 +2281,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31125 | loss 5.7439 | text 2.2002 audio 5.3039 | grad 2.645 | 1.46s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.16 | nonfinite=0 +step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2335,8 +2308,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31150 | loss 5.7190 | text 2.1559 audio 5.2878 | grad 1.900 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.82 | nonfinite=0 +step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2362,8 +2335,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31175 | loss 5.7093 | text 2.1537 audio 5.2786 | grad 2.073 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.91 | nonfinite=0 +step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2389,8 +2362,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31200 | loss 5.7151 | text 2.1503 audio 5.2850 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.18 | nonfinite=0 +step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2416,8 +2389,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31225 | loss 5.7159 | text 2.1445 audio 5.2870 | grad 1.819 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.79 | nonfinite=0 +step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2443,12 +2416,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31250 | loss 5.7228 | text 2.1591 audio 5.2910 | grad 2.032 | 1.47s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.90 | nonfinite=0 - running validation at step 31250... - val/composite=3.0194 (text=0.6437 audio=4.6031) val/loss=4.7319 cb0_acc=38.7% text_acc=90.7% - [val] per-codebook acc: cb0=38.7% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0163) +step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 + running validation at step 32250... + val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2474,8 +2456,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31275 | loss 5.7330 | text 2.1806 audio 5.2969 | grad 1.797 | 2.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.81 | nonfinite=0 +step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2501,8 +2483,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31300 | loss 5.7208 | text 2.1705 audio 5.2867 | grad 2.195 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.01 | nonfinite=0 +step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2528,8 +2510,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31325 | loss 5.7276 | text 2.1491 audio 5.2978 | grad 2.566 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.95 model_audio_embed=0.06 projection=0.05 text_embed=0.18 | spike L=+0.00 G=1.17 | nonfinite=0 +step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2555,8 +2537,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31350 | loss 5.7069 | text 2.1178 audio 5.2834 | grad 1.952 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=-0.00 G=0.88 | nonfinite=0 +step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2582,8 +2564,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31375 | loss 5.7308 | text 2.1203 audio 5.3067 | grad 1.813 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 +step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2609,8 +2591,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31400 | loss 5.7433 | text 2.1584 audio 5.3116 | grad 2.645 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.23 | nonfinite=0 +step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2636,8 +2618,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31425 | loss 5.7413 | text 2.1406 audio 5.3132 | grad 1.860 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.84 | nonfinite=0 +step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2663,8 +2645,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31450 | loss 5.7094 | text 2.1542 audio 5.2786 | grad 1.536 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.71 | nonfinite=0 +step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2690,8 +2672,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31475 | loss 5.7086 | text 2.1526 audio 5.2781 | grad 3.518 | 1.59s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.67 | nonfinite=0 +step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2717,16 +2699,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31500 | loss 5.7234 | text 2.1505 audio 5.2933 | grad 2.082 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 31500... - val/composite=3.0129 (text=0.6302 audio=4.6013) val/loss=4.7273 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% +step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 + running validation at step 32500... + val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_97c5nxlf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b51f5-564c334001068ae16e82ce0f;aafdec3d-7141-4994-a73d-192be2632797) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -2757,8 +2739,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31525 | loss 5.7418 | text 2.1563 audio 5.3106 | grad 2.122 | 21.38s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 +step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2784,8 +2766,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31550 | loss 5.7494 | text 2.1284 audio 5.3237 | grad 2.364 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.06 | nonfinite=0 +step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2811,8 +2793,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31575 | loss 5.7309 | text 2.1538 audio 5.3001 | grad 1.760 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 +step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2838,8 +2820,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31600 | loss 5.7379 | text 2.1287 audio 5.3122 | grad 3.343 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.09 | spike L=+0.00 G=1.53 | nonfinite=0 +step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2865,8 +2847,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31625 | loss 5.6686 | text 2.1377 audio 5.2410 | grad 2.997 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.11 | spike L=-0.01 G=1.30 | nonfinite=0 +step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2892,8 +2874,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31650 | loss 5.6969 | text 2.1637 audio 5.2641 | grad 1.750 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 +step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2919,8 +2901,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31675 | loss 5.7165 | text 2.1522 audio 5.2861 | grad 2.261 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 + [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 +step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2946,8 +2929,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31700 | loss 5.6886 | text 2.1403 audio 5.2606 | grad 2.322 | 1.57s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.01 | nonfinite=0 +step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -2973,8 +2956,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31725 | loss 5.7274 | text 2.1259 audio 5.3022 | grad 1.478 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.04 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.64 | nonfinite=0 +step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3000,12 +2983,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31750 | loss 5.6918 | text 2.1276 audio 5.2663 | grad 2.754 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.23 lora=0.97 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=-0.00 G=1.24 | nonfinite=0 - running validation at step 31750... - val/composite=3.0149 (text=0.6321 audio=4.6034) val/loss=4.7299 cb0_acc=39.0% text_acc=90.9% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0129) +step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 + running validation at step 32750... + val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3031,8 +3023,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31775 | loss 5.7190 | text 2.1568 audio 5.2876 | grad 1.540 | 2.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.68 | nonfinite=0 +step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3058,8 +3050,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31800 | loss 5.7086 | text 2.1397 audio 5.2807 | grad 1.333 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.60 | nonfinite=0 +step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3085,8 +3077,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31825 | loss 5.7059 | text 2.1289 audio 5.2801 | grad 1.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3112,8 +3104,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31850 | loss 5.7276 | text 2.1494 audio 5.2978 | grad 1.604 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.08 text_embed=0.09 | spike L=+0.00 G=0.77 | nonfinite=0 +step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3139,9 +3131,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 3.4x (grad 6.92 vs EMA 2.02) step 31875 -step 31875 | loss 5.7635 | text 2.1532 audio 5.3329 | grad 6.916 | 1.58s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.10 lora=0.99 model_audio_embed=0.05 projection=0.02 text_embed=0.08 | spike L=+0.01 G=3.42 | nonfinite=0 +step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3167,8 +3158,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31900 | loss 5.7007 | text 2.1354 audio 5.2736 | grad 1.531 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.10 text_embed=0.12 | spike L=-0.00 G=0.61 | nonfinite=0 +step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3194,8 +3185,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31925 | loss 5.7284 | text 2.1511 audio 5.2981 | grad 2.452 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.02 | nonfinite=0 +step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3221,8 +3212,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31950 | loss 5.7026 | text 2.1339 audio 5.2759 | grad 2.128 | 1.54s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=0.88 | nonfinite=0 +step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3248,8 +3239,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 31975 | loss 5.7325 | text 2.1323 audio 5.3060 | grad 3.383 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.19 lora=0.98 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.00 G=1.42 | nonfinite=0 +step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3275,15 +3266,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32000 | loss 5.7196 | text 2.1309 audio 5.2935 | grad 2.305 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.93 | nonfinite=0 - running validation at step 32000... - val/composite=3.0141 (text=0.6317 audio=4.6024) val/loss=4.7287 cb0_acc=38.9% text_acc=90.9% - [val] per-codebook acc: cb0=38.9% cb1=20.0% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 8/10 (best 3.0129) +step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G + [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 + running validation at step 33000... + val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% + [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_y3j7s14l/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 +[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3294,17 +3285,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_032000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5701-676c84df289d499c023dd0e8;593c1b3a-6ef8-4474-8a06-3d170a535244) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3318,9 +3308,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32025 | loss 5.7313 | text 2.1524 audio 5.3008 | grad 1.656 | 20.67s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3345,9 +3335,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32050 | loss 5.7048 | text 2.1554 audio 5.2737 | grad 2.361 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3372,9 +3362,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32075 | loss 5.7039 | text 2.0979 audio 5.2843 | grad 1.845 | 1.48s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3399,9 +3389,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32100 | loss 5.7270 | text 2.1393 audio 5.2991 | grad 1.905 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3426,9 +3416,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32125 | loss 5.7158 | text 2.1071 audio 5.2944 | grad 1.418 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.62 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3453,9 +3443,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32150 | loss 5.7424 | text 2.1768 audio 5.3070 | grad 2.442 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3480,9 +3470,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32175 | loss 5.7048 | text 2.1509 audio 5.2746 | grad 2.600 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3507,9 +3497,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32200 | loss 5.6916 | text 2.1584 audio 5.2599 | grad 2.043 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.90 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3534,9 +3524,9 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32225 | loss 5.7350 | text 2.1529 audio 5.3044 | grad 2.626 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3561,16 +3551,17 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32250 | loss 5.6589 | text 2.1243 audio 5.2341 | grad 1.595 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.42 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.70 | nonfinite=0 - running validation at step 32250... - val/composite=3.0107 (text=0.6328 audio=4.5959) val/loss=4.7225 cb0_acc=39.1% text_acc=91.0% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.0% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.7% +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G + [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 + running validation at step 33250... + val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_2q_qpjcc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5a53-7aa975cd701e1c3a179e1ee4;c16d8895-718f-427d-aa46-6b75ac134eac) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -3601,8 +3592,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32275 | loss 5.7120 | text 2.1521 audio 5.2816 | grad 2.076 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3628,8 +3619,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32300 | loss 5.7328 | text 2.1731 audio 5.2982 | grad 2.765 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.26 | nonfinite=0 +step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3655,8 +3646,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32325 | loss 5.6889 | text 2.1296 audio 5.2630 | grad 3.661 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.18 lora=0.98 model_audio_embed=0.05 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.62 | nonfinite=0 +step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3682,8 +3673,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32350 | loss 5.7429 | text 2.1437 audio 5.3142 | grad 1.754 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.06 projection=0.08 text_embed=0.18 | spike L=+0.01 G=0.73 | nonfinite=0 +step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3709,8 +3700,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32375 | loss 5.7029 | text 2.1066 audio 5.2816 | grad 1.824 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3736,8 +3727,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32400 | loss 5.7371 | text 2.1651 audio 5.3041 | grad 2.395 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.05 | nonfinite=0 +step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3763,8 +3754,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32425 | loss 5.7261 | text 2.1371 audio 5.2987 | grad 2.860 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.11 | spike L=+0.00 G=1.25 | nonfinite=0 +step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3790,8 +3781,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32450 | loss 5.7025 | text 2.1555 audio 5.2714 | grad 2.142 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.91 | nonfinite=0 +step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3817,8 +3808,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32475 | loss 5.7278 | text 2.1393 audio 5.3000 | grad 1.822 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=+0.00 G=0.78 | nonfinite=0 +step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3844,16 +3835,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32500 | loss 5.7113 | text 2.1342 audio 5.2845 | grad 2.222 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.98 | nonfinite=0 - running validation at step 32500... - val/composite=3.0101 (text=0.6280 audio=4.5982) val/loss=4.7238 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 + running validation at step 33500... + val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_lc3w192r/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b5dc0-53129883635583d7158de6f9;688fe6af-c937-462d-86e6-91c9a1e6d783) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -3884,8 +3875,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32525 | loss 5.7083 | text 2.1319 audio 5.2819 | grad 2.125 | 21.39s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=0.94 | nonfinite=0 +step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3911,8 +3902,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32550 | loss 5.7155 | text 2.1400 audio 5.2875 | grad 2.616 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=1.16 | nonfinite=0 +step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3938,8 +3929,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32575 | loss 5.7057 | text 2.1315 audio 5.2794 | grad 1.909 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.83 | nonfinite=0 +step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3965,8 +3956,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32600 | loss 5.7114 | text 2.1253 audio 5.2864 | grad 1.935 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.86 | nonfinite=0 +step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -3992,8 +3983,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32625 | loss 5.7150 | text 2.1387 audio 5.2872 | grad 1.556 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.00 G=0.70 | nonfinite=0 +step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4019,8 +4010,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32650 | loss 5.7460 | text 2.1485 audio 5.3163 | grad 1.521 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=+0.01 G=0.71 | nonfinite=0 +step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4046,9 +4037,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 6.5x (grad 13.57 vs EMA 2.09) step 32675 -step 32675 | loss 5.7083 | text 2.1197 audio 5.2843 | grad 13.572 | 1.52s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.05 lora=1.00 model_audio_embed=0.05 projection=0.01 text_embed=0.06 | spike L=-0.00 G=6.49 | nonfinite=0 +step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4074,8 +4064,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32700 | loss 5.7349 | text 2.1457 audio 5.3057 | grad 2.156 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.13 | spike L=+0.00 G=0.67 | nonfinite=0 +step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4101,8 +4091,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32725 | loss 5.6968 | text 2.1348 audio 5.2698 | grad 1.604 | 1.56s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.51 | nonfinite=0 +step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4128,21 +4118,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32750 | loss 5.7333 | text 2.1561 audio 5.3021 | grad 1.545 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.46 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.09 | spike L=+0.00 G=0.52 | nonfinite=0 - running validation at step 32750... - val/composite=3.0069 (text=0.6297 audio=4.5917) val/loss=4.7177 cb0_acc=39.1% text_acc=91.1% - [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_70qjzrkf/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b612d-0d3151834f43bd51736c7f2d;ea00c642-bdf8-4474-948f-a84d24d9a120) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 + running validation at step 33750... + val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4168,8 +4149,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32775 | loss 5.7253 | text 2.1631 audio 5.2927 | grad 1.639 | 21.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.58 | nonfinite=0 +step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4195,8 +4176,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32800 | loss 5.7059 | text 2.1316 audio 5.2796 | grad 2.172 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.15 | spike L=-0.00 G=0.80 | nonfinite=0 +step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4222,8 +4203,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32825 | loss 5.6958 | text 2.1395 audio 5.2679 | grad 1.854 | 1.51s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 +step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4249,8 +4230,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32850 | loss 5.7124 | text 2.1521 audio 5.2819 | grad 1.945 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.75 | nonfinite=0 +step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4276,8 +4257,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32875 | loss 5.7092 | text 2.1398 audio 5.2812 | grad 2.402 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=-0.00 G=0.95 | nonfinite=0 +step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4303,8 +4284,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32900 | loss 5.7017 | text 2.1311 audio 5.2754 | grad 1.630 | 1.55s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.65 | nonfinite=0 +step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4330,8 +4311,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32925 | loss 5.7284 | text 2.1699 audio 5.2945 | grad 2.090 | 1.53s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 +step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4357,8 +4338,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32950 | loss 5.7212 | text 2.1425 audio 5.2927 | grad 3.371 | 1.49s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=+0.00 G=1.41 | nonfinite=0 +step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4384,8 +4365,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 32975 | loss 5.7070 | text 2.1172 audio 5.2836 | grad 1.944 | 1.45s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.78 | nonfinite=0 +step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4411,15 +4392,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33000 | loss 5.7159 | text 2.1396 audio 5.2880 | grad 1.531 | 1.50s/step | peak -1.0G | host_rss 45.3G - [diag] gradnorm depth=0.41 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.17 | spike L=+0.00 G=0.63 | nonfinite=0 - running validation at step 33000... - val/composite=3.0095 (text=0.6220 audio=4.6011) val/loss=4.7255 cb0_acc=38.9% text_acc=91.1% - [val] per-codebook acc: cb0=38.9% cb1=20.1% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0069) -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. +step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G + [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 + running validation at step 34000... + val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% + [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_dqysxqvc/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 +[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4430,21 +4420,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_033000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b649a-59cd28081005181a07e291fa;3ccd8853-dd52-43bb-9394-e4cffa93737b) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4454,8 +4444,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33025 | loss 5.7435 | text 2.1748 audio 5.3085 | grad 1.711 | 20.72s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.01 G=0.73 | nonfinite=0 +step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4481,8 +4471,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33050 | loss 5.7479 | text 2.1749 audio 5.3129 | grad 2.267 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.14 | spike L=+0.01 G=1.00 | nonfinite=0 +step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4508,8 +4498,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33075 | loss 5.6842 | text 2.1272 audio 5.2587 | grad 3.121 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.01 G=1.37 | nonfinite=0 +step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4535,8 +4525,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33100 | loss 5.7066 | text 2.1162 audio 5.2834 | grad 2.275 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.31 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.16 | spike L=-0.00 G=0.96 | nonfinite=0 +step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4562,8 +4552,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33125 | loss 5.7258 | text 2.1497 audio 5.2958 | grad 2.791 | 1.52s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.19 | nonfinite=0 +step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4589,8 +4579,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33150 | loss 5.6867 | text 2.1251 audio 5.2617 | grad 2.156 | 1.53s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.13 | spike L=-0.01 G=0.90 | nonfinite=0 +step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4616,8 +4606,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33175 | loss 5.7465 | text 2.1708 audio 5.3123 | grad 3.098 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=+0.01 G=1.31 | nonfinite=0 +step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4643,8 +4633,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33200 | loss 5.6825 | text 2.1098 audio 5.2605 | grad 2.620 | 1.47s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.13 | spike L=-0.01 G=1.07 | nonfinite=0 +step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4670,8 +4660,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33225 | loss 5.7178 | text 2.1299 audio 5.2918 | grad 2.059 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.84 | nonfinite=0 +step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4697,21 +4687,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33250 | loss 5.7431 | text 2.1444 audio 5.3142 | grad 2.975 | 1.50s/step | peak -1.0G | host_rss 45.8G - [diag] gradnorm depth=0.22 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=+0.01 G=1.23 | nonfinite=0 - running validation at step 33250... - val/composite=3.0059 (text=0.6220 audio=4.5951) val/loss=4.7195 cb0_acc=39.0% text_acc=91.1% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ymnypmm2/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b67f1-30da824202ed69fc19290dd3;819c5ecd-b13d-42a5-a250-43e1632cc59f) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 + running validation at step 34250... + val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% + [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4737,8 +4718,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33275 | loss 5.7334 | text 2.1417 audio 5.3051 | grad 2.408 | 21.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=0.97 | nonfinite=0 +step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4764,8 +4745,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33300 | loss 5.7199 | text 2.1622 audio 5.2874 | grad 2.667 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=+0.00 G=1.08 | nonfinite=0 +step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4791,8 +4772,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33325 | loss 5.7126 | text 2.1546 audio 5.2817 | grad 2.010 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 +step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4818,8 +4799,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33350 | loss 5.7029 | text 2.1469 audio 5.2736 | grad 1.681 | 1.46s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.69 | nonfinite=0 +step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4845,8 +4826,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33375 | loss 5.6882 | text 2.1282 audio 5.2626 | grad 2.886 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.22 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.22 | nonfinite=0 +step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4872,8 +4853,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33400 | loss 5.6880 | text 2.1483 audio 5.2584 | grad 1.699 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.70 | nonfinite=0 +step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4899,8 +4880,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33425 | loss 5.7187 | text 2.1416 audio 5.2904 | grad 2.010 | 1.47s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.86 | nonfinite=0 +step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4926,8 +4907,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33450 | loss 5.7143 | text 2.1262 audio 5.2891 | grad 3.706 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.07 projection=0.04 text_embed=0.16 | spike L=+0.00 G=1.60 | nonfinite=0 +step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4953,8 +4934,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33475 | loss 5.7084 | text 2.1426 audio 5.2798 | grad 2.644 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.08 | nonfinite=0 +step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -4980,21 +4961,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33500 | loss 5.7220 | text 2.1670 audio 5.2886 | grad 2.575 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.04 | nonfinite=0 - running validation at step 33500... - val/composite=3.0006 (text=0.6133 audio=4.5922) val/loss=4.7149 cb0_acc=39.0% text_acc=91.4% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_j0mrl85n/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b6b5c-127a9b1067da909d31675cc1;34eee324-3c86-4a8d-8a80-d0538e4113dd) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 + running validation at step 34500... + val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% + [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5020,8 +4992,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33525 | loss 5.7135 | text 2.1668 audio 5.2801 | grad 1.389 | 21.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.56 | nonfinite=0 +step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5047,8 +5019,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33550 | loss 5.7341 | text 2.1403 audio 5.3060 | grad 1.642 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.69 | nonfinite=0 +step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5074,8 +5046,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33575 | loss 5.7011 | text 2.1289 audio 5.2754 | grad 1.863 | 1.50s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.81 | nonfinite=0 +step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5101,8 +5073,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33600 | loss 5.7024 | text 2.1479 audio 5.2728 | grad 1.537 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.68 | nonfinite=0 +step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5128,8 +5100,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33625 | loss 5.7159 | text 2.1464 audio 5.2866 | grad 1.842 | 1.48s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=+0.00 G=0.84 | nonfinite=0 +step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5155,8 +5127,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33650 | loss 5.6905 | text 2.1480 audio 5.2609 | grad 1.867 | 1.60s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=0.87 | nonfinite=0 +step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5182,8 +5154,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33675 | loss 5.6960 | text 2.1260 audio 5.2708 | grad 3.509 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.12 | spike L=-0.00 G=1.65 | nonfinite=0 +step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5209,8 +5181,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33700 | loss 5.6934 | text 2.1037 audio 5.2727 | grad 2.820 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.06 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 +step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5236,8 +5208,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33725 | loss 5.6974 | text 2.0929 audio 5.2788 | grad 1.813 | 1.58s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.78 | nonfinite=0 +step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5263,12 +5235,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33750 | loss 5.6610 | text 2.1089 audio 5.2392 | grad 2.502 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.01 G=1.10 | nonfinite=0 - running validation at step 33750... - val/composite=3.0042 (text=0.6186 audio=4.5947) val/loss=4.7184 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 3.0006) +step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 + running validation at step 34750... + val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% + [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5294,8 +5266,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33775 | loss 5.7118 | text 2.1578 audio 5.2802 | grad 2.755 | 2.36s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.09 | spike L=+0.00 G=1.20 | nonfinite=0 +step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5321,8 +5293,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33800 | loss 5.7103 | text 2.1690 audio 5.2765 | grad 1.722 | 1.56s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.74 | nonfinite=0 +step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5348,8 +5320,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33825 | loss 5.6946 | text 2.1025 audio 5.2741 | grad 2.228 | 1.51s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=0.98 | nonfinite=0 +step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5375,8 +5347,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33850 | loss 5.7031 | text 2.1177 audio 5.2796 | grad 1.597 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.70 | nonfinite=0 +step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5402,8 +5374,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33875 | loss 5.7057 | text 2.1149 audio 5.2827 | grad 1.535 | 1.52s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.00 G=0.70 | nonfinite=0 +step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5429,8 +5401,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33900 | loss 5.7093 | text 2.1131 audio 5.2867 | grad 2.293 | 1.55s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.00 G=1.07 | nonfinite=0 +step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5456,8 +5428,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33925 | loss 5.7178 | text 2.1343 audio 5.2910 | grad 2.053 | 1.57s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.95 | nonfinite=0 +step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5483,8 +5455,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33950 | loss 5.7401 | text 2.1343 audio 5.3133 | grad 3.863 | 1.54s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.17 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.15 | spike L=+0.01 G=1.80 | nonfinite=0 +step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5510,8 +5482,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 33975 | loss 5.7176 | text 2.1229 audio 5.2931 | grad 2.129 | 1.53s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.92 | nonfinite=0 +step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5537,16 +5509,43 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34000 | loss 5.6863 | text 2.1245 audio 5.2614 | grad 3.301 | 1.49s/step | peak -1.0G | host_rss 45.9G - [diag] gradnorm depth=0.21 lora=0.97 model_audio_embed=0.06 projection=0.04 text_embed=0.13 | spike L=-0.00 G=1.44 | nonfinite=0 - running validation at step 34000... - val/composite=2.9981 (text=0.6123 audio=4.5887) val/loss=4.7112 cb0_acc=39.2% text_acc=91.4% - [val] per-codebook acc: cb0=39.2% cb1=20.4% cb2=17.0% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% +step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 + [audio-demo] ar_cb0_acc=12.0% (35s) + running validation at step 35000... + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  + + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s + New Data Upload : | | 0.00B / 0.00B, ???B/s + ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB + Processing Files (0 / 0) : | | 0.00B / 0.00B + New Data Upload : | | 0.00B / 0.00B  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  + + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s + New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s + ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB + pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 + val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_svr4iuqv/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7066-54ca26647ada8ba96b65eca4;963ada0a-af52-4cfb-851f-0368c724f72f) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -5554,7 +5553,7 @@ Make sure your token has the correct permissions. * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_mfze1yuj/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5565,19 +5564,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_034000 +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7245-375335441723bd3941902801;ffc4ad61-7145-484a-84ff-c0139e1eb24f) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7c65-1c38aecc5ca4acdb5008e0b4;48e1227e-d30a-4783-b035-8b02985e5e1d) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5589,11 +5585,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34025 | loss 5.7225 | text 2.1311 audio 5.2963 | grad 2.120 | 39.94s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.92 model_audio_embed=0.06 projection=0.06 text_embed=0.19 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35025 | loss 5.7000 | text 2.1331 audio 5.2734 | grad 2.588 | 41.18s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.18 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5616,11 +5612,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34050 | loss 5.6833 | text 2.1355 audio 5.2562 | grad 2.113 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35050 | loss 5.7063 | text 2.1069 audio 5.2850 | grad 2.479 | 1.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5643,11 +5639,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34075 | loss 5.6950 | text 2.1184 audio 5.2713 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + [ALERT] grad spike 3.0x (grad 6.76 vs EMA 2.25) step 35075 +step 35075 | loss 5.7092 | text 2.1201 audio 5.2851 | grad 6.761 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.11 lora=0.98 model_audio_embed=0.06 projection=0.02 text_embed=0.12 | spike L=+0.00 G=3.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5670,11 +5667,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34100 | loss 5.6984 | text 2.1317 audio 5.2720 | grad 2.787 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.22 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35100 | loss 5.7286 | text 2.1354 audio 5.3015 | grad 2.020 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5697,11 +5694,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34125 | loss 5.7281 | text 2.1302 audio 5.3021 | grad 2.069 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35125 | loss 5.6800 | text 2.1305 audio 5.2539 | grad 1.847 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5724,11 +5721,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34150 | loss 5.7037 | text 2.1209 audio 5.2796 | grad 3.054 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.21 lora=0.96 model_audio_embed=0.05 projection=0.04 text_embed=0.15 | spike L=-0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35150 | loss 5.7182 | text 2.1559 audio 5.2870 | grad 1.641 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5751,11 +5748,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34175 | loss 5.6899 | text 2.1143 audio 5.2671 | grad 2.012 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.13 | spike L=-0.00 G=0.84 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35175 | loss 5.6631 | text 2.1155 audio 5.2400 | grad 2.109 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5778,11 +5775,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34200 | loss 5.7261 | text 2.1484 audio 5.2965 | grad 1.367 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.50 lora=0.85 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=+0.00 G=0.58 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35200 | loss 5.7188 | text 2.1798 audio 5.2828 | grad 2.145 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5805,11 +5802,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34225 | loss 5.6907 | text 2.1303 audio 5.2646 | grad 1.792 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35225 | loss 5.7170 | text 2.1077 audio 5.2955 | grad 1.866 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.78 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5832,15 +5829,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34250 | loss 5.7166 | text 2.1371 audio 5.2892 | grad 1.868 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.16 | spike L=+0.00 G=0.85 | nonfinite=0 - running validation at step 34250... - val/composite=3.0032 (text=0.6142 audio=4.5959) val/loss=4.7187 cb0_acc=39.0% text_acc=91.2% - [val] per-codebook acc: cb0=39.0% cb1=20.1% cb2=16.9% cb3=11.0% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9981) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35250 | loss 5.6647 | text 2.0979 audio 5.2451 | grad 1.703 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.73 | nonfinite=0 + running validation at step 35250... + val/composite=2.9985 (text=0.6053 audio=4.5939) val/loss=4.7150 cb0_acc=39.1% text_acc=91.6% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.6% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9973) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5863,11 +5860,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34275 | loss 5.6904 | text 2.1224 audio 5.2659 | grad 3.344 | 2.44s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.19 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.54 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35275 | loss 5.7344 | text 2.1597 audio 5.3025 | grad 4.716 | 2.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=+0.01 G=2.07 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5890,11 +5887,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34300 | loss 5.7227 | text 2.1574 audio 5.2913 | grad 2.031 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.17 | spike L=+0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35300 | loss 5.6869 | text 2.1172 audio 5.2634 | grad 1.765 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.70 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5917,11 +5914,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34325 | loss 5.6939 | text 2.1397 audio 5.2660 | grad 1.870 | 1.61s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.83 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35325 | loss 5.7376 | text 2.1312 audio 5.3113 | grad 1.650 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5944,11 +5941,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34350 | loss 5.6886 | text 2.1227 audio 5.2641 | grad 1.419 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.64 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35350 | loss 5.6878 | text 2.1130 audio 5.2652 | grad 1.950 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5971,11 +5968,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34375 | loss 5.7424 | text 2.1397 audio 5.3145 | grad 1.599 | 1.59s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=+0.01 G=0.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35375 | loss 5.6788 | text 2.1222 audio 5.2544 | grad 1.724 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -5998,11 +5995,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34400 | loss 5.7078 | text 2.1227 audio 5.2832 | grad 2.389 | 1.53s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.14 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35400 | loss 5.6788 | text 2.1020 audio 5.2584 | grad 1.961 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6025,11 +6022,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34425 | loss 5.7407 | text 2.1707 audio 5.3066 | grad 1.427 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35425 | loss 5.6969 | text 2.0995 audio 5.2770 | grad 1.766 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6052,11 +6049,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34450 | loss 5.7077 | text 2.1706 audio 5.2736 | grad 2.570 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.25 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35450 | loss 5.6848 | text 2.0913 audio 5.2665 | grad 1.521 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.69 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6079,11 +6076,11 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34475 | loss 5.6641 | text 2.1020 audio 5.2437 | grad 1.802 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.47 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.10 | spike L=-0.01 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35475 | loss 5.7112 | text 2.1194 audio 5.2873 | grad 2.124 | 1.60s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6106,12 +6103,24 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34500 | loss 5.7111 | text 2.1372 audio 5.2836 | grad 1.723 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=+0.00 G=0.83 | nonfinite=0 - running validation at step 34500... - val/composite=3.0024 (text=0.6199 audio=4.5907) val/loss=4.7147 cb0_acc=39.1% text_acc=91.3% - [val] per-codebook acc: cb0=39.1% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.7% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9981) +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +step 35500 | loss 5.7002 | text 2.0979 audio 5.2806 | grad 2.371 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 + running validation at step 35500... + val/composite=2.9923 (text=0.5962 audio=4.5896) val/loss=4.7089 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_3crsay5y/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b816c-2712d01710663cc45f59dc84;4de1714a-3953-4c64-bba2-5dfc7dfb76bb) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6137,8 +6146,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34525 | loss 5.7139 | text 2.1525 audio 5.2835 | grad 1.805 | 2.35s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.14 | spike L=+0.00 G=0.89 | nonfinite=0 +step 35525 | loss 5.7157 | text 2.1252 audio 5.2907 | grad 1.712 | 21.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6164,8 +6173,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34550 | loss 5.7045 | text 2.1328 audio 5.2779 | grad 1.541 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.00 G=0.77 | nonfinite=0 +step 35550 | loss 5.7160 | text 2.1306 audio 5.2898 | grad 2.455 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6191,8 +6200,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34575 | loss 5.6994 | text 2.1397 audio 5.2715 | grad 2.426 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.23 | nonfinite=0 +step 35575 | loss 5.7021 | text 2.1011 audio 5.2819 | grad 2.139 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6218,8 +6227,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34600 | loss 5.7011 | text 2.1191 audio 5.2773 | grad 1.647 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.39 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 +step 35600 | loss 5.6976 | text 2.1119 audio 5.2752 | grad 1.692 | 1.59s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6245,8 +6254,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34625 | loss 5.7109 | text 2.1229 audio 5.2864 | grad 4.351 | 1.56s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.04 projection=0.03 text_embed=0.09 | spike L=+0.00 G=2.20 | nonfinite=0 +step 35625 | loss 5.7269 | text 2.1576 audio 5.2954 | grad 1.948 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.93 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6272,8 +6281,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34650 | loss 5.7072 | text 2.1289 audio 5.2814 | grad 2.504 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.13 | nonfinite=0 +step 35650 | loss 5.7038 | text 2.0968 audio 5.2844 | grad 2.009 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6299,8 +6308,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34675 | loss 5.7400 | text 2.1441 audio 5.3112 | grad 2.248 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.00 | nonfinite=0 +step 35675 | loss 5.6964 | text 2.1176 audio 5.2729 | grad 1.781 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.86 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6326,8 +6335,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34700 | loss 5.6857 | text 2.1158 audio 5.2625 | grad 2.498 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 +step 35700 | loss 5.7023 | text 2.1459 audio 5.2731 | grad 1.470 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6353,8 +6362,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34725 | loss 5.7169 | text 2.1167 audio 5.2936 | grad 2.229 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=+0.00 G=0.98 | nonfinite=0 +step 35725 | loss 5.6936 | text 2.1079 audio 5.2720 | grad 1.577 | 1.61s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6380,12 +6389,21 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34750 | loss 5.7057 | text 2.1438 audio 5.2769 | grad 2.059 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=0.91 | nonfinite=0 - running validation at step 34750... - val/composite=3.0000 (text=0.6135 audio=4.5911) val/loss=4.7138 cb0_acc=39.0% text_acc=91.5% - [val] per-codebook acc: cb0=39.0% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 7/10 (best 2.9981) +step 35750 | loss 5.7018 | text 2.1473 audio 5.2724 | grad 3.195 | 1.56s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.64 | nonfinite=0 + running validation at step 35750... + val/composite=2.9894 (text=0.5988 audio=4.5832) val/loss=4.7030 cb0_acc=39.1% text_acc=91.8% + [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. + warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") +[ckpt] uploading /tmp/ckpt_q4bfdk_b/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b84eb-23e8713c6407ba4a163897ad;da59c5d6-9f19-4363-b138-8063d8495773) + +403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. +Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. +Make sure your token has the correct permissions. + * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6411,8 +6429,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34775 | loss 5.7248 | text 2.1412 audio 5.2966 | grad 1.818 | 2.31s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.35 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.18 | spike L=+0.00 G=0.81 | nonfinite=0 +step 35775 | loss 5.6928 | text 2.1394 audio 5.2649 | grad 2.077 | 21.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6438,8 +6456,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34800 | loss 5.6919 | text 2.1225 audio 5.2674 | grad 2.318 | 1.57s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.05 | nonfinite=0 +step 35800 | loss 5.7009 | text 2.1499 audio 5.2709 | grad 3.615 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.75 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6465,8 +6483,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34825 | loss 5.7131 | text 2.1577 audio 5.2816 | grad 1.742 | 1.58s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.79 | nonfinite=0 +step 35825 | loss 5.7072 | text 2.1152 audio 5.2842 | grad 1.769 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.79 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6492,8 +6510,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34850 | loss 5.7182 | text 2.1306 audio 5.2921 | grad 1.742 | 1.52s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.80 | nonfinite=0 +step 35850 | loss 5.7080 | text 2.1325 audio 5.2815 | grad 1.590 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6519,8 +6537,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34875 | loss 5.7273 | text 2.1472 audio 5.2978 | grad 1.850 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 +step 35875 | loss 5.7191 | text 2.1272 audio 5.2937 | grad 2.377 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6546,8 +6564,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34900 | loss 5.6855 | text 2.1320 audio 5.2591 | grad 2.839 | 1.55s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.35 | nonfinite=0 +step 35900 | loss 5.6961 | text 2.0964 audio 5.2768 | grad 1.901 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6573,8 +6591,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34925 | loss 5.6774 | text 2.1182 audio 5.2537 | grad 1.635 | 1.48s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.75 | nonfinite=0 +step 35925 | loss 5.6746 | text 2.1152 audio 5.2516 | grad 1.636 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6600,8 +6618,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34950 | loss 5.7159 | text 2.1427 audio 5.2874 | grad 1.754 | 1.54s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.83 | nonfinite=0 +step 35950 | loss 5.6555 | text 2.1133 audio 5.2328 | grad 1.572 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.76 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6627,8 +6645,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 34975 | loss 5.6886 | text 2.1512 audio 5.2584 | grad 3.280 | 1.51s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.05 projection=0.04 text_embed=0.09 | spike L=-0.00 G=1.58 | nonfinite=0 +step 35975 | loss 5.6890 | text 2.1208 audio 5.2648 | grad 2.115 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6654,51 +6672,15 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35000 | loss 5.7111 | text 2.1238 audio 5.2864 | grad 2.058 | 1.50s/step | peak -1.0G | host_rss 46.4G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=+0.00 G=0.94 | nonfinite=0 - [audio-demo] ar_cb0_acc=12.0% (35s) - running validation at step 35000... - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  - - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB  Processing Files (1 / 1) : 100%|██████████| 230kB / 230kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...es/step_035000/source.wav: 100%|██████████| 230kB / 230kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, ???B/s - New Data Upload : | | 0.00B / 0.00B, ???B/s - ...step_035000/target_gt.wav: 100%|██████████| 192kB / 192kB - Processing Files (0 / 0) : | | 0.00B / 0.00B - New Data Upload : | | 0.00B / 0.00B  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  - - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s  Processing Files (1 / 1) : 100%|██████████| 192kB / 192kB, 18.6kB/s - New Data Upload : 100%|██████████| 192kB / 192kB, 18.6kB/s - ...step_035000/generated.wav: 100%|██████████| 192kB / 192kB - pushed 3 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/samples/step_035000 - val/composite=2.9973 (text=0.6081 audio=4.5901) val/loss=4.7118 cb0_acc=39.2% text_acc=91.6% - [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=16.9% cb3=11.1% cb4=8.3% cb5=6.8% cb6=5.8% cb7=5.6% -/opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. - warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ib4oqihz/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7a89-48f990b947c53fdd7df45d51;504d5845-4577-45c5-ab47-e5b9f376f1b4) - -403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. -Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. -Make sure your token has the correct permissions. - * new best val — saved to gs:/tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +step 36000 | loss 5.7153 | text 2.1097 audio 5.2934 | grad 2.235 | 1.49s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.10 | nonfinite=0 + running validation at step 36000... + val/composite=2.9897 (text=0.5974 audio=4.5846) val/loss=4.7041 cb0_acc=39.2% text_acc=91.8% + [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9894) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_ij5q_kus/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 +[ckpt] uploading /tmp/ckpt_skck_i9h/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3 [ckpt-index] log failed: Cannot read the W&B step in shared mode. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6709,22 +6691,22 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_035000 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b7c65-1c38aecc5ca4acdb5008e0b4;48e1227e-d30a-4783-b035-8b02985e5e1d) +[ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8865-1ec56561651c1a56354c727e;d2e4e91e-8960-46ef-b922-28020e2e7473) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. Make sure your token has the correct permissions. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. + pushed 1 file(s) -> https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3/tree/main/logs sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6733,8 +6715,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35025 | loss 5.7000 | text 2.1331 audio 5.2734 | grad 2.588 | 41.18s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.12 | spike L=-0.00 G=1.18 | nonfinite=0 +step 36025 | loss 5.6831 | text 2.1034 audio 5.2624 | grad 1.684 | 20.86s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.06 projection=0.08 text_embed=0.15 | spike L=-0.00 G=0.82 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6760,8 +6742,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35050 | loss 5.7063 | text 2.1069 audio 5.2850 | grad 2.479 | 1.48s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.11 | nonfinite=0 +step 36050 | loss 5.7093 | text 2.0998 audio 5.2893 | grad 1.919 | 1.59s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6787,9 +6769,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. - [ALERT] grad spike 3.0x (grad 6.76 vs EMA 2.25) step 35075 -step 35075 | loss 5.7092 | text 2.1201 audio 5.2851 | grad 6.761 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.11 lora=0.98 model_audio_embed=0.06 projection=0.02 text_embed=0.12 | spike L=+0.00 G=3.00 | nonfinite=0 +step 36075 | loss 5.6986 | text 2.1149 audio 5.2756 | grad 2.041 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.14 | spike L=+0.00 G=1.02 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6815,8 +6796,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35100 | loss 5.7286 | text 2.1354 audio 5.3015 | grad 2.020 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=+0.00 G=0.75 | nonfinite=0 +step 36100 | loss 5.7108 | text 2.1152 audio 5.2878 | grad 1.744 | 1.50s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.87 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6842,8 +6823,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35125 | loss 5.6800 | text 2.1305 audio 5.2539 | grad 1.847 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.34 lora=0.92 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.70 | nonfinite=0 +step 36125 | loss 5.6697 | text 2.1321 audio 5.2433 | grad 2.215 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.01 G=1.12 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6869,8 +6850,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35150 | loss 5.7182 | text 2.1559 audio 5.2870 | grad 1.641 | 1.51s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.42 lora=0.89 model_audio_embed=0.06 projection=0.08 text_embed=0.13 | spike L=+0.00 G=0.64 | nonfinite=0 +step 36150 | loss 5.6836 | text 2.1104 audio 5.2616 | grad 1.771 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=-0.00 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6896,8 +6877,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35175 | loss 5.6631 | text 2.1155 audio 5.2400 | grad 2.109 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.06 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.86 | nonfinite=0 +step 36175 | loss 5.7177 | text 2.1067 audio 5.2964 | grad 1.322 | 1.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.49 lora=0.86 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.67 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6923,8 +6904,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35200 | loss 5.7188 | text 2.1798 audio 5.2828 | grad 2.145 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.88 | nonfinite=0 +step 36200 | loss 5.6609 | text 2.1179 audio 5.2373 | grad 2.159 | 1.55s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.30 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.01 G=1.13 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6950,8 +6931,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35225 | loss 5.7170 | text 2.1077 audio 5.2955 | grad 1.866 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.78 | nonfinite=0 +step 36225 | loss 5.7207 | text 2.1407 audio 5.2925 | grad 2.000 | 1.50s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=1.03 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -6977,12 +6958,12 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35250 | loss 5.6647 | text 2.0979 audio 5.2451 | grad 1.703 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.38 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.11 | spike L=-0.01 G=0.73 | nonfinite=0 - running validation at step 35250... - val/composite=2.9985 (text=0.6053 audio=4.5939) val/loss=4.7150 cb0_acc=39.1% text_acc=91.6% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.8% cb6=5.8% cb7=5.6% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9973) +step 36250 | loss 5.6867 | text 2.0964 audio 5.2674 | grad 2.338 | 1.47s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.09 | spike L=-0.00 G=1.20 | nonfinite=0 + running validation at step 36250... + val/composite=2.9903 (text=0.5982 audio=4.5851) val/loss=4.7047 cb0_acc=39.2% text_acc=91.8% + [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 8/10 (best 2.9894) sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7008,8 +6989,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35275 | loss 5.7344 | text 2.1597 audio 5.3025 | grad 4.716 | 2.47s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.15 lora=0.98 model_audio_embed=0.05 projection=0.03 text_embed=0.11 | spike L=+0.01 G=2.07 | nonfinite=0 +step 36275 | loss 5.7501 | text 2.1488 audio 5.3204 | grad 1.751 | 2.30s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.01 G=0.88 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7035,8 +7016,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35300 | loss 5.6869 | text 2.1172 audio 5.2634 | grad 1.765 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.00 G=0.70 | nonfinite=0 +step 36300 | loss 5.6584 | text 2.0785 audio 5.2427 | grad 1.793 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.01 G=0.91 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7062,8 +7043,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35325 | loss 5.7376 | text 2.1312 audio 5.3113 | grad 1.650 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.40 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.01 G=0.67 | nonfinite=0 +step 36325 | loss 5.7092 | text 2.1386 audio 5.2815 | grad 1.547 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.80 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7089,8 +7070,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35350 | loss 5.6878 | text 2.1130 audio 5.2652 | grad 1.950 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.82 | nonfinite=0 +step 36350 | loss 5.7118 | text 2.1361 audio 5.2846 | grad 2.371 | 1.60s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.24 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7116,8 +7097,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35375 | loss 5.6788 | text 2.1222 audio 5.2544 | grad 1.724 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.74 | nonfinite=0 +step 36375 | loss 5.7245 | text 2.1202 audio 5.3005 | grad 1.915 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7143,8 +7124,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35400 | loss 5.6788 | text 2.1020 audio 5.2584 | grad 1.961 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.39 lora=0.90 model_audio_embed=0.05 projection=0.07 text_embed=0.15 | spike L=-0.00 G=0.86 | nonfinite=0 +step 36400 | loss 5.7019 | text 2.1270 audio 5.2765 | grad 2.564 | 1.51s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.15 | spike L=+0.00 G=1.32 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7170,8 +7151,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35425 | loss 5.6969 | text 2.0995 audio 5.2770 | grad 1.766 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.45 lora=0.88 model_audio_embed=0.05 projection=0.10 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36425 | loss 5.7086 | text 2.1082 audio 5.2870 | grad 1.945 | 1.58s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7197,8 +7178,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35450 | loss 5.6848 | text 2.0913 audio 5.2665 | grad 1.521 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.69 | nonfinite=0 +step 36450 | loss 5.6757 | text 2.1256 audio 5.2505 | grad 1.849 | 1.57s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.38 lora=0.91 model_audio_embed=0.05 projection=0.07 text_embed=0.16 | spike L=-0.00 G=0.92 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7224,8 +7205,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35475 | loss 5.7112 | text 2.1194 audio 5.2873 | grad 2.124 | 1.60s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.04 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.00 | nonfinite=0 +step 36475 | loss 5.6761 | text 2.1255 audio 5.2510 | grad 1.958 | 1.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.98 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7251,16 +7232,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35500 | loss 5.7002 | text 2.0979 audio 5.2806 | grad 2.371 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 - running validation at step 35500... - val/composite=2.9923 (text=0.5962 audio=4.5896) val/loss=4.7089 cb0_acc=39.1% text_acc=91.8% - [val] per-codebook acc: cb0=39.1% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.3% cb5=6.9% cb6=5.8% cb7=5.7% +step 36500 | loss 5.6908 | text 2.0905 audio 5.2727 | grad 1.988 | 1.54s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.13 | spike L=-0.00 G=1.00 | nonfinite=0 + running validation at step 36500... + val/composite=2.9857 (text=0.5922 audio=4.5813) val/loss=4.6998 cb0_acc=39.4% text_acc=91.8% + [val] per-codebook acc: cb0=39.4% cb1=20.1% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_3crsay5y/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_pnu0pz72/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b816c-2712d01710663cc45f59dc84;4de1714a-3953-4c64-bba2-5dfc7dfb76bb) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b8d52-3c934e2072b2c29c7b9ed4a4;c8972a1a-b255-4d8a-a62a-f52dd68711e3) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -7291,8 +7272,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35525 | loss 5.7157 | text 2.1252 audio 5.2907 | grad 1.712 | 21.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.10 | spike L=+0.00 G=0.80 | nonfinite=0 +step 36525 | loss 5.6839 | text 2.1440 audio 5.2551 | grad 2.686 | 21.43s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.24 lora=0.96 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.35 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7318,8 +7299,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35550 | loss 5.7160 | text 2.1306 audio 5.2898 | grad 2.455 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.10 | spike L=+0.00 G=1.17 | nonfinite=0 +step 36550 | loss 5.6939 | text 2.1047 audio 5.2730 | grad 2.494 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.26 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.21 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7345,8 +7326,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35575 | loss 5.7021 | text 2.1011 audio 5.2819 | grad 2.139 | 1.57s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.10 | spike L=-0.00 G=1.00 | nonfinite=0 +step 36575 | loss 5.6677 | text 2.0906 audio 5.2496 | grad 1.994 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.95 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7372,8 +7353,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35600 | loss 5.6976 | text 2.1119 audio 5.2752 | grad 1.692 | 1.59s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36600 | loss 5.7210 | text 2.1266 audio 5.2957 | grad 2.306 | 1.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.29 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.09 | spike L=+0.01 G=1.10 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7399,8 +7380,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35625 | loss 5.7269 | text 2.1576 audio 5.2954 | grad 1.948 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.05 projection=0.07 text_embed=0.12 | spike L=+0.00 G=0.93 | nonfinite=0 +step 36625 | loss 5.7019 | text 2.1429 audio 5.2733 | grad 1.521 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.72 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7426,8 +7407,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35650 | loss 5.7038 | text 2.0968 audio 5.2844 | grad 2.009 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.37 lora=0.91 model_audio_embed=0.06 projection=0.07 text_embed=0.14 | spike L=-0.00 G=0.97 | nonfinite=0 +step 36650 | loss 5.7157 | text 2.1187 audio 5.2920 | grad 1.489 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.46 lora=0.87 model_audio_embed=0.05 projection=0.09 text_embed=0.12 | spike L=+0.00 G=0.73 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7453,8 +7434,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35675 | loss 5.6964 | text 2.1176 audio 5.2729 | grad 1.781 | 1.56s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.35 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.86 | nonfinite=0 +step 36675 | loss 5.6895 | text 2.1200 audio 5.2655 | grad 1.530 | 1.58s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.14 | spike L=-0.00 G=0.77 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7480,8 +7461,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35700 | loss 5.7023 | text 2.1459 audio 5.2731 | grad 1.470 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.72 | nonfinite=0 +step 36700 | loss 5.6871 | text 2.1240 audio 5.2623 | grad 1.735 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=0.89 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7507,8 +7488,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35725 | loss 5.6936 | text 2.1079 audio 5.2720 | grad 1.577 | 1.61s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.44 lora=0.88 model_audio_embed=0.05 projection=0.09 text_embed=0.13 | spike L=-0.00 G=0.79 | nonfinite=0 +step 36725 | loss 5.6945 | text 2.1135 audio 5.2718 | grad 2.142 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.11 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7534,16 +7515,16 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35750 | loss 5.7018 | text 2.1473 audio 5.2724 | grad 3.195 | 1.56s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.10 | spike L=-0.00 G=1.64 | nonfinite=0 - running validation at step 35750... - val/composite=2.9894 (text=0.5988 audio=4.5832) val/loss=4.7030 cb0_acc=39.1% text_acc=91.8% - [val] per-codebook acc: cb0=39.1% cb1=20.3% cb2=17.1% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% +step 36750 | loss 5.6832 | text 2.1276 audio 5.2577 | grad 1.583 | 1.55s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.44 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=-0.00 G=0.81 | nonfinite=0 + running validation at step 36750... + val/composite=2.9854 (text=0.5932 audio=4.5801) val/loss=4.6988 cb0_acc=39.2% text_acc=92.0% + [val] per-codebook acc: cb0=39.2% cb1=20.2% cb2=17.0% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.8% /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_q4bfdk_b/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val +[ckpt] uploading /tmp/ckpt_jwbvue5g/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val [ckpt] upload complete: gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/best_by_val -[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b84eb-23e8713c6407ba4a163897ad;da59c5d6-9f19-4363-b138-8063d8495773) +[ckpt] WARNING: hub push failed: (Request ID: Root=1-6a5b90c5-6a65d78d5cb61e8b6beabdcb;44000a67-8fdf-4e95-91a1-63030a9bba23) 403 Forbidden: Private repository storage limit reached, please upgrade your plan to increase your private storage limit. Cannot access content at: https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3.git/info/lfs/objects/batch. @@ -7574,8 +7555,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35775 | loss 5.6928 | text 2.1394 audio 5.2649 | grad 2.077 | 21.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.06 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.00 | nonfinite=0 +step 36775 | loss 5.7233 | text 2.1374 audio 5.2958 | grad 2.439 | 21.48s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.10 | spike L=+0.01 G=1.28 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7601,8 +7582,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35800 | loss 5.7009 | text 2.1499 audio 5.2709 | grad 3.615 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.20 lora=0.97 model_audio_embed=0.04 projection=0.04 text_embed=0.08 | spike L=-0.00 G=1.75 | nonfinite=0 +step 36800 | loss 5.7377 | text 2.1306 audio 5.3116 | grad 2.060 | 1.45s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.07 text_embed=0.11 | spike L=+0.01 G=1.05 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7628,8 +7609,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35825 | loss 5.7072 | text 2.1152 audio 5.2842 | grad 1.769 | 1.54s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.36 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=+0.00 G=0.79 | nonfinite=0 +step 36825 | loss 5.7254 | text 2.1173 audio 5.3019 | grad 2.148 | 1.52s/step | peak -1.0G | host_rss 46.9G + [diag] gradnorm depth=0.31 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=+0.00 G=1.09 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7655,8 +7636,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35850 | loss 5.7080 | text 2.1325 audio 5.2815 | grad 1.590 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.43 lora=0.89 model_audio_embed=0.05 projection=0.09 text_embed=0.11 | spike L=+0.00 G=0.73 | nonfinite=0 +step 36850 | loss 5.6654 | text 2.0803 audio 5.2493 | grad 1.928 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.35 lora=0.92 model_audio_embed=0.05 projection=0.07 text_embed=0.14 | spike L=-0.01 G=0.97 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7682,8 +7663,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35875 | loss 5.7191 | text 2.1272 audio 5.2937 | grad 2.377 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.28 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=+0.00 G=1.12 | nonfinite=0 +step 36875 | loss 5.7116 | text 2.1150 audio 5.2885 | grad 1.605 | 1.45s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.05 projection=0.08 text_embed=0.12 | spike L=+0.00 G=0.81 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7709,8 +7690,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35900 | loss 5.6961 | text 2.0964 audio 5.2768 | grad 1.901 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.33 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.12 | spike L=-0.00 G=0.89 | nonfinite=0 +step 36900 | loss 5.6732 | text 2.0938 audio 5.2545 | grad 1.952 | 1.52s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.37 lora=0.92 model_audio_embed=0.04 projection=0.07 text_embed=0.11 | spike L=-0.00 G=1.00 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7736,8 +7717,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35925 | loss 5.6746 | text 2.1152 audio 5.2516 | grad 1.636 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.40 lora=0.91 model_audio_embed=0.04 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.77 | nonfinite=0 +step 36925 | loss 5.6833 | text 2.0830 audio 5.2667 | grad 2.282 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.30 lora=0.95 model_audio_embed=0.04 projection=0.06 text_embed=0.11 | spike L=-0.00 G=1.17 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7763,8 +7744,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35950 | loss 5.6555 | text 2.1133 audio 5.2328 | grad 1.572 | 1.58s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.41 lora=0.90 model_audio_embed=0.05 projection=0.08 text_embed=0.10 | spike L=-0.01 G=0.76 | nonfinite=0 +step 36950 | loss 5.6657 | text 2.0712 audio 5.2514 | grad 2.144 | 1.46s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.32 lora=0.94 model_audio_embed=0.05 projection=0.06 text_embed=0.11 | spike L=-0.01 G=1.08 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7790,8 +7771,8 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 35975 | loss 5.6890 | text 2.1208 audio 5.2648 | grad 2.115 | 1.55s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.06 text_embed=0.12 | spike L=-0.00 G=1.05 | nonfinite=0 +step 36975 | loss 5.6745 | text 2.0822 audio 5.2581 | grad 2.598 | 1.49s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.25 lora=0.96 model_audio_embed=0.04 projection=0.05 text_embed=0.11 | spike L=-0.00 G=1.30 | nonfinite=0 sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. @@ -7817,13 +7798,13 @@ sys:1: UserWarning: Full backward hook is firing when gradients are computed wit sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. sys:1: UserWarning: Full backward hook is firing when gradients are computed with respect to module outputs since no inputs require gradients. See https://docs.pytorch.org/docs/main/generated/torch.nn.Module.html#torch.nn.Module.register_full_backward_hook for more details. -step 36000 | loss 5.7153 | text 2.1097 audio 5.2934 | grad 2.235 | 1.49s/step | peak -1.0G | host_rss 46.9G - [diag] gradnorm depth=0.27 lora=0.95 model_audio_embed=0.05 projection=0.05 text_embed=0.11 | spike L=+0.00 G=1.10 | nonfinite=0 - running validation at step 36000... - val/composite=2.9897 (text=0.5974 audio=4.5846) val/loss=4.7041 cb0_acc=39.2% text_acc=91.8% - [val] per-codebook acc: cb0=39.2% cb1=20.1% cb2=17.0% cb3=11.2% cb4=8.4% cb5=6.9% cb6=5.9% cb7=5.7% - [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9894) +step 37000 | loss 5.6766 | text 2.0915 audio 5.2583 | grad 2.042 | 1.47s/step | peak -1.0G | host_rss 47.0G + [diag] gradnorm depth=0.34 lora=0.93 model_audio_embed=0.04 projection=0.07 text_embed=0.09 | spike L=-0.00 G=0.99 | nonfinite=0 + running validation at step 37000... + val/composite=2.9927 (text=0.5950 audio=4.5912) val/loss=4.7101 cb0_acc=39.3% text_acc=91.9% + [val] per-codebook acc: cb0=39.3% cb1=20.2% cb2=16.9% cb3=11.1% cb4=8.4% cb5=6.9% cb6=5.8% cb7=5.7% + [early-stop] no val improvement (>0.0); patience 9/10 (best 2.9854) /opt/tinyaya/.venv/lib/python3.12/site-packages/peft/utils/save_and_load.py:356: UserWarning: Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`. warnings.warn("Setting `save_embedding_layers` to `True` as embedding layers found in `target_modules`.") -[ckpt] uploading /tmp/ckpt_skck_i9h/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_036000 +[ckpt] uploading /tmp/ckpt_w5h8lisx/* -> gs://tinyaya-stage2-eu/checkpoints/stage2-v6e16-mh-v03-r2/step_037000 pushed to https://huggingface.co/tiny-aya-translate/tr-hi-s2st-v0.3