the-jb's picture
Upload folder using huggingface_hub
ab6553c verified
Raw
History Blame Contribute Delete
12 kB
[2025-05-03 01:02:31,956][model][INFO] - Setting pad_token as eos token: <|eot_id|>
[2025-05-03 01:02:31,959][evaluator][INFO] - Evaluations stored in the experiment directory: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals
[2025-05-03 01:02:31,960][evaluator][INFO] - ***** Running TOFU evaluation suite *****
[2025-05-03 01:02:31,961][evaluator][INFO] - Fine-grained evaluations will be saved to: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-03 01:02:31,961][evaluator][INFO] - Aggregated evaluations will be summarised in: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_SUMMARY.json
[2025-05-03 01:02:35,578][metrics][INFO] - Loading evaluations from saves/eval/tofu_Llama-3.1-8B-Instruct_retain90/TOFU_EVAL.json
[2025-05-03 01:02:35,589][metrics][INFO] - Evaluating forget_Q_A_PARA_Prob
[2025-05-03 01:02:49,166][metrics][INFO] - Loading evaluations from saves/eval/tofu_Llama-3.1-8B-Instruct_retain90/TOFU_EVAL.json
[2025-05-03 01:02:49,178][metrics][INFO] - Evaluating forget_Q_A_PERT_Prob
[2025-05-03 01:03:43,684][metrics][INFO] - Loading evaluations from saves/eval/tofu_Llama-3.1-8B-Instruct_retain90/TOFU_EVAL.json
[2025-05-03 01:03:43,695][metrics][INFO] - Evaluating forget_truth_ratio
[2025-05-03 01:03:43,696][metrics][INFO] - Loading evaluations from saves/eval/tofu_Llama-3.1-8B-Instruct_retain90/TOFU_EVAL.json
[2025-05-03 01:03:43,705][metrics][INFO] - Evaluating forget_quality
[2025-05-03 01:03:43,706][evaluator][INFO] - Result for metric forget_quality: 0.09352461409425458
[2025-05-03 01:03:45,427][metrics][INFO] - Evaluating forget_Q_A_Prob
[2025-05-03 01:03:55,895][evaluator][INFO] - Result for metric forget_Q_A_Prob: 0.07000148265040025
[2025-05-03 01:03:58,128][metrics][INFO] - Evaluating forget_Q_A_ROUGE
[2025-05-03 01:05:04,124][evaluator][INFO] - Result for metric forget_Q_A_ROUGE: 0.28561408364618274
[2025-05-03 01:05:06,555][metrics][INFO] - Evaluating retain_Q_A_Prob
[2025-05-03 01:05:17,815][metrics][INFO] - Evaluating retain_Q_A_ROUGE
[2025-05-03 01:05:51,740][metrics][INFO] - Evaluating retain_Q_A_PARA_Prob
[2025-05-03 01:06:03,901][metrics][INFO] - Evaluating retain_Q_A_PERT_Prob
[2025-05-03 01:06:53,625][metrics][INFO] - Evaluating retain_Truth_Ratio
[2025-05-03 01:06:56,023][metrics][INFO] - Evaluating ra_Q_A_Prob
[2025-05-03 01:06:59,254][metrics][INFO] - Evaluating ra_Q_A_PERT_Prob
[2025-05-03 01:07:04,050][metrics][INFO] - Evaluating ra_Q_A_Prob_normalised
[2025-05-03 01:07:05,784][metrics][INFO] - Evaluating ra_Q_A_ROUGE
[2025-05-03 01:07:13,999][metrics][INFO] - Skipping ra_Truth_Ratio's precompute ra_Q_A_Prob, already evaluated.
[2025-05-03 01:07:13,999][metrics][INFO] - Skipping ra_Truth_Ratio's precompute ra_Q_A_PERT_Prob, already evaluated.
[2025-05-03 01:07:13,999][metrics][INFO] - Evaluating ra_Truth_Ratio
[2025-05-03 01:07:15,786][metrics][INFO] - Evaluating wf_Q_A_Prob
[2025-05-03 01:07:18,696][metrics][INFO] - Evaluating wf_Q_A_PERT_Prob
[2025-05-03 01:07:23,782][metrics][INFO] - Evaluating wf_Q_A_Prob_normalised
[2025-05-03 01:07:25,545][metrics][INFO] - Evaluating wf_Q_A_ROUGE
[2025-05-03 01:07:36,402][metrics][INFO] - Skipping wf_Truth_Ratio's precompute wf_Q_A_Prob, already evaluated.
[2025-05-03 01:07:36,402][metrics][INFO] - Skipping wf_Truth_Ratio's precompute wf_Q_A_PERT_Prob, already evaluated.
[2025-05-03 01:07:36,402][metrics][INFO] - Evaluating wf_Truth_Ratio
[2025-05-03 01:07:36,402][metrics][INFO] - Evaluating model_utility
[2025-05-03 01:07:36,403][evaluator][INFO] - Result for metric model_utility: 0.6140003792455105
[2025-05-03 01:07:39,695][metrics][INFO] - Loading evaluations from saves/eval/tofu_Llama-3.1-8B-Instruct_retain90/TOFU_EVAL.json
[2025-05-03 01:07:39,708][metrics][INFO] - Evaluating mia_min_k
[2025-05-03 01:07:57,083][metrics][INFO] - Loading evaluations from saves/eval/tofu_Llama-3.1-8B-Instruct_retain90/TOFU_EVAL.json
[2025-05-03 01:07:57,094][metrics][INFO] - Evaluating privleak
[2025-05-03 01:07:57,094][evaluator][INFO] - Result for metric privleak: 40.278085859319106
[2025-05-03 01:07:58,893][metrics][INFO] - Evaluating extraction_strength
[2025-05-03 01:08:07,928][evaluator][INFO] - Result for metric extraction_strength: 0.06814847463944672
[2025-05-09 18:55:32,743][model][INFO] - Setting pad_token as eos token: <|eot_id|>
[2025-05-09 18:55:32,747][evaluator][INFO] - Evaluations stored in the experiment directory: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals
[2025-05-09 18:55:32,748][evaluator][INFO] - Loading existing evaluations from saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-09 18:55:32,801][evaluator][INFO] - ***** Running TOFU evaluation suite *****
[2025-05-09 18:55:32,801][evaluator][INFO] - Fine-grained evaluations will be saved to: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-09 18:55:32,801][evaluator][INFO] - Aggregated evaluations will be summarised in: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_SUMMARY.json
[2025-05-09 18:55:32,801][evaluator][INFO] - Skipping forget_quality, already evaluated.
[2025-05-09 18:55:32,802][evaluator][INFO] - Result for metric forget_quality: 0.09352461409425458
[2025-05-09 18:55:32,810][evaluator][INFO] - Skipping forget_Q_A_Prob, already evaluated.
[2025-05-09 18:55:32,810][evaluator][INFO] - Result for metric forget_Q_A_Prob: 0.07000148265040025
[2025-05-09 18:55:32,812][evaluator][INFO] - Skipping forget_Q_A_ROUGE, already evaluated.
[2025-05-09 18:55:32,812][evaluator][INFO] - Result for metric forget_Q_A_ROUGE: 0.28561408364618274
[2025-05-09 18:55:32,814][evaluator][INFO] - Skipping model_utility, already evaluated.
[2025-05-09 18:55:32,814][evaluator][INFO] - Result for metric model_utility: 0.6140003792455105
[2025-05-09 18:55:32,815][evaluator][INFO] - Skipping privleak, already evaluated.
[2025-05-09 18:55:32,815][evaluator][INFO] - Result for metric privleak: 40.278085859319106
[2025-05-09 18:55:32,816][evaluator][INFO] - Skipping extraction_strength, already evaluated.
[2025-05-09 18:55:32,816][evaluator][INFO] - Result for metric extraction_strength: 0.06814847463944672
[2025-05-09 18:59:59,459][model][INFO] - Setting pad_token as eos token: <|eot_id|>
[2025-05-09 18:59:59,461][evaluator][INFO] - Evaluations stored in the experiment directory: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals
[2025-05-09 18:59:59,462][evaluator][INFO] - Loading existing evaluations from saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-09 18:59:59,551][evaluator][INFO] - ***** Running TOFU evaluation suite *****
[2025-05-09 18:59:59,551][evaluator][INFO] - Fine-grained evaluations will be saved to: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-09 18:59:59,551][evaluator][INFO] - Aggregated evaluations will be summarised in: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_SUMMARY.json
[2025-05-09 18:59:59,551][evaluator][INFO] - Skipping forget_quality, already evaluated.
[2025-05-09 18:59:59,551][evaluator][INFO] - Result for metric forget_quality: 0.09352461409425458
[2025-05-09 18:59:59,559][evaluator][INFO] - Skipping forget_Q_A_Prob, already evaluated.
[2025-05-09 18:59:59,559][evaluator][INFO] - Result for metric forget_Q_A_Prob: 0.07000148265040025
[2025-05-09 18:59:59,562][evaluator][INFO] - Skipping forget_Q_A_ROUGE, already evaluated.
[2025-05-09 18:59:59,562][evaluator][INFO] - Result for metric forget_Q_A_ROUGE: 0.28561408364618274
[2025-05-09 18:59:59,565][evaluator][INFO] - Skipping model_utility, already evaluated.
[2025-05-09 18:59:59,565][evaluator][INFO] - Result for metric model_utility: 0.6140003792455105
[2025-05-09 18:59:59,568][evaluator][INFO] - Skipping privleak, already evaluated.
[2025-05-09 18:59:59,568][evaluator][INFO] - Result for metric privleak: 40.278085859319106
[2025-05-09 18:59:59,570][evaluator][INFO] - Skipping extraction_strength, already evaluated.
[2025-05-09 18:59:59,570][evaluator][INFO] - Result for metric extraction_strength: 0.06814847463944672
[2025-05-09 20:32:18,055][model][INFO] - Setting pad_token as eos token: <|eot_id|>
[2025-05-09 20:32:18,060][evaluator][INFO] - Evaluations stored in the experiment directory: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals
[2025-05-09 20:32:18,062][evaluator][INFO] - Loading existing evaluations from saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-09 20:32:18,114][evaluator][INFO] - ***** Running TOFU evaluation suite *****
[2025-05-09 20:32:18,114][evaluator][INFO] - Fine-grained evaluations will be saved to: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-09 20:32:18,114][evaluator][INFO] - Aggregated evaluations will be summarised in: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_SUMMARY.json
[2025-05-09 20:32:18,114][evaluator][INFO] - Skipping forget_quality, already evaluated.
[2025-05-09 20:32:18,114][evaluator][INFO] - Result for metric forget_quality: 0.09352461409425458
[2025-05-09 20:32:18,350][evaluator][INFO] - Skipping forget_Q_A_Prob, already evaluated.
[2025-05-09 20:32:18,350][evaluator][INFO] - Result for metric forget_Q_A_Prob: 0.07000148265040025
[2025-05-09 20:32:18,361][evaluator][INFO] - Skipping forget_Q_A_ROUGE, already evaluated.
[2025-05-09 20:32:18,361][evaluator][INFO] - Result for metric forget_Q_A_ROUGE: 0.28561408364618274
[2025-05-09 20:32:18,364][evaluator][INFO] - Skipping model_utility, already evaluated.
[2025-05-09 20:32:18,364][evaluator][INFO] - Result for metric model_utility: 0.6140003792455105
[2025-05-09 20:32:18,367][evaluator][INFO] - Skipping privleak, already evaluated.
[2025-05-09 20:32:18,367][evaluator][INFO] - Result for metric privleak: 40.278085859319106
[2025-05-09 20:32:18,369][evaluator][INFO] - Skipping extraction_strength, already evaluated.
[2025-05-09 20:32:18,369][evaluator][INFO] - Result for metric extraction_strength: 0.06814847463944672
[2025-05-10 18:48:58,416][model][INFO] - Setting pad_token as eos token: <|eot_id|>
[2025-05-10 18:48:58,419][evaluator][INFO] - Evaluations stored in the experiment directory: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals
[2025-05-10 18:48:58,420][evaluator][INFO] - Loading existing evaluations from saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-10 18:48:58,465][evaluator][INFO] - ***** Running TOFU evaluation suite *****
[2025-05-10 18:48:58,465][evaluator][INFO] - Fine-grained evaluations will be saved to: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_EVAL.json
[2025-05-10 18:48:58,465][evaluator][INFO] - Aggregated evaluations will be summarised in: saves/unlearn/tofu_Llama-3.1-8B-Instruct_forget10_NPO/evals/TOFU_SUMMARY.json
[2025-05-10 18:48:58,465][evaluator][INFO] - Skipping forget_quality, already evaluated.
[2025-05-10 18:48:58,466][evaluator][INFO] - Result for metric forget_quality: 0.09352461409425458
[2025-05-10 18:48:58,534][evaluator][INFO] - Skipping forget_Q_A_Prob, already evaluated.
[2025-05-10 18:48:58,534][evaluator][INFO] - Result for metric forget_Q_A_Prob: 0.07000148265040025
[2025-05-10 18:48:58,543][evaluator][INFO] - Skipping forget_Q_A_ROUGE, already evaluated.
[2025-05-10 18:48:58,543][evaluator][INFO] - Result for metric forget_Q_A_ROUGE: 0.28561408364618274
[2025-05-10 18:48:58,559][evaluator][INFO] - Skipping model_utility, already evaluated.
[2025-05-10 18:48:58,559][evaluator][INFO] - Result for metric model_utility: 0.6140003792455105
[2025-05-10 18:48:58,565][evaluator][INFO] - Skipping privleak, already evaluated.
[2025-05-10 18:48:58,565][evaluator][INFO] - Result for metric privleak: 40.278085859319106
[2025-05-10 18:48:58,568][evaluator][INFO] - Skipping extraction_strength, already evaluated.
[2025-05-10 18:48:58,568][evaluator][INFO] - Result for metric extraction_strength: 0.06814847463944672