diff --git "a/artifacts/chunks.json" "b/artifacts/chunks.json" new file mode 100644--- /dev/null +++ "b/artifacts/chunks.json" @@ -0,0 +1,17966 @@ +[ + { + "text": "Milad Saeedi\nUniversity of Toronto, Canada\n(437) 326-7991 | Email |LinkedIn Profile |Google Scholar | Git Hub | Hugging Face\nSUMMARY\nGeospatial Data Science | Machine Learning Engineer | Spatial AI & Computer Vision | High-\nDimensional Spatial Data | Generative AI & LLMs | End-to-End ML Workflows\nApplied Machine Learning Scientist with PhD-level expertise in Machine Learning, Computer Vision, Spatial AI, Multimodal\nLearning, and Generative AI. Experienced in designing end-to-end AI systems for large-scale geospatial, visual, and sensor\ndata, from data engineering and feature extraction to model development, fine-tuning, evaluation, and deployment.\nSkilled in deep learning (PyTorch, TensorFlow, CNNs, Transformers, and LLMs), with a proven record of delivering\nproduction-oriented AI solutions and high-impact research.\nTECHNICAL SKILLS\nProgramming: Python (PyTorch, Scikit-learn, NumPy, Pandas, XGBoost), SQL, R, C++", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "0-0", + "page": 1, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "LLMs), with a proven record of delivering\nproduction-oriented AI solutions and high-impact research.\nTECHNICAL SKILLS\nProgramming: Python (PyTorch, Scikit-learn, NumPy, Pandas, XGBoost), SQL, R, C++\nMachine Learning: Supervised & Unsupervised Learning, Regression, Classification, Clustering, Feature Engineering,\nModel Selection, Ensemble Learning, Model Evaluation & Interpretability\nDeep Learning & Computer Vision: CNNs, RNNs, LSTMs, GRUs, Transformers, Transfer Learning, Object Detection\n(YOLO), Image Classification, Semantic Segmentation, Representation Learning, VAEs, GANs\nGenerative AI & LLMs: Hugging Face Transformers, Encoder-only (BERT), Decoder-only (GPT), Encoder–Decoder Models\n(T5, BART), Chat Templates, Embeddings, Prompt Engineering, Supervised Fine-Tuning (SFT), LoRA (PEFT), TRL\n(SFTTrainer), GRPO, Agentic AI\nData Engineering: ETL pipelines, Spark, large-scale data processing, geospatial & spatiotemporal data pipelines, GPS–\nimage–sensor data integration", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "0-1", + "page": 1, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "T), LoRA (PEFT), TRL\n(SFTTrainer), GRPO, Agentic AI\nData Engineering: ETL pipelines, Spark, large-scale data processing, geospatial & spatiotemporal data pipelines, GPS–\nimage–sensor data integration\nCloud, Platforms & Tools: AWS (SageMaker, S3, Lambda), Databricks, Git, Linux, MATLAB\nApplied ML (Geospatial Data): GeoPandas, ArcGIS, QGIS, rasterio, feature and model selection under autocorrelation,\nspatial & spatiotemporal cross-validation, objective-aware validation, high-dimensional raster & spatial datasets\nVisualization: Matplotlib, Contextily, Seaborn, ggplot2, GIS mapping, SHAP interpretation\nAREAS OF EXPERTISE\nMachine Learning | Deep Learning | Computer Vision | Generative AI & LLMs | Multimodal AI (Imagery, GPS & Sensor\nData)| Geospatial AI | End-to-End ML Pipelines | Model Development & Evaluation | Scalable AI Systems | Spatial &\nSpatiotemporal Modeling | Large-Scale Data Processing\nPROFESSIONAL WORK EXPERIENCE", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "0-2", + "page": 1, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "or\nData)| Geospatial AI | End-to-End ML Pipelines | Model Development & Evaluation | Scalable AI Systems | Spatial &\nSpatiotemporal Modeling | Large-Scale Data Processing\nPROFESSIONAL WORK EXPERIENCE\nMachine Learning Engineer / Data Scientist – University of Toronto, Toronto, Canada | Sep 2021 – Present\n•\nDesigned a geospatial ML framework for robust and transferable spatial predictions | Paper (ES&T, 2026)\no Designed a target-oriented validation framework for autocorrelated geospatial data by aligning model\nevaluation with spatial structure, improving model reliability and regional transferability.\no Reduced spatial generalization error by 63% (MAPE: 216% → 79%) by integrating autocorrelation-aware\nfeature engineering, feature selection, and structured cross-validation into the modeling pipeline.\n•\nCity-scale geospatial ML using courier-truck mobile sensing | Paper (TRR, 2025)\no Built scalable data pipelines and XGBoost models on 1M+ mobile sensing records, improving city-scale", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "0-3", + "page": 1, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "ling pipeline.\n•\nCity-scale geospatial ML using courier-truck mobile sensing | Paper (TRR, 2025)\no Built scalable data pipelines and XGBoost models on 1M+ mobile sensing records, improving city-scale\npollution prediction and spatial generalization.\n•\nMultimodal geospatial modeling from 360° imagery and air quality data | Paper (submitted)\no Developed a YOLOv8 pipeline to extract fine-grained vehicle categories from 360° imagery and integrated\nmultimodal traffic features into ML models, improving causal interpretation and model transferability.\n1", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "0-4", + "page": 1, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "•\nApplied geospatial ML to large-scale spatial datasets to generate actionable insights for decision-making\no Applied geospatial ML to large-scale spatial datasets, integrating multimodal data to extract patterns\nand generate actionable insights for decision-making, contributing to 6 peer-reviewed publications.\n•\nApplied Deep Learning & Production ML Workflows:\no Applied end-to-end ML workflows on high-dimensional and time-series data, performing feature\nengineering and developing supervised and unsupervised models for data-driven predictions.\no Built scalable ML pipelines in Databricks (PySpark), enabling feature engineering and training of\nproduction-ready classification models.\no Designed and implemented PyTorch-based neural networks for binary and multiclass classification on\nstructured and image datasets, optimizing network architectures, loss functions, and training workflows\nfor robust generalization.", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "1-0", + "page": 2, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "ch-based neural networks for binary and multiclass classification on\nstructured and image datasets, optimizing network architectures, loss functions, and training workflows\nfor robust generalization.\no Developed and fine-tuned CNN-based models with transfer learning for image classification by\nleveraging pre-trained architectures, optimizing training pipelines, and applying them to hand gesture\nrecognition datasets, achieving improved accuracy and efficient model convergence on test data.\no Developed PyTorch-based generative deep learning models, including autoencoders and conditional\nGANs, for image reconstruction, synthesis, and colorization by implementing end-to-end training\npipelines and evaluating alternative network architectures and training strategies.\no Built and optimized RNN, LSTM, and GRU models for time-series and textual sequence data (e.g.,\ntweets) by developing PyTorch training pipelines, achieving strong predictive performance and stable\nconvergence.", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "1-1", + "page": 2, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "optimized RNN, LSTM, and GRU models for time-series and textual sequence data (e.g.,\ntweets) by developing PyTorch training pipelines, achieving strong predictive performance and stable\nconvergence.\no Designed and deployed a YOLO-based traffic light detection system for real-time inference.\n•\nLLM & Generative AI (Hugging Face):\no Developed and published transformer- and LLM-based applications on the Hugging Face Hub, including\ntext classification, question answering, translation, summarization, token classification, causal language\nmodeling, supervised fine-tuning (SFT), LoRA (PEFT), GRPO, and AI agents.\nTeaching Assistant – University of Toronto | Sep 2021 – Dec 2025\n● Delivered lectures in Machine Learning and Data Mining; supported students in applied modeling and\nquantitative problem solving.\nAir Quality Analyst — Tehran Air Quality Control Company| Jan 2017 – Dec 2017\n● Developed GIS-based emission inventories and performed spatial analysis of pollution patterns to support", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "1-2", + "page": 2, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "oblem solving.\nAir Quality Analyst — Tehran Air Quality Control Company| Jan 2017 – Dec 2017\n● Developed GIS-based emission inventories and performed spatial analysis of pollution patterns to support\nurban planning initiatives.\nResearch Assistant — Sharif University of Technology (M.Sc.) | Sep 2014– Dec 2016\n● Developed GIS-based emission inventories and applied numerical air-quality models to analyze spatiotemporal\npollution patterns for municipal planning and policy support.\nEDUCATION\nPh.D., Civil & Mineral Engineering, University of Toronto (2021–Present),\n•\nFocus: Research in geospatial machine learning, computer vision, and large-scale spatial modeling using\nmultimodal data (360° imagery, GPS, environmental sensors).\nM.Sc., Mechanical Engineering, Sharif University of Technology (2014–2017)\nB.Sc., Mechanical Engineering, Isfahan University of Technology (2010–2014)\nAWARDS & SCHOLARSHIPS", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "1-3", + "page": 2, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "S, environmental sensors).\nM.Sc., Mechanical Engineering, Sharif University of Technology (2014–2017)\nB.Sc., Mechanical Engineering, Isfahan University of Technology (2010–2014)\nAWARDS & SCHOLARSHIPS\nRichard Soberman Graduate Student Fellowship (2022, 2024, 2025 | Heavy Construction Association of Toronto\nGraduate Scholarship in Intelligent Transportation Systems (2023) | Dr. Mazen Hassounah Graduate Scholarship in\nMass Events Transportation and Crowd Management (2022)\nPUBLICATIONS\n● 15 peer-reviewed publications, including 5 first-author papers, in machine learning, statistical modeling,\ntransportation and travel survey analytics, computer vision, and large-scale spatial modeling.\n● 7+ presentations at major conferences and symposia, including the Transportation Research Board (TRB) Annual\nMeeting.\n2", + "source": "Milad Saeedi - GEO.pdf", + "file_type": "pdf", + "chunk_id": "1-4", + "page": 2, + "document_type": "cv_geospatial", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - GEO.pdf", + "section": null + }, + { + "text": "Milad Saeedi\nUniversity of Toronto, Canada\n(437) 326-7991 | Email |LinkedIn Profile |Google Scholar |Git Hub |Hugging Face\nSUMMARY\nMachine Learning | Generative AI & LLMs | Spatial AI & Computer Vision | Multimodal\nLearning | End-to-End AI Systems\nApplied Machine Learning Scientist with PhD-level expertise in machine learning, deep learning, computer vision,\ngenerative AI, and multimodal AI. Experienced building end-to-end AI systems for geospatial, visual, and language data,\nwith hands-on expertise in modern deep learning, natural language processing, and large language model development.\nProven ability to translate research into scalable, real-world AI solutions.\nTECHNICAL SKILLS\nProgramming: Python (NumPy, Pandas, Scikit-learn, XGBoost, PyTorch, TensorFlow), SQL, R, C++\nMachine Learning: Supervised & Unsupervised Learning, Regression, Classification, Clustering, Feature Engineering,\nModel Selection, Ensemble Learning, Model Evaluation & Interpretability", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "2-0", + "page": 1, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "SQL, R, C++\nMachine Learning: Supervised & Unsupervised Learning, Regression, Classification, Clustering, Feature Engineering,\nModel Selection, Ensemble Learning, Model Evaluation & Interpretability\nDeep Learning: CNNs, RNNs, LSTMs, GRUs, Transformers, Transfer Learning, Object Detection (YOLO), Image\nClassification, Semantic Segmentation, Representation Learning, VAEs, GANs\nGenerative AI & LLM: Hugging Face Transformers, Encoder-only (BERT), Decoder-only (GPT), Encoder–Decoder Models\n(T5, BART), Chat Templates, Embeddings, Prompt Engineering, Supervised Fine-Tuning (SFT), LoRA (PEFT), TRL\n(SFTTrainer), GRPO, Agentic AI\nData Engineering: ETL pipelines, Spark, large-scale data processing, GPS–image–sensor synchronization and processing\nCloud, Platforms & Tools: AWS (SageMaker, S3, Lambda), Databricks, Git, Linux, MATLAB\nApplied ML (Geospatial Data): GeoPandas, ArcGIS, QGIS, rasterio, spatial feature engineering, location intelligence", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "2-1", + "page": 1, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "oud, Platforms & Tools: AWS (SageMaker, S3, Lambda), Databricks, Git, Linux, MATLAB\nApplied ML (Geospatial Data): GeoPandas, ArcGIS, QGIS, rasterio, spatial feature engineering, location intelligence\nVisualization: Matplotlib, Seaborn, ggplot2, GIS mapping, SHAP interpretation\nAREAS OF EXPERTISE\nMachine Learning | Deep Learning | Computer Vision | Generative AI & LLMs | Scalable ML Pipelines | Distributed Data\nProcessing (Spark) | Feature Engineering | Model Validation & Generalization\nPROFESSIONAL WORK EXPERIENCE\nResearch Scientist / Machine Learning Engineer– University of Toronto, Toronto, Canada | Sep 2021 – Present\n•\nDeveloped a framework to bridge the gap between data reproduction and prediction| Paper (ES&T, 2026)\no Designed a novel target-oriented framework to address overfitting caused by spatial autocorrelation by\naligning model validation with data structure, improving model reliability and transferability.", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "2-2", + "page": 1, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "Designed a novel target-oriented framework to address overfitting caused by spatial autocorrelation by\naligning model validation with data structure, improving model reliability and transferability.\no Reduced high spatial generalization error caused by spatial autocorrelation by coupling feature selection\nand hyperparameter tuning with structured cross validation, decreasing MAPE from 216% to 79%.\n•\nMobile air pollution mapping using courier-truck sensors| Paper (TRR, 2025)\no Built scalable data pipelines and XGBoost models on 1M+ mobile sensing records, improving city-scale\npollution prediction and spatial generalization.\n•\nMultimodal modeling using 360° imagery and sensor data | Paper (submitted)\no Developed a YOLOv8 pipeline to extract fine-grained vehicle categories from 360° imagery and integrated\nmultimodal traffic features into ML models, improving causal interpretation and model transferability.\n•", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "2-3", + "page": 1, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "a YOLOv8 pipeline to extract fine-grained vehicle categories from 360° imagery and integrated\nmultimodal traffic features into ML models, improving causal interpretation and model transferability.\n•\nApplied data science methods to city-scale exposure and environmental health modeling\no Conducted data-driven analysis of mobile-sensor and land-use datasets to investigate disparities, wildfire\nimpacts, exposure-health relationships, and backcasting, contributing to 6 peer-reviewed publications.\n•\nApplied Deep Learning & Production ML Workflows:\n1", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "2-4", + "page": 1, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "o Applied end-to-end data science workflows to address real-world engineering and business problems\nusing high-dimensional, time-series, and retail transactional data with customer response (promotion)\nlabels by performing feature engineering and developing supervised and unsupervised ML models (e.g.,\nregression, trees, PCA), enabling accurate prediction and data-driven decision-making.\no Built ML pipelines in Databricks (PySpark) for retail transaction and promotion response modeling,\nperforming feature engineering and training classification models to enable scalable, production-ready\npredictions.\no Designed and implemented PyTorch-based neural networks for binary and multiclass classification tasks\nby developing forward/backpropagation workflows, tuning architectures and loss functions, and applying\nthem to structured and image datasets, achieving stable convergence and strong generalization\nperformance on validation and test data.", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "3-0", + "page": 2, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "workflows, tuning architectures and loss functions, and applying\nthem to structured and image datasets, achieving stable convergence and strong generalization\nperformance on validation and test data.\no Developed and fine-tuned CNN-based models with transfer learning for image classification by leveraging\npre-trained architectures, optimizing training pipelines, and applying them to hand gesture recognition\ndatasets, achieving improved accuracy and efficient model convergence on test data.\no Developed generative deep learning models (Autoencoders, VAEs, GANs, and cGANs) in PyTorch for\nrepresentation learning, image generation, image-to-image translation (image colorization), and\nevaluating model robustness to adversarial attacks.\no Built and optimized RNN, LSTM, and GRU models for time-series and textual sequence data (tweets) by\ndeveloping PyTorch training pipelines, achieving strong predictive performance and stable convergence.", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "3-1", + "page": 2, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "lt and optimized RNN, LSTM, and GRU models for time-series and textual sequence data (tweets) by\ndeveloping PyTorch training pipelines, achieving strong predictive performance and stable convergence.\no Designed and deployed a YOLO-based traffic light detection system for real-time inference.\n•\nLLM & Generative AI:\no Developed and published transformer- and LLM-based applications on the Hugging Face Hub, including\ntext classification, question answering, translation, summarization, token classification, causal language\nmodeling, supervised fine-tuning (SFT), LoRA (PEFT), GRPO, and AI agents.\nTeaching Assistant – University of Toronto, Toronto, Canada| Sep 2021 – Dec 2025\n● Delivered lectures and led tutorials in Machine Learning, Data Mining, and quantitative engineering courses;\nmentored students in algorithm implementation and model evaluation.\nData Scientist — Environmental & Geospatial Data — Tehran Air Quality Control Company| Jan 2017 – Dec 2020", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "3-2", + "page": 2, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "ive engineering courses;\nmentored students in algorithm implementation and model evaluation.\nData Scientist — Environmental & Geospatial Data — Tehran Air Quality Control Company| Jan 2017 – Dec 2020\n● Analyzed large-scale spatiotemporal air quality data to support municipal decision-making by developing GIS-\nbased data pipelines and applying numerical and statistical models, generating actionable insights for urban\npollution management and policy planning.\nResearch Assistant — Sharif University of Technology (M.Sc.), Iran | Sep 2014 – Dec 2016\n● Conducted computational modeling of transport and dispersion processes using traffic, GPS, and meteorological\ndatasets.\nEDUCATION\nPh.D., Civil & Mineral Engineering, University of Toronto (2021–Present),\n•\nFocus: Machine learning, computer vision, and large-scale geospatial modeling using multimodal data.\nM.Sc., Mechanical Engineering, Sharif University of Technology (2014–2017)", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "3-3", + "page": 2, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "ronto (2021–Present),\n•\nFocus: Machine learning, computer vision, and large-scale geospatial modeling using multimodal data.\nM.Sc., Mechanical Engineering, Sharif University of Technology (2014–2017)\nB.Sc., Mechanical Engineering, Isfahan University of Technology (2010–2014)\nPUBLICATIONS\n● 15 peer-reviewed publications, including 5 first-author papers, in applied machine learning, data science,\ncomputer vision, large-scale data analysis, and numerical modeling.\n● 7+ presentations at major conferences and symposia, including the Transportation Research Board (TRB) Annual\nMeeting.\nAWARDS & SCHOLARSHIPS\nRichard Soberman Graduate Student Fellowship (2022, 2024, 2025 | Heavy Construction Association of Toronto\nGraduate Scholarship in Intelligent Transportation Systems (2023) | Dr. Mazen Hassounah Graduate Scholarship in\nMass Events Transportation and Crowd Management (2022)\n2", + "source": "Milad Saeedi - ML.pdf", + "file_type": "pdf", + "chunk_id": "3-4", + "page": 2, + "document_type": "cv_machine_learning", + "page_header": null, + "repository": null, + "relative_path": "Milad Saeedi - ML.pdf", + "section": null + }, + { + "text": "Bridging the Gap Between Data Reproduction and Prediction: The\nImpact of Feature Selection and Cross-Validation Strategies on\nPrediction of Ambient Ultrafine Particles Collected with Mobile\nMonitoring\n\nMilad Saeedi, Jad Zalzal, Arman Ganji, Junshi Xu, Sebastian D. Goodfellow, and Marianne Hatzopoulou*\n\nCite This: https://doi.org/10.1021/acs.est.5c12601\nRead Online\n\nACCESS\nMetrics & More\nArticle Recommendations\n*\nsı\nSupporting Information\n\nABSTRACT: Reliable exposure assessment is vital for epidemio-\nlogical research, but weaknesses in land-use regression (LUR)\nmodels undermine its validity. Using mobile ultrafine particle\n(UFP) data in Toronto, we compared LUR models trained under\nrandom, spatial, temporal, and spatiotemporal cross-validation\n(CV), with and without forward feature selection (FFS). Model\nhyperparameters and feature subsets were optimized within each\nCV scheme. Spatial CV folds were designed at fine scales to reflect\nUFP autocorrelation.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "4-0", + "page": 1, + "document_type": "research_paper", + "page_header": "pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "d without forward feature selection (FFS). Model\nhyperparameters and feature subsets were optimized within each\nCV scheme. Spatial CV folds were designed at fine scales to reflect\nUFP autocorrelation. Each approach was evaluated on a hold-out\ntest set, across CV schemes, and against independent stationary\nbackyard measurements. Models based on spatiotemporal CV\ncoupled with FFS were able to reduce overfitting, improve\ngeneralization, and produce stable exposure surfaces. These\nsurfaces avoided the spatial artifacts and exaggerated variable effects typically seen in models trained with random CV. Models\ntuned with random CV overfit, performed poorly on independent samples, and were sensitive to outliers. The average percentage\nerror (APE) decreased from ∼217% for a model with random-CV to ∼79% with spatiotemporal CV and FFS. Our findings\ndemonstrate that proper alignment of model design with the data’s spatiotemporal structure and modeling objective ensures", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "4-1", + "page": 1, + "document_type": "research_paper", + "page_header": "pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "r a model with random-CV to ∼79% with spatiotemporal CV and FFS. Our findings\ndemonstrate that proper alignment of model design with the data’s spatiotemporal structure and modeling objective ensures\nreliability, minimizes data reproduction, and enables true prediction.\nKEYWORDS: UFP, mobile monitoring, land-use regression, cross-validation, feature selection, spatiotemporal\n\nSee https://pubs.acs.org/sharingguidelines for options on how to legitimately share published articles.\n\nDownloaded via UNIV OF TORONTO on February 5, 2026 at 17:16:06 (UTC).\n\n1. INTRODUCTION\n\nAir pollution is a major public health concern, especially in urban\nareas.1−3 Accurately assessing exposure necessitates advanced\nmeasurement and modeling methods, including mobile\nmonitoring, remote sensing, land-use regression (LUR), and\nmachine learning (ML).4,5 LUR modeling is widely used to\nestimate air pollution exposure in urban areas by leveraging\ngeographic and environmental predictors to map pollutant", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "4-2", + "page": 1, + "document_type": "research_paper", + "page_header": "pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "se regression (LUR), and\nmachine learning (ML).4,5 LUR modeling is widely used to\nestimate air pollution exposure in urban areas by leveraging\ngeographic and environmental predictors to map pollutant\nconcentrations at fine scales. Historically, LUR relied on data\nfrom a limited number of fixed monitoring stations,6,7 which\nconstrained its ability to capture localized variations.8 Recent\nprogress has broadened LUR’s scope, incorporating advanced\ndata collection methods and diverse predictors. Machine\nlearning approaches further enhance predictive accuracy.9,10\n\nMobile monitoring enhances exposure assessment for\nspatially varying air pollutants by covering many more locations\nand capturing the large spatial variability characteristic of urban\nareas.11 When enabled by portable low-cost sensors on vehicles,\nit also offers high-resolution measurements over wide areas with\nfewer devices.12 Mobile air quality measurements are often\nconducted at specific times and under varying meteorological", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "4-3", + "page": 1, + "document_type": "research_paper", + "page_header": "pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "s on vehicles,\nit also offers high-resolution measurements over wide areas with\nfewer devices.12 Mobile air quality measurements are often\nconducted at specific times and under varying meteorological\n\n© XXXX American Chemical Society\n\nA\n\nconditions, making it challenging to distinguish true spatial\npollution patterns from short-term fluctuations caused by factors\nlike traffic, weather, or localized emission sources.13,14 Ultrafine\nparticles (UFP) differ markedly from other pollutants such as\nPM2.5 or NO2 in both their physical behavior and chemical\nreactivity.15 Emitted primarily from traffic sources, UFPs have a\nhigh surface-area-to-mass ratio that enhances their volatility and\ncapacity to adsorb or desorb semivolatile compounds, resulting\nin rapid transformation and growth into larger particles in the\natmosphere.16 Because of their short atmospheric lifetime and\nstrong dependence on local traffic and meteorological\nconditions, UFPs exhibit steep spatial gradients and high", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "4-4", + "page": 1, + "document_type": "research_paper", + "page_header": "pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "o larger particles in the\natmosphere.16 Because of their short atmospheric lifetime and\nstrong dependence on local traffic and meteorological\nconditions, UFPs exhibit steep spatial gradients and high\nvariability at fine spatial scales.17,18 As a result, the collected data\nset represents an episodic sample in time (a series of short-term\n\nReceived:\nSeptember 8, 2025\nRevised:\nJanuary 28, 2026\nAccepted:\nJanuary 29, 2026", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "4-5", + "page": 1, + "document_type": "research_paper", + "page_header": "pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "snapshots) with extensive spatial coverage. This trade-off\ncomplicates the disentanglement of spatial vs temporal\nvariation.19−21 Variations in data collection frequency and\nsampling routes can introduce biases, potentially affecting the\nreliability of LUR models trained on mobile data.13,22,23 A model\nmight simply “memorize” the specific conditions of the sampling\nperiod (a phenomenon we term “data reproduction” defined as\noverfitting to the specific spatiotemporal structure of the training\ndata) rather than learning generalizable patterns (prediction).\nFor instance, Blanco et al. (2023)24 demonstrated that different\nsampling designs can produce significantly different exposure\nsurfaces, even within the same study area. Their findings indicate\nthat campaigns with temporal restrictions (limited to business\nhours, rush hours, or specific seasons) reduced model\nperformance and produced different spatial surfaces. While\nthey suggest improving sampling strategies to enhance data", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-0", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ctions (limited to business\nhours, rush hours, or specific seasons) reduced model\nperformance and produced different spatial surfaces. While\nthey suggest improving sampling strategies to enhance data\nrepresentativeness, refining modeling approaches might also be\nnecessary to better account for spatial and temporal variability\nand ensure more reliable exposure assessments.\n\nModel selection and validation are crucial for building reliable\nLUR models for air pollution.5 A common approach for tuning a\nmodel is random cross-validation (CV). Random CV is\nstraightforward but often overestimates performance in spatially\nor temporally structured data, as nearby locations or time\nperiods can leak between training and test sets, inflating\nperformance metrics.25−27 This limitation arises because\nconventional random CV assumes that observations are\nindependent and identically distributed, a core assumption\nunderpinning its validity as an unbiased estimator of out-of-", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-1", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "limitation arises because\nconventional random CV assumes that observations are\nindependent and identically distributed, a core assumption\nunderpinning its validity as an unbiased estimator of out-of-\nsample error.28−30 In spatial or temporal data, this assumption is\nviolated, observations are not independent, leading to\ninformation leakage between folds and overly optimistic\nperformance estimates.31−33 The spatial and temporal autocor-\nrelation in air pollution data may render random splitting\ninappropriate, risking overfitting to noncausal predictors and\nover optimistic performance metrics. Therefore, more advanced\nstrategies are needed to avoid overfitting.34,35 Roberts et al.\n(2017)36 recommend block CV for model selection, ensuring\ndata are partitioned by spatial or temporal structures. Blocking\nby location or time provides a stringent test of whether a model\ncan predict in completely unseen regions or periods. Indeed,\nusing appropriate CV helps select models and hyperparameters", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-2", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "es. Blocking\nby location or time provides a stringent test of whether a model\ncan predict in completely unseen regions or periods. Indeed,\nusing appropriate CV helps select models and hyperparameters\nthat avoid overfitting to spurious correlations present in training\ndata.36,37\n\nThe accuracy of LUR models depends on predictor variables,\nwhich capture built environment and other factors influencing\nair pollution levels.38,39 It is common to explore a wide array of\npotential predictor variables when constructing LUR mod-\nels.40−43 However, excessive inclusion of predictor variables can\nlead to overfitting, reducing model generalizability and\npredictive performance.34,44,45 As a result, feature selection is a\ncritical step in LUR. Various feature selection methods,\nincluding forward stepwise linear regression46−48 and\nLASSO,49 have been used to improve LUR performance and\nreduce overfitting. Forward feature selection (FFS) is an", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-3", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "rious feature selection methods,\nincluding forward stepwise linear regression46−48 and\nLASSO,49 have been used to improve LUR performance and\nreduce overfitting. Forward feature selection (FFS) is an\niterative procedure that adds variables sequentially based on\ntheir contribution to model performance. In each step, variables\nare evaluated for significance and multicollinearity, excluding\nredundant or highly correlated land-use predictors. However,\nfeature selection methods such as FFS cannot remove spurious\npredictors under random CV, where spatially dependent\nobservations appear in both training and test sets.44,50,51 This\ncan inflate performance metrics, leading to overfitted models\n\nand artifactual spatial patterns that fail to generalize across new\nregions or time periods. More broadly, these risks stem from\nreliance on purely statistical criteria that may omit causal\nrelationships. This leads to selecting variables that appear", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-4", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "alize across new\nregions or time periods. More broadly, these risks stem from\nreliance on purely statistical criteria that may omit causal\nrelationships. This leads to selecting variables that appear\nsignificant by chance but do not generalize in autocorrelated\ndata, generating misleading models.51−54 As the number of\ncandidate variables increases, the likelihood of selecting\nvariables that appear statistically significant by chance but do\nnot contribute to generalizable predictions also increases.55\n\nStudies demonstrate that selected variables may be significant by\nchance but fail to generalize in spatially autocorrelated data.56,57\n\nTo address these limitations, this study explores several\nmodeling frameworks to evaluate the approaches that can best\nmitigate overfitting in mobile monitoring data. For example, in\none framework, we examine the integration of FFS with\nspatiotemporal CV to assess how aligning feature selection", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-5", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "proaches that can best\nmitigate overfitting in mobile monitoring data. For example, in\none framework, we examine the integration of FFS with\nspatiotemporal CV to assess how aligning feature selection\nand hyperparameter tuning with the data’s spatial and temporal\nstructure influences model generalization.\n\nRobust models for data with spatial and temporal\nautocorrelation require CV strategies for hyperparameter tuning\nand feature selection to ensure generalizability and minimize\noverfitting.37,58 Meyer et al. (2019)44 showed that combining\nspatial CV with FFS significantly improved model robustness by\navoiding spurious predictors in spatially correlated data.\nSimilarly, Meyer et al. (2018)52 integrated FFS with a target-\noriented spatiotemporal CV approach for temperature and soil-\ncontent data from stationary stations, systematically removing\nmisleading variables and achieving more generalizable pre-\ndictions. Baumberger et al. (2024)59 highlighted the importance", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-6", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ure and soil-\ncontent data from stationary stations, systematically removing\nmisleading variables and achieving more generalizable pre-\ndictions. Baumberger et al. (2024)59 highlighted the importance\nof spatiotemporal CV, FFS, and targeted hyperparameter tuning\nin a high-resolution environmental model using stationary\nstation records for soil temperature and soil moisture.\n\nIn this study, we develop LUR models using Extreme\nGradient Boosting (XGBoost) to predict UFP concentrations\nbased on mobile monitoring data collected in Toronto. We\nsystematically compare seven modeling approaches to assess the\nimpact of hyperparameter tuning and feature selection on model\ngeneralizability. In addition, we develop four additional models\nto assess the robustness of modeling approaches under different\noutlier-handling methods. We tested CV strategies for tuning\nand feature selection and assessed model robustness under\ndifferent outlier treatments. This allowed us to compare", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-7", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "roaches under different\noutlier-handling methods. We tested CV strategies for tuning\nand feature selection and assessed model robustness under\ndifferent outlier treatments. This allowed us to compare\nmodeling approaches and understand how they influence\nperformance when dealing with mobile monitoring data. This\nstudy contributes to the development of machine learning-based\nexposure assessment models, providing insights into optimizing\nfeature selection, hyperparameter tuning, and CV strategies to\nimprove the reliability of mobile air pollution estimates. Finally,\nwe compare model predictions against continuously recorded\ndata from stationary sensors.\n\n2. MATERIALS AND METHODS\n\n2.1. Study Area and Data Collection\n\nIn Toronto, UFP concentrations are primarily influenced by traffic\nemissions, particularly from diesel vehicles and congestion along major\nroadways, as well as localized industrial activities.60 These sources\ncreate strong spatial gradients and diurnal variation, with higher", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-8", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": ", particularly from diesel vehicles and congestion along major\nroadways, as well as localized industrial activities.60 These sources\ncreate strong spatial gradients and diurnal variation, with higher\nconcentrations during rush hours and lower levels at night and on\nweekends.5 Dilution and dispersion in green and open areas such as\nparks and vegetated corridors further reduce local UFP levels through\nenhanced atmospheric mixing and particle deposition.61 Between April\n7 and June 23, 2021, we conducted mobile monitoring in Toronto to\ncapture spatial variations in UFP concentrations between 8 AM and 7\n\nB", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "5-9", + "page": 2, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "PM over 34 days. Data were recorded at one-second intervals using the\nUrbanScanner platform, which integrates a UFP DiscMini, wind\nanemometer, and GPS, along routes covering major roads, residential,\nindustrial, and highways as described by (Ganji et al., 202312). In\naddition, 12 DiscMini and 5 UFP Partector instruments were deployed\nin residential backyards from July 9 to 30, 2021, providing continuous\nUFP measurements over a 3 week period (Figure 1). Backyard\n\nFigure 1. Study area in Toronto with mobile monitoring routes and\nstationary backyard sensor locations.\n\nDiscMini sensors recorded UFP concentrations at 10 s intervals, while\nPartector sensors collected data every second. To align with the mobile\nmonitoring campaign, we used data recorded by these sensors between\n8 AM and 7 PM. The DiscMini instruments were factory-calibrated by\nthe manufacturer prior to deployment, following procedures consistent\nwith previous UrbanScanner campaigns (Ganji et al., 2020;62 Xu et al.,\n202239).", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "6-0", + "page": 3, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "he DiscMini instruments were factory-calibrated by\nthe manufacturer prior to deployment, following procedures consistent\nwith previous UrbanScanner campaigns (Ganji et al., 2020;62 Xu et al.,\n202239). The performance of this class of diffusion-charger-based UFP\nsensors has been independently validated in laboratory studies; for\nexample, Mills et al. (2013)63 compared a DiSCmini-type instrument\nwith reference CPC and SMPS instruments, reporting strong\nagreement (R2 > 0.9) across 10−300 nm and concentrations up to\n106 particles cm−3.\n\n2.2. Predictor Variables\nTo ensure spatial consistency, we defined 100 m distance points along\nthe mobile monitoring routes. At each of these points, air pollution\nmeasurements [49] were aggregated and a median value reported.64\n\nSpatial buffers were then created around these 100 m interval points to\nextract land use and traffic-related variables within predefined buffer\nzones. To model UFP concentrations, we incorporated a compre-", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "6-1", + "page": 3, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "l buffers were then created around these 100 m interval points to\nextract land use and traffic-related variables within predefined buffer\nzones. To model UFP concentrations, we incorporated a compre-\nhensive set of 194 predictor variables, categorized into land use, traffic-\nrelated, distance-based, and temporally changing variables (Table S1).\n2.3. Model Development\nWe used XGBoost, a gradient boosting decision tree algorithm, to\ndevelop the LUR models for UFP. Outlier removal by road type was\napplied using the Interquartile Range (IQR) method, where outliers\nwere identified separately for each road type and removed if they fell\noutside Q1 −1.5 × IQR and Q3 + 1.5 × IQR. Finally, normalization\nwas applied to standardize continuous predictors. From this data set, we\ncreated seven models to evaluate different CV strategies for\nhyperparameter tuning and feature selection. In addition, we explored\ntwo alternative outlier-handling techniques: removing values above the", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "6-2", + "page": 3, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "eated seven models to evaluate different CV strategies for\nhyperparameter tuning and feature selection. In addition, we explored\ntwo alternative outlier-handling techniques: removing values above the\n95th percentile or retaining all spikes, by building four additional\nmodels following two modeling frameworks: random CV alone, and\nspatiotemporal CV combined with feature selection.\n\nWe used 80% of the data points as training set for hyperparameter\ntuning and feature selection. All model development and validation\nprocedures were implemented in Python using the scikit-learn and\nXGBoost libraries.\n\nWe defined four CV schemes for model tuning, each using K-fold CV\non the training data set.\n\n• Random CV (RCV): Data were randomly split into training and\nvalidation (10-folds), ignoring spatial or temporal structure.\n• Spatial CV (SCV): We generated 400 geographic clusters using\nthe 100 m-aggregated training records (see Figure S1a). These", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "6-3", + "page": 3, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "raining and\nvalidation (10-folds), ignoring spatial or temporal structure.\n• Spatial CV (SCV): We generated 400 geographic clusters using\nthe 100 m-aggregated training records (see Figure S1a). These\nclusters were randomly divided into 10 folds for spatial cross-\nvalidation (SCV), ensuring each cluster served as the validation\nregion once (Figure S1a). A 300 m dead buffer was imposed\naround validation clusters to limit spatial leakage36 (step-by-step\ndetails are provided in the SI).\n• Temporal CV (TCV): We partitioned the data by sampling day.\nWe sorted the mobile data chronologically and created eight\n\nFigure 2. Conceptual diagram of the framework integrating FFS and STCV, representing the iterative workflow from data partitioning and\nhyperparameter tuning to FFS combined with STCV and final model evaluation.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "6-4", + "page": 3, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "Table 1. Summary of Model Configurations\n\nfeature\nselection\noutlier handling\nnote\n\ncross-validation for\nhyperparameter tuning\n\nmodel name\n\nRCV_HP\nRandom CV\nNo\nby road-type IQR-\n\nbased removal\n\nhyperparameter tuning with random CV.\nall features are used as inputs.\nSCV_HP\nSpatial CV\nNo\nby road-type IQR-\n\nbased removal\n\nhyperparameter tuning with spatial CV.\nall features are used as inputs.\nTCV_HP\nTemporal CV\nNo\nby road-type IQR-\n\nbased removal\n\nuses day-wise splits to handle temporal variability.\nhyperparameter tuning done under temporal CV while using all features.\nSTCV_HP\nSpatiotemporal CV\nNo\nby road-type IQR-\n\nbased removal\n\nSCV_FFS\nSpatial CV\nYes\nby road-type IQR-\n\nbased removal\n\nhyperparameter tuning done under spatiotemporal CV using all features.\n\nforward feature selection integrated with spatial CV.\nhyperparameter tuning with spatial CV.\nTCV_FFS\nTemporal CV\nYes\nby road-type IQR-\n\nbased removal\n\nFfs integrated with temporal CV.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-0", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "all features.\n\nforward feature selection integrated with spatial CV.\nhyperparameter tuning with spatial CV.\nTCV_FFS\nTemporal CV\nYes\nby road-type IQR-\n\nbased removal\n\nFfs integrated with temporal CV.\nhyperparameter tuning with temporal CV while using selected features.\nSTCV_FFS\nSpatiotemporal CV\nYes\nby road-type IQR-\n\nbased removal\n\nFFS integrated with Spatiotemporal CV.\nhyperparameter tuning with temporal CV while using selected features.\nRCV_HP _95\nRandom CV\nNo\nremove top 5% outliers\nthis model is similar to RCV_HP.\ntests the effect of removing extreme spikes (above the 95th percentile)\n\nwhile using RCV_HP framework.\nSTCV_FFS_95\nSpatiotemporal CV\nYes\nremove top 5% outliers\nthis model is similar to STCV_FFS.\ntests the effect of removing extreme spikes (above the 95th percentile)\n\nwhile using STCV_FFS framework.\nRCV_HP_noOut\nRandom CV\nNo\nno outlier removal\nthis model is similar to RCV_HP.\ntests the effect of not removing extreme spikes while using RCV_HP\n\nframework.\nSTCV_FFS_noOut", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-1", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "e using STCV_FFS framework.\nRCV_HP_noOut\nRandom CV\nNo\nno outlier removal\nthis model is similar to RCV_HP.\ntests the effect of not removing extreme spikes while using RCV_HP\n\nframework.\nSTCV_FFS_noOut\nSpatiotemporal CV\nYes\nno outlier removal\nthis model is similar to STCV_FFS.\ntests the effect of not removing extreme spikes while using STCV_FFS\n\nframework.\nNote: Model names indicate the cross-validation (CV) strategy used for hyperparameter tuning (RCV = Random CV, STCV = Spatiotemporal\nCV), followed by whether forward feature selection (FFS) was applied, and then followed by outlier-handling approach. Models with suffixes like\n“_95” or “_noOut” test the effect of removing or retaining outliers.\n\nfolds, each corresponding to a subset of days (Figure S2). In\neach fold, models were trained on seven-eighths of the days and\nvalidated on the remaining one-eighth (validation involved\npredicting data from dates entirely unseen in training) (step-by-\nstep details are provided in the SI).", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-2", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ned on seven-eighths of the days and\nvalidated on the remaining one-eighth (validation involved\npredicting data from dates entirely unseen in training) (step-by-\nstep details are provided in the SI).\n• Spatiotemporal CV (STCV): We combined spatial and\ntemporal blocking by holding out data that were independent\nin both space and time (step-by-step details are provided in the\nSI). In our implementation, we ensured that no location from a\ncertain set of clusters on certain days appeared in training if it\nwas in the validation set. Essentially, we withheld entire day-of-\ncluster combinations. This STCV approach is very stringent: for\nexample, the model might be trained on data from most\nlocations and days but tested on data from a particular group of\nlocations on specific days (neither those locations nor dates\nappear in training). Because our data set is large, we achieved\nSTCV by first clustering locations (as in SCV) and then in each", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-3", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "group of\nlocations on specific days (neither those locations nor dates\nappear in training). Because our data set is large, we achieved\nSTCV by first clustering locations (as in SCV) and then in each\nfold selecting a subset of clusters and specific days to hold out,\nsuch that validation data are separated in both dimensions. This\napproach aligns with the “target-oriented” validation concept\nintroduced by Meyer et al. (2018).52\n\nFor each CV scheme, we conducted hyperparameter tuning (HP) by\nBayesian optimization using only the training folds and evaluated\nperformance on the validation fold in each iteration. Hyperparameters\nsuch as tree depth, learning rate, and regularization parameters were\noptimized based on the average validation performance across the folds.\n\nIn addition to the full-feature models, we built models with FFS\nunder the three structured CV frameworks (see SI for details). We\nemployed a forward stepwise selection approach: we added predictors", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-4", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "In addition to the full-feature models, we built models with FFS\nunder the three structured CV frameworks (see SI for details). We\nemployed a forward stepwise selection approach: we added predictors\none at a time, in each step including the variable that most improved the\nCV mean squared error (MSE) under the given CV scheme. This\nprocess was repeated until adding any remaining variable did not\nincrease the MSE. The feature selection was performed within each CV\nstrategy (see SI for details). This ensured that selected predictors\n\ncontributed to generalizable skill rather than fitting coincidences of the\ntraining data.44,52 Figure 2 presents a conceptual diagram of the STCV\ncoupled with FFS modeling framework, illustrating the iterative\nworkflow from spatiotemporal fold definition and hyperparameter\ntuning to feature selection and final model evaluation. While this\nschematic highlights the methodological innovation specific to the", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-5", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "e\nworkflow from spatiotemporal fold definition and hyperparameter\ntuning to feature selection and final model evaluation. While this\nschematic highlights the methodological innovation specific to the\nSTCV−FFS framework, the same forward feature selection workflow\nwas also applied under the spatial and temporal CV schemes, with fold\ndefinitions differing across frameworks.\n\nUltimately, we trained and evaluated 11 models (Table 1) to\nexamine how modeling frameworks (including hyperparameter tuning\nand feature selection) and outlier treatment influence model perform-\nance and robustness. First, we evaluated seven modeling frameworks:\nRCV_HP, SCV_HP, TCV_HP, STCV_HP, SCV_FFS, TCV_FFS,\nand STCV_FFS which differ in their hyperparameter tuning and feature\nselection processes. Next, to examine how outlier handling and\nsampling changes affect the RCV_HP and STCV_FFS frameworks,\nfour additional models were developed: RCV_HP_95,\nSTCV_FFS_95, RCV_HP_noOut, STCV_FFS_noOut. Notably,", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-6", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": ", to examine how outlier handling and\nsampling changes affect the RCV_HP and STCV_FFS frameworks,\nfour additional models were developed: RCV_HP_95,\nSTCV_FFS_95, RCV_HP_noOut, STCV_FFS_noOut. Notably,\nRCV_HP, RCV_HP_95, and RCV_HP_noOut all follow the same\nrandom CV approach but vary in how they address outliers, while\nSTCV_FFS, STCV_FFS_95, and STCV_FFS_noOut share the same\nspatiotemporal CV and feature selection strategy, they differ by their\noutlier-removal methods. The subscript 95 represents the removal of\nvalues in the top 5% or 95th percentile and noOut refers to no outlier-\nremoval. Models without either of these suffixes follow the baseline\noutlier-handling approach, which applies IQR-based removal by road\ntype.\n2.4. Model Implementation and Evaluation\n\nAppropriate selection of evaluation metrics is essential to accurately\nreflect model performance and generalizability in air pollution exposure\nassessment studies.35,36,44 Given the spatial and temporal complexities", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-7", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ction of evaluation metrics is essential to accurately\nreflect model performance and generalizability in air pollution exposure\nassessment studies.35,36,44 Given the spatial and temporal complexities\ninvolved in mobile monitoring data sets, we employed two distinct\nevaluation frameworks to check generalizability and overfitting of the", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "7-8", + "page": 4, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "Figure 3. Predicted versus actual UFP concentrations on the 20% random holdout test set for each model. The dashed line denotes perfect 1:1\nagreement.\n\nmodels. First, to assess overfitting and performance on unseen data, we\nused a 20% random holdout sample excluded from training, tuning, and\nfeature selection. We computed median concentrations to mitigate\ntransient spikes. This data set, consistently used across all models,\nprovided a standardized comparison basis, reporting coefficient of\ndetermination (R2), Mean Absolute Error (MAE), and Root Mean\nSquared Error (RMSE) as evaluation metrics (see SI for details).\nSecond, to assess whether a modeling framework improved general-\nization, we reallocated all data into random, spatial, temporal, and\nspatiotemporal folds. The tuning and feature selection process relied on\na specific data split, whereas final split for checking the models were\ndeveloped using the entire data set with different cluster allocations.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "8-0", + "page": 5, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "lds. The tuning and feature selection process relied on\na specific data split, whereas final split for checking the models were\ndeveloped using the entire data set with different cluster allocations.\nFigure S1b illustrates an example of a single fold from the 10-fold spatial\nCV used during model tuning and feature selection, while Figure S3\nshows an example of a single fold from the 10-fold spatial CV used\nduring the evaluation phase. Once each model had undergone tuning\nand feature selection with its designated strategy, we tested it using all\nfour newly defined CV methods in a comparable manner (e.g.,\nconsistent cluster assignments for SCV and STCV, day-based splits for\nTCV). This ensures a direct comparison of performance while\nminimizing confounding effects from varying fold assignments (see SI\nfor details). Evaluating each model under these four CV schemes\nenables a nuanced assessment of performance across varying levels of", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "8-1", + "page": 5, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "imizing confounding effects from varying fold assignments (see SI\nfor details). Evaluating each model under these four CV schemes\nenables a nuanced assessment of performance across varying levels of\n\ngeographic and temporal separation between training and validation\nsubsets.\n\nGiven our modeling context, the interpretation of CV metrics,\nespecially R2, requires caution. Mobile monitoring data aggregated from\nlimited visits per location is susceptible to episodic pollution spikes and\ntemporal variability, potentially inflating prediction errors and lowering\nR2, especially under rigorous CV schemes (SCV, TCV, STCV).\nConsequently, while R2 is essential for seeing the improvement in\ngeneralization, it does not necessarily represent the model’s capability\nto accurately estimate average pollutant levels over a given period. A\nmodel evaluated under TCV, for example, might exhibit low R2 due to\ntemporal variability but still accurately capture average spatial pollution", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "8-2", + "page": 5, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "stimate average pollutant levels over a given period. A\nmodel evaluated under TCV, for example, might exhibit low R2 due to\ntemporal variability but still accurately capture average spatial pollution\npatterns when evaluated on continuous data representative of the given\nperiod.23,25,47,65 In this study, CV performance metrics (e.g., R2, MSE)\nwere primarily used for hyperparameter tuning and feature selection\nand then each model was tested across four CV. Overall, by integrating\nrigorous CV for evaluation (to inform robust model development and\nprevent overfitting) and an internal holdout sample evaluation, we aim\nto provide a balanced, transparent, and realistic assessment of each\nmodel’s capability to generalize.\n\nSensitivity analysis was conducted using SHAP-based feature\nimportance plots, which reveal how each predictor influences UFP.\nBy averaging the absolute SHAP values, we identified the most critical\n\nE", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "8-3", + "page": 5, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "Figure 4. Mean APE (%) of UFP predictions compared with stationary measurements across all models.\n\npredictors. Finally, we generated a 50 m resolution UFP map for\nToronto by applying the final XGBoost model to a regular grid,\nincorporating land-use, traffic, and distance-based predictors. This\napproach captures spatial distributions and allows visual comparisons\nacross different model predictions.\n\nTo independently explore model performance, we performed\nexternal validation against a stationary network of 17 UFP fixed\nmonitors installed in backyards of private homes throughout the city of\nToronto (12 DiscMini and 5 Partector) and which continuously\nmeasured UFP from July 9 to July 30, 2021, while our mobile campaign\nconcluded on June 23, 2021. Despite nonoverlapping periods,\nmeteorological and emission conditions were sufficiently similar,\nenabling the backyard data to serve as a valuable continuous external\nholdout sample. Model performance was evaluated by computing the", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "9-0", + "page": 6, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "eteorological and emission conditions were sufficiently similar,\nenabling the backyard data to serve as a valuable continuous external\nholdout sample. Model performance was evaluated by computing the\nAbsolute Percentage Error (APE) between predicted and observed\nmean UFP concentrations at each site (APE; see SI for the equation).\n\n3. RESULTS\n\n3.1. Measured Concentrations\nThe mobile monitoring campaign yielded an extensive data set of UFP\nmeasurements (over 440,000 valid observations after quality control).\nConcentrations varied greatly from day to day, reflecting changing\ntraffic, meteorology, and local emissions (Figure S4). To further\ncharacterize the data set, we first examined the distribution of UFP\nconcentrations after applying IQR-based filtering by road type. This\nmethod excluded road-specific statistical outliers and resulted in a mean\nconcentration of 16,817 particles/cm3 and a median of 13,856\nparticles/cm3. We also assessed the impact of two alternative outlier", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "9-1", + "page": 6, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "d excluded road-specific statistical outliers and resulted in a mean\nconcentration of 16,817 particles/cm3 and a median of 13,856\nparticles/cm3. We also assessed the impact of two alternative outlier\nstrategies: retaining all values without removal and excluding the top 5%\nof concentrations (95th percentile removal). The unfiltered data set\nyielded the highest mean (23,906 particles/cm3) and a median of\n14,777 particles/cm3, indicating a right-skewed distribution influenced\nby extreme values. The 95th percentile removal approach produced\nsimilar summary statistics to the IQR method, with a mean of 16,816\nparticles/cm3 and a median of 14,017 particles/cm3. These results,\nsummarized in Table S2, highlight how different filtering strategies can\nshape the overall distribution of mobile UFP data.\n\nMeasurements from the backyard sensors over the three-week\nmonitoring period (Figure S5) reveal substantial variability in UFP\nconcentrations across locations, likely reflecting differences in", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "9-2", + "page": 6, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ata.\n\nMeasurements from the backyard sensors over the three-week\nmonitoring period (Figure S5) reveal substantial variability in UFP\nconcentrations across locations, likely reflecting differences in\n\nproximity to emission sources such as roads, buildings, or vegetation.\nSummary statistics for each monitor (Table S3) further highlight this\nheterogeneity: median UFP concentrations ranged from as low as 2915\nparticles/cm3 to as high as 9018 particles/cm3. Compared to mobile\nmonitoring data, with a median concentration of approximately 13,856\nparticles/cm3 (after IQR-based filtering), stationary measurements\nexhibited lower medians and reduced day-to-day variability.\n\n3.2. Impact of Modeling Framework\n\nSeven modeling frameworks were evaluated to assess the impact of\nhyperparameter tuning and feature selection combined with CV\nstrategies on performance, generalization, and overfitting. Table S4\ncompares 80% training and 20% holdout performance, highlighting", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "9-3", + "page": 6, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ct of\nhyperparameter tuning and feature selection combined with CV\nstrategies on performance, generalization, and overfitting. Table S4\ncompares 80% training and 20% holdout performance, highlighting\nfitting, overfitting, and how CV strategies and FFS affect robustness.\nThe RCV_HP model exhibits a near-perfect training R2 of 0.999, with\nlow RMSE (400) and MAE (261). However, its test R2 declines to\n0.715 (RMSE = 6692, MAE = 3869) which is a sign of overfitting.\nModels tuned using structured CV approaches\u0001such as SCV_HP,\nTCV_HP, and STCV_HP\u0001show more moderate training R2 values\n(0.544−0.825) and achieve test set R2 values between 0.510 and 0.660.\nIn addition to the models without feature selection, Table S4 shows\nhow incorporating FFS combined with different CV strategies impacts\nmodel performance. For instance, the SCV_FFS model yields a training\nR2 of 0.957 and a test R2 of 0.717. Meanwhile, models with feature\nselection merged and tuned under Temporal (TCV_FFS) and", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "9-4", + "page": 6, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "impacts\nmodel performance. For instance, the SCV_FFS model yields a training\nR2 of 0.957 and a test R2 of 0.717. Meanwhile, models with feature\nselection merged and tuned under Temporal (TCV_FFS) and\nSpatiotemporal (STCV_FFS) CV schemes exhibit lower overall R2\nvalues on both training and test sets: 0.413 and 0.441 in training, and\n0.398 and 0.450 in testing, indicating that overfitting was effectively\navoided.\n\nFigure 3 presents scatter plots of the predicted versus actual UFP\nconcentrations for all the seven modeling frameworks using the 20%\nrandom holdout internal test set. The red dashed line represents perfect\nprediction. Models using Random CV tend to have points clustered\nnear this line, indicating high accuracy in a similar data split; however,\nthis may hide overfitting to local patterns. In contrast, models using\nspatial, temporal, or spatiotemporal CV and especially those with\nforward feature selection, show a wider spread in the data. This spread", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "9-5", + "page": 6, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "hide overfitting to local patterns. In contrast, models using\nspatial, temporal, or spatiotemporal CV and especially those with\nforward feature selection, show a wider spread in the data. This spread\nreflects the challenge of predicting pollution in different locations and\ntimes. Lower R2 values in these cases do not automatically mean the\nmodels perform poorly, they reflect the limited visits per location and\nthe episodic nature of mobile monitoring.\n\nF", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "9-6", + "page": 6, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "Figure 5. Predicted UFP exposure surfaces based on all models, highlighting the spatial variability associated with different cross-validation, feature\nselection, and outlier treatment strategies.\n\nTable S5 presents the performance of seven XGBoost-based models\nunder four CV approaches: Random, Spatial, Temporal, and\nSpatiotemporal. Each model’s performance is summarized by three\nmetrics: R2, RMSE, and MAE. When using Random CV, RCV_HP\nattains an R2 of 0.735, with an RMSE of 6245 and an MAE of 3656,\nwhile SCV_FFS records R2 = 0.717, RMSE = 6406, MAE = 3964.\nModels such as TCV_FFS and STCV_FFS show lower R2 values\n(0.389 and 0.413, respectively) yet vary in their RMSE and MAE.\nUnder Spatial CV, RCV_HP R2 is 0.355, while SCV_FFS reaches\n0.418. Similarly, TCV_FFS and STCV_FFS have R2 values of 0.265\nand 0.278, respectively. Under Temporal CV, the R2 values typically\ndrop further, reflecting the challenge of day-to-day variability in", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "10-0", + "page": 7, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "0.418. Similarly, TCV_FFS and STCV_FFS have R2 values of 0.265\nand 0.278, respectively. Under Temporal CV, the R2 values typically\ndrop further, reflecting the challenge of day-to-day variability in\nmeteorology, traffic flows, and emission sources that can dominate\nmobile monitoring data. For instance, RCV_HP experiences a stark\n\ndecrease from 0.355 in Spatial CV to a mere 0.035 in Temporal CV,\nindicating its inability to generalize across distinct sampling days.\nModels specifically incorporating feature selection and/or tuned for\ntemporal splits like TCV_FFS, achieve a higher R2 of 0.173. Some\nmodels, such as SCV_FFS, score as low as 0.054 in R2. RMSE and MAE\nsimilarly show larger absolute values for most models under this\npartitioning. Under spatiotemporal CV evaluation, RCV_HP has an R2\n\nof 0.105, and STCV_FFS achieves 0.186.\n\n3.3. Impact of Outliers on Predictions\n\nTable S6 presents the training and test set performance of four", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "10-1", + "page": 7, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": ". Under spatiotemporal CV evaluation, RCV_HP has an R2\n\nof 0.105, and STCV_FFS achieves 0.186.\n\n3.3. Impact of Outliers on Predictions\n\nTable S6 presents the training and test set performance of four\nadditional models differing by outlier handling (none or 95th-percentile\nremoval) and modeling framework (random vs spatiotemporal CV with\nfeature selection). Comparing the metrics for this table with the results\n\nG", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "10-2", + "page": 7, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "from Table S4 shows that the models tuned under Random CV often\nexhibit near-perfect R2 on the training set yet experience marked\ndeclines on the test set, suggesting overfitting when extreme values are\neither retained or filtered. Models that have feature selection combined\nwith spatiotemporal CV tend to show lower training R2 but more\nbalanced performance on the test set, indicating better generalization.\nThe four additional models have also been tested under four different\nCV strategies. The results of this evaluation (Table S7) reveal that\nmodeling frameworks with feature selection in conjunction with proper\nCV can improve generalization across time and space. Although models\ntuned under random CV might yield deceptively high R2, imposing\nstricter splits along spatial and temporal dimensions forces the models\nto learn predictors that hold up under varied conditions.\n\nBy comparing scatter plots in Figure 3 and Figure S6, we observe two", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-0", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "stricter splits along spatial and temporal dimensions forces the models\nto learn predictors that hold up under varied conditions.\n\nBy comparing scatter plots in Figure 3 and Figure S6, we observe two\ndistinct groups of models with similar behavior. The models\nSTCV_FFS, STCV_FFS_95, and STCV_FFS_noOut exhibit similar\nscatter patterns, suggesting consistent predictive structure even under\ndifferent outlier-handling strategies. In contrast, the RCV_HP,\nRCV_HP_95, and RCV_HP_noOut models display another shared\npattern reinforcing the idea that model structure and CV strategy have a\nmore pronounced influence on predictions than minor variations in\noutlier removal.\n\n3.4. Comparing with Stationary Measurements\n\nFigure 4 summarizes mean Absolute Percentage Error (APE) values by\ncomparing average of stationary measurements with the predictions of\nthe models. For instance, RCV_HP shows notably high mean APE:\n134%. Meanwhile, TCV_FFS and STCV_FFS markedly reduce these\nerrors.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-1", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "y\ncomparing average of stationary measurements with the predictions of\nthe models. For instance, RCV_HP shows notably high mean APE:\n134%. Meanwhile, TCV_FFS and STCV_FFS markedly reduce these\nerrors. TCV_FFS achieves 65%, while STCV_FFS 78%. Models\nincorporating temporal or spatiotemporal CV, particularly when\ncombined with FFS, appear to track average of continuous fixed-site\nsensor data more reliably than those relying on random or purely spatial\nCV for tuning and feature selection. Additionally, comparing models\nwith different outlier-handling approaches reveals that RCV_HP and its\nvariations (RCV_HP_95 and RCV_HP_noOut) can produce\nsubstantially higher APEs (greater than 130%) indicating that models\ntuned with random CV do not align well with three-week average\nconcentrations from stationary sensors. In contrast, models using\nspatiotemporal CV with feature selection (STCV_FFS,\nSTCV_FFS_95, or STCV_FFS_noOut) consistently achieve lower", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-2", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ith three-week average\nconcentrations from stationary sensors. In contrast, models using\nspatiotemporal CV with feature selection (STCV_FFS,\nSTCV_FFS_95, or STCV_FFS_noOut) consistently achieve lower\nAPE, suggesting they may better capture average pollution trends across\nthe city, rather than being overly influenced by episodic mobile\nmeasurements.\n\n3.5. Selected Predictors and Development of Exposure\nSurfaces\n\nSHAP analysis was applied to each of the XGBoost models to quantify\nthe influence of individual predictors on UFP concentrations. In models\nwithout feature selection (e.g., RCV_HP, SCV_HP, TCV_HP,\nSTCV_HP), the SHAP plots (Figure S7) highlight a wide range of\npredictors, prominently featuring traffic-related variables (such as\ndistance to highways, road area across multiple buffers, and AADT) and\nseveral meteorological variables. In contrast, models incorporating\nforward feature selection (SCV_FFS, TCV_FFS, STCV_FFS,\nSTCV_FFS_95, and STCV_FFS_noOut) retain fewer predictors,", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-3", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ers, and AADT) and\nseveral meteorological variables. In contrast, models incorporating\nforward feature selection (SCV_FFS, TCV_FFS, STCV_FFS,\nSTCV_FFS_95, and STCV_FFS_noOut) retain fewer predictors,\nemphasizing more robust and interpretable features such as proximity\nto major roads, road density, and specific land-use categories. Among\nthe spatiotemporal CV models, STCV_FFS retains 17 predictors,\nSTCV_FFS_95 includes 20, and STCV_FFS_noOut reduces the\ncount to just 9, demonstrating the effect of both CV strategy and outlier\ntreatment on model parsimony and interpretability.\n\nFigure 5 presents the exposure surfaces. Models relying on Random\nCV (RCV_HP), Spatial CV for hyperparameter tuning (SCV_HP),\nSpatial CV with feature selection (SCV_FFS), and Temporal CV\nwithout feature selection (TCV_HP) exhibit abrupt “lines” where UFP\nlevels shift dramatically, resulting in concentration surfaces that appear\nless representative of known spatial trends in Toronto. These models", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-4", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ure selection (TCV_HP) exhibit abrupt “lines” where UFP\nlevels shift dramatically, resulting in concentration surfaces that appear\nless representative of known spatial trends in Toronto. These models\nalso show abrupt pollutant changes that deviate from expected averages,\nresulting in exposure surfaces with artifacts and indicating overfitting\nand a failure to reproduce realistic spatial gradients. In contrast, the\n\nmodel using spatiotemporal CV, even without feature selection\n(STCV_HP), produces spatial patterns that better follow traffic-related\nvariations, with elevated UFP levels along major roadways and lower\nconcentrations in parks and open areas. The TCV_FFS and\nSTCV_FFS models refine these trends further, assigning lower UFP\nvalues to vegetated areas and higher concentrations to major corridors.\nTheir predicted gradients capture the expected UFP behavior near\nroads more accurately and appear more physically consistent and better", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-5", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "o vegetated areas and higher concentrations to major corridors.\nTheir predicted gradients capture the expected UFP behavior near\nroads more accurately and appear more physically consistent and better\naligned with average observations over the study period.\n\nFigure 5 also includes variations of these models under different\noutlier treatments. For Random CV (RCV_HP, RCV_HP_95,\nRCV_HP_noOut), altering the outlier threshold leads to noticeably\ndifferent spatial surfaces, further illustrating the instability of random\nCV-based models. In contrast, STCV_FFS models (STCV_FFS,\nSTCV_FFS_95, STCV_FFS_noOut) produce consistently smooth\nand interpretable concentration maps regardless of the outlier\nthreshold, demonstrating that combining spatiotemporal CV with\nfeature selection improves robustness to data variability and enhances\nspatial reliability. Residual spatial autocorrelation was evaluated using\nglobal Moran’s I on model residuals, with results reported in the", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-6", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ection improves robustness to data variability and enhances\nspatial reliability. Residual spatial autocorrelation was evaluated using\nglobal Moran’s I on model residuals, with results reported in the\nSupporting Information (Table S8).\n\n4. DISCUSSION\n\nIn this study, two internal evaluation approaches were used: a\n20% random holdout and four cross-validation schemes, which\nwere applied to assess each LUR model using mobile monitoring\ndata. Because mobile data are collected over short periods and\ncan exhibit high temporal and spatial variability, these internal\nevaluations may not fully capture true long-term pollution levels.\nIn addition, structured CV can impose overly stringent spatial\nand temporal constraints, creating test scenarios that are more\nchallenging than typical real-world conditions and may not\naccurately reflect a model’s true predictive power. Therefore, the\nmetrics derived from mobile data under different CV schemes", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-7", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "that are more\nchallenging than typical real-world conditions and may not\naccurately reflect a model’s true predictive power. Therefore, the\nmetrics derived from mobile data under different CV schemes\nshould not be interpreted as indicators of each model’s ability to\npredict longer-term concentrations, which is the ultimate goal of\nmodeling. Instead, they should be used to compare robustness,\noverfitting, and improvement in generalization among the\nmodels. Accordingly, the reported R2 values should be\ninterpreted in a comparative sense, highlighting relative model\nstability across validation strategies rather than as absolute\nmeasures of predictive accuracy.\n\n4.1. Impact of Cross Validation for Hyperparameter Tuning\n\nThe choice of CV for tuning and feature selection had a large\nimpact on model performance and generalizability. Models\ntuned with conventional random CV appeared to fit the training\ndata extremely well but failed to generalize to unseen space or\ntime.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-8", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ad a large\nimpact on model performance and generalizability. Models\ntuned with conventional random CV appeared to fit the training\ndata extremely well but failed to generalize to unseen space or\ntime. For instance, RCV_HP achieved an almost perfect fit to its\ntraining data (R2 ≈0.999) but saw its R2 drop to ∼0.71 on a\nhold-out test set. This sharp decline, alongside a jump in error\nmeasures, illustrates severe overfitting when spatial and\ntemporal autocorrelations are ignored during hyperparameter\ntuning. In contrast, models tuned under structured CV scheme\nyielded more moderate training fits (training R2 ∼0.54 to 0.83)\nand comparable test performance (test R2 ∼0.51 to 0.66).\n\nComparing the performance of models without feature\nselection under different CV methods indicate that using\nconventional random CV for tuning inflates model performance\nbecause it mixes spatially correlated data within both training\nand test sets. For instance, models like RCV_HP exhibit very", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-9", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "te that using\nconventional random CV for tuning inflates model performance\nbecause it mixes spatially correlated data within both training\nand test sets. For instance, models like RCV_HP exhibit very\nhigh R2 values of 0.735 under random CV, yet their performance\nsharply declines when evaluated with structured CV approaches,\ndropping to as low as 0.035 under Temporal CV. On the other\nhand, using more strict splits for hyperparameter tuning (e.g.,\n\nH", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "11-10", + "page": 8, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "STCV_HP) improve generalization across sampling days and\nlocations. Models tuned with temporal or spatiotemporal CV\nalign closely with the central tendency of aggregated mobile\nmeasurements at locations with similar traffic patterns, rather\nthan reproducing short-term peaks from transient events.\nMoreover, the prediction power of models tuned with spatially\nand temporally blocked CV was evident when evaluating on\ntruly independent continuous data.\n\n4.2. Feature Selection Combined with CV Methods and\nPredictor Importance\n\nFFS further enhanced model generalizability and interpret-\nability, especially when coupled with structured CV. Without\nfeature selection, even a spatially blocked model could retain\nspurious predictors and overfit certain conditions. For example,\nthe Spatial CV model with all features (SCV_HP) and even with\nFFS (SCV_FFS) showed relatively high training R2 (∼0.95)\nalongside a lower test R2 (∼0.72), suggesting some overfitting\npersisted.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-0", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "xample,\nthe Spatial CV model with all features (SCV_HP) and even with\nFFS (SCV_FFS) showed relatively high training R2 (∼0.95)\nalongside a lower test R2 (∼0.72), suggesting some overfitting\npersisted. This outcome indicates that spatial blocking alone, or\nspatial blocking plus FFS, may not fully prevent overfitting if\ntemporal variability is unaccounted for. In contrast, introducing\ntemporal separation during feature selection and tuning had a\nclear regularizing effect: the temporally blocked (TCV_FFS)\nand spatiotemporal (STCV_FFS) models ended up with much\nlower training R2 (≈0.41 to 0.44) that were nearly matched by\ntheir test R2 (≈0.40 to 0.45). The very small training/testing\nperformance gap for STCV_FFS implies that overfitting was\nreduced, yielding a model that prioritizes generalizable\npredictors. Such conservative models might initially appear to\nperform worse (because they do not chase every idiosyncrasy in\nthe training data), but they in fact proved more robust on", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-1", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "generalizable\npredictors. Such conservative models might initially appear to\nperform worse (because they do not chase every idiosyncrasy in\nthe training data), but they in fact proved more robust on\nindependent data and produced more credible spatial\npredictions. Combining feature selection with temporal or\nspatiotemporal CV narrows the models to predictors that reflect\nstable patterns rather than short-lived spikes.\n\nOur study focuses on generalizable prediction, aiming to\nestimate exposure surfaces (long-term average concentrations)\nacross new locations and periods rather than interpolate within\nthe sampled domain. Unlike spatiotemporal interpolation\nmethods such as kriging, which deliberately exploit spatial\ndependence as the predictive signal,66 our objective is to learn\ntransferable relationships between predictors and pollution\nlevels. In this context, spatial autocorrelation can inflate cross-\nvalidation performance, nudging models to reproduce patterns", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-2", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "s to learn\ntransferable relationships between predictors and pollution\nlevels. In this context, spatial autocorrelation can inflate cross-\nvalidation performance, nudging models to reproduce patterns\ninstead of capturing generalizable processes.36,51,52 The robust-\nness of models with feature selection combined with\nspatiotemporal CV was evident when evaluating on truly\nindependent data. The random CV model showed poor\nagreement, whereas the spatiotemporal CV with feature\nselection had an error nearly half as large. In fact, a model\ntrained with random CV and no outlier filtering completely\nfailed to capture average pollution levels (overall APE ≈217%),\nwhereas the STCV_FFS, STCV_FFS_95, and STCV_FFS_no-\nOut maintained error around 78%−79% even with different\noutlier treatment. This contrast in external validation perform-\nance confirms that models tuned with naive random resampling\nwere overfit to peculiarities of the training campaign, whereas", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-3", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "different\noutlier treatment. This contrast in external validation perform-\nance confirms that models tuned with naive random resampling\nwere overfit to peculiarities of the training campaign, whereas\nincorporating feature selection and spatiotemporal CV yielded\nmodels that better generalize to independent data.\n\nThe model interpretation using SHAP plots provides insight\ninto how feature selection using different CV methods narrowed\nthe predictors down to the most important ones. The overfit\nrandom-CV model that used the full feature set (RCV_HP)\n\neffectively relied on all 194 candidate features, making it difficult\nto interpret and indicating the model was fitting many spurious\nrelationships. In contrast, the models with FFS combined\nspatiotemporal CV retained a much smaller subset of variables.\nThe SHAP analysis consistently identified a few dominant\ndrivers of UFP across these models including traffic and road\ninfrastructure such as road area within 50 m, distance to the", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-4", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "bset of variables.\nThe SHAP analysis consistently identified a few dominant\ndrivers of UFP across these models including traffic and road\ninfrastructure such as road area within 50 m, distance to the\nnearest highway, and traffic volume (AADT). This further\ndemonstrates that the STCV framework successfully identified\nthese variables as sources of nongeneralizable, spurious\ncorrelation. Even meteorological or site-specific variables,\nwhich appeared in some less constrained models, were largely\nabsent or deemphasized in the STCV_FFS models, likely\nbecause those factors do not generalize spatially if they were tied\nto specific days or locations. The fact that the same core features\nemerged under different outlier-handling scenarios suggests that\nthese predictors represent enduring signals rather than artifacts\nof preprocessing choice.\n4.3. Impact of Outlier Treatment\nHandling of extreme UFP readings (“outliers”) had a notable\neffect on models tuned with random CV but had minimal impact", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-5", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ather than artifacts\nof preprocessing choice.\n4.3. Impact of Outlier Treatment\nHandling of extreme UFP readings (“outliers”) had a notable\neffect on models tuned with random CV but had minimal impact\non those built with the spatiotemporal CV + feature selection\nframework. When using random CV for model tuning, the\ninclusion or exclusion of transient extreme spikes significantly\naltered both error metrics and predicted spatial patterns,\nindicating a tendency to “chase” these peaks. In contrast, the\nmodels developed with spatiotemporal CV and feature selection\ndemonstrated a remarkable robustness to outlier handling. The\nSTCV_FFS models performed nearly identically regardless of\nwhether we kept all spikes or filtered them out. For these models,\nwhen predictions were compared with stationary measurements,\nthe errors remained consistently low across all outlier treat-\nments, indicating that the model had learned the underlying\nsignal in the data. Despite this improved performance, the", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-6", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "ary measurements,\nthe errors remained consistently low across all outlier treat-\nments, indicating that the model had learned the underlying\nsignal in the data. Despite this improved performance, the\nSTCV_FFS, STCV_FFS_95, and STCV_FFS_noOut models\nstill tended to overpredict relative to stationary backyard\nmonitors, resulting in a mean absolute percentage error of\nabout 78%−79%. This discrepancy is expected because mobile\nmonitoring measurements were collected directly on roadways,\nwhereas stationary monitors are typically located in residential\nbackyards farther from emission sources. UFP concentrations\nare known to decrease steeply with distance from traffic, often\ndropping by 50%−70% within the first 100−300 m and\napproaching background levels beyond 300 m.67,68 Therefore,\nhigher predicted values compared to backyard observations are\nphysically consistent with the well-established near-road decay\nbehavior of UFPs. The robustness of the STCV_FFS framework", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-7", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": ",68 Therefore,\nhigher predicted values compared to backyard observations are\nphysically consistent with the well-established near-road decay\nbehavior of UFPs. The robustness of the STCV_FFS framework\nto outliers is driven by its ability to remove predictors that do not\nmeaningfully contribute to persistent patterns, while retaining\nthose linked to stable sources such as traffic and road networks.\nThis property is especially important for exposure mapping: it\nmeans the model can incorporate genuine high-pollution events\n(e.g., data recorded while passing a truck or during high traffic\nhours) into the predictions without being misled by them. This\nguarantees that true extreme pollution events are kept and\ncorrectly shown as hotspots if they are part of consistent\npatterns, such as persistently high concentrations along highway\ninterchanges.\n4.4. Modeling Recommendations\nBased on our findings and previous studies, we propose a set of", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-8", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "hey are part of consistent\npatterns, such as persistently high concentrations along highway\ninterchanges.\n4.4. Modeling Recommendations\nBased on our findings and previous studies, we propose a set of\nbest practices for developing generalizable models using\nautocorrelated air quality data collected via mobile modeling.", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "12-9", + "page": 9, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "Define the Modeling Objective: Clearly distinguish between\ninterpolation (within-sample prediction) and generalization\n(prediction at new locations or times), as this determines how\ntraining and validation data are split.\n\nAlign Cross-Validation with the Modeling Goal and Data\nStructure: The structure of CV and the definition of folds should\nreflect both the study objective and the data’s spatiotemporal\ncorrelation structure.\n\nDesign Folds Based on Autocorrelation Range: Spatial folds\nand cluster sizes should reflect the data’s autocorrelation range,\nwith buffer zones applied around validation clusters to prevent\nspatial leakage.\n\nIntegrate CV Into Feature Selection and Tuning: Perform\nfeature selection and hyperparameter tuning within the\nstructured CV to prevent overfitting and the selection of\nspurious predictors tied to specific sampling patterns.\n\nPrioritize External and Visual Validation: For mobile data,\nwhen the goal is to estimate average surfaces over time, rely on", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-0", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "the selection of\nspurious predictors tied to specific sampling patterns.\n\nPrioritize External and Visual Validation: For mobile data,\nwhen the goal is to estimate average surfaces over time, rely on\nindependent continuous measurements and visual inspection to\nconfirm expected spatial trends and detect artifacts. Internal CVs\nusing mobile data should not be interpreted as performance\nindicators. When comparing several models, internal CV metrics\nshould be interpreted comparatively to assess robustness and\noverfitting. Random CV reflects the model’s fit within the\nsampled data, while structured CV (spatial, temporal, or\nspatiotemporal) provides a more realistic measure of general-\nization. Models that maintain more balanced metrics across\nthese CV schemes generally demonstrate stronger real-world\nreliability.\n\n4.5. Future Directions and Implications\n\nFuture work should investigate the transferability of the feature\nselection approach combined with spatiotemporal CV by", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-1", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "rate stronger real-world\nreliability.\n\n4.5. Future Directions and Implications\n\nFuture work should investigate the transferability of the feature\nselection approach combined with spatiotemporal CV by\napplying the model to different time periods and additional\nurban areas. Testing the model’s performance on multiyear\nmobile monitoring data sets in diverse cities or time periods\nwould help determine whether the identified predictors remain\nrobust over time and across various urban contexts. Future work\ncould also test the applicability of this framework to other\npollutants such as NO2 or PM2.5, which spatial and temporal\ncharacteristics differ from UFPs. For such pollutants, the spatial\nfolds in blocked cross-validation should be defined according to\nthe range and structure of their respective autocorrelation\npatterns. Another promising direction is to integrate supple-\nmentary data sets, such as satellite-derived imagery or detailed", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-2", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "according to\nthe range and structure of their respective autocorrelation\npatterns. Another promising direction is to integrate supple-\nmentary data sets, such as satellite-derived imagery or detailed\nmeteorological data, to further refine the model’s predictive\ncapabilities and capture finer-scale variations. Moreover,\nexploring alternative feature selection techniques, or hybrid\nmethods that combine statistical and domain-expertise-driven\napproaches, may yield further improvements in model perform-\nance and interpretability, ultimately leading to more reliable air\npollution exposure assessments.\n\nThis study demonstrates that aligning model development\nwith the spatiotemporal structure of mobile air pollution data\nand the objective of modeling are critical for reliable predictions.\nAmong the tested frameworks, models using spatiotemporal CV\nwith forward feature selection performed best, avoiding\noverfitting and producing stable, interpretable UFP maps. In", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-3", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "reliable predictions.\nAmong the tested frameworks, models using spatiotemporal CV\nwith forward feature selection performed best, avoiding\noverfitting and producing stable, interpretable UFP maps. In\ncontrast, conventional random CV overlooked spatial−tempo-\nral autocorrelation, encouraging data reproduction and unstable\npredictions. By focusing on a small set of meaningful predictors\nand using rigorous CV, we improved both robustness and\nexternal validity. Feature selection and objective-oriented CV\n\nare essential for building exposure models from mobile data,\noffering a reliable path toward high-resolution pollution\nmapping in urban environments.\n■ASSOCIATED CONTENT\n*\nsı Supporting Information\nThe Supporting Information is available free of charge at\nhttps://pubs.acs.org/doi/10.1021/acs.est.5c12601.\n\nSupplementary methods, predictor details, detailed cross-\nvalidation and feature selection workflows, outlier\nanalyses, SHAP plots, and additional figures and tables\n(PDF)", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-4", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "i/10.1021/acs.est.5c12601.\n\nSupplementary methods, predictor details, detailed cross-\nvalidation and feature selection workflows, outlier\nanalyses, SHAP plots, and additional figures and tables\n(PDF)\n■AUTHOR INFORMATION\nCorresponding Author\n\nMarianne Hatzopoulou −Department of Civil and Mineral\n\nEngineering, University of Toronto, Toronto, ON M5S 1A4,\nCanada;\norcid.org/0000-0002-7107-7086; Phone: +1-\n416-978-0864; Email: marianne.hatzopoulou@utoronto.ca\n\nAuthors\n\nMilad Saeedi −Department of Civil and Mineral Engineering,\n\nUniversity of Toronto, Toronto, ON M5S 1A4, Canada;\n\norcid.org/0000-0002-5681-1342\nJad Zalzal −Department of Civil and Mineral Engineering,\n\nUniversity of Toronto, Toronto, ON M5S 1A4, Canada;\n\norcid.org/0000-0002-0444-1124\nArman Ganji −Department of Civil and Mineral Engineering,\n\nUniversity of Toronto, Toronto, ON M5S 1A4, Canada;\n\norcid.org/0000-0001-9462-5324\nJunshi Xu −Department of Geography, University of Hong\n\nKong, Pok Fu Lam 999077, Hong Kong;", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-5", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "nd Mineral Engineering,\n\nUniversity of Toronto, Toronto, ON M5S 1A4, Canada;\n\norcid.org/0000-0001-9462-5324\nJunshi Xu −Department of Geography, University of Hong\n\nKong, Pok Fu Lam 999077, Hong Kong;\norcid.org/0000-\n0003-2834-2291\nSebastian D. Goodfellow −Department of Civil and Mineral\n\nEngineering, University of Toronto, Toronto, ON M5S 1A4,\nCanada\n\nComplete contact information is available at:\nhttps://pubs.acs.org/10.1021/acs.est.5c12601\n\nNotes\nThe authors declare no competing financial interest.\n■ACKNOWLEDGMENTS\nThis study was funded by a grant from the Natural Sciences and\nEngineering Research Council (NSERC) of Canada (RGPIN-\n2022-04597). During manuscript preparation, the authors used\nOpenAI’s ChatGPT to improve clarity and grammar. After using\nthis tool, the authors reviewed and edited the text and take full\nresponsibility for the content.\n■REFERENCES\n\n(1) Lloyd, M.; Olaniyan, T.; Ganji, A.; Xu, J.; Venuta, A.; Simon, L.;", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-6", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "mmar. After using\nthis tool, the authors reviewed and edited the text and take full\nresponsibility for the content.\n■REFERENCES\n\n(1) Lloyd, M.; Olaniyan, T.; Ganji, A.; Xu, J.; Venuta, A.; Simon, L.;\nZhang, M.; Saeedi, M.; Yamanouchi, S.; Wang, A.; Schmidt, A.; Chen,\nH.; Villeneuve, P.; Apte, J.; Lavigne, E.; Burnett, R. T.; Tjepkema, M.;\nHatzopoulou, M.; Weichenthal, S. Airborne Nanoparticle Concen-\ntrations Are Associated with Increased Mortality Risk in Canada’s Two\nLargest Cities. Am. J. Respir. Crit. Care Med. 2024, 210, 1338.\n\n(2) Lloyd, M.; Olaniyan, T.; Ganji, A.; Xu, J.; Simon, L.; Zhang, M.;\nSaeedi, M.; Yamanouchi, S.; Wang, A.; Burnett, R. T.; Tjepkema, M.;\nHatzopoulou, M.; Weichenthal, S. Airborne Ultrafine Particle\nConcentrations and Brain Cancer Incidence in Canada’s Two Largest\nCities. Environ. Int. 2024, 193, 109088.\n\nJ", + "source": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "file_type": "pdf", + "chunk_id": "13-7", + "page": 10, + "document_type": "research_paper", + "page_header": "Environmental Science & Technology | pubs.acs.org/est | Article", + "repository": null, + "relative_path": "bridging-the-gap-between-data-reproduction-and-prediction-the-impact-of-feature-selection-and-cross-validation.pdf", + "section": null + }, + { + "text": "From Data Reproduction to Prediction in Mobile Air Pollution\nMapping\nby\nMilad Saeedi\nA thesis submitted in conformity with the requirements\nfor the degree of Doctor of Philosophy\nDepartment of Civil & Mineral Engineering\nUniversity of Toronto\n© Copyright 2026 by Milad Saeedi", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "14-0", + "page": 1, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "From Data Reproduction to Prediction in Mobile Air Pollution Mapping\nMilad Saeedi\nDoctor of Philosophy\nDepartment of Civil & Mineral Engineering\nUniversity of Toronto\n2026\nAbstract\nTraffic-related air pollution (TRAP) exhibits pronounced spatiotemporal variability within cities,\nwith steep near-road gradients that are often missed by conventional fixed-site monitoring networks.\nMobile monitoring platforms equipped with high–time-resolution sensors offer a promising means\nof capturing this fine-scale variability and supporting high-resolution mapping of urban air qual-\nity. In this thesis, mobile observations of particulate matter (PM2.5), ultrafine particles (UFP),\nand black carbon (BC) are collected across the urban road network and integrated with spatial\npredictors derived from geographic information systems (GIS), meteorological variables, and tem-\nporally varying covariates using land use regression (LUR)–based modeling frameworks to generate\nfine-scale exposure surfaces.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "15-0", + "page": 2, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "geographic information systems (GIS), meteorological variables, and tem-\nporally varying covariates using land use regression (LUR)–based modeling frameworks to generate\nfine-scale exposure surfaces. However, since mobile air-pollution data are highly structured in space\nand time, strong model performance can arise from preprocessing decisions, validation choices, or\nother methodological artifacts rather than from relationships that generalize—reflecting the idea\nthat extensive data manipulation may lead to results that appear to support a desired conclusion.\nIn the first contribution, the thesis establishes a practical framework for producing city-wide\nPM2.5 exposure surfaces from mobile measurements. The analysis demonstrates that low-cost mo-\nbile sensing can extend beyond route-level observations to generate coherent urban concentration\npatterns, particularly when models incorporate predictors that reflect both near-road influences\nand broader background conditions.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "15-1", + "page": 2, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "nd route-level observations to generate coherent urban concentration\npatterns, particularly when models incorporate predictors that reflect both near-road influences\nand broader background conditions. Comparisons with regulatory monitoring stations indicate that\nthe resulting surfaces provide a credible representation of spatial variability suitable for exposure-\noriented applications.\nIn the second contribution, the thesis demonstrates that the credibility of exposure surfaces de-\npends strongly on the modeling framework, particularly for highly variable pollutants. The analysis\nexplicitly accounts for data leakage and autocorrelation, and recognizes that different modeling goals\nmay require different validation strategies and modeling frameworks, showing that the appropriate\nframework depends on the intended use of the surface (e.g., reproducing observed measurements\nversus producing a product that generalizes to genuinely unobserved locations or conditions). Be-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "15-2", + "page": 2, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "opriate\nframework depends on the intended use of the surface (e.g., reproducing observed measurements\nversus producing a product that generalizes to genuinely unobserved locations or conditions). Be-\ncause mobile-monitoring observations are autocorrelated, frameworks built around naive random", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "15-3", + "page": 2, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "resampling can allow information leakage between training and testing data, leading to models that\nappear highly accurate yet largely reproduce structured artifacts rather than transferable relation-\nships. In contrast, autocorrelation-aware frameworks that enforce separation in space and/or time\nand align feature selection and hyperparameter tuning with the same data splits reduce leakage and\noverfitting and yield exposure surfaces that are more stable, interpretable, and physically plausible.\nThese frameworks are also less sensitive to analytical choices such as outlier handling, supporting\nmore defensible surfaces.\nThe third contribution investigates how traffic composition and its interannual changes influence\nurban air pollution by integrating mobile air-quality monitoring with image-based traffic sensing.\nUFP and black carbon (BC) are measured during summer campaigns in 2021 and 2023 using a mobile\nplatform equipped with sensors and a rooftop 360◦camera.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "16-0", + "page": 3, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ity monitoring with image-based traffic sensing.\nUFP and black carbon (BC) are measured during summer campaigns in 2021 and 2023 using a mobile\nplatform equipped with sensors and a rooftop 360◦camera. A deep learning object-detection model\nis used to quantify multiple vehicle classes from imagery, and LUR models are developed to predict\nspatial distributions of both pollutant levels and vehicle-type indicators for each year. Comparing\nthe resulting surfaces enables identification of persistent and emerging hotspots, evaluation of year-\nto-year shifts in traffic-related pollution, and assessment of how fleet composition contributes to\nobserved spatial patterns beyond what is captured by aggregate traffic indicators alone.\nOverall, this dissertation demonstrates an integrated approach for high-resolution TRAP map-\nping that combines mobile sensing, rigorous model evaluation under autocorrelation, and image-\nderived traffic characterization.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "16-1", + "page": 3, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tion demonstrates an integrated approach for high-resolution TRAP map-\nping that combines mobile sensing, rigorous model evaluation under autocorrelation, and image-\nderived traffic characterization.\nThe work contributes practical guidance for producing credible\nexposure surfaces from mobile monitoring data and highlights how composition-aware traffic infor-\nmation can strengthen interpretation of urban pollution patterns and support more reliable exposure\nassessment.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "16-2", + "page": 3, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "List of Abbreviations\nAADT Annual Average Daily Traffic. 2, 14, 25, 26, 33, 35, 57, 61, 67, 69, 70, 72, 74, 77, 79, 123,\n129\nANN Artificial Neural Network. 17, 29\nAPE Average Percentage Error. x, 46, 54, 57, 58, 61, 69, 72, 76, 114, 125\nBC Black Carbon. 1–3, 6–17, 25, 26, 29, 30, 64–72, 75, 76, 78, 79, 81, 83, 84, 122, 126, 129\nCV Cross-Validation. x, 20, 21, 35, 38, 39, 41, 44–63, 108, 109, 111, 113, 116–119\nEMME A Multimodal Transportation Planning Software Package. 33\nFFS Forward Feature Selection. x, 18, 46, 48, 51, 52, 55, 57, 60, 61, 68, 72, 75, 117–119, 121\nGBDT Gradient Boosting Decision Tree. 29\nGIS Geographic Information System. ix, 1, 12–14, 16, 24, 67\nGPS Global Positioning System. 122\nGTAModel Greater Toronto Area travel demand model. 35\nIQR Interquartile Range. 50, 52, 54, 115, 116\nLASSO Least Absolute Shrinkage and Selection Operator. 18, 29, 48\nLOOCV Leave-One-Out Cross-Validation. 29\nLUR Land Use Regression.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "17-0", + "page": 13, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ravel demand model. 35\nIQR Interquartile Range. 50, 52, 54, 115, 116\nLASSO Least Absolute Shrinkage and Selection Operator. 18, 29, 48\nLOOCV Leave-One-Out Cross-Validation. 29\nLUR Land Use Regression. ix, 1, 12, 13, 15–18, 26, 28–30, 33, 35, 38, 41, 44–48, 50, 59, 64–70, 72,\n75–79, 83, 114, 123, 125\nLUR-K Land Use Regression–Kriging. 29\nMAE Mean Absolute Error. 19, 52, 55, 69, 71, 76, 113, 116, 118, 124, 125\nML Machine Learning. ix, 1, 12, 13, 16–18, 22, 46, 49, 81\nMSE Mean Squared Error. 53\nNO Nitric Oxide. 29, 33, 36\nNO2 Nitrogen Dioxide. 31, 33, 36, 47, 63, 65, 123\nNOx Nitrogen Oxides. 6, 33, 36, 79\nPM Particulate Matter. 6, 31", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "17-1", + "page": 13, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "PM10 Particulate Matter smaller than 10 microns. 31\nPM2.5 Particulate Matter smaller than 2.5 microns. x, 1, 6–10, 15, 28, 30–45, 47, 63, 65, 79, 81–83,\n123\nR2 Coefficient of Determination. viii, x, 19, 29, 35, 37–41, 44, 49, 52, 53, 55–57, 60, 69, 71, 74, 76,\n77, 107, 109, 112, 113, 116, 118, 124\nRCV Random Cross-Validation. ix, 20–22, 29, 45, 113, 121\nRMSE Root Mean Squared Error. 19, 52, 55, 69, 71, 76, 113, 116, 118, 124, 125\nSCV Spatial Cross-Validation. ix, 20–22, 50, 51, 53, 109–113, 121\nSHAP SHapley Additive exPlanations. 17, 38, 54, 57, 61, 69, 72, 119, 120, 129\nSTCV spatiotemporal cross-validation. ix, x, 20, 22, 48, 51–53, 61, 68, 72, 75, 77, 110–113, 121\nTCV Temporal Cross-Validation. ix, 20, 22, 53, 110–113, 121\nTRAP Traffic Related Air Pollution. 1, 81\nUFP Ultrafine Particles. x, 1–3, 6–17, 25, 26, 29, 46–50, 54–59, 61–72, 75, 76, 78, 79, 81–84, 107,\n109, 114–116, 118–122, 125–127, 129\nUrbanScanner Mobile sensing platform for air pollution and imagery. 49, 122", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "18-0", + "page": 14, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "e Particles. x, 1–3, 6–17, 25, 26, 29, 46–50, 54–59, 61–72, 75, 76, 78, 79, 81–84, 107,\n109, 114–116, 118–122, 125–127, 129\nUrbanScanner Mobile sensing platform for air pollution and imagery. 49, 122\nXGBoost Extreme Gradient Boosting. 17, 18, 28–30, 34, 35, 39, 48, 50, 54, 55, 57, 67–69, 108,\n109, 112", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "18-1", + "page": 14, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Author Contributions\nThis dissertation includes three manuscripts completed with co-authors, two of which have been pub-\nlished in peer-reviewed journals and one of which is intended for submission. The three manuscripts\nwere completed with co-authors; details of author contributions are provided below.\nChapter 3: “Urban Air Pollution Data Collection, Mapping, and Prediction Using Mo-\nbile Sensors Installed on Courier Trucks.”\nThis manuscript was completed by myself as first author, with Junshi Xu, Usman Ahmed, Matthew\nRoorda, and Marianne Hatzopoulou as co-authors.\nI was responsible for the overall study de-\nsign, data processing, model development, analysis, interpretation of results, and preparation of\nthe manuscript. The mobile monitoring campaign was conducted in collaboration with Junshi Xu,\nwith sensor installation and operational support provided by Usman Ahmed. Matthew Roorda and\nMarianne Hatzopoulou contributed intellectually to the study, provided guidance, comments, and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "19-0", + "page": 15, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Junshi Xu,\nwith sensor installation and operational support provided by Usman Ahmed. Matthew Roorda and\nMarianne Hatzopoulou contributed intellectually to the study, provided guidance, comments, and\nedited the manuscript.\nChapter 4: “Bridging the Gap Between Data Reproduction and Prediction: The Im-\npact of Feature Selection and Cross-Validation Strategies on Prediction of Ambient\nUltrafine Particles Collected with Mobile Monitoring.”\nThis manuscript was completed by myself as first author, with Jad Zalzal, Arman Ganji, Junshi Xu,\nSebastian D. Goodfellow, and Marianne Hatzopoulou as co-authors. I was responsible for develop-\ning the methodological framework, conducting the full analysis, interpreting the results, and leading\nthe manuscript write-up. Jad Zalzal contributed extensively through conceptual discussions and\nmethodological development. Arman Ganji and Junshi Xu provided guidance related to the mobile\nmonitoring data and study design. Sebastian D.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "19-1", + "page": 15, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "contributed extensively through conceptual discussions and\nmethodological development. Arman Ganji and Junshi Xu provided guidance related to the mobile\nmonitoring data and study design. Sebastian D. Goodfellow and Marianne Hatzopoulou contributed\nintellectually, provided critical feedback, and edited the manuscript.\nChapter 5: “Post-COVID Traffic Composition: Why Are Reductions in Ultrafine Par-\nticles and Black Carbon Stalling?”\n(Intended for submission to Environmental Pollution).\nThis manuscript was completed by myself as first author, with Junshi Xu, Jad Zalzal, Arman Ganji,\nand Marianne Hatzopoulou as co-authors. I was responsible for the air pollution modeling, data\nanalysis, interpretation of results, and preparation of the manuscript. Junshi Xu contributed to\nthe development and application of the image-based vehicle detection models. Jad Zalzal provided\nmethodological guidance and contributed to the analytical framework. Arman Ganji contributed to", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "19-2", + "page": 15, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ted to\nthe development and application of the image-based vehicle detection models. Jad Zalzal provided\nmethodological guidance and contributed to the analytical framework. Arman Ganji contributed to\nthe design and execution of the mobile monitoring campaigns. Marianne Hatzopoulou contributed\nintellectually, provided guidance, comments, and edited the manuscript.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "19-3", + "page": 15, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Publication Details\nChapters 3, 4, and 5 of this dissertation are adapted from peer-reviewed journal articles that have\nbeen published or submitted for publication. The details of these publications are listed below.\nChapter 3\nMilad Saeedi, Junshi Xu, Usman Ahmed, Matthew Roorda, and Marianne Hatzopoulou. Urban\nAir Pollution Data Collection, Mapping, and Prediction Using Mobile Sensors Installed on Courier\nTrucks. Transportation Research Record, 2025.\nChapter 4\nMilad Saeedi, Jad Zalzal, Arman Ganji, Junshi Xu, Sebastian D. Goodfellow, and Marianne Hat-\nzopoulou. Bridging the Gap Between Data Reproduction and Prediction: The Impact of Feature\nSelection and Cross-Validation Strategies on Prediction of Ambient Ultrafine Particles Collected\nwith Mobile Monitoring. Environmental Science & Technology, 2026.\nChapter 5\nMilad Saeedi, Junshi Xu, Jad Zalzal, Arman Ganji, and Marianne Hatzopoulou. Post-COVID Traffic", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "20-0", + "page": 16, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ltrafine Particles Collected\nwith Mobile Monitoring. Environmental Science & Technology, 2026.\nChapter 5\nMilad Saeedi, Junshi Xu, Jad Zalzal, Arman Ganji, and Marianne Hatzopoulou. Post-COVID Traffic\nComposition: Why Are Reductions in Ultrafine Particles and Black Carbon Stalling? Manuscript\nready for submission to Environmental Pollution.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "20-1", + "page": 16, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Chapter 1\nIntroduction and research objectives\n1.1\nBackground and motivation\nUrban air pollution remains a major environmental and public health concern, particularly in cities\nwhere dense transportation networks coincide with high population exposure [1, 2]. A substantial\nbody of epidemiological research has linked long-term exposure to Traffic Related Air Pollution\n(TRAP) with adverse cardiovascular and respiratory outcomes, increased mortality, and broader\nsocietal health burdens [3]. Among traffic-related pollutants, Ultrafine Particles (UFP) and Black\nCarbon (BC) are of particular concern due to their strong association with vehicular combustion\nemissions, their pronounced spatial heterogeneity at fine spatial scales, and their ability to penetrate\ndeeply into the respiratory system and contribute to oxidative stress, inflammation, and cardiovascu-\nlar and respiratory diseases[4]. Unlike regulated pollutants such as Particulate Matter smaller than", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "21-0", + "page": 17, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ply into the respiratory system and contribute to oxidative stress, inflammation, and cardiovascu-\nlar and respiratory diseases[4]. Unlike regulated pollutants such as Particulate Matter smaller than\n2.5 microns (PM2.5), which often reflect regional background conditions, UFP and BC exhibit steep\nconcentration gradients near roadways and respond rapidly to local traffic activity, fleet composition,\nand driving dynamics [5].\nAccurate characterization of population exposure to these pollutants therefore requires modeling\napproaches capable of resolving spatial variability at the scale of individual road segments and neigh-\nborhoods [6]. Conventional fixed-site monitoring networks, while essential for regulatory compliance\nand long-term trend analysis, are typically too sparse to capture the fine-scale spatial contrasts\nassociated with traffic emissions [7]. This limitation has motivated the growing use of mobile mon-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "21-1", + "page": 17, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and long-term trend analysis, are typically too sparse to capture the fine-scale spatial contrasts\nassociated with traffic emissions [7]. This limitation has motivated the growing use of mobile mon-\nitoring platforms equipped with compact, high–time-resolution sensors, which enable dense spatial\nsampling of urban environments within relatively short time periods [8, 9]. Mobile monitoring has\ndemonstrated strong potential for revealing near-road pollution gradients and localized hotspots\nthat are not captured by stationary networks [10].\nHowever, translating mobile measurements into reliable, city-wide exposure surfaces presents\nsubstantial methodological challenges. Mobile datasets are inherently structured in space and time,\nwith observations clustered along road networks, repeated across limited sampling days, and influ-\nenced by transient traffic and meteorological conditions [8]. Spatial air pollution models developed", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "21-2", + "page": 17, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "with observations clustered along road networks, repeated across limited sampling days, and influ-\nenced by transient traffic and meteorological conditions [8]. Spatial air pollution models developed\nusing such data must therefore balance the ability to reproduce observed concentrations with the\ncapacity to generalize to unmeasured locations and conditions. Land Use Regression (LUR) models\nand Machine Learning (ML)–enhanced variants have emerged as widely used tools for this purpose,\nas they provide a flexible framework for relating measured concentrations to spatial predictors de-\nrived from Geographic Information System (GIS), traffic indicators, and environmental covariates\n[11, 12].\nDespite their widespread adoption, recent studies have highlighted important limitations in how\nspatial air pollution models are commonly trained, validated, and interpreted. In particular, val-\nidation strategies that ignore spatial or temporal dependence can lead to information leakage and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "21-3", + "page": 17, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ow\nspatial air pollution models are commonly trained, validated, and interpreted. In particular, val-\nidation strategies that ignore spatial or temporal dependence can lead to information leakage and\noverly optimistic estimates of predictive performance [13]. Models may appear highly accurate when", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "21-4", + "page": 17, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "evaluated using random data splits, yet fail to generalize to genuinely unobserved locations or time\nperiods [14]. At the same time, the increasing use of flexible machine learning models introduces\nadditional challenges related to interpretability and robustness, including the risk that models ex-\nploit spurious correlations or artifacts in structured environmental data rather than learning stable\nemission–concentration relationships [15].\nAccurate representation of traffic activity represents a further critical constraint in modeling\ntraffic-related air pollution. Many studies rely on aggregate traffic indicators such as Annual Average\nDaily Traffic (AADT), which provide long-term averages but do not capture short-term variation\nin vehicle mix or reflect conditions encountered during mobile monitoring campaigns [16].\nYet\nemissions of BC and UFP are highly sensitive to traffic composition, with heavy-duty and diesel-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "22-0", + "page": 18, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ariation\nin vehicle mix or reflect conditions encountered during mobile monitoring campaigns [16].\nYet\nemissions of BC and UFP are highly sensitive to traffic composition, with heavy-duty and diesel-\npowered vehicles contributing disproportionately to local pollutant levels [17]. As a result, spatial\nvariation in fleet composition can strongly influence observed pollution patterns even where total\ntraffic volumes are similar [18].\nRecent advances in computer vision offer new opportunities to address this gap by enabling di-\nrect, data-driven characterization of traffic composition from street-level and 360-degree imagery\n[19]. Deep learning–based object detection models, such as those in the YOLO family, allow auto-\nmated detection and classification of vehicles at high spatial and temporal resolution [20]. When\nintegrated with mobile monitoring platforms, these methods provide an independent and physi-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "22-1", + "page": 18, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ow auto-\nmated detection and classification of vehicles at high spatial and temporal resolution [20]. When\nintegrated with mobile monitoring platforms, these methods provide an independent and physi-\ncally interpretable means of capturing traffic activity along sampled road networks, complementing\ntraditional traffic datasets and improving representation of emission sources [21].\nMotivated by these challenges and opportunities, this thesis investigates the integration of mobile\nair pollution monitoring, robust spatial modeling frameworks, and image-based traffic composition\nestimation to improve high-resolution prediction and interpretation of urban air pollution. Rather\nthan focusing solely on predictive accuracy, the work emphasizes methodological rigor, appropriate\nvalidation under spatial and temporal dependence, and interpretability of model behavior.\nBy\nsystematically examining data preprocessing strategies, modeling approaches, validation designs,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "22-2", + "page": 18, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "propriate\nvalidation under spatial and temporal dependence, and interpretability of model behavior.\nBy\nsystematically examining data preprocessing strategies, modeling approaches, validation designs,\nand traffic representation, this research aims to advance exposure assessment for traffic-related air\npollutants and support more reliable inference in complex urban environments.\n1.2\nProblem statement\nHigh-resolution exposure assessment in cities requires concentration surfaces that capture fine-scale\nvariability driven by transportation activity, yet producing reliable and interpretable maps remains\nchallenging when measurements are collected using mobile platforms and models are trained on\nhighly structured spatiotemporal data. In practice, exposure models must simultaneously (i) lever-\nage the dense spatial coverage of mobile monitoring, (ii) avoid overfitting and information leakage\ncaused by spatial/temporal dependence, and (iii) represent traffic emissions in a way that is phys-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "22-3", + "page": 18, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "age the dense spatial coverage of mobile monitoring, (ii) avoid overfitting and information leakage\ncaused by spatial/temporal dependence, and (iii) represent traffic emissions in a way that is phys-\nically meaningful and interpretable (e.g., accounting for fleet composition rather than only aggre-\ngate volume). This thesis addresses this overarching problem through three complementary studies\n(Chapters 3–5) that investigate how mobile data collection design, modeling/validation strategy,\nand traffic characterization influence the reliability, generalizability, and interpretability of urban\nair-pollution prediction surfaces.\n1. A fundamental problem is how to transform route-limited and temporally uneven mobile mea-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "22-4", + "page": 18, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "surements into stable, city-wide concentration surfaces suitable for exposure assessment. This\nincludes selecting appropriate preprocessing and spatial aggregation, identifying predictors\nthat reflect emissions and dispersion environments, and developing models that remain robust\ndespite heterogeneous sampling density and limited temporal representativeness.\n2. For highly variable pollutants measured with mobile platforms, standard model-development\npipelines (hyperparameter tuning, feature selection, and evaluation under random splits) can\nexploit spatial and temporal dependence, leading to information leakage and inflated perfor-\nmance. The resulting models may appear accurate while failing to generalize to genuinely\nunobserved roads, neighborhoods, or sampling periods. The problem is to design modeling\nframeworks workflows that are explicitly aligned with the intended prediction task under spa-\ntiotemporal dependence.\n3.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "23-0", + "page": 19, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ed roads, neighborhoods, or sampling periods. The problem is to design modeling\nframeworks workflows that are explicitly aligned with the intended prediction task under spa-\ntiotemporal dependence.\n3. A persistent limitation in traffic-related pollution modeling is inadequate representation of\ntraffic emissions. Aggregate traffic indicators often fail to capture variation in fleet composi-\ntion, even though heavy-duty and diesel-powered vehicles can contribute disproportionately to\npollutants such as BC and UFP. The problem is to obtain spatially and temporally aligned\ntraffic-composition information at scale and integrate it into modeling frameworks in a phys-\nically interpretable way that improves explanation and prediction of localized pollution pat-\nterns.\nIn summary, this thesis develops an integrated framework for high-resolution urban air-pollution\nmapping that reconciles the strengths of mobile monitoring with the challenges posed by structured\nspatiotemporal data.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "23-1", + "page": 19, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "thesis develops an integrated framework for high-resolution urban air-pollution\nmapping that reconciles the strengths of mobile monitoring with the challenges posed by structured\nspatiotemporal data. By aligning data collection, modeling, validation, and traffic characterization\nwith the intended prediction task, the work aims to produce exposure surfaces that are credible,\ntransferable across space and time, and useful for exposure science and mitigation planning.\n1.3\nObjectives and research questions\n1.3.1\nOverall objective\nThe overall objective of this thesis is to develop and evaluate an integrated framework for high-\nresolution urban air pollution mapping using mobile monitoring data, with an emphasis on (i)\nproducing spatially detailed concentration surfaces for traffic-related pollutants, (ii) ensuring scien-\ntifically valid model evaluation under spatiotemporal dependence, and (iii) improving interpretation", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "23-2", + "page": 19, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ing spatially detailed concentration surfaces for traffic-related pollutants, (ii) ensuring scien-\ntifically valid model evaluation under spatiotemporal dependence, and (iii) improving interpretation\nof observed and predicted pollution patterns through physically meaningful traffic characterization,\nincluding traffic composition derived from imagery.\n1.3.2\nResearch objectives and questions\nThis research aims to develop robust and generalizable frameworks for predicting traffic-related\nair pollutants using mobile monitoring data. The specific objectives and corresponding research\nquestions are as follows:\n1. Develop high-resolution exposure surfaces from mobile monitoring data. How can\nroute-constrained and temporally uneven mobile measurements be transformed into stable,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "23-3", + "page": 19, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "city-wide concentration surfaces that capture fine-scale spatial variability relevant for exposure\nassessment?\n2. Establish modeling and validation frameworks that ensure generalization under\nspatial and temporal dependence. How does a modeling framework that aligns model\nselection, feature selection, and objective-aware validation with the underlying spatial and\ntemporal structure of the data affect model performance, robustness, and transferability to\ngenuinely unobserved locations and conditions?\n3. Incorporate traffic composition using computer vision to improve model inter-\npretability and performance. To what extent does integrating high-resolution traffic com-\nposition derived from street-level imagery improve model interpretability and the ability to\nexplain spatial variability in traffic-related pollutants, including temporal changes such as\nthose observed during and after COVID-19?\n1.4\nResearch significance", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "24-0", + "page": 20, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "l interpretability and the ability to\nexplain spatial variability in traffic-related pollutants, including temporal changes such as\nthose observed during and after COVID-19?\n1.4\nResearch significance\nThis dissertation advances high-resolution exposure assessment for traffic-related air pollutants by\naddressing key methodological and data limitations that affect the credibility, interpretability, and\npredictive usefulness of urban air pollution maps. Using mobile monitoring data and spatial model-\ning, the work develops fine-scale concentration surfaces for particulate matter (PM2.5), black carbon\n(BC), and ultrafine particles (UFP), with a particular focus on near-road environments where spatial\nvariability is pronounced and population exposure is highly heterogeneous. The research demon-\nstrates the feasibility of producing coherent, city-wide exposure surfaces from mobile measurements", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "24-1", + "page": 20, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tial\nvariability is pronounced and population exposure is highly heterogeneous. The research demon-\nstrates the feasibility of producing coherent, city-wide exposure surfaces from mobile measurements\nwhile explicitly confronting the additional challenges posed by highly variable pollutants such as BC\nand UFP.\nA central contribution of this research is its emphasis on moving beyond data reproduction\ntoward defensible spatial prediction, while strengthening rigor in spatial model development and\ninterpretation. Rather than prioritizing agreement with observed measurements under convenient\nvalidation schemes, the dissertation reframes model performance in terms of generalization to gen-\nuinely unobserved locations, road segments, and conditions. By explicitly accounting for spatial\nand temporal dependence in mobile monitoring data, this work shows that commonly used random\nresampling approaches can introduce information leakage and overestimate model performance. In", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "24-2", + "page": 20, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "for spatial\nand temporal dependence in mobile monitoring data, this work shows that commonly used random\nresampling approaches can introduce information leakage and overestimate model performance. In\ncontrast, the proposed modeling and validation frameworks are aligned with realistic prediction ob-\njectives and are complemented by diagnostic and interpretability analyses that reduce the risk of\nflexible statistical and machine-learning models exploiting spurious correlations or structured ar-\ntifacts in environmental data. Together, these contributions support more transparent inference,\nenhance confidence in the physical plausibility of predicted surfaces, and improve the credibility and\ntransferability of derived exposure maps.\nIn addition, the work significantly improves representation of traffic composition for understand-\ning exposure surfaces by moving beyond coarse, long-term traffic indicators toward detailed char-\nacterization of traffic composition.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "24-3", + "page": 20, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "y improves representation of traffic composition for understand-\ning exposure surfaces by moving beyond coarse, long-term traffic indicators toward detailed char-\nacterization of traffic composition.\nThrough computer-vision–based vehicle detection applied to\nstreet-level or 360-degree imagery, the dissertation provides a scalable and data-driven pathway to\nderive spatially explicit and temporally aligned indicators of fleet mix. This enhances the physical", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "24-4", + "page": 20, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "interpretability of modeled concentration patterns and supports attribution of localized pollution\nvariability to plausible emission sources, particularly in contexts where heavy-duty and diesel vehicles\ndisproportionately influence BC and UFP levels.\nCollectively, these contributions improve the scientific credibility and practical utility of high-\nresolution air pollution predictions for exposure science and public health research. More reliable and\ntransferable concentration surfaces can reduce exposure misclassification in epidemiological studies,\nsupport identification of localized hotspots and traffic-affected communities, and inform targeted\nmitigation strategies and urban planning decisions aimed at reducing inequities in traffic-related air\npollution exposure.\n1.5\nStructure and overview of chapters\nThis thesis is organized into six chapters. Chapter 1 introduces the research context, motivates the\nwork, and defines the problem addressed in this dissertation.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "25-0", + "page": 21, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ".5\nStructure and overview of chapters\nThis thesis is organized into six chapters. Chapter 1 introduces the research context, motivates the\nwork, and defines the problem addressed in this dissertation. It then outlines the overall objective,\nspecific objectives, and research questions that guide the thesis. Chapter 2 provides the scientific and\nmethodological background required for the remainder of the dissertation, including traffic-related air\npollution, mobile monitoring, spatial air pollution modeling frameworks, predictor development, and\nkey issues in model evaluation, validation, and interpretation under spatial and temporal dependence.\nChapters 3–5 present the three core studies that form the dissertation. Chapter 3 presents the first\nstudy, which develops a high-resolution modeling framework for estimating traffic-related pollutant\nconcentrations using mobile monitoring measurements and spatial predictors, and evaluates the\nresulting exposure surfaces.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "25-1", + "page": 21, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "a high-resolution modeling framework for estimating traffic-related pollutant\nconcentrations using mobile monitoring measurements and spatial predictors, and evaluates the\nresulting exposure surfaces. Chapter 4 presents the second study, which focuses on methodological\naspects of model development and evaluation, emphasizing generalization, robust validation, and\nthe implications of spatial and temporal dependence for performance assessment and interpretation.\nChapter 5 presents the third study, which advances traffic characterization by estimating traffic\ncomposition using computer vision on street-level or 360-degree imagery and examines how this\ninformation can support interpretation and improve representation of traffic-related emission drivers\nwithin spatial air pollution modeling.\nFinally, Chapter 6 concludes the dissertation by summarizing the main findings across the three\nstudies, highlighting the methodological and practical contributions to high-resolution exposure", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "25-2", + "page": 21, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "eling.\nFinally, Chapter 6 concludes the dissertation by summarizing the main findings across the three\nstudies, highlighting the methodological and practical contributions to high-resolution exposure\nassessment, discussing key limitations, and outlining directions for future research.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "25-3", + "page": 21, + "document_type": "thesis", + "page_header": "CHAPTER 1. INTRODUCTION AND RESEARCH OBJECTIVES", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Chapter 2\nLiterature review\nAir quality is a critical environmental concern with significant implications for ecosystems and human\nhealth worldwide. Exposure to ambient air pollutants affects millions of people and has been linked\nto a wide range of adverse health outcomes. Among these pollutants, Particulate Matter (PM)—\nincluding UFP, BC, and PM2.5—as well as Nitrogen Oxides (NOx) and ozone, are of particular\nconcern due to their prevalence and toxicity. UFP (particles smaller than 100 nm) and BC exhibit\ndistinct physical and chemical behaviors compared to larger particulate matter fractions, as they can\npenetrate deep into the respiratory system and, in the case of UFP, translocate into the bloodstream,\nthereby posing elevated risks to human health [22, 23, 2, 24].\nPM2.5 is one of the most extensively studied air pollutants due to its widespread occurrence, abil-\nity to remain suspended in the atmosphere, and well-established associations with adverse health\noutcomes.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "26-0", + "page": 22, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "is one of the most extensively studied air pollutants due to its widespread occurrence, abil-\nity to remain suspended in the atmosphere, and well-established associations with adverse health\noutcomes. PM2.5 originates from both primary emissions, such as combustion processes and mechan-\nical abrasion, and secondary formation through atmospheric chemical reactions involving gaseous\nprecursors including sulfur dioxide, nitrogen oxides, and volatile organic compounds. Due to its size,\nPM2.5 can penetrate deep into the lower respiratory tract and has been linked to cardiovascular and\nrespiratory morbidity, premature mortality, and systemic inflammation [1, 25]. Unlike UFP and BC,\nwhich often exhibit strong local spatial variability driven by traffic activity, PM2.5 typically reflects\na combination of local sources and regional background contributions, resulting in smoother spatial\ngradients at the urban scale. This characteristic has made PM2.5 a central pollutant in epidemio-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "26-1", + "page": 22, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s\na combination of local sources and regional background contributions, resulting in smoother spatial\ngradients at the urban scale. This characteristic has made PM2.5 a central pollutant in epidemio-\nlogical studies and regulatory frameworks, while also posing challenges for source attribution and\nhigh-resolution exposure assessment in complex urban environments[26, 27, 28].\nUnderstanding the spatiotemporal distribution of air pollutant concentrations is essential for\ndeveloping effective mitigation strategies and for generating accurate exposure surfaces (i.e., air\npollution maps) used in long-term health assessments [6, 29].\nWithin urban environments, air\npollution levels exhibit pronounced variability across both space and time, particularly for primary\npollutants such as UFP and BC, which are strongly influenced by local emission sources. Black\ncarbon, also known as soot, is widely recognized as a marker of diesel exhaust and traffic-related\ncombustion emissions [4].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "26-2", + "page": 22, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "UFP and BC, which are strongly influenced by local emission sources. Black\ncarbon, also known as soot, is widely recognized as a marker of diesel exhaust and traffic-related\ncombustion emissions [4].\nA central challenge in epidemiological studies investigating the long-\nterm health impacts of air pollution is the accurate characterization of population exposure. This\nrequires reliable methods to measure, model, and predict pollutant concentrations at fine spatial\nand temporal scales that reflect real-world human activity patterns [30].\n2.1\nTraffic-related air pollution\nAir pollution originates from a variety of natural and anthropogenic sources; however, road traffic\nis a key contributor to air quality degradation in urban environments [31, 32].\nIn urban areas,\nlarge numbers of vehicles and the close proximity of roadways to human activity result in higher\npollutant concentrations near roads, making traffic-related air pollution (TRAP) an important focus", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "26-3", + "page": 22, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "of environmental and public health research [33, 34]. Road traffic emissions generate a complex\nmixture of gaseous pollutants and particulate matter that are released directly from motor vehicles\nor formed through atmospheric transformation processes. This mixture includes primary pollutants\nemitted during fuel combustion as well as secondary pollutants produced through chemical reactions\ninvolving traffic-related precursors in the atmosphere [35].\nAmong traffic-related pollutants, BC and UFP are of particular interest due to their strong as-\nsociation with vehicular combustion and their pronounced spatial and temporal variability [27, 36].\nUnlike regulated pollutants such as PM2.5, which often reflect broader regional background condi-\ntions, BC and UFP respond rapidly to changes in traffic activity, fleet composition, and driving\nconditions. These characteristics make them effective tracers of local traffic emissions while simul-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "27-0", + "page": 23, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tions, BC and UFP respond rapidly to changes in traffic activity, fleet composition, and driving\nconditions. These characteristics make them effective tracers of local traffic emissions while simul-\ntaneously posing challenges for accurate characterization using conventional fixed-site monitoring\napproaches [37, 38].\nTraffic-related emissions originate from both exhaust and non-exhaust sources associated with\nvehicle operation and roadway use[39]. Exhaust emissions result from fuel combustion in vehicle\nengines and are the primary source of BC and UFP, particularly under high-load or transient driving\nconditions[40]. Non-exhaust sources, including brake wear, tire wear, and the resuspension of road\ndust, contribute primarily to coarse and fine particulate matter but may also influence local particle\nconcentrations in heavily trafficked environments[41].\nThe relative contribution of these sources", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "27-1", + "page": 23, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ust, contribute primarily to coarse and fine particulate matter but may also influence local particle\nconcentrations in heavily trafficked environments[41].\nThe relative contribution of these sources\ndepends on vehicle technology, fleet composition, roadway characteristics, and driving behavior[42,\n43].\nVehicle type and fuel technology play a central role in shaping traffic-related emission profiles[44].\nHeavy-duty vehicles, particularly those powered by diesel engines, emit disproportionately higher\nlevels of BC and UFP compared to light-duty gasoline vehicles[45, 46]. Consequently, traffic com-\nposition, specifically the proportion of heavy-duty and diesel vehicles, can strongly influence local\npollutant concentrations even when overall traffic volumes are moderate[47]. Emission rates are\nfurther affected by driving dynamics such as congestion, frequent acceleration and deceleration, and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "27-2", + "page": 23, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "al\npollutant concentrations even when overall traffic volumes are moderate[47]. Emission rates are\nfurther affected by driving dynamics such as congestion, frequent acceleration and deceleration, and\nstop-and-go traffic, which are commonly associated with localized pollution hotspots near intersec-\ntions, highway ramps, and signalized road segments[48, 49, 50].\nThe spatial distribution of BC and UFP in urban environments is characterized by pronounced\nheterogeneity and steep near-roadway gradients. Concentrations typically peak adjacent to traffic\ncorridors and decrease rapidly with increasing distance from the roadway, with the most pronounced\ndeclines often occurring within the first few hundred meters. These spatial patterns are influenced by\ntraffic intensity, vehicle composition, roadway configuration, and features of the built environment\nsuch as building density, street orientation, and street-canyon geometry, which affect airflow and\npollutant dispersion[51, 52].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "27-3", + "page": 23, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "position, roadway configuration, and features of the built environment\nsuch as building density, street orientation, and street-canyon geometry, which affect airflow and\npollutant dispersion[51, 52].\nThis strong fine-scale variability has important implications for exposure assessment. Fixed-site\nmonitoring networks, which are generally sparse and designed to represent neighborhood- or city-\nscale conditions, often fail to capture localized exposure contrasts associated with traffic emissions.\nAs a result, individuals living, working, or commuting within the same urban area may experience\nsubstantially different exposure levels depending on proximity to traffic sources. This limitation has\nmotivated the development and application of high-resolution spatial approaches to better charac-\nterize traffic-related pollution patterns and associated health risks[7, 10, 53].\nA substantial body of epidemiological evidence has linked exposure to traffic-related air pollu-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "27-4", + "page": 23, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tion with adverse health outcomes, including impacts on cardiovascular and respiratory systems,\npregnancy outcomes, and overall mortality. BC is commonly used as a marker of traffic-related\nparticulate matter and has been associated with inflammatory responses, oxidative stress, and car-\ndiovascular dysfunction[3, 54]. Collectively, this literature underscores the importance of accurately\ncharacterizing the fine-scale spatial variability of traffic-related air pollution to support robust ex-\nposure assessment and inform mitigation strategies.\n2.2\nPhysical and chemical properties\nAirborne particulate pollutants differ substantially in their physical and chemical properties, which\ninfluence their atmospheric behavior, spatial distribution, and relevance for exposure assessment.\nKey characteristics include particle size, surface area, chemical composition, atmospheric lifetime,\nand susceptibility to dispersion and transformation processes. These properties play a central role in", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "28-0", + "page": 24, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "characteristics include particle size, surface area, chemical composition, atmospheric lifetime,\nand susceptibility to dispersion and transformation processes. These properties play a central role in\ndetermining how pollutants are emitted, transported, and measured within urban environments[55,\n56].\nPM2.5 represents a heterogeneous mixture of solid and liquid components. Owing to their rela-\ntively larger size compared to ultrafine particles, PM2.5 particles typically exhibit longer atmospheric\nlifetimes and are influenced by both local emission sources and regional background contributions.\nAs a result, PM2.5 concentrations often display smoother spatial gradients at the urban scale and\nare less sensitive to short-term fluctuations in traffic activity. Chemically, PM2.5 includes a range\nof constituents such as sulfates, nitrates, organic carbon, elemental carbon, metals, and crustal ma-\nterial, reflecting contributions from combustion processes, secondary atmospheric formation, and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "28-1", + "page": 24, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "of constituents such as sulfates, nitrates, organic carbon, elemental carbon, metals, and crustal ma-\nterial, reflecting contributions from combustion processes, secondary atmospheric formation, and\nmechanical sources[57, 58].\nUFP are characterized by extremely high number concentrations and large surface area relative\nto their mass. These properties make UFP highly reactive and sensitive to atmospheric processes\nsuch as coagulation, condensation, and rapid dilution following emission[59, 60]. Due to their short\natmospheric lifetimes, UFP concentrations typically exhibit strong spatial and temporal variability,\nwith pronounced gradients near emission sources. In urban environments, UFP are closely associated\nwith combustion-related activities, particularly road traffic, and respond rapidly to changes in vehicle\nactivity and driving conditions[61, 62].\nBC is a carbonaceous component of particulate matter formed through incomplete combustion\nof fossil fuels and biomass.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "28-2", + "page": 24, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "d respond rapidly to changes in vehicle\nactivity and driving conditions[61, 62].\nBC is a carbonaceous component of particulate matter formed through incomplete combustion\nof fossil fuels and biomass.\nUnlike PM2.5 mass, BC is commonly quantified based on its light-\nabsorbing properties and serves as a tracer of combustion-related emissions, especially from diesel-\npowered vehicles. BC particles often reside within the fine and ultrafine size fractions and exhibit\nstrong spatial contrasts near traffic corridors. Compared to UFP, BC tends to be more stable in\nthe atmosphere, allowing it to persist over longer distances while still retaining sensitivity to local\nemission sources[63, 62, 64].\nThe contrasting physical and chemical characteristics of PM2.5, BC, and UFP have important\nimplications for exposure assessment and modeling. While PM2.5 mass concentrations are effec-\ntive for capturing broader spatial patterns and long-term average exposure, BC and UFP provide", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "28-3", + "page": 24, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "important\nimplications for exposure assessment and modeling. While PM2.5 mass concentrations are effec-\ntive for capturing broader spatial patterns and long-term average exposure, BC and UFP provide\nenhanced sensitivity to local combustion sources and near-roadway variability. Consequently, in-\ntegrating multiple particulate metrics is often necessary to capture the full range of spatial scales\nrelevant to urban air pollution exposure.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "28-4", + "page": 24, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "2.3\nMeasurement methods\nAccurate characterization of particulate air pollution depends on measurement techniques that are\ncompatible with the physical and chemical properties of the target pollutants. Because PM2.5, BC,\nand UFP differ substantially in particle size, mass contribution, and optical behavior, they cannot be\nquantified using a single measurement approach. Distinct monitoring methods have therefore been\ndeveloped, each with specific strengths and limitations that influence spatial resolution, temporal\ncoverage, and suitability for exposure assessment[65].\nMonitoring of PM2.5 is primarily based on mass concentration measurements and forms the\nfoundation of regulatory air quality monitoring worldwide. Reference and equivalent methods such\nas gravimetric samplers, beta attenuation monitors, and tapered element oscillating microbalances\nprovide robust long-term measurements and support epidemiological analyses and regulatory com-\npliance.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "29-0", + "page": 25, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s gravimetric samplers, beta attenuation monitors, and tapered element oscillating microbalances\nprovide robust long-term measurements and support epidemiological analyses and regulatory com-\npliance. However, these instruments are typically deployed at a limited number of fixed locations\nand are designed to capture neighborhood- or city-scale conditions rather than fine-scale spatial\nvariability. Optical particle sensors have increasingly been used to supplement fixed-site networks\nand enable denser spatial coverage, although their measurements require careful calibration due to\nsensitivity to particle composition and environmental conditions[65, 66, 67, 68].\nUFP contribute negligibly to total particulate mass despite dominating particle number concen-\ntrations. Consequently, UFP cannot be reliably characterized using mass-based monitoring tech-\nniques and are instead measured using number-based approaches that are sensitive to nanometer-\nscale particles[69, 70].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "29-1", + "page": 25, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "equently, UFP cannot be reliably characterized using mass-based monitoring tech-\nniques and are instead measured using number-based approaches that are sensitive to nanometer-\nscale particles[69, 70]. Condensation particle counters (CPCs) are the most widely used instruments\nfor measuring total UFP number concentration. In these instruments, sampled air is exposed to\na supersaturated vapor (commonly butanol or water), causing nanometer-sized particles to grow\ninto larger droplets through condensation. The enlarged particles are then detected and counted\noptically as they pass through a laser beam, allowing direct quantification of particle number concen-\ntration[71]. Mobility-based instruments provide additional information on particle size by classifying\nUFP according to their electrical mobility. Particles are first electrically charged and then passed\nthrough an electric field, where their trajectories depend on particle size and charge. By selectively", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "29-2", + "page": 25, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "according to their electrical mobility. Particles are first electrically charged and then passed\nthrough an electric field, where their trajectories depend on particle size and charge. By selectively\ntransmitting particles of specific mobilities to a detector, these instruments generate size-resolved\nnumber distributions. Together, CPCs and mobility-based techniques enable detailed characteriza-\ntion of UFP number concentrations and size spectra[72].\nBecause UFP undergo rapid atmospheric processes such as coagulation, condensation, and di-\nlution following emission, measurements are highly sensitive to temporal averaging and sampling\nlocation. These characteristics make UFP instruments particularly effective for capturing short-\nterm emission dynamics and near-source variability, while also contributing to their limited use in\nlong-term regulatory monitoring due to cost, operational complexity, and data intensity[73].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "29-3", + "page": 25, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "short-\nterm emission dynamics and near-source variability, while also contributing to their limited use in\nlong-term regulatory monitoring due to cost, operational complexity, and data intensity[73].\nIn addition to condensation (CPCs) and mobility-based techniques, diffusion-charging–based\nportable sensors have been widely adopted for mobile and personal exposure assessment of UFP. In\nthese instruments, aerosol particles are exposed to a unipolar corona charger, where nanometer-sized\nparticles acquire electrical charge proportional to their surface area. Excess ions are subsequently\nremoved using an ion trap, and the charged aerosol is directed through multiple electrometer stages.\nBy measuring the electrical current associated with particle deposition in diffusion and filter stages,\nthese sensors estimate UFP number concentration and, in some designs, an effective mean particle", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "29-4", + "page": 25, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "diameter.\nOwing to their compact size, low power consumption, and high temporal resolution,\ndiffusion-charging sensors are particularly well suited for mobile monitoring applications, although\nthey do not provide fully size-resolved particle distributions[74].\nMeasurement of BC most commonly exploits its strong light-absorbing properties and is typi-\ncally performed using optical or thermal-based techniques[75]. Continuous BC monitoring is widely\nconducted using optical absorption instruments, such as aethalometers, which estimate BC concen-\ntration by measuring the attenuation of light transmitted through particles collected on a filter. As\nBC accumulates on the filter, increased light absorption is interpreted as higher BC mass concentra-\ntion, allowing high temporal resolution measurements that are particularly well suited for capturing\nshort-term variability associated with combustion-related emissions[76].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "30-0", + "page": 26, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "higher BC mass concentra-\ntion, allowing high temporal resolution measurements that are particularly well suited for capturing\nshort-term variability associated with combustion-related emissions[76].\nIn addition to optical methods, BC can also be quantified using thermal or thermal–optical\ntechniques applied to filter samples. These approaches separate carbonaceous aerosol components\nbased on their volatility at different temperatures, enabling distinction between elemental carbon\nand organic carbon fractions. While thermal-based methods provide detailed characterization of\ncarbonaceous particles, they are typically limited to offline analysis and lower temporal resolution\ncompared to optical instruments[77].\nOptical BC instruments are commonly deployed in both fixed-site and mobile monitoring ap-\nplications due to their portability and high time resolution. However, BC measurements can be", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "30-1", + "page": 26, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "nstruments[77].\nOptical BC instruments are commonly deployed in both fixed-site and mobile monitoring ap-\nplications due to their portability and high time resolution. However, BC measurements can be\ninfluenced by particle mixing state, coating effects, and instrument-specific assumptions related to\nlight absorption properties[77]. As a result, careful interpretation is required when comparing BC\nconcentrations across instruments, environments, or studies[78, 79]. Despite these limitations, BC\nmonitoring provides a valuable proxy for combustion-related particulate pollution and complements\nmass-based PM2.5 and number-based UFP measurements in exposure assessment.\n2.4\nTraffic composition and pollutant emissions\nTraffic composition plays a central role in determining both the magnitude and spatial variability\nof traffic-related air pollutant emissions in urban environments. Although total traffic volume is", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "30-2", + "page": 26, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s\nTraffic composition plays a central role in determining both the magnitude and spatial variability\nof traffic-related air pollutant emissions in urban environments. Although total traffic volume is\ncommonly used as a proxy for emission intensity, pollutant emissions vary substantially across vehicle\nclasses due to differences in engine technology, fuel type, vehicle mass, and operating conditions.\nAs a result, aggregate traffic metrics alone may provide an incomplete representation of emission\nsources and associated air quality impacts[80, 81, 17].\nEmission factors for traffic-related pollutants differ markedly by vehicle class. Heavy-duty ve-\nhicles, particularly those powered by diesel engines, generally exhibit substantially higher emission\nfactors for pollutants such as BC and UFP than light-duty gasoline vehicles. These differences reflect\nvariations in combustion processes, after-treatment technologies, and engine operating regimes, im-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "30-3", + "page": 26, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tors for pollutants such as BC and UFP than light-duty gasoline vehicles. These differences reflect\nvariations in combustion processes, after-treatment technologies, and engine operating regimes, im-\nplying that even modest fractions of heavy-duty vehicles can contribute disproportionately to total\nemissions[40, 82, 83].\nEmpirical studies have shown that the spatial distribution of BC and UFP in urban environments\nis closely associated with freight traffic, buses, and other heavy-duty vehicle activity.\nElevated\nconcentrations are frequently observed along truck corridors, bus routes, and near logistics hubs,\nhighlighting the localized influence of vehicle classes with high emission intensities[84, 18, 85].\nThese effects are particularly pronounced in near-road environments, where emissions occur in", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "30-4", + "page": 26, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "close proximity to human activity and traffic composition strongly modulates observed pollutant\nlevels. Vehicle-class–specific information therefore provides a more direct link between traffic activity\nand combustion-related pollutant concentrations than aggregate traffic indicators. Incorporating\nhigh-resolution traffic composition data improves identification of dominant emission drivers and\nsupports interpretation of fine-scale spatial and temporal variability in urban air pollution.\n2.5\nTemporal changes in urban air pollution\nUrban air pollution exhibits substantial temporal variability across multiple time scales, driven by\nchanges in emission patterns, meteorological conditions, and broader societal factors. Understanding\ntemporal dynamics is essential for interpreting observed concentration differences across monitoring\ncampaigns, assessing long-term trends, and ensuring that predictive models capture variability be-\nyond purely spatial contrasts.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "31-0", + "page": 27, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "for interpreting observed concentration differences across monitoring\ncampaigns, assessing long-term trends, and ensuring that predictive models capture variability be-\nyond purely spatial contrasts. Temporal changes are particularly relevant in studies that integrate\ndata collected across different years, seasons, or operating conditions[86, 87, 88, 89].\nYear-to-year variations in air pollutant concentrations have been observed in many urban environ-\nments, reflecting changes in traffic activity, fleet composition, emission regulations, and background\natmospheric conditions. The COVID-19 pandemic represents a notable example of how abrupt soci-\netal changes can influence urban air quality. Restrictions on mobility, changes in commuting behav-\nior, and reductions in traffic volume during lockdown periods led to measurable changes in pollutant\nconcentrations in many cities. Traffic-related pollutants exhibited varying responses depending on", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "31-1", + "page": 27, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ior, and reductions in traffic volume during lockdown periods led to measurable changes in pollutant\nconcentrations in many cities. Traffic-related pollutants exhibited varying responses depending on\nlocal conditions, pollutant type, and the relative contribution of different emission sources. These\ndisruptions underscore the sensitivity of urban air pollution to changes in traffic activity and empha-\nsize the need to consider extraordinary temporal events when interpreting multi-year datasets[90,\n91].\nSeasonal variability is another key driver of temporal changes in urban air pollution. Differences\nbetween summer and winter conditions influence both emission processes and atmospheric behavior.\nSeasonal changes in traffic patterns, such as increased idling during colder months or altered travel\nbehavior, can affect emission rates. In addition, seasonal variation in atmospheric stability, mixing\nheight, and photochemical activity contributes to differences in observed concentrations.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "31-2", + "page": 27, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "d travel\nbehavior, can affect emission rates. In addition, seasonal variation in atmospheric stability, mixing\nheight, and photochemical activity contributes to differences in observed concentrations. For pol-\nlutants such as BC and UFP, seasonal contrasts may be pronounced due to changes in combustion\nefficiency, cold-start emissions, and dispersion conditions[92, 93].\nMeteorological factors play a central role in shaping temporal variability in air pollutant concen-\ntrations. Wind speed and direction influence pollutant dispersion and transport, while temperature\naffects emission rates, atmospheric chemistry, and boundary-layer dynamics. Low wind speeds and\nstable atmospheric conditions are often associated with pollutant accumulation, whereas higher wind\nspeeds promote dilution. Temperature can also indirectly affect concentrations through its influence\non traffic behavior and energy use. As a result, meteorological variability must be carefully con-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "31-3", + "page": 27, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "speeds promote dilution. Temperature can also indirectly affect concentrations through its influence\non traffic behavior and energy use. As a result, meteorological variability must be carefully con-\nsidered when comparing observations across time or when developing models intended to generalize\nacross different temporal conditions[94, 95, 96].\nOverall, temporal changes in urban air pollution arise from the combined effects of evolving\nemission sources, seasonal processes, meteorological variability, and episodic events. These consider-\nations are particularly important for studies that integrate mobile monitoring data collected under\nvarying temporal conditions or aim to assess changes in pollution patterns over time.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "31-4", + "page": 27, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "2.6\nAir pollution modeling approaches\nA wide range of modeling approaches has been developed to estimate spatial and spatiotemporal\nvariability in ambient air pollutant concentrations, particularly in urban environments where direct\nmonitoring data are spatially sparse.\nThese models aim to translate limited measurements into\ncontinuous pollution surfaces that support exposure assessment, epidemiological analyses, and urban\nplanning applications. Broadly, spatial air pollution modeling approaches can be categorized into\ndeterministic, statistical, and data-driven frameworks, each characterized by distinct assumptions,\ndata requirements, and modeling objectives[65].\nDeterministic models, including dispersion and chemical transport models, explicitly represent\nthe physical and chemical processes governing pollutant emission, transport, transformation, and\nremoval. These approaches rely on detailed emissions inventories, meteorological inputs, and at-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "32-0", + "page": 28, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "esent\nthe physical and chemical processes governing pollutant emission, transport, transformation, and\nremoval. These approaches rely on detailed emissions inventories, meteorological inputs, and at-\nmospheric chemistry parameterizations, and are commonly applied for regional-scale air quality\nassessment, regulatory analysis, and policy evaluation. Although physically interpretable, deter-\nministic models are computationally demanding and often operate at spatial resolutions that are\ninsufficient to capture fine-scale intra-urban variability, particularly in near-road environments[97].\nStatistical modeling approaches empirically relate observed pollutant concentrations to spatial\npredictors describing land use, traffic, and environmental characteristics. These models are generally\ncomputationally efficient and well suited for urban-scale exposure assessment, with LUR representing\na commonly used framework within this class of models. In parallel, data-driven approaches based on", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "32-1", + "page": 28, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "omputationally efficient and well suited for urban-scale exposure assessment, with LUR representing\na commonly used framework within this class of models. In parallel, data-driven approaches based on\nML have gained increasing attention in spatial air pollution modeling. These methods offer enhanced\nflexibility for capturing non-linear relationships and complex interactions among predictors, and can\nbe applied either independently or as extensions of traditional statistical frameworks[98].\nWithin this broader modeling landscape, LUR and ML-based approaches have emerged as par-\nticularly relevant for characterizing traffic-related pollutants such as BC and UFP, which exhibit\nstrong local gradients and are difficult to represent using regional-scale deterministic models. The\nfollowing sections describe these two modeling paradigms in greater detail.\nFigure 2.1 illustrates a conceptual framework for spatial air pollution modeling based on LUR\nand ML–enhanced LUR approaches.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "32-2", + "page": 28, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ollowing sections describe these two modeling paradigms in greater detail.\nFigure 2.1 illustrates a conceptual framework for spatial air pollution modeling based on LUR\nand ML–enhanced LUR approaches. In this framework, air quality measurements collected from\nmobile sensing platforms to capture fine-scale spatial variability. These observations are integrated\nwith spatial predictors derived from GIS layers, including land use, traffic, and built-environment\ncharacteristics, as well as stationary air quality and meteorological variables that influence pollutant\ndispersion. Together, these inputs are used to train statistical or machine learning models that\ngenerate high-resolution air pollution maps for urban exposure assessment.\n2.6.1\nMobile monitoring of air pollutants\nAir pollution monitoring has traditionally relied on stationary monitoring networks designed to\nprovide long-term, continuous measurements of ambient pollutant concentrations. These stations", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "32-3", + "page": 28, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "llutants\nAir pollution monitoring has traditionally relied on stationary monitoring networks designed to\nprovide long-term, continuous measurements of ambient pollutant concentrations. These stations\nare typically equipped with high-precision instruments and are strategically located to represent\nregional background conditions, urban background levels, or specific source-influenced environments.\nStationary monitors play a critical role in regulatory compliance, trend analysis, and epidemiological\nstudies that require long temporal coverage. However, due to their limited spatial density, stationary\nnetworks often fail to capture fine-scale spatial variability in pollutant concentrations, particularly", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "32-4", + "page": 28, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Air quality data from\nmobile sensors or\nmonitoring network\nModel\n(Sampling)\n+\nUrban Scanner\nAir quality sensor\nAir Pollution Map\nAir quality from\nstationary\nstations\n(Temporal\nfeature)\nLand use, traffic, etc.\nClimate and\n(Spatial features)\nmeteorological\ncondition\n(Temporal feature)\nFigure 2.1: Conceptual framework for LUR and machine learning–enhanced LUR modeling of ur-\nban air pollution. Air quality measurements collected from mobile sensing platforms and stationary\nmonitoring networks provide complementary information on spatial and temporal variability. These\nobservations are combined with spatial predictors derived from GIS layers (e.g., land use, traffic, and\nbuilt-environment characteristics) and meteorological variables to inform LUR or ML-based LUR\nmodels. The trained models are used to generate high-resolution air pollution maps for exposure\nassessment.\nfor traffic-related pollutants that exhibit strong gradients over short distances[99, 100].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "33-0", + "page": 29, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "models. The trained models are used to generate high-resolution air pollution maps for exposure\nassessment.\nfor traffic-related pollutants that exhibit strong gradients over short distances[99, 100].\nThis limitation is especially pronounced for pollutants such as UFP and BC, whose concentrations\ncan vary substantially within tens to hundreds of meters due to localized emission sources and\ndispersion conditions[101]. As a result, exposure estimates derived solely from stationary monitoring\ndata may not accurately reflect the conditions experienced by individuals in proximity to roadways\nor within complex urban environments. These challenges have motivated the increasing use of mobile\nmonitoring approaches to complement traditional stationary measurements[102].\nAdvances in sensor miniaturization and data logging have facilitated the integration of multi-\nple pollutant sensors on mobile platforms, including vehicles, bicycles, and pedestrians[101]. By", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "33-1", + "page": 29, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "[102].\nAdvances in sensor miniaturization and data logging have facilitated the integration of multi-\nple pollutant sensors on mobile platforms, including vehicles, bicycles, and pedestrians[101]. By\nsampling pollutants while traversing urban road networks, mobile monitoring enables the character-\nization of spatial variability at much finer resolutions than those achievable with stationary networks\nalone. This approach is particularly well suited for mapping traffic-related air pollution, identifying\npollution hotspots, and evaluating exposure patterns along transportation corridors[103].\nCompared to stationary monitors, mobile monitoring offers several key advantages.\nFirst, it\nallows for high spatial coverage within relatively short time periods, enabling detailed assessment of\nintra-urban variability. Second, mobile measurements can be directly linked to roadway characteris-\ntics, traffic conditions, and surrounding land use, facilitating the investigation of source–concentration", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "33-2", + "page": 29, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "rban variability. Second, mobile measurements can be directly linked to roadway characteris-\ntics, traffic conditions, and surrounding land use, facilitating the investigation of source–concentration\nrelationships. Third, mobile platforms can be deployed flexibly and repeatedly, allowing researchers\nto target specific areas of interest or to capture variability across different times of day, seasons, or\ntraffic conditions[104].\nMobile monitoring has been widely applied in previous urban air quality studies to characterize\nspatial patterns of traffic-related pollutants. Numerous campaigns have demonstrated the ability\nof mobile measurements to reveal sharp concentration gradients near roadways, identify previously", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "33-3", + "page": 29, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "unrecognized hotspots, and improve estimates of population exposure. These studies have employed\na range of platforms and sensor technologies, from research-grade instruments to increasingly so-\nphisticated low-cost sensors, reflecting ongoing advancements in mobile sensing capabilities[66, 9].\nDespite its advantages, mobile monitoring also presents several challenges related to data quality\nand interpretation. Measurements collected from moving platforms are subject to rapid fluctuations\ndriven by transient emission events, vehicle dynamics, and changing meteorological conditions. As\na result, mobile datasets often exhibit higher variability than stationary measurements and require\ncareful processing, aggregation, and quality control. Instrument response time, sensor calibration,\nand synchronization with location data are particularly important considerations, as delays or inac-\ncuracies can introduce spatial misalignment or bias in estimated concentrations[102, 105].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "34-0", + "page": 30, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ration,\nand synchronization with location data are particularly important considerations, as delays or inac-\ncuracies can introduce spatial misalignment or bias in estimated concentrations[102, 105].\nAdditionally, mobile monitoring campaigns typically provide limited temporal coverage at any\ngiven location compared to stationary monitors. This characteristic necessitates thoughtful study\ndesign and appropriate analytical methods to ensure that observed spatial patterns are represen-\ntative and not driven by short-term temporal effects. Repeated sampling, temporal normalization,\nand integration with stationary monitoring data are commonly used strategies to address these\nlimitations[8, 106, 107].\nOverall, mobile monitoring represents a powerful complement to traditional stationary air quality\nmonitoring. When combined with appropriate data processing and modeling approaches, mobile\nmeasurements enable high-resolution characterization of urban air pollution and provide valuable", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "34-1", + "page": 30, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "air quality\nmonitoring. When combined with appropriate data processing and modeling approaches, mobile\nmeasurements enable high-resolution characterization of urban air pollution and provide valuable\ninsights into the spatial structure of traffic-related emissions. These capabilities are particularly\nimportant for pollutants such as UFP and BC, for which fine-scale spatial variability plays a critical\nrole in determining population exposure and associated health risks.\n2.6.2\nPredictors for spatial air pollution modeling\nBuilding on the modeling frameworks described above, this section reviews the main categories of\npredictors commonly used to characterize spatial variability in urban air pollution. Accurate pre-\ndiction of pollutant concentrations relies on predictors that represent emission sources, dispersion\nprocesses, and characteristics of the built and natural environment. In spatial air pollution mod-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "34-2", + "page": 30, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "diction of pollutant concentrations relies on predictors that represent emission sources, dispersion\nprocesses, and characteristics of the built and natural environment. In spatial air pollution mod-\nels, predictors are typically derived from GIS, transportation datasets, meteorological observations,\nremote sensing products, and increasingly, image-based data[11, 108, 109].\nTraffic-related predictors are among the most influential variables in spatial air pollution mod-\neling, particularly for pollutants such as BC and UFP that are strongly associated with vehicular\nemissions. Common traffic indicators include road length by functional class, distance to major\nroadways, intersection density, and traffic volume metrics such as AADT. These variables are typi-\ncally calculated within multiple buffer distances around monitoring locations to capture the spatial\ninfluence of traffic emissions at different scales. While AADT provides a useful measure of long-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "34-3", + "page": 30, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "cally calculated within multiple buffer distances around monitoring locations to capture the spatial\ninfluence of traffic emissions at different scales. While AADT provides a useful measure of long-\nterm traffic intensity, it often represents average conditions and may not fully capture short-term\nvariability, fleet composition, or congestion patterns relevant for near-road exposure assessment[16,\n110, 111].\nLand-use and built-environment predictors provide complementary information related to source\nactivity and dispersion conditions. Variables such as residential, commercial, industrial, and green\nspace coverage reflect differences in emission density and human activity patterns, while indicators\nof urban form, including building density, street configuration, and land-use mix, influence airflow", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "34-4", + "page": 30, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and pollutant dispersion. Population density is also commonly included as a proxy for anthropogenic\nactivity and potential exposure. Beyond area-based representations of land use and traffic intensity,\nspatial air pollution models frequently incorporate proximity- and point-based predictors to represent\nlocalized emission sources and infrastructure[65, 111, 112].\nMeteorological predictors play an important role in modulating observed pollutant concentrations\nby influencing atmospheric dispersion, dilution, and transformation processes. Variables such as\nwind speed, wind direction, temperature, atmospheric stability, and mixing height are frequently\nincorporated either directly or indirectly in spatial models. While meteorological conditions are\noften treated as temporal covariates, they can also be used to normalize mobile monitoring data or\nsupport the interpretation of spatial patterns derived from short-term sampling campaigns[112, 111,\n113].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "35-0", + "page": 31, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "treated as temporal covariates, they can also be used to normalize mobile monitoring data or\nsupport the interpretation of spatial patterns derived from short-term sampling campaigns[112, 111,\n113].\nRemote sensing and geospatial datasets have expanded the range of predictors available for spatial\nair pollution modeling. Satellite-derived products, including land cover classifications, vegetation\nindices, surface temperature, and nighttime light intensity, provide consistent spatial coverage and\nhave been used to capture regional background contributions and urbanization patterns.\nThese\npredictors are particularly valuable in areas with limited ground-based monitoring or incomplete\nland-use data[114].\nMore recently, image-based predictors have emerged as a promising extension to traditional\npredictor sets, enabling improved characterization of traffic composition and local urban features\nthat are difficult to capture using conventional datasets[19, 115, 116].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "35-1", + "page": 31, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ng extension to traditional\npredictor sets, enabling improved characterization of traffic composition and local urban features\nthat are difficult to capture using conventional datasets[19, 115, 116].\nThe relevance and importance of individual predictors vary by pollutant, spatial scale, and\nmodeling approach. Pollutants dominated by regional background contributions, such as PM2.5,\ntend to be more strongly associated with land-use and remote sensing variables, whereas traffic-\nrelated pollutants such as BC and UFP exhibit stronger relationships with traffic and road network\ncharacteristics.\nConsequently, effective spatial air pollution modeling often requires integrating\nmultiple predictor types to capture both local emission sources and broader spatial context. Careful\npredictor selection and evaluation are therefore essential to balance model interpretability, predictive\nperformance, and robustness across different urban environments[12, 117].\n2.6.3", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "35-2", + "page": 31, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ntext. Careful\npredictor selection and evaluation are therefore essential to balance model interpretability, predictive\nperformance, and robustness across different urban environments[12, 117].\n2.6.3\nLinear-based land use regression (LUR) models\nLinear-based LUR models have been widely applied to estimate spatial variability in ambient air pol-\nlutant concentrations, particularly in urban environments where direct monitoring data are spatially\nsparse. These models relate measured pollutant concentrations to spatial predictors that character-\nize surrounding land use, transportation infrastructure, and other features of the built environment.\nBy leveraging readily available geographic and environmental data, linear-based LUR models pro-\nvide a practical and interpretable framework for generating high-resolution pollution surfaces and\nestimating population exposure[66].\nClassic linear-based LUR approaches are typically implemented using linear regression, in which", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "35-3", + "page": 31, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ble framework for generating high-resolution pollution surfaces and\nestimating population exposure[66].\nClassic linear-based LUR approaches are typically implemented using linear regression, in which\npollutant concentrations measured at monitoring locations are regressed against spatial predictors\nderived within predefined buffer distances. Predictor selection is a central component of this frame-\nwork and is commonly used to balance interpretability and predictive performance.\nThe trans-\nparency and simplicity of linear-based LUR models have made them particularly well suited for\nepidemiological studies that require interpretable exposure estimates[11, 12].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "35-4", + "page": 31, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "A defining feature of linear-based LUR modeling is the construction of spatial predictors that\nrepresent emission sources and dispersion environments. Predictors are typically derived from GIS-\nbased datasets and calculated across multiple buffer distances to capture spatial influence at different\nscales. Road network characteristics, such as road length by functional class, intersection density, and\ntraffic volume within buffers, play a particularly important role in modeling traffic-related pollutants.\nLand-use variables, including residential, commercial, industrial, and green space coverage, provide\ncomplementary information related to source activity and dispersion processes[111, 110].\nDespite their widespread application, traditional linear-based LUR models exhibit several well-\ndocumented limitations. One key challenge is their reliance on relatively small numbers of stationary\nmonitoring locations, which may inadequately capture fine-scale spatial variability, particularly for", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "36-0", + "page": 32, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "mented limitations. One key challenge is their reliance on relatively small numbers of stationary\nmonitoring locations, which may inadequately capture fine-scale spatial variability, particularly for\npollutants with strong local gradients. In addition, linear model structures may struggle to represent\nnon-linear relationships, interaction effects, and complex dependencies among spatial predictors.\nTemporal variability is often treated implicitly, further limiting the ability of classic LUR models to\nrepresent dynamic pollution processes[118, 119].\nLimitations also arise from the spatial resolution and content of conventional GIS-derived pre-\ndictors. Predictors prepared at relatively coarse spatial scales may exhibit mismatches with high-\nresolution monitoring data, introducing spatial imprecision in estimated concentrations[120]. More-\nover, traditional GIS variables often fail to capture fine-scale, street-level features such as traffic com-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "36-1", + "page": 32, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "n monitoring data, introducing spatial imprecision in estimated concentrations[120]. More-\nover, traditional GIS variables often fail to capture fine-scale, street-level features such as traffic com-\nposition, roadside infrastructure, and localized built-environment characteristics that can strongly\ninfluence pollutant concentrations[121]. These constraints highlight the need for more flexible and\nspatially precise predictor representations and the integration of alternative data sources to better\nresolve fine-scale variability[122].\nThese challenges are particularly relevant for pollutants such as UFP and BC, which exhibit\nhigh spatial variability and strong sensitivity to local traffic conditions. While linear-based LUR\nmodels have been successfully applied to estimate spatial patterns of BC in many urban settings,\ntheir application to UFP has proven more challenging due to the dynamic nature of particle num-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "36-2", + "page": 32, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ed LUR\nmodels have been successfully applied to estimate spatial patterns of BC in many urban settings,\ntheir application to UFP has proven more challenging due to the dynamic nature of particle num-\nber concentrations and their sensitivity to meteorological and traffic-related factors. Nevertheless,\nseveral studies have demonstrated that linear-based LUR models can provide useful estimates of\nUFP spatial variability when supported by high-resolution monitoring data and carefully selected\npredictors[5, 123, 124].\nOverall, linear-based LUR models remain a foundational tool in spatial air pollution research,\noffering a balance between interpretability, data availability, and spatial resolution. However, their\nreliance on linear regression limits flexibility in representing non-linear effects, interaction structures,\nand correlated predictors commonly encountered in spatial datasets[11, 125]. These limitations have", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "36-3", + "page": 32, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "on linear regression limits flexibility in representing non-linear effects, interaction structures,\nand correlated predictors commonly encountered in spatial datasets[11, 125]. These limitations have\nmotivated the development of extended modeling strategies that move beyond conventional linear\nregression while retaining the conceptual foundation of land use regression modeling.\n2.6.4\nMachine learning in air pollution prediction\nRecent advances in ML have led to the increasing adoption of ML-based LUR models for air pol-\nlution prediction, particularly in urban environments characterized by complex emission patterns\nand strong spatial heterogeneity. Within this context, machine learning is primarily used to extend\ntraditional LUR frameworks by enabling non-linear relationships and interactions among land-use\nand traffic-related predictors to be learned directly from data. As a result, ML methods have been", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "36-4", + "page": 32, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "increasingly applied as flexible alternatives to linear regression within the LUR paradigm, rather\nthan as purely standalone modeling approaches[126, 127, 118].\nA variety of ML algorithms have been incorporated into ML-based LUR models, including tree-\nbased ensemble methods such as Extreme Gradient Boosting (XGBoost) and Random Forests, as\nwell as neural network approaches such as Artificial Neural Network (ANN). Tree-based models\nare particularly well suited for spatial air pollution modeling due to their ability to handle high-\ndimensional predictor spaces, capture non-linear effects, and accommodate complex interactions\nwithout requiring explicit functional specification. Neural network models offer additional flexibility,\nbut typically require larger training datasets and careful tuning to achieve stable and generalizable\nperformance within LUR-style applications[128, 129, 86].\nAn important consideration in ML-based LUR modeling is the management of high-dimensional", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "37-0", + "page": 33, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "d careful tuning to achieve stable and generalizable\nperformance within LUR-style applications[128, 129, 86].\nAn important consideration in ML-based LUR modeling is the management of high-dimensional\npredictor spaces, which can influence both predictive performance and generalization. While ML-\nenhanced LUR models provide increased flexibility compared to linear regression–based approaches,\nthis flexibility introduces challenges related to model robustness and interpretability[128, 12]. In-\nterpretability has therefore emerged as a critical consideration in the application of ML-enhanced\nLUR models for environmental exposure assessment. In contrast to traditional LUR models, which\noffer transparent coefficient-based interpretation, many ML methods are often viewed as black-box\nmodels.\nTo address this limitation, post hoc interpretability techniques such as feature impor-\ntance metrics and SHapley Additive exPlanations (SHAP) values have been increasingly adopted to", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "37-1", + "page": 33, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ack-box\nmodels.\nTo address this limitation, post hoc interpretability techniques such as feature impor-\ntance metrics and SHapley Additive exPlanations (SHAP) values have been increasingly adopted to\nsupport insight into the relative contributions of predictors and the drivers of predicted pollutant\nconcentrations[130, 131]. Despite their advantages, ML-based LUR models present challenges when\napplied to spatial air pollution data. In particular, the risk of overfitting is heightened when models\nare trained on spatially clustered observations, underscoring the importance of validation strategies\nthat explicitly account for spatial dependence[14, 13].\nOverall, ML-enhanced LUR models represent a powerful extension of traditional LUR frame-\nworks, retaining their conceptual foundation while improving the ability to capture non-linear effects\nand fine-scale spatial variability. When combined with appropriate validation and interpretability", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "37-2", + "page": 33, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "rks, retaining their conceptual foundation while improving the ability to capture non-linear effects\nand fine-scale spatial variability. When combined with appropriate validation and interpretability\nstrategies, these approaches offer a robust pathway for advancing spatial prediction of traffic-related\npollutants such as UFP and BC.\nFeature selection in machine learning–based LUR models\nIn machine learning, feature selection is a key modeling step used to improve predictive performance,\ncomputational efficiency, and interpretability by identifying a subset of informative predictors from\na potentially high-dimensional feature space.\nAlthough many ML algorithms can accommodate\nlarge numbers of predictors and complex interactions, excessive dimensionality can increase training\ntime, amplify sensitivity to noise, and reduce generalization, particularly when training data are\nlimited. Feature selection is therefore commonly employed to control model complexity and mitigate", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "37-3", + "page": 33, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "g\ntime, amplify sensitivity to noise, and reduce generalization, particularly when training data are\nlimited. Feature selection is therefore commonly employed to control model complexity and mitigate\noverfitting in ML applications[132, 133, 134, 13]. Figure 2.2 summarizes the primary motivations\nfor feature selection in machine learning.\nIn the context of ML-based LUR models for spatial air pollution prediction, feature selection\nplays an especially important role due to the high dimensionality and strong correlation structure of\nspatial predictors. Predictor sets commonly include variables derived from multiple buffer distances,\nland-use categories, transportation networks, meteorological conditions, and emerging data sources", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "37-4", + "page": 33, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 2.2: Illustration of the primary motivations for feature selection in machine learning.\nsuch as remote sensing and imagery. While this diversity of predictors enhances model flexibility, it\nalso increases the risk of redundancy, multicollinearity, and spurious associations, particularly when\nthe number of monitoring locations is limited relative to the number of candidate predictors[13, 135,\n136].\nFeature selection methods used in air pollution modeling can be broadly categorized into filter-\nbased, embedded, and wrapper-based approaches. Filter-based methods apply statistical criteria,\nsuch as correlation thresholds or univariate associations with pollutant concentrations, to remove\nweak or redundant predictors prior to model training.\nEmbedded methods perform feature se-\nlection as part of the model fitting process; for example, regularization techniques such as Least\nAbsolute Shrinkage and Selection Operator (LASSO) penalize model complexity by shrinking coef-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "38-0", + "page": 34, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "e se-\nlection as part of the model fitting process; for example, regularization techniques such as Least\nAbsolute Shrinkage and Selection Operator (LASSO) penalize model complexity by shrinking coef-\nficients of less informative predictors toward zero. Tree-based ensemble models, including Random\nForests and XGBoost, provide intrinsic measures of feature importance based on split frequency or\ninformation gain, enabling identification of predictors that contribute most strongly to model per-\nformance. Wrapper-based methods explicitly evaluate subsets of predictors by iteratively training\nmodels and assessing predictive performance under a specified validation scheme. Forward Feature\nSelection (FFS), for example, incrementally adds predictors based on their contribution to model\nimprovement. Although computationally more demanding, wrapper-based approaches allow feature\nselection to be directly aligned with the modeling objective and validation strategy, making them", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "38-1", + "page": 34, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "to model\nimprovement. Although computationally more demanding, wrapper-based approaches allow feature\nselection to be directly aligned with the modeling objective and validation strategy, making them\nparticularly suitable for spatial air pollution applications where overfitting is a central concern[137,\n138, 136].\nIn spatial air pollution modeling, many predictors are highly correlated due to overlapping buffer\nsizes, shared spatial patterns, or common underlying processes (e.g., road length and traffic volume).\nIncluding large numbers of correlated predictors can inflate model complexity without providing\nadditional explanatory power, leading to unstable predictions and reduced generalizability. Feature\nselection therefore serves as a critical mechanism for identifying the most informative predictors and\nstabilizing ML-based LUR models in the presence of strong spatial dependence[136, 13, 14].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "38-2", + "page": 34, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "2.6.5\nModel evaluation\nModel evaluation is a fundamental component of air pollution modeling, as it determines the extent\nto which predicted concentrations accurately represent observed conditions and generalize beyond\nthe training data. In exposure modeling applications, evaluation results directly influence confidence\nin downstream analyses, including spatial mapping, epidemiological inference, and policy-relevant\ninterpretation.\nConsequently, careful consideration of evaluation metrics, validation design, and\ndiagnostic tools is essential[139, 9, 12].\nQuantitative evaluation of air pollution models is commonly performed using summary statistics\nthat compare predicted and observed concentrations. Let yi denote the observed pollutant concen-\ntration at location i, ˆyi the corresponding model prediction, ¯y the mean of observed values, and n\nthe number of evaluation samples.\nThe Coefficient of Determination (R2) is widely used to assess the proportion of variance in the", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "39-0", + "page": 35, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "orresponding model prediction, ¯y the mean of observed values, and n\nthe number of evaluation samples.\nThe Coefficient of Determination (R2) is widely used to assess the proportion of variance in the\nobservations explained by the model and is defined as\nPn\ni=1(yi −ˆyi)2\nR2 = 1 −\ni=1(yi −¯y)2 .\n(2.1)\nPn\nWhile R2 provides a useful measure of goodness-of-fit, it does not directly quantify the magnitude\nof prediction errors and can be sensitive to the validation strategy employed[11].\nError-based metrics provide complementary information by quantifying the absolute deviation\nbetween predictions and observations. The Mean Absolute Error (MAE) is defined as\nn\nMAE = 1\nX\n|yi −ˆyi| ,\n(2.2)\nn\ni=1\nand represents the average magnitude of prediction error, with equal weighting applied to all devi-\nations. The Root Mean Squared Error (RMSE) is given by\nv\nn\nu\nt 1\nX\nu\n(yi −ˆyi)2,\n(2.3)\nRMSE =\nn\ni=1\nwhich places greater emphasis on larger errors and is therefore more sensitive to outliers. Both", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "39-1", + "page": 35, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "i-\nations. The Root Mean Squared Error (RMSE) is given by\nv\nn\nu\nt 1\nX\nu\n(yi −ˆyi)2,\n(2.3)\nRMSE =\nn\ni=1\nwhich places greater emphasis on larger errors and is therefore more sensitive to outliers. Both\nMAE and RMSE are expressed in the same units as the pollutant concentration and are commonly\nreported alongside R2 to provide a more complete assessment of model performance[140].\nBeyond numerical performance measures, model evaluation must be aligned with the intended\nuse of the model. Models designed to reproduce observed concentrations at monitoring locations\nmay prioritize goodness-of-fit, whereas models intended for spatial prediction at unmonitored loca-\ntions require strong generalization capability. In spatial air pollution modeling, this distinction is\ncritical, as high performance under inappropriate validation schemes can mask poor predictive abil-\nity in independent spatial domains.This challenge is amplified when models are trained using mobile", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "39-2", + "page": 35, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "critical, as high performance under inappropriate validation schemes can mask poor predictive abil-\nity in independent spatial domains.This challenge is amplified when models are trained using mobile\nmonitoring data, which often consist of dense, spatially clustered observations collected along road\nnetworks and exhibit strong spatial and temporal autocorrelation. Standard evaluation approaches\nmay therefore overemphasize short-term fluctuations or densely sampled areas, rather than meaning-\nful spatial contrasts. Aggregation strategies, such as averaging measurements at the road-segment", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "39-3", + "page": 35, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "level or over defined temporal windows, play an important role in stabilizing evaluation metrics and\nimproving representativeness. For these reasons, model evaluation should extend beyond summary\nstatistics to include diagnostic analyses that examine residual distributions, spatial structure of er-\nrors, and consistency with known physical processes. Such diagnostics are particularly important\nwhen short-term or spatially intensive data are used to support prediction of long-term average\nconcentration surfaces, as conventional metrics may reflect the model’s ability to reproduce sampled\nobservations rather than its suitability for long-term exposure assessment. Diagnostic analyses can\nhelp identify systematic biases related to temporal representativeness, sampling design, and aggre-\ngation choices, and provide insight into whether model outputs are consistent with their intended\napplication[8, 14, 13, 141, 142].\nCross-validation", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "40-0", + "page": 36, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "al representativeness, sampling design, and aggre-\ngation choices, and provide insight into whether model outputs are consistent with their intended\napplication[8, 14, 13, 141, 142].\nCross-validation\nCross-Validation (CV) is a standard framework for estimating predictive performance by repeatedly\npartitioning data into training and testing subsets. Its primary objective is to approximate how\na model will perform on unseen data, thereby supporting model comparison, hyperparameter tun-\ning, and feature selection. Conventional CV approaches, such as Random Cross-Validation (RCV),\nimplicitly assume that observations are independent and identically distributed. While this assump-\ntion is reasonable for many classical machine learning tasks, it is frequently violated in datasets\nexhibiting structured dependence in space, time, or both[143].\nWhen observations are correlated, random partitioning can result in training and testing sets that", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "40-1", + "page": 36, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "it is frequently violated in datasets\nexhibiting structured dependence in space, time, or both[143].\nWhen observations are correlated, random partitioning can result in training and testing sets that\nshare information through proximity in space, time, or other latent structures. This phenomenon,\noften referred to as information leakage, leads to systematically optimistic performance estimates, as\nmodels are evaluated on data that are not truly independent of the training samples. Roberts et al.\n[13] demonstrated that such violations can severely bias model assessment and obscure overfitting,\nparticularly in spatial, temporal, and spatiotemporal prediction problems[141]. Figure 2.3 illustrates\nthis issue by contrasting RCV and Spatial Cross-Validation (SCV). Under random partitioning,\ntraining and testing samples are intermingled across the study area, whereas SCV enforces geographic\nseparation between folds, reducing information leakage and yielding more realistic estimates of spatial", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "40-2", + "page": 36, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ing and testing samples are intermingled across the study area, whereas SCV enforces geographic\nseparation between folds, reducing information leakage and yielding more realistic estimates of spatial\ngeneralization performance.\nTo address these limitations, alternative validation strategies explicitly align data partitioning\nwith the intended prediction target. Temporal Cross-Validation (TCV) enforces separation in time\nby withholding entire periods during testing, enabling evaluation of a model’s ability to generalize\nto unseen temporal conditions. SCV enforces separation in space by withholding geographically\ndistinct regions or locations, ensuring that predictions are evaluated at spatial domains not repre-\nsented during training. More recently, spatiotemporal cross-validation (STCV) strategies have been\ndeveloped to jointly address dependence in both space and time. These approaches restrict overlap", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "40-3", + "page": 36, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "e-\nsented during training. More recently, spatiotemporal cross-validation (STCV) strategies have been\ndeveloped to jointly address dependence in both space and time. These approaches restrict overlap\nbetween training and testing data in both dimensions, for example by withholding entire locations\nduring specific time periods[13]. Meyer et al. [14] showed that performance estimates obtained under\nRCV can differ substantially from those obtained using target-oriented strategies such as SCV, TCV,\nor STCV, revealing overfitting that would otherwise remain undetected. Figure 2.4 summarizes how\ndifferent CV strategies can be constructed by enforcing separation across observations, space, time,\nor their combination.\nBy visualizing the structure of training and testing partitions, the figure\nhighlights how increasingly restrictive validation schemes reduce dependence between training and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "40-4", + "page": 36, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 2.3:\nConceptual comparison of RCV and SCV. In random cross-validation (top row),\ntraining and testing samples are interspersed across the study area, increasing the risk of information\nleakage due to spatial dependence. In spatial cross-validation (bottom row), testing data are withheld\nfrom geographically distinct regions, providing a more realistic assessment of spatial generalization\nperformance[144].\nevaluation data[13, 14].\nOverall, these developments highlight that CV is not merely a technical implementation detail,\nbut a fundamental modeling decision. In the presence of spatial, temporal, or spatiotemporal depen-\ndence, target-oriented validation strategies are essential for obtaining reliable performance estimates\nand for avoiding misleading conclusions about model generalizability.\n2.7\nChallenges in spatial air pollution modeling\nDespite substantial advances in sensing technologies and modeling methodologies, spatial air pol-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "41-0", + "page": 37, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "misleading conclusions about model generalizability.\n2.7\nChallenges in spatial air pollution modeling\nDespite substantial advances in sensing technologies and modeling methodologies, spatial air pol-\nlution modeling faces several fundamental challenges that complicate inference, evaluation, and\ninterpretation. These challenges arise primarily from the structured nature of environmental data,\nthe complexity of urban emission processes, and the increasing reliance on flexible machine learning\nmodels. In particular, spatial autocorrelation, information leakage, and limitations in model inter-\npretability and generalization represent critical issues that must be carefully addressed to ensure\nscientifically valid and policy-relevant predictions[145, 146].\n2.7.1\nSpatial autocorrelation and data leakage\nSpatial air pollution data are inherently autocorrelated: observations collected at nearby locations", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "41-1", + "page": 37, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ly valid and policy-relevant predictions[145, 146].\n2.7.1\nSpatial autocorrelation and data leakage\nSpatial air pollution data are inherently autocorrelated: observations collected at nearby locations\ntend to exhibit similar pollutant concentrations due to shared emission sources, meteorological condi-\ntions, and characteristics of the built environment. While this spatial and temporal structure reflects\nmeaningful environmental processes, it violates the assumption of independent and identically dis-\ntributed observations that underpins many standard statistical and machine learning evaluation\nprocedures[13, 141].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "41-2", + "page": 37, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Random\n3-fold\nCV\n3-fold CV\nspatial\nTemporal\n3-fold CV\nTemporal\n3-fold CV\nSpatio\nFigure 2.4: Schematic illustration of cross-validation strategies under RCV, SCV, TCV, and STCV.\nColored blocks represent distinct folds, illustrating how training and testing sets are constructed by\nenforcing separation across observations, space, time, or both[13, 14].\nWhen spatial autocorrelation is not explicitly accounted for, model development and evaluation\ncan suffer from information leakage. In particular, commonly used random data partitioning strate-\ngies may assign spatially proximate observations to both training and testing sets. As a result,\nmodels may be evaluated on data that are not truly independent of the training samples, leading to\noverly optimistic estimates of predictive performance. This issue is especially pronounced in dense\nurban datasets and in mobile monitoring campaigns, where observations are often clustered along", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "42-0", + "page": 38, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "to\noverly optimistic estimates of predictive performance. This issue is especially pronounced in dense\nurban datasets and in mobile monitoring campaigns, where observations are often clustered along\nroad networks or within specific neighborhoods[147, 141, 148].\nThe consequences of spatial leakage extend beyond inflated performance metrics. Models trained\nand evaluated under inappropriate validation schemes may appear to perform well while failing to\ngeneralize to genuinely unobserved locations. Such failures can undermine the reliability of predicted\npollution surfaces and compromise downstream applications, including exposure assessment and\nepidemiological analysis. Addressing spatial autocorrelation therefore requires validation strategies\nthat explicitly align with the intended prediction task and enforce appropriate separation between\ntraining and evaluation data[13, 149, 15].\n2.7.2\nModel interpretation, generalization, and the Clever Hans effect", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "42-1", + "page": 38, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ly align with the intended prediction task and enforce appropriate separation between\ntraining and evaluation data[13, 149, 15].\n2.7.2\nModel interpretation, generalization, and the Clever Hans effect\nIn parallel with challenges related to spatial dependence, the increasing use of machine learning\n(ML) models in environmental prediction raises important concerns regarding model interpretability\nand generalization. While ML algorithms provide substantial flexibility for capturing non-linear\nrelationships and complex interactions, this flexibility also increases the risk that models rely on\nspurious patterns that do not correspond to underlying physical or environmental processes [15,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "42-2", + "page": 38, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "150, 151].\nThis risk is closely related to the so-called Clever Hans effect, named after a horse in the early 20th\ncentury that appeared capable of performing arithmetic but was later shown to respond instead to\nsubtle, unintentional cues from its handler (Figure 2.5). In modern machine learning, the Clever Hans\neffect refers to situations in which a model produces seemingly accurate predictions by exploiting\nunintended correlations or artifacts in the data rather than learning the causal relationships of\ninterest [152]. Because such behavior is not readily revealed by standard evaluation metrics, it poses\na particular challenge for model interpretation and generalization.\nFigure 2.5: Historical illustration of Clever Hans[153], a horse that appeared to perform arithmetic\ntasks in the early 20th century. Subsequent investigation revealed that Hans was responding to\nsubtle, unintentional cues from human observers rather than performing genuine calculations.\nLapuschkin et al.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "43-0", + "page": 39, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "in the early 20th century. Subsequent investigation revealed that Hans was responding to\nsubtle, unintentional cues from human observers rather than performing genuine calculations.\nLapuschkin et al. [152] provide a compelling demonstration of the Clever Hans effect across\nmultiple machine learning applications. Using explanation techniques such as layer-wise relevance\npropagation, they show that models with strong test-set performance can base their predictions on\nsuperficial or irrelevant features. In one illustrative example (Figure 2.6), a classifier trained to\nrecognize horses in images from the PASCAL VOC dataset relied primarily on a small source tag\nembedded in the images, rather than on the visual characteristics of the horse itself. When this\ntag was removed, model performance collapsed, while inserting the tag into images of other objects\ninduced systematic misclassification, revealing that the learned decision logic was fundamentally\nnon-generalizable.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "43-1", + "page": 39, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "model performance collapsed, while inserting the tag into images of other objects\ninduced systematic misclassification, revealing that the learned decision logic was fundamentally\nnon-generalizable.\nIn spatial air pollution modeling, analogous Clever Hans–type behavior may arise when mod-\nels exploit artifacts introduced by sampling design, predictor construction, spatial clustering, or\ncorrelated covariates. For example, models may implicitly encode location-specific identifiers, road-\nnetwork geometry, or campaign-specific characteristics that correlate with pollutant concentrations\nbut do not represent causal emission or dispersion mechanisms.\nWithout careful interpretation,\nsuch models may yield plausible spatial predictions while failing to generalize beyond the conditions\nunder which they were trained.\nThese concerns highlight the importance of complementing predictive performance metrics with", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "43-2", + "page": 39, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "sible spatial predictions while failing to generalize beyond the conditions\nunder which they were trained.\nThese concerns highlight the importance of complementing predictive performance metrics with\ninterpretability analyses that explicitly examine which information a model uses to generate predic-\ntions. Interpretation tools, sensitivity analyses, and spatial diagnostics can help determine whether\nlearned relationships align with known physical processes or instead reflect Clever Hans–type short-\ncuts. Reliable spatial prediction therefore requires not only accurate models, but models whose", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "43-3", + "page": 39, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 2.6: Illustration of the Clever Hans effect from Lapuschkin et al. [152]. A classifier trained\nto recognize horses relies primarily on a spurious source tag present in the image rather than on the\nvisual features of the horse.\ninternal decision logic is consistent with the intended environmental inference and exposure assess-\nment objectives.\nThis broader risk in data-driven modeling arises when flexible models capture spurious patterns\ndriven by data structure rather than underlying processes, leading to misleadingly strong perfor-\nmance that does not generalize. In the context of spatial air pollution modeling, excessive flexibility\nin preprocessing and tuning, when combined with non-independent validation or limited validation\ndata, can produce models that appear highly accurate yet fail to generalize beyond the sampled\ndomain. Such behavior reflects not genuine predictive skill, but the imposition of researcher expec-\ntations onto structured and correlated data.\n2.8", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "44-0", + "page": 40, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ccurate yet fail to generalize beyond the sampled\ndomain. Such behavior reflects not genuine predictive skill, but the imposition of researcher expec-\ntations onto structured and correlated data.\n2.8\nImage-based traffic detection and composition\nRecent advances in computer vision and imaging technologies have enabled the use of visual data\nas a powerful source of information for characterizing traffic conditions in urban environments. In\ncontrast to traditional air pollution modeling approaches that rely on aggregated or static traffic\nindicators, image-based methods allow direct observation of vehicles and road activity at fine spatial\nand temporal scales. As a result, imagery has become an increasingly important data source for\nestimating traffic volume and traffic composition in complex urban settings[154, 155].\nImagery used for traffic analysis can originate from a range of platforms, including fixed roadside", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "44-1", + "page": 40, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "a source for\nestimating traffic volume and traffic composition in complex urban settings[154, 155].\nImagery used for traffic analysis can originate from a range of platforms, including fixed roadside\ncameras, aerial and satellite systems, and increasingly, street-level and 360-degree mobile imaging\nplatforms. Street-level imagery provides detailed views of vehicles, lane structure, and surrounding\nbuilt environments, while mobile imaging systems offer continuous spatial coverage along road net-\nworks. When combined with precise location information, these data enable spatially explicit char-\nacterization of traffic patterns that are difficult to capture using conventional GIS-based datasets[19,\n156, 155, 157, 158].\nThe automated detection and classification of vehicles in images has been greatly advanced by the\ndevelopment of deep learning–based object detection models. Early traffic monitoring approaches", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "44-2", + "page": 40, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "158].\nThe automated detection and classification of vehicles in images has been greatly advanced by the\ndevelopment of deep learning–based object detection models. Early traffic monitoring approaches\nrelied on handcrafted visual descriptors and classical classifiers, which often exhibited limited ro-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "44-3", + "page": 40, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "bustness under real-world conditions involving variable lighting, occlusion, and heterogeneous vehicle\ntypes. Modern convolutional neural network–based detectors address these limitations by learning\nhierarchical visual representations directly from data, enabling more reliable vehicle detection and\nclassification across diverse urban environments[159, 156, 155, 157, 158].\nAmong contemporary object detection frameworks, the You Only Look Once (YOLO) family of\nmodels has become particularly prominent in traffic-related applications.Figure 2.7 illustrates the\ngeneral architecture and detection pipeline of a YOLO-based object detection model. YOLO models\nperform object localization and classification in a single forward pass, enabling efficient processing of\nlarge image datasets while maintaining high detection accuracy. Successive versions of YOLO have\nintroduced architectural improvements that enhance detection of small objects, improve robustness", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "45-0", + "page": 41, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "g of\nlarge image datasets while maintaining high detection accuracy. Successive versions of YOLO have\nintroduced architectural improvements that enhance detection of small objects, improve robustness\nunder complex visual conditions, and support multi-class vehicle classification. These characteristics\nmake YOLO-based models well suited for large-scale extraction of traffic composition from street-\nlevel and mobile imagery[20, 159].\nFigure 2.7: Schematic overview of the YOLO object detection architecture.\nYOLO and related deep learning models have been widely applied to traffic monitoring tasks,\nincluding vehicle counting, estimation of traffic density, and classification of vehicles by type. Of\nparticular relevance for environmental applications is the ability to distinguish between passenger\ncars, light-duty vehicles, buses, and heavy-duty trucks, as different vehicle classes exhibit markedly\ndifferent emission profiles.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "45-1", + "page": 41, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "nmental applications is the ability to distinguish between passenger\ncars, light-duty vehicles, buses, and heavy-duty trucks, as different vehicle classes exhibit markedly\ndifferent emission profiles. Image-based estimation of traffic composition therefore provides a more\nphysically meaningful representation of emission sources than aggregate traffic metrics alone[19,\n160].\nIn air pollution research, image-derived traffic information has increasingly been used to support\ninterpretation of observed concentration patterns and to improve representation of traffic-related\nemissions. Vehicle-type–specific counts obtained from imagery can serve as direct proxies for emis-\nsion intensity, especially for pollutants such as BC and UFP that are strongly influenced by fleet\ncomposition. Compared to model-based or average traffic indicators, image-based traffic detection\noffers improved sensitivity to local traffic conditions and spatial heterogeneity[161, 21].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "45-2", + "page": 41, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "by fleet\ncomposition. Compared to model-based or average traffic indicators, image-based traffic detection\noffers improved sensitivity to local traffic conditions and spatial heterogeneity[161, 21].\nA key advantage of image-based traffic detection lies in its ability to capture dynamic and context-\nspecific traffic characteristics.\nTraditional traffic indicators, such as AADT, typically represent\nlong-term averages and may not reflect short-term variability in traffic flow, congestion, or fleet\nmix. In contrast, imagery collected during monitoring campaigns can provide temporally aligned", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "45-3", + "page": 41, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "traffic composition information, enabling closer linkage between observed pollutant concentrations\nand contemporaneous emission sources[160, 21, 162].\nOverall, the use of computer vision for traffic detection and composition estimation represents\na complementary data acquisition strategy that enhances traditional traffic datasets without in-\ntroducing opaque image-to-pollution mappings. By transforming raw visual data into structured\ntraffic composition variables, image-based traffic detection supports integration with established air\npollution modeling frameworks while preserving interpretability and physical relevance.\n2.9\nGaps in existing literature\nDespite extensive research on traffic-related air pollution, several important gaps remain in the\ncurrent literature, particularly with respect to high-resolution exposure assessment, representation\nof traffic composition, and model generalization.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "46-0", + "page": 42, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "pollution, several important gaps remain in the\ncurrent literature, particularly with respect to high-resolution exposure assessment, representation\nof traffic composition, and model generalization.\nFirst, although numerous studies have documented the strong influence of traffic on urban air\npollution, most spatial modeling frameworks rely on aggregate traffic indicators such as road length\nor AADT. These metrics do not explicitly capture traffic composition, despite well-established dif-\nferences in emission profiles between light-duty and heavy-duty vehicles. As a result, many existing\nmodels are limited in their ability to distinguish emission hotspots driven by freight and service vehi-\ncles from those dominated by passenger cars, especially in near-road and street-level environments.\nSecond, while mobile monitoring has substantially improved the characterization of fine-scale\nspatial variability in pollutants such as BC and UFP, integration of mobile measurements with", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "46-1", + "page": 42, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "vironments.\nSecond, while mobile monitoring has substantially improved the characterization of fine-scale\nspatial variability in pollutants such as BC and UFP, integration of mobile measurements with\nvehicle-type–specific traffic information remains limited. Traditional traffic datasets often lack suffi-\ncient spatial resolution or vehicle-class detail to fully exploit the potential of mobile air quality data.\nConsequently, the link between observed pollutant patterns and underlying traffic composition is\noften inferred indirectly rather than quantified explicitly.\nThird, recent advances in computer vision have enabled automated vehicle detection from street-\nlevel imagery, yet these methods have primarily been applied in transportation engineering and urban\nanalytics rather than in air pollution exposure modeling. Existing air quality studies that incorporate\nimagery typically use visual data as generic predictors or contextual indicators, with relatively few", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "46-2", + "page": 42, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ics rather than in air pollution exposure modeling. Existing air quality studies that incorporate\nimagery typically use visual data as generic predictors or contextual indicators, with relatively few\nstudies transforming imagery into physically interpretable traffic composition variables that can be\ndirectly integrated into established modeling frameworks such as LUR.\nFourth, temporal changes in traffic composition and their implications for air pollution patterns\nremain insufficiently explored. While several studies have examined year-to-year or seasonal changes\nin pollutant concentrations, fewer have jointly modeled changes in both air pollution and vehicle-\ntype–specific traffic activity using consistent spatial frameworks. This limits understanding of how\nshifts in fleet composition, such as post-pandemic changes in freight and delivery traffic, translate\ninto evolving pollution hotspots.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "46-3", + "page": 42, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "consistent spatial frameworks. This limits understanding of how\nshifts in fleet composition, such as post-pandemic changes in freight and delivery traffic, translate\ninto evolving pollution hotspots.\nFinally, the increasing use of flexible machine learning models in spatial air pollution modeling\nhas raised concerns regarding overfitting, spatial data leakage, and model interpretability. Many\nstudies report strong predictive performance using random validation strategies that do not ade-\nquately account for spatial or spatiotemporal dependence, potentially overstating model generaliza-\ntion. There remains a need for modeling frameworks that combine high-resolution data sources with\nvalidation strategies explicitly designed to mitigate spatial and temporal leakage, while preserving", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "46-4", + "page": 42, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "interpretability and physical relevance.\nCollectively, these gaps highlight the need for integrated approaches that (i) explicitly represent\ntraffic composition at high spatial resolution, (ii) leverage mobile monitoring data to capture fine-\nscale pollutant variability, (iii) use image-based methods to generate interpretable traffic indicators\nrather than opaque predictors, (iv) examine temporal changes in both traffic activity and air pollu-\ntion, and (v) employ robust modeling and validation strategies that prioritize generalization. The\nstudies presented in this dissertation directly address these gaps.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "47-0", + "page": 43, + "document_type": "thesis", + "page_header": "CHAPTER 2. LITERATURE REVIEW", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Chapter 3\nUrban air pollution data collection, mapping, and\nprediction using mobile sensors installed on courier\ntrucks\n3.1\nChapter overview\nPM2.5 has significant impacts on human health, making it essential to understand its spatial and\ntemporal variations. This study focuses on developing LUR models and improving their perfor-\nmance in predicting PM2.5 concentrations in an urban setting. In this study, air quality data were\ncollected using a sensor on a courier truck in downtown Toronto. XGBoost, a machine learning\nalgorithm, was employed to address limitations in traditional linear regression-based LUR models,\nincorporating predictors such as land use, meteorology, and emissions to build robust models. A\ntotal of 27 models were trained, with varying road segment lengths, predictors, and outlier treatment\nthresholds. Three models tested the impact of road segment length on model predictions. Eight", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "48-0", + "page": 44, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ". A\ntotal of 27 models were trained, with varying road segment lengths, predictors, and outlier treatment\nthresholds. Three models tested the impact of road segment length on model predictions. Eight\nmodels examined the effect of removing outliers with different thresholds, revealing that appropri-\nate thresholds improve accuracy. Ten models assessed the addition of emission and traffic data,\nwhich did not enhance performance, likely due to overlapping effects with other predictors. In six\nmodels, time-variant predictors such as time of day, month, humidity, wind speed, temperature,\nand pollutant concentrations from stationary stations were included. Adding these predictors sig-\nnificantly improved model performance, highlighting the complex relationships in LUR models for\nPM2.5 predictions and offering valuable insights for air quality assessment.\n3.2\nIntroduction\nAir quality stands as a critical environmental concern, impacting millions of individuals worldwide.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "48-1", + "page": 44, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "for\nPM2.5 predictions and offering valuable insights for air quality assessment.\n3.2\nIntroduction\nAir quality stands as a critical environmental concern, impacting millions of individuals worldwide.\nAmong the air pollutants, PM2.5 severely affects human health[163].Understanding spatial or spa-\ntiotemporal pollutant concentrations plays a vital role in implementing effective mitigation strategies\nand in developing accurate exposure surfaces to support the analysis of long-term health effect [6].\nReference air quality stations provide reliable air pollutant concentration information, but they have\nlimitations due to data being available at limited points[163]. To overcome these limitations, re-\nsearchers use models such as geostatistical interpolation, photo-chemical dispersion, and LUR to\nassess air pollution exposure and its associated health effects in urban areas[164]. While geostatis-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "48-2", + "page": 44, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "searchers use models such as geostatistical interpolation, photo-chemical dispersion, and LUR to\nassess air pollution exposure and its associated health effects in urban areas[164]. While geostatis-\ntical interpolation and dispersion models are valuable, their resolution is often coarse, making them\nless suitable for capturing small-scale urban variability that is crucial in exposure assessment[8]. In\ncontrast, LUR models have become popular for quantifying intraurban variation in air pollutant\nconcentrations, offering simplicity and the ability to predict fine-scale pollution variations that can\nreduce exposure measurement errors[8, 165]. The main components of LUR models include monitor-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "48-3", + "page": 44, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ing data, geographic predictors, model development, and validation[165, 166]. Traditionally, LUR\nmodels have relied on stationary air quality monitors, but these methods fail to capture small-scale\nspatial variations in pollutant concentrations. Mobile monitoring offers an alternative with better\nspatial coverage, though it presents challenges in addressing temporal variability[166, 105].\nCriticism of linear regression based LUR models includes limited flexibility and difficulty in\nincorporating highly correlated predictors [167]. Studies have explored machine learning models,\nsuch as neural networks, support vector machine, and tree-based models, which improve accuracy\nand handle multicollinearity[167, 168]. Predictors for LUR models include land use, meteorology,\npollutant emissions, and built environment characteristics. Common predictor variables are traffic,\npopulation density, land use, topography, and location[127].\nUnderstanding the impact of such", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "49-0", + "page": 45, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "logy,\npollutant emissions, and built environment characteristics. Common predictor variables are traffic,\npopulation density, land use, topography, and location[127].\nUnderstanding the impact of such\npredictors is crucial for improving models and capturing spatiotemporal variation. Studies have us\ned various methods to assess LUR model performance. Common approaches include cross-validation\ntechniques such as Leave-One-Out Cross-Validation (LOOCV), where each data point is tested once\nwhile the model is trained on the rest. RCV evaluates model error by leaving out data points within a\nspecific region and predicting concentrations there[11].Some studies use external data for validation,\nnoting that a high training R2 may not correspond to a high test-set R2 when estimating long-term\naverage exposure[169].\nAcross the literature, different sets of LUR models, predictors, and postprocessing of data have\nbeen used to improve predicted long-term or short-term exposure surfaces.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "49-1", + "page": 45, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "g-term\naverage exposure[169].\nAcross the literature, different sets of LUR models, predictors, and postprocessing of data have\nbeen used to improve predicted long-term or short-term exposure surfaces. In a study by Messier\net al.[166], Nitric Oxide (NO) and BC measured by mobile monitoring and a land use Land Use\nRegression–Kriging (LUR-K) was implemented on part of a dataset to predict concentration of pol-\nlutants at unobserved locations. Kerckhoffs et al.[128] compared the performance of a simple linear\nregression, several algorithms based on regularization, and several machine learning algorithms (ran-\ndom forest, boosting, and bagging) to predict long-term UFP. They showed that machine learning\nmodels had better performance for their data. In another study, Kerckhoffs et al.[170] implemented\nthree models (supervised stepwise regression, LASSO, and random forest) to predict long-term av-\nerage UFP by using a deconvoluting method to separate local UFP contributions from background", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "49-2", + "page": 45, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "lemented\nthree models (supervised stepwise regression, LASSO, and random forest) to predict long-term av-\nerage UFP by using a deconvoluting method to separate local UFP contributions from background\nconcentrations. In another study by Liu et al. [167], a linear regression was compared withGradient\nBoosting Decision Tree (GBDT) models.The study showed that the GBDT model captured non-\nlinear relationships and explained more of the variations in the pollutant concentrations [167, 170].\nIn recent years, more complex machine learning algorithms such as ANN have been used for predic-\ntion. In a study by Wang et al. [171], the performance of two machine learning approaches, ANN\nand gradient boost, were compared with simple linear regression LUR using five data segmentation\nschemes. Machine learning methods exhibited superior performance over a simple LUR model. In\nthe studies conducted by Ren et al. [172] and Wong et al. [118], various machine learning methods", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "49-3", + "page": 45, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "entation\nschemes. Machine learning methods exhibited superior performance over a simple LUR model. In\nthe studies conducted by Ren et al. [172] and Wong et al. [118], various machine learning methods\nand simple LUR models were assessed. Both studies arrived at the conclusion that the XGBoost\nalgorithm outperformed the other approaches in terms of performance.\nLUR models are widely used in air quality research to estimate pollutant concentrations in areas\nlacking direct monitoring data. By incorporating spatial predictors such as land use types, traffic\ndensity, population distribution, and meteorological data, LUR models can effectively predict pollu-\ntion levels beyond measurement locations (whether fixed or mobile), including off-road environments.\nFor instance, Padhi et al.[173] discusses the application of LUR models in air pollution assessment,\nhighlighting their effectiveness in capturing spatial variability.\nSimilarly, research by Hankey et", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "49-4", + "page": 45, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "al.[174] demonstrates the successful modeling of on-road particulate air pollution, including particle\nnumber, BC, and PM2.5, using LUR techniques based on mobile monitoring data. Their findings\nhighlight the effectiveness of LUR models in capturing spatial variability of pollutants across urban\nsettings. While uncertainties may arise when extrapolating to areas far from monitoring points, LUR\nmodels have been validated in diverse settings, demonstrating their robustness in capturing spatial\ntrends across urban and suburban environments. These strengths make LUR models valuable for\nunderstanding exposure in off-road locations.\nIn this study, we used the XGBoost algorithm to develop multiple models with the aim of\npredicting concentrations of PM2.5. To gather the data for developing the models, a sensor was\ninstalled on a courier truck operating in Toronto, Canada.\nThe primary focus of this research", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "50-0", + "page": 46, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "the aim of\npredicting concentrations of PM2.5. To gather the data for developing the models, a sensor was\ninstalled on a courier truck operating in Toronto, Canada.\nThe primary focus of this research\nis to investigate the possibility of enhancing prediction ability of the model by employing diverse\npredictors and addressing outliers. By constructing distinct models for varied sets of input data as\npredictors, we conducted comparisons to understand the impact of these predictors and modifications\non the overall performance of the model. Additionally, the feasibility of utilizing delivery vehicle\nroutes for constructing a LUR model to predict exposure surfaces was assessed. The study outcomes\ncan be leveraged to investigate the potential of a fleet serving as a mobile monitoring platform.\nBeyond contributing to the methodological development of exposure surfaces, this study offers\nvaluable insights for urban planning, public health, and community-level air quality management.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "50-1", + "page": 46, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "latform.\nBeyond contributing to the methodological development of exposure surfaces, this study offers\nvaluable insights for urban planning, public health, and community-level air quality management.\nPrevious research has demonstrated the value of LUR models in identifying urban air pollution\nhotspots and supporting interventions to mitigate health risks associated with air pollution. For\ninstance, Hoek et al.[11] reviewed the application of LUR models in air pollution epidemiology,\nhighlighting their effectiveness in exposure assessment and their potential to inform public health\npolicies. Similarly, Jerrett et al.[175] used LUR models to examine the relationship between air pollu-\ntion and mortality in urban areas, underscoring their role in exposure analysis and public health risk\nassessment. Furthermore, Lu et al.[176] discussed the application of LUR models for identifying air\npollution hotspots in urban environments, emphasizing their value in supporting population health", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "50-2", + "page": 46, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "k\nassessment. Furthermore, Lu et al.[176] discussed the application of LUR models for identifying air\npollution hotspots in urban environments, emphasizing their value in supporting population health\nby guiding urban planning decisions. By using low-cost mobile sensors and implementing advanced\nLUR modeling techniques, our findings allow for a more nuanced understanding of PM2.5 variability\nacross urban areas, which can help inform effective strategies to reduce population exposure and\npromote healthier urban spaces.\n3.3\nMethods\n3.3.1\nStudy area\nThis study is set in the city of Toronto, located on the north shore of Lake Ontario in southeastern\nCanada and has a population of over 2.7 million, making it the largest city in the country. Toronto\nexperiences a temperate climate with distinct seasons. This city has a relatively flat topography,\nespecially in the downtown core and along the lakeshore. Air quality in Toronto can vary depending", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "50-3", + "page": 46, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "o\nexperiences a temperate climate with distinct seasons. This city has a relatively flat topography,\nespecially in the downtown core and along the lakeshore. Air quality in Toronto can vary depending\non several factors, including weather, traffic, industrial activity, seasonal influences, and proximity\nto major sources like Toronto’s municipal expressways and provincial highways or its two airports:\nToronto Pearson International Airport located west of the city and Billy Bishop Toronto City Air-\nport on the Toronto Islands.\nAmong all the sources, traffic emissions are a major local source", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "50-4", + "page": 46, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "of air pollution in Toronto. This includes emissions from cars, trucks, buses, and other forms of\ntransportation.\n3.3.2\nData collection and preprocessing\nWe conducted a mobile monitoring campaign in Toronto by installing a sensor on a truck operated by\nPurolator, a courier and package delivery company in Toronto, to capture spatiotemporal variation\nof ambient PM2.5. The mobile monitoring campaign took place between June 21, 2022, to January\n11, 2023. The sensor, manufactured by Geotab, a telematics and fleet management company, uses\na Honeywell HPMA 11S50 low-cost sensor. The HPMA 11S50 sensor is capable of measuring ambi-\nent PM concentrations. It can detect and quantify real-time PM2.5 and Particulate Matter smaller\nthan 10 microns (PM10). In fact, the sensor is designed to capture a broad spectrum of particulate\nconcentrations, suitable for urban air quality monitoring. It responds rapidly to changes in PM2.5\nlevels, ensuring near real-time data collection.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "51-0", + "page": 47, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "designed to capture a broad spectrum of particulate\nconcentrations, suitable for urban air quality monitoring. It responds rapidly to changes in PM2.5\nlevels, ensuring near real-time data collection. The sensor operates reliably across a wide range of\ntemperatures and humidity conditions. It was factory-calibrated to ensure accurate and consistent\nmeasurements[177, 178]. This sensor is integrated into a Geotab unit which also includes a SGX Sen-\nsortech MiCS-4514 to collect Nitrogen Dioxide (NO2) concentrations and Bosch Sensortech BME280\nto collect temperature, humidity, and atmospheric pressure[168]. Moreover, Geotab reports a wide\nrange of vehicle telematics data such as GPS location, fuel level, and other on-board diagnostics\n(OBD). The sensor reports a record whenever a property changes, for example whenever PM2.5 con-\ncentration changes, a new record is reported. The instruments were calibrated by the manufacturer\nprior to installation.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "51-1", + "page": 47, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "r reports a record whenever a property changes, for example whenever PM2.5 con-\ncentration changes, a new record is reported. The instruments were calibrated by the manufacturer\nprior to installation. To minimize the risk of self-pollution from the truck’s exhaust emissions, the\nPM2.5 sensor was strategically placed with the airflow intake positioned on the top of the truck’s\nright window, facing outward. This positioning was chosen to avoid direct exposure to exhaust gases\nand ensure that the sensor primarily measured ambient air quality rather than pollutants emitted\nby the vehicle itself. To assess the reliability of the sensor readings over time, we compared PM2.5\nconcentrations recorded by the mobile sensor with data from a near-reference monitoring station\nlocated along College Street, a busy arterial road on the campus of the University of Toronto during\nperiods when the truck operated near this station (within a 200-meter radius). This station uses a", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "51-2", + "page": 47, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "located along College Street, a busy arterial road on the campus of the University of Toronto during\nperiods when the truck operated near this station (within a 200-meter radius). This station uses a\nThermo Scientific 5030 SHARP Synchronized Hybrid Ambient Real-time Particulate Monitor; its\nlocation is presented in Figure 3.1. This comparison helped verify the consistency and accuracy of\nthe mobile measurements, providing validation for the sensor’s stability and accuracy throughout\nthe study period.\nThe concentration of PM2.5 was recorded for a total of 56,577 data points while the truck was in\noperation. The truck was operating mostly between 7 am and 7 pm and 20% of the data points were\nrecorded during the morning (7 am to 11 am), 41% midday (11 am to 3 pm), 31% in the afternoon\n(3–7 pm), and the remaining were recorded outside those hours. The truck mostly operated in the\ndowntown area. During the campaign, the truck was operational for approximately 300 hours over", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "51-3", + "page": 47, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "fternoon\n(3–7 pm), and the remaining were recorded outside those hours. The truck mostly operated in the\ndowntown area. During the campaign, the truck was operational for approximately 300 hours over\nthe six-month period, this was due to variable daily operational hours of the Purolator delivery\nvehicle, which did not consistently follow an 8-hour schedule. However, the PM2.5 sensor was active\nfor only 157 hours. This discrepancy arises from instances where drivers inadvertently forget to turn\non the sensor. These factors led to reduced data collection time over the six-month period. Another\nobservation from our data analysis is the presence of records for 5,767 distinct minutes (equivalent\nto 96 hours). This indicates that the sensor didn’t continuously report data throughout its 157-hour", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "51-4", + "page": 47, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 3.1: Map featuring the locations of monitoring stations in Toronto. The near-reference\nmonitoring stations (Hanlan’s Point Station and University of Toronto Station) are pinpointed along\nwith the four regulatory monitoring stations operated by the Ontario Ministry of Conservation and\nParks (Toronto Downtown Station, Toronto East Station, Toronto North Station, and Toronto West\nStation).\nactive time. The gaps in the data correspond to moments when the sensor did not report new values\ndue to the absence of significant concentration changes.\nEach data point consists of a pair of x, y coordinates associated with a PM2.5 concentration.\nFor this reason, we averaged all observations corresponding with each road segment, defining those\nsegments using different methods. In the first method, the observations were aggregated for each\nroad segment (defined from one intersection to another) from a road network shapefile for Toronto.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "52-0", + "page": 48, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ose\nsegments using different methods. In the first method, the observations were aggregated for each\nroad segment (defined from one intersection to another) from a road network shapefile for Toronto.\nThen the center of each road was identified, and buffers were created around each point to compute\nmodel predictors. As a result, all data points were assigned to approximately 1,050 road segments.\nIn the second and third methods, segments were defined every 50-meter and 100-meter intervals,\nwith the average concentration computed for all readings associated with these segments. With a\n50-meter aggregation, the total number of segments is 1,673, and with a 100-meter aggregation, it\nis 705. These three methods were used to evaluate the impact of spatial aggregation. We used the\n50-meter segment as the base case, against which all models will be compared since it refers to the\nhighest resolution. These three methods of aggregation have been employed by other studies[8, 166,\n179].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "52-1", + "page": 48, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "3.3.3\nGeneration of predictor variables\nSeveral LUR models were developed using different sets of predictors to evaluate the impact of\nchoosing predictors on the performance of LUR models.\nWe divided our predictors into three\ncategories including: specific to a location predictor, traffic related predictors, as well as weather\nand air quality related predictors which can change temporally. To calculate predictors, varying\nbuffer sizes (50, 100, 200, 250, 300, 500, 1000, 2000 meters) were created around each segment mid-\npoint (based on the three methods for data aggregation). The calculation of predictors involved the\ncalculation of area, length, or number of land use types and traffic related variables within each of\nthese buffers, computation of distances to important places in the city from the buffers’ centers, and\ntemporally changing variables (weather and regional air quality).\nPredictors that accurately explain emissions or traffic can affect LUR models and their effec-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "53-0", + "page": 49, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "he city from the buffers’ centers, and\ntemporally changing variables (weather and regional air quality).\nPredictors that accurately explain emissions or traffic can affect LUR models and their effec-\ntiveness in estimating pollution concentrations[167]. In this study, LUR models were created to\nassess the contrasting effects of utilizing two data sources for traffic and emissions. In one model,\ntraffic volume and emissions for NOx from the GTAModel, an activity-based model for the Greater\nToronto Area which uses the A Multimodal Transportation Planning Software Package (EMME)\ntraffic assignment package, were included as predictors[180, 181]. Then, four models were developed\nby including AADT data instead of emission data. The AADT data used in this study was devel-\noped by Ganji et al.[162] by using traffic count data to calculate AADT instead of a travel demand\nmodel. The AADT data was improved by Ganji et al.[182] using aerial imagery. Four models were", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "53-1", + "page": 49, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "devel-\noped by Ganji et al.[162] by using traffic count data to calculate AADT instead of a travel demand\nmodel. The AADT data was improved by Ganji et al.[182] using aerial imagery. Four models were\ntrained with AADT data from four different years (2017, 2018, 2019, 2020).\nTo address temporal variability of PM2.5 concentrations, several models were trained by adding\nnew predictors. For this purpose, we utilized data from two near-reference monitoring stations,\none situated at the University of Toronto and the other one near Lake Ontario, at Hanlan’s point,\nsouth of the city. These monitoring stations provided environmental variables including humidity,\nwind speed, temperature, regional PM2.5, and gaseous pollutants such as NOx, NO, and NO2. We\nalso added time of day as three categories (morning, mid-day, and afternoon), and day of week and\nmonth for examining the possibility of capturing daily, weekly and seasonal trends in PM2.5. All", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "53-2", + "page": 49, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ". We\nalso added time of day as three categories (morning, mid-day, and afternoon), and day of week and\nmonth for examining the possibility of capturing daily, weekly and seasonal trends in PM2.5. All\nthe predictors, their categories, and type of measurement are detailed in Table 3.1.\nTable 3.1: Summary of predictors categorized into land use-related variables (e.g., area, population,\nproximity), traffic-related variables (e.g., road lengths, traffic counts, emissions), and temporally\nvarying variables (e.g., time of day, month, weather conditions).\nAttribute\nType of measurement\nLand use related variables\nParking lot area\nArea within buffer normalized by buffer area\nCommercial area\nArea within buffer normalized by buffer area\nGovernmental area\nArea within buffer normalized by buffer area\nIndustrial area\nArea within buffer normalized by buffer area\nOpen area\nArea within buffer normalized by buffer area\nResidential area\nArea within buffer normalized by buffer area\nWaterbody area", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "53-3", + "page": 49, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "r area\nIndustrial area\nArea within buffer normalized by buffer area\nOpen area\nArea within buffer normalized by buffer area\nResidential area\nArea within buffer normalized by buffer area\nWaterbody area\nArea within buffer normalized by buffer area", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "53-4", + "page": 49, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Attribute (continued)\nType of measurement (continued)\nParks\nArea within buffer normalized by buffer area\nRoad area\nArea within buffer normalized by buffer area\nDistance to Lake Ontario\nDistance in meters\nPopulation\nPopulation within buffer normalized by area\nDistance to Billy Bishop Toronto City Air-\nDistance in meters\nport\nAccommodation points\nCount within buffer\nChimneys\nCount within buffer\nGas stations\nCount within buffer\nRestaurants\nCount within buffer\nTraffic related variables\nLength of highways within buffer\nLength in meters\nLength of railways within buffer\nLength in meters\nLength of major roads within buffer\nLength in meters\nLength of bus routes within buffer\nLength in meters\nTotal road length within buffer\nLength in meters\nDistance to nearest highway\nDistance in meters\nDistance to nearest major road\nDistance in meters\nDistance to nearest rail line\nDistance in meters\nNumber of intersections\nCount within buffer\nNumber of traffic signals\nCount within buffer\nNumber of bus stops", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "54-0", + "page": 50, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "to nearest major road\nDistance in meters\nDistance to nearest rail line\nDistance in meters\nNumber of intersections\nCount within buffer\nNumber of traffic signals\nCount within buffer\nNumber of bus stops\nCount within buffer\nNOx emissions (daily and yearly)\nEmissions normalized by buffer area\nTraffic volume (daily and yearly)\nTraffic count normalized by buffer area\nAADT (multiple years)\nTraffic count normalized by buffer area\nTemporally changing variables\nTime of day\nCategorical (morning, afternoon, evening)\nMonth\nMonth of data collection\nWeekday\nDay of week\nStation-based pollutant concentrations\nNO, NO2, NOx, PM2.5\nWeather conditions\nWind speed, relative humidity, temperature\n3.3.4\nLand-use regression modelling\nIn this study, the XGBoost model, which is a tree-based machine learning method, was trained using\ndifferent sets of predictors to predict long-term PM2.5 exposure surfaces and to study the impact\nof using different sets of predictors on the model performance.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "54-1", + "page": 50, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ine learning method, was trained using\ndifferent sets of predictors to predict long-term PM2.5 exposure surfaces and to study the impact\nof using different sets of predictors on the model performance. The XGBoost model can capture\nnon-linear relations between predictors [171, 172]. The importance of predictors can be extracted\nfrom this model using different feature importance measures. In this study, The Python libraries", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "54-2", + "page": 50, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "scikit-learn and XGBoost were used [183]. The grid search method was used to tune our models.\nAfter tuning and choosing the best model, we used a feature importance method to assess the impact\nof different features on model output and predictive performance.\nTo assess model accuracy, the dataset was partitioned into training and testing sets, with 20%\nallocated for testing. The model was trained and evaluated at the same time using a 10-fold CV\napproach in combination with grid search for hyperparameter tuning. In fact, by using this method,\nthe training data set was divided into 10 subsets, and the model was trained and evaluated 10 times.\nEach time, one of the 10-folds is used for testing the model and other folds for training. The model\nperformance was averaged for all these 10 models to provide an overall assessment of its accuracy\nand generalization. After finding the best models and their hyperparameters, the models were tested", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "55-0", + "page": 51, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "l\nperformance was averaged for all these 10 models to provide an overall assessment of its accuracy\nand generalization. After finding the best models and their hyperparameters, the models were tested\nby using the 20% test set which the model had not seen during training and tuning. The best R2\nvalues for the test set and 10-fold CV were compared for the different models.\nFinally, we conducted a validation analysis by comparing the model’s predicted PM2.5 concen-\ntrations with the average of measurements during sampling period from four regulatory monitoring\nstations in Toronto, operated by the Ontario Ministry of Conservation and Parks (excluding the\nnear-reference station used for sensor comparison); the locations of these stations are identified\nin Figure 3.1.\nWe computed the accuracy of model predictions, calculated as Accuracy (%) =\n\u0010\n\u0011\n1 −|Measured Value−Modeled Value|\n100 ×\n.\nMeasured Value\nThe impact of deleting outliers on XGBoost performance was evaluated.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "55-1", + "page": 51, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "puted the accuracy of model predictions, calculated as Accuracy (%) =\n\u0010\n\u0011\n1 −|Measured Value−Modeled Value|\n100 ×\n.\nMeasured Value\nThe impact of deleting outliers on XGBoost performance was evaluated. To evaluate the effect\nof deleting outliers on the performance of LUR models, we compared our models when we delete\noutliers larger and less than specific percentiles. We deleted PM2.5 concentrations when they were\nlarger than 99.7th, 99th, and 95th percentiles and less than 5th (all zero records), 10th, and 20th.\nWe trained eight models by using various combinations of these thresholds, removing data exceeding\nor falling below the specified values as necessary and for some models, we exclusively utilized either\nthe upper or lower threshold for data removal. Three models involved the removal of data greater\nthan the 99th, 99.7th, and 95th percentiles (referred as Outlier 1, 2, 3). Another set of three models", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "55-2", + "page": 51, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "the upper or lower threshold for data removal. Three models involved the removal of data greater\nthan the 99th, 99.7th, and 95th percentiles (referred as Outlier 1, 2, 3). Another set of three models\nfocused on removal of data greater than the 95th percentile and less than the 5th percentile, greater\nthan the 95th percentile and less than the 10th percentile, and greater than the 95th percentile\nand less than the 20th percentile (referred as Outlier 4, 5, 6, respectively). Additionally, one model\nexclusively removed data falling below the 5th percentile (all zeros) (Outlier 7), while another model\nremoved data greater than the 99.7th and less than the 5th percentiles (Outlier 8).\nTo assess the influence of incorporating emission and traffic count predictors on the performance\nof the model, we extended our analysis. In one model, we incorporated both emission and traffic\ncount data from the Greater Toronto Area travel demand model (GTAModel). Additionally, we", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "55-3", + "page": 51, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "e performance\nof the model, we extended our analysis. In one model, we incorporated both emission and traffic\ncount data from the Greater Toronto Area travel demand model (GTAModel). Additionally, we\nincluded traffic count data as AADT for four different years (2017, 2018, 2019, 2020) into four\nseparate models. These predictor sets were integrated not only into the base case model, but also\ninto the model configuration where we removed data points less than 5th percentile (remove all\nzeros) and all data points greater than the 95th percentile. We also tested the impact of adding\nemission and traffic count predictors to the Outlier 4 model.\nTo handle and capture temporal variation in the concentration of the pollutants in the models, we\nhave assessed several models with different predictors. In one model (named Model time) time of day\nrepresented as three categorical variables (morning, midday, afternoon) was considered. In another", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "55-4", + "page": 51, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "have assessed several models with different predictors. In one model (named Model time) time of day\nrepresented as three categorical variables (morning, midday, afternoon) was considered. In another\nmodel (named as day month) day of week (Monday to Sunday) and month of (1-12) were added as", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "55-5", + "page": 51, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "predictors. Moreover, 4 different models (named as Station 1 to 4) were developed by adding data\nfrom stationary stations as predictors. In the Station 1 model, hourly average relative humidity,\nwind speed and weather temperature from two stations (Hanlan’s Point station and University of\nToronto station) were added to the base case as predictors; in the Station 2 model, hourly average\nrelative humidity, wind speed and ambient temperature along with NO, NO2, NOx and PM2.5\nconcentrations from the two stations were added as predictors; in the Station 3 model, all the\npredictors from Station 2 model using data less than 95th percentile and higher than 5th percentile\nwere used; and for Station 4 model, we added time of day as three categories to the setting of\nStation 3 model.\nIncluding the different methods of data aggregation, use of predictors, and treatment of outliers,\nwe end up with 27 models, summarized in Figure 3.2.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "56-0", + "page": 52, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ee categories to the setting of\nStation 3 model.\nIncluding the different methods of data aggregation, use of predictors, and treatment of outliers,\nwe end up with 27 models, summarized in Figure 3.2.\nFigure 3.2: Overview of the 27 models categorized by key factors and represented in distinct colors:\nmodels evaluating (1) road segment length (orange), (2) outlier treatment (green), (3) emission and\ntraffic data incorporation (blue), (4) combined emission/traffic data with outlier treatment (red),\nand (5) temporally changing predictors (yellow).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "56-1", + "page": 52, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "3.4\nResults\n3.4.1\nSensor performance\nThe comparison between the PM2.5 values recorded by the mobile sensor and the regulatory moni-\ntoring station yielded an R2 value of 0.63. This suggests reasonable alignment between the mobile\nand stationary measurements, though variability is still present. A significant factor influencing this\ncorrelation is the number of records per minute and the proximity of the truck to the station, as only\nrecords collected within a 200-meter radius of the station were included in this comparison. Min-\nutes with a higher density of second-by-second records tend to show averaged readings closer to the\nstation’s data, as more frequent measurements reduce random fluctuations (Figure 3.3). Conversely,\nminutes with fewer records or greater distances may be more impacted by transient conditions,\nleading to less consistent alignment with the stationary measurements.\nFigure 3.3: Scatter plot of PM2.5 concentrations recorded by the mobile sensor versus the University", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "57-0", + "page": 53, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ed by transient conditions,\nleading to less consistent alignment with the stationary measurements.\nFigure 3.3: Scatter plot of PM2.5 concentrations recorded by the mobile sensor versus the University\nof Toronto monitoring station, with color indicating the number of measurements per minute for the\nsensor installed on the truck. Station records represent average PM2.5 concentration per minute.\nThe dashed line (y = x) represents perfect agreement.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "57-1", + "page": 53, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "3.4.2\nDescriptive statistics\nOur data analysis shows an average PM2.5 concentration of 9.56 µg/m³ with a standard deviation\nof 14.34 µg/m³, ranging from zero to a maximum of 1096 µg/m³, suggesting potential measurement\nerrors at high levels. The high standard deviation underscores the variability in air quality within\ndowntown Toronto, reinforcing the importance of effective analysis to capture variability. The re-\nported average concentration is within the range commonly observed in urban environments and is\nbelow widely referenced short-term guideline values. However, this does not imply the absence of\nhealth risks. A substantial body of evidence indicates that long-term exposure to PM2.5, even at\nrelatively low concentrations, is associated with adverse health outcomes, including cardiovascular\nand respiratory diseases, particularly under repeated exposure in urban settings. The extremely high", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "58-0", + "page": 54, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "elatively low concentrations, is associated with adverse health outcomes, including cardiovascular\nand respiratory diseases, particularly under repeated exposure in urban settings. The extremely high\nmaximum value observed in this dataset is likely influenced by measurement artifacts or transient\nlocal emission events and should not be interpreted as representative of typical exposure conditions.\nFigure 3.4a illustrates the box plot for PM2.5 concentrations for each sampling day. The data reveals\nsignificant variation between and within days, without a clear trend between summer and fall.\nFigure 3.4b displays the distribution of average PM2.5 concentrations on road segments. This\nfigure shows the complex spatial variability of PM2.5, which may be influenced by both spatial\nand temporal factors, considering that not all segments were sampled simultaneously. Based on\nFigure 3.4b, the concentration of PM2.5 in areas closer to the Toronto downtown core exhibit elevated", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "58-1", + "page": 54, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ial\nand temporal factors, considering that not all segments were sampled simultaneously. Based on\nFigure 3.4b, the concentration of PM2.5 in areas closer to the Toronto downtown core exhibit elevated\nconcentrations when compared to regions farther from the city center, such as the eastern part of\nthe city. Figure 3.4a and Figure 3.4b provide evidence that pollution levels change within short\ndistances and time periods, affected by weather, local emissions, and the layout of the city. These\nsmall-scale urban variabilities can be captured using LUR models[165].\n3.4.3\nBase model\nThe base case uses a 50-meter road segment aggregation, chosen as the benchmark for comparison.\nThis fine resolution increases data points for training and captures spatial variability. It serves as\na foundation for evaluating other aggregation methods. The R2 for the base-case model, yielded\na value of 0.249 for the test dataset and 0.251 for the 10-fold CV. The SHAP summary plot for", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "58-2", + "page": 54, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "serves as\na foundation for evaluating other aggregation methods. The R2 for the base-case model, yielded\na value of 0.249 for the test dataset and 0.251 for the 10-fold CV. The SHAP summary plot for\nthe base model (Figure 3.5) highlights key predictors influencing PM2.5 concentrations, such as road\narea in small buffers, proximity to major roads, and intersection density. These predictors align with\nthe concentration sensitivity to localized factors, demonstrating its ability to capture the intricate\nimpacts of urban infrastructure and traffic emissions on air quality.\n3.4.4\nVariations on base model\nThe results of the 27 different models tested and which included 3 data aggregation methods, 8\noutlier treatment methods, 5 local traffic emissions data sources, 5 local traffic emissions data sources\napplied with an outlier treatment method, and 6 models with temporal variables are presented in\nFigure 3.6. The R2 values for the test set and 10-fold CV for these models are reported.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "58-3", + "page": 54, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s data sources\napplied with an outlier treatment method, and 6 models with temporal variables are presented in\nFigure 3.6. The R2 values for the test set and 10-fold CV for these models are reported.\nThe two segment aggregation methods (100-meter and original segments) were compared with\nbase case to determine their impact on model performance. The 100-meter aggregation method per-\nformed slightly better than both the 50-meter (base case) and the original road segment aggregation\nin terms of model accuracy. However, the original road segment aggregation, while not as accurate", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "58-4", + "page": 54, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "(a)\n(b)\nFigure 3.4: (a) Distribution of PM2.5 measurements and (b) average PM2.5 concentration per road\nsegment, highlighting spatial variability across Toronto.\nas the 100-meter approach, provided more realistic representation of the spatial variability inher-\nent to road networks. These results highlight the importance of testing various spatial aggregation\nresolutions, as each method balances accuracy with the ability to represent spatial structures and\npatterns effectively.\nWe tested the impact of different threshold values for outlier removal on the performance of\nour XGBoost models for predicting PM2.5 concentration. The threshold values are highlighted in\nFigure 3.4a to illustrate the range of data points affected by each outlier removal method. In the\nfirst three models (Outlier 1, Outlier 2, and Outlier 3), where we employed outlier removal criteria\ntargeting data points greater than the 99th, 99.7th, and 95th percentiles, the model’s performance", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "59-0", + "page": 55, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "first three models (Outlier 1, Outlier 2, and Outlier 3), where we employed outlier removal criteria\ntargeting data points greater than the 99th, 99.7th, and 95th percentiles, the model’s performance\nexhibited a slight improvement compared to the base case. However, in the third model utilizing\nthe 95th percentile threshold, there is a smaller change in the R2 of the 10-fold CV. This can be\nthe result of removing data points for some days with high concentration that do not qualify as\noutliers for that day. For the models where we removed datapoints less than the 10th percentile\nor 20th percentile threshold (Outlier 5 and Outlier 6), the R2 decreased, possibly because correctly\nrecorded data points (not true outliers) may have been removed. As shown in Figure 3.4a, these\ntwo methods (removing datapoints less than the 10th percentile or 20th) may result in the removal", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "59-1", + "page": 55, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 3.5: SHAP summary plot for top 30 features of the base model ranked by their importance.\nEvery point represents a SHAP value, reflecting the influence of a feature on the model output for\na specific observation. The color of each dot indicates the value of the corresponding independent\nfeature, ranging from low (blue) to high (red). Higher SHAP values for features suggest elevated\nlevels of PM2.5.\nof most records for some days. In our data, the 5th percentile of the data is equal to zero. For\nthree models, we used the 5th percentile and lower threshold to remove data points. When we just\nuse lower threshold (Outlier 7 model), the R2 diminished slightly. The results for ‘Outlier 4’ and\n‘Outlier 8’ models where we had upper thresholds for 95th and 99.7th percentile, respectively, and\nzeros were removed shows that the R2 for these two models were slightly better than other ones.\nHowever, the model results show that removing outliers does not significantly improve performance", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "60-0", + "page": 56, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ", and\nzeros were removed shows that the R2 for these two models were slightly better than other ones.\nHowever, the model results show that removing outliers does not significantly improve performance\nin predicting PM2.5 concentrations.\nThe results of the models that we created by adding local traffic and emission predictors show\nno improvement, possibly because the impact of traffic had been captured by predictors that we had\nin the base case model such as road density and road typology.\nThe results of the models with temporally varying predictors show that by adding day of week\nand month as predictor, the performance of the model improved significantly. By adding humidity,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "60-1", + "page": 56, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "temperature, and wind speed, the R2 increased from 0.26 to 0.69 which shows the importance of\nthese variables that can capture temporal variability of the model for predicting PM2.5 and by adding\nconcentration of pollutants from stationary stations, the R2 increased to 0.753. Moreover, the results\nfor Station 3 and Station 4 model demonstrate that by treating outliers and adding weather and\nregional air quality predictors to the model (Station 2), the performance of LUR models for PM2.5\ncan be improved substantially.\nFigure 3.6:\nComparison of the 27 tested models, incorporating variations in data aggregation\nmethods, outlier treatments, local traffic and emissions data, and temporal predictors. Each bar\nrepresents the R2 values for both the test dataset and 10-fold CV.\nThe best-performing model, ”Station 4” includes additional features with larger buffer sizes of\n500m, 1000m, and 2000m to capture regional impacts.\nThese larger buffers provide a broader", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "61-0", + "page": 57, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "spatial context that complements the finer-scale variability captured by smaller buffers, resulting\nin a balanced representation of both local and regional influences on PM2.5 concentrations. This\napproach also mitigates concerns about spatial autocorrelation by capturing regional trends and\naccounting for spatial dependencies in the model. The results of this approach, shown in Figure 3.7,\nrepresent the PM2.5 concentration surface for afternoon hours (3 pm to 7 pm), where the inclusion\nof larger buffer sizes highlights regional pollution patterns across the study area. In Figure 3.7, the\nPM2.5 surface using the ”Station 4” model is used to predict PM2.5 for the entire city, including\nregions beyond the area that the data was originally collected (beyond downtown area). As expected,\nhigher concentrations are predicted along major roads and highways due to increased emissions. This\nfigure demonstrates that the model effectively predicts pollutant concentrations in locations outside", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "62-0", + "page": 58, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "gher concentrations are predicted along major roads and highways due to increased emissions. This\nfigure demonstrates that the model effectively predicts pollutant concentrations in locations outside\nthe original data collection area.\nFigure 3.7: Predicted PM2.5 concentration for the City of Toronto using the ”Station 4” model\nduring afternoon hours (3 pm to 7 pm).\n3.4.5\nValidation against observations at reference stations\nThe results of the comparison of the Station 4 model predictions against observations at the 4\nreference stations operated by the Ministry of Environment, Conservation, and Parks are presented\nin Figure 3.8. They demonstrate reasonable consistency between the predicted and observed values\nacross different times of day and locations.\nFor instance, in the afternoon, the accuracy of the\nmodel’s predictions ranged from 84.6% in Toronto Downtown to 95.5% in Toronto North. Similar", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "62-1", + "page": 58, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "patterns were observed during the morning and midday periods, with accuracies exceeding 88% for\nmost comparisons. This consistency highlights the model’s ability to extend its predictions beyond\nthe immediate road network, leveraging broad spatial predictors such as land use and meteorology to\ncapture regional trends effectively. While limitations remain in areas far removed from roadways, the\nupdated surface incorporates larger buffer sizes (500m, 1000m, and 2000m), enhancing the model’s\ncapability to estimate off-road concentrations more reliably. These findings support the robustness\nof our approach for predicting spatial pollution patterns across diverse urban areas.\n(a)\n(b)\nFigure 3.8: Validation of the Station 4 model against measurements at four regulatory monitoring\nstations in Toronto.\nPanel (a) shows the observed (measured) and predicted (modeled) PM2.5\nconcentrations (in µg/m³) across stations for each time of day (morning, midday, afternoon). Panel", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "63-0", + "page": 59, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "onitoring\nstations in Toronto.\nPanel (a) shows the observed (measured) and predicted (modeled) PM2.5\nconcentrations (in µg/m³) across stations for each time of day (morning, midday, afternoon). Panel\n(b) presents the calculated accuracy (%) of the model predictions.\n3.5\nDiscussion and conclusion\nA mobile monitoring campaign, utilizing a sensor-equipped truck operated by a courier company,\ndemonstrates a novel approach to capture spatiotemporal variations in PM2.5 concentration. The\nstudy’s timeframe, from June 21, 2022, to January 11, 2023, allows for a nuanced analysis of seasonal", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "63-1", + "page": 59, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "influences on air quality. This study investigates factors influencing PM2.5 concentration predictions\nusing LUR models by developing several different models using different sets of predictors. The\nfindings highlight the complex interplay of factors influencing PM2.5 levels, including meteorologi-\ncal conditions, traffic emissions, and land use variables. Considering temporal factors significantly\nenhanced the model’s performance, as evident from the comparison of the ”day month” model with\nthe base case. The incorporation of day of the week and month as predictors resulted in a substantial\nimprovement in the R2 for both the test set and the 10-fold CV. This improvement suggests that\ndistinct patterns exist throughout the week, possibly influenced by varying human activities, traffic\ndensity, and meteorological conditions. Moreover, this improvement suggests that accounting for\ntemporal variations, such as day-of-week patterns and seasonal differences, contributes valuable in-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "64-0", + "page": 60, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ic\ndensity, and meteorological conditions. Moreover, this improvement suggests that accounting for\ntemporal variations, such as day-of-week patterns and seasonal differences, contributes valuable in-\nformation to the model. For instance, weekdays might exhibit different pollution patterns compared\nto weekends, and seasonal variations could be crucial in understanding the impact of weather con-\nditions on PM2.5 concentrations. This finding aligns with previous research emphasizing the impact\nof temporal variations on air quality and emphasizes the need for comprehensive spatiotemporal\nmodels [184, 185].\nThe study also investigated the impact of different spatial aggregation methods on model per-\nformance, providing insights into the impact of aggregation method on model prediction.\nThe\ncomparison of three aggregation methods—road segment shapefile, 50-meter, and 100-meter dis-\ntanced points showed that the 100-meter resolution performed slightly better. The results imply", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "64-1", + "page": 60, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ction.\nThe\ncomparison of three aggregation methods—road segment shapefile, 50-meter, and 100-meter dis-\ntanced points showed that the 100-meter resolution performed slightly better. The results imply\nthat choosing an appropriate spatial aggregation method is crucial for achieving a more accurate\nrepresentation of local air quality patterns. So, for developing an LUR model, testing several reso-\nlutions might help to improve the model’s accuracy which is aligned with the finding of the study\nby Hankey et al. [174].\nThe study evaluates the impact of outlier removal on model performance, revealing trade-offs as-\nsociated with different thresholds. Models that deleted outliers (errors in recording) while retaining\nnon-outlier records demonstrated improvement in the performance compared to models with exten-\nsive data removal. For instance, the ”Outlier 4” model, which focused on removing data greater", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "64-2", + "page": 60, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "retaining\nnon-outlier records demonstrated improvement in the performance compared to models with exten-\nsive data removal. For instance, the ”Outlier 4” model, which focused on removing data greater\nthan the 95th percentile while keeping lower concentrations, exhibited a slightly better R2 for both\nthe test set and 10-fold CV. This indicates that careful outlier handling, preserving valid data points,\ncontributes to a more robust and accurate LUR model as it has been shown in previous studies [186,\n187]. However, this study shows that it is essential to strike a balance, as overly aggressive outlier\nremoval may lead to the loss of crucial information and compromise model performance.\nThe inclusion of predictors related to emissions and traffic counts presented mixed results. While\nthe addition of predictors from EMME and AADT data did not show significant improvement in the\nbase case model, it is crucial to note that the base case model already included relevant predictors", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "64-3", + "page": 60, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ile\nthe addition of predictors from EMME and AADT data did not show significant improvement in the\nbase case model, it is crucial to note that the base case model already included relevant predictors\nrelated to land use, traffic, and location. The lack of substantial improvement suggests that the\nexisting set of predictors in the base case model adequately captured the spatial variability caused\nby traffic and emissions in PM2.5 concentrations. This finding underscores the importance of carefully\nselecting predictors based on the study context and existing knowledge of local sources of pollution.\nThe models incorporating data from stationary monitoring stations as predictors exhibited a\nremarkable increase in performance. The introduction of temporal predictors, such as time of day\nand month, as well as data from stationary monitoring stations, enhances the ability of models\nto capture temporal variation. The marked improvement in R2 values for models incorporating", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "64-4", + "page": 60, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "temporal variables highlights the significance of considering time-related factors in predicting PM2.5\nconcentrations.\nIn this study, model performance was evaluated using CV due to the limited availability of inde-\npendent datasets for validation, which is consistent with common practice in air pollution modeling.\nWhile this approach provides a useful estimate of predictive accuracy, the use of random CV (RCV)\nintroduces important limitations in the context of mobile monitoring data. Specifically, because such\ndata exhibit strong spatial and temporal autocorrelation, randomly partitioning observations into\ntraining and testing subsets can result in samples that are geographically and temporally proximate.\nThis lack of independence can lead to information leakage, producing overly optimistic performance\nestimates and favoring models that reproduce localized patterns rather than generalize to truly\nunseen locations or conditions.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "65-0", + "page": 61, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "an lead to information leakage, producing overly optimistic performance\nestimates and favoring models that reproduce localized patterns rather than generalize to truly\nunseen locations or conditions.\nThe implications of this limitation are examined in detail in Chapter 4, where alternative vali-\ndation frameworks that explicitly account for spatial and temporal dependence are developed and\nevaluated. In that chapter, model selection, feature selection, and hyperparameter tuning are system-\natically aligned with structured CV strategies to provide more realistic assessments of generalization\nperformance and to reduce the risk of overfitting associated with dependence in the data.\nDespite these limitations, the results of this study demonstrate that incorporating temporal,\nspatial, and outlier-related considerations can improve the development of accurate LUR models\nfor predicting PM2.5 concentrations. In particular, the integration of innovative mobile monitoring", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "65-1", + "page": 61, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ral,\nspatial, and outlier-related considerations can improve the development of accurate LUR models\nfor predicting PM2.5 concentrations. In particular, the integration of innovative mobile monitoring\ntechniques, careful predictor selection, and systematic exploration of outlier impacts contributes\nto enhanced model performance.\nThese findings highlight the importance of adopting a multi-\ndimensional modeling approach that accounts for both spatial and temporal variability in urban\nenvironments.\nFinally, while the inclusion of pollutant concentrations from stationary monitoring stations im-\nproves model performance by capturing regional background conditions, these predictors rely on\nexternal data sources that may not be available in all application contexts. As a result, models\nincorporating such variables are conditional on the availability of monitoring infrastructure and may\nhave reduced transferability to regions lacking comparable data. Future work should therefore fo-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "65-2", + "page": 61, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "incorporating such variables are conditional on the availability of monitoring infrastructure and may\nhave reduced transferability to regions lacking comparable data. Future work should therefore fo-\ncus on incorporating independent datasets and developing validation strategies that better reflect\nreal-world deployment scenarios to further assess model robustness and generalizability.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "65-3", + "page": 61, + "document_type": "thesis", + "page_header": "CHAPTER 3. URBAN AIR POLLUTION DATA COLLECTION, MAPPING, AND ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Chapter 4\nBridging the gap between data reproduction and\nprediction: The impact of feature selection and cross-\nvalidation strategies on prediction of ambient ultra-\nfine particles collected with mobile monitoring\n4.1\nChapter overview\nReliable exposure assessment is vital for epidemiological research, but weaknesses in LUR models\nundermine its validity. Using mobile UFP data in Toronto, we compared LUR models trained under\nrandom, spatial, temporal, and spatiotemporal CV, with and without FFS. Model hyperparameters\nand feature subsets were optimized within each CV scheme. Considering UFP’s short atmospheric\npersistence and sharp spatial gradients, spatial CV folds were designed at fine spatial scales to re-\nflect their autocorrelation structure. Each approach was evaluated on a hold-out test set, across\nCV schemes, and against independent stationary backyard measurements. The modeling framework", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "66-0", + "page": 62, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s to re-\nflect their autocorrelation structure. Each approach was evaluated on a hold-out test set, across\nCV schemes, and against independent stationary backyard measurements. The modeling framework\nintegrating spatiotemporal CV with FFS to enhance hyperparameter tuning and predictor selection\nwas designed with the primary objective of capturing the spatial trends in UFP. This approach ef-\nfectively reduced overfitting, improved generalization, and produced stable exposure surfaces. These\nsurfaces avoided the spatial artifacts and exaggerated variable effects typically seen in models trained\nwith random CV. Models tuned with conventional random CV overfit, performed poorly on indepen-\ndent samples, and were highly sensitive to outliers. For example, a model trained under random CV\nhad an Average Percentage Error (APE) of ∼217% against an independent dataset, whereas spa-\ntiotemporal CV with FFS reduced APE to ∼79%. Our findings demonstrate that proper alignment", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "66-1", + "page": 62, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "under random CV\nhad an Average Percentage Error (APE) of ∼217% against an independent dataset, whereas spa-\ntiotemporal CV with FFS reduced APE to ∼79%. Our findings demonstrate that proper alignment\nof model design with the data’s spatiotemporal structure and modeling objective ensures reliability,\nminimizes data reproduction, and enables true prediction.\n4.2\nIntroduction\nAir pollution is a major public health concern, especially in urban areas [188, 189, 190]. Accurately\nassessing exposure necessitates advanced measurement and modeling methods, including mobile\nmonitoring, remote sensing, LUR, and ML [191, 86]. LUR modeling is widely used to estimate air\npollution exposure in urban areas by leveraging geographic and environmental predictors to map\npollutant concentrations at fine scales. Historically, LUR relied on data from a limited number of\nfixed monitoring stations [119, 192], which constrained its ability to capture localized variations\n[184].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "66-2", + "page": 62, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tant concentrations at fine scales. Historically, LUR relied on data from a limited number of\nfixed monitoring stations [119, 192], which constrained its ability to capture localized variations\n[184]. Recent progress has broadened LUR’s scope, incorporating advanced data collection methods\nand diverse predictors. ML approaches further enhance predictive accuracy [193, 194].\nMobile monitoring enhances exposure assessment for spatially varying air pollutants by covering", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "66-3", + "page": 62, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "many more locations and capturing the large spatial variability characteristic of urban areas [195].\nWhen enabled by portable low-cost sensors on vehicles, it also offers high-resolution measurements\nover wide areas with fewer devices [196]. Mobile air quality measurements are often conducted at\nspecific times and under varying meteorological conditions, making it challenging to distinguish true\nspatial pollution patterns from short-term fluctuations caused by factors like traffic, weather, or\nlocalized emission sources [197, 198]. UFP differ markedly from other air pollutants such as PM2.5\nand NO2 in both their physical behavior and chemical reactivity [199]. glsUFP are of particular con-\ncern due to their ability to penetrate deep into the respiratory system and enter the bloodstream.\nUnlike regulated pollutants such as , there is no established safe exposure threshold, and health\nrisks increase with concentration. Both short-term spikes and long-term cumulative exposure are", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "67-0", + "page": 63, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ream.\nUnlike regulated pollutants such as , there is no established safe exposure threshold, and health\nrisks increase with concentration. Both short-term spikes and long-term cumulative exposure are\nrelevant, as transient peaks can trigger acute responses while repeated exposure is associated with\ncardiovascular and respiratory diseases. This highlights the importance of capturing both localized\nvariability and broader spatial patterns in exposure models. Emitted primarily from traffic-related\nsources, UFP have a high surface-area-to-mass ratio, which enhances their volatility and their abil-\nity to adsorb and desorb semi-volatile compounds. This property facilitates rapid transformation\nprocesses, including coagulation and growth into larger particles in the atmosphere [200]. Due to\ntheir short atmospheric lifetime and strong dependence on local traffic activity and meteorological\nconditions, UFP exhibit steep spatial gradients and pronounced variability at fine spatial scales [201,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "67-1", + "page": 63, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "r short atmospheric lifetime and strong dependence on local traffic activity and meteorological\nconditions, UFP exhibit steep spatial gradients and pronounced variability at fine spatial scales [201,\n202]. As a result, the collected dataset represents an episodic sample in time (a series of short-term\nsnapshots) with extensive spatial coverage. This trade-off complicates the disentanglement of spatial\nvs. temporal variation [186, 203, 204]. Variations in data collection frequency and sampling routes\ncan introduce biases, potentially affecting the reliability of LUR models trained on mobile data [205,\n8, 206]. A model might simply “memorize” the specific conditions of the sampling period (a phe-\nnomenon we term ’data reproduction’ defined as overfitting to the specific spatiotemporal structure\nof the training data) rather than learning generalizable patterns (prediction). For instance, Blanco\net al.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "67-2", + "page": 63, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "we term ’data reproduction’ defined as overfitting to the specific spatiotemporal structure\nof the training data) rather than learning generalizable patterns (prediction). For instance, Blanco\net al. (2023) [206] demonstrated that different sampling designs can produce significantly different\nexposure surfaces, even within the same study area. Their findings indicate that campaigns with\ntemporal restrictions (limited to business hours, rush hours, or specific seasons) reduced model per-\nformance and produced different spatial surfaces. While they suggest improving sampling strategies\nto enhance data representativeness, refining modeling approaches might also be necessary to better\naccount for spatial and temporal variability and ensure more reliable exposure assessments.\nModel selection and validation are crucial for building reliable LUR models for air pollution\n[191]. A common approach for tuning a model is random CV. Random CV is straightforward but", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "67-3", + "page": 63, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "e assessments.\nModel selection and validation are crucial for building reliable LUR models for air pollution\n[191]. A common approach for tuning a model is random CV. Random CV is straightforward but\noften overestimates performance in spatially or temporally structured data, as nearby locations or\ntime periods can leak between training and test sets, inflating performance metrics [207, 149, 208].\nThis limitation arises because conventional random CV assumes that observations are independent\nand identically distributed, a core assumption underpinning its validity as an unbiased estimator\nof out-of-sample error [209, 210, 211].\nIn spatial or temporal data, this assumption is violated,\nobservations are not independent, leading to information leakage between folds and overly optimistic\nperformance estimates[212, 213, 214]. The spatial and temporal autocorrelation in air pollution data\nmay render random splitting inappropriate, risking overfitting to non-causal predictors and over", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "67-4", + "page": 63, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "c\nperformance estimates[212, 213, 214]. The spatial and temporal autocorrelation in air pollution data\nmay render random splitting inappropriate, risking overfitting to non-causal predictors and over\noptimistic performance metrics. Therefore, more advanced strategies are needed to avoid overfitting\n[215, 172]. Roberts et al. (2017) [13] recommend block CV for model selection, ensuring data are", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "67-5", + "page": 63, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "partitioned by spatial or temporal structures. Blocking by location or time provides a stringent test\nof whether a model can predict in completely unseen regions or periods. Indeed, using appropriate\nCV helps select models and hyperparameters that avoid overfitting to spurious correlations present\nin training data [13, 216].\nThe accuracy of LUR models depends on predictor variables, which capture built environment\nand other factors influencing air pollution levels [122, 19]. It is common to explore a wide array of\npotential predictor variables when constructing LUR models [179, 105, 29, 111]. However, excessive\ninclusion of predictor variables can lead to overfitting, reducing model generalizability and predictive\nperformance [15, 217, 218]. As a result, feature selection is a critical step in LUR. Various feature\nselection methods, including forward stepwise linear regression [218, 219, 166] and LASSO [170] have\nbeen used to improve LUR performance and reduce overfitting.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "68-0", + "page": 64, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "critical step in LUR. Various feature\nselection methods, including forward stepwise linear regression [218, 219, 166] and LASSO [170] have\nbeen used to improve LUR performance and reduce overfitting. FFS is an iterative procedure that\nadds variables sequentially based on their contribution to model performance. In each step, variables\nare evaluated for significance and multicollinearity, excluding redundant or highly correlated land-\nuse predictors. However, feature selection methods such as FFS cannot remove spurious predictors\nunder random CV, where spatially dependent observations appear in both training and test sets\n[15, 220, 221]. This can inflate performance metrics, leading to overfitted models and artefactual\nspatial patterns that fail to generalize across new regions or time periods. More broadly, these risks\nstem from reliance on purely statistical criteria that may omit causal relationships. This leads to", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "68-1", + "page": 64, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tial patterns that fail to generalize across new regions or time periods. More broadly, these risks\nstem from reliance on purely statistical criteria that may omit causal relationships. This leads to\nselecting variables that appear significant by chance but do not generalize in autocorrelated data,\ngenerating misleading models [221, 14, 222, 223]. As the number of candidate variables increases, the\nlikelihood of selecting variables that appear statistically significant by chance but do not contribute\nto generalizable predictions also increases [224]. Studies demonstrate that selected variables may be\nsignificant by chance but fail to generalize in spatially autocorrelated data [225, 226]. To address\nthese limitations, this study explores several modeling frameworks to evaluate the approaches that\ncan best mitigate overfitting in mobile monitoring data. For example, in one framework, we examine", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "68-2", + "page": 64, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s\nthese limitations, this study explores several modeling frameworks to evaluate the approaches that\ncan best mitigate overfitting in mobile monitoring data. For example, in one framework, we examine\nthe integration of FFS with STCV to assess how aligning feature selection and hyperparameter tuning\nwith the data’s spatial and temporal structure influences model generalization.\nRobust models for data with spatial and temporal autocorrelation require CV strategies for\nhyperparameter tuning and feature selection to ensure generalizability and minimize overfitting [13,\n227]. Meyer et al. (2019) [15] showed that combining spatial CV with FFS significantly improved\nmodel robustness by avoiding spurious predictors in spatially correlated data. Similarly, Meyer et\nal. (2018) [14] integrated FFS with a target-oriented spatiotemporal CV approach for temperature\nand soil-content data from stationary stations, systematically removing misleading variables and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "68-3", + "page": 64, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "er et\nal. (2018) [14] integrated FFS with a target-oriented spatiotemporal CV approach for temperature\nand soil-content data from stationary stations, systematically removing misleading variables and\nachieving more generalizable predictions. Baumberger et al. (2024) [228] highlighted the importance\nof spatiotemporal CV, FFS, and targeted hyperparameter tuning in a high-resolution environmental\nmodel using stationary station records for soil temperature and soil moisture.\nIn this study, we develop LUR models using XGBoost to predict UFP concentrations based on\nmobile monitoring data collected in Toronto. We systematically compare seven modeling approaches\nto assess the impact of hyperparameter tuning and feature selection on model generalizability. In\naddition, we develop four additional models to assess the robustness of modeling approaches under\ndifferent outlier-handling methods. We tested CV strategies for tuning and feature selection and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "68-4", + "page": 64, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ". In\naddition, we develop four additional models to assess the robustness of modeling approaches under\ndifferent outlier-handling methods. We tested CV strategies for tuning and feature selection and\nassessed model robustness under different outlier treatments. This allowed us to compare modeling\napproaches and understand how they influence performance when dealing with mobile monitoring", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "68-5", + "page": 64, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "data. This study contributes to the development of ML-based exposure assessment models, providing\ninsights into optimizing feature selection, hyperparameter tuning, and CV strategies to improve\nthe reliability of mobile air pollution estimates.\nFinally, we compare model predictions against\ncontinuously recorded data from stationary sensors.\n4.3\nMethods\n4.3.1\nStudy area and data collection\nIn Toronto, UFP concentrations are primarily influenced by traffic emissions, particularly from diesel\nvehicles and congestion along major roadways, as well as localized industrial activities[229]. These\nsources create strong spatial gradients and diurnal variation, with higher concentrations during\nrush hours and lower levels at night and on weekends[86]. Dilution and dispersion in green and\nopen areas such as parks and vegetated corridors further reduce local UFP levels through enhanced\natmospheric mixing and particle deposition[230]. Between April 7 and June 23, 2021, we conducted", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "69-0", + "page": 65, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and\nopen areas such as parks and vegetated corridors further reduce local UFP levels through enhanced\natmospheric mixing and particle deposition[230]. Between April 7 and June 23, 2021, we conducted\nmobile monitoring in Toronto to capture spatial variations in UFP concentrations between 8 AM and\n7 PM over 34 days. Data were recorded at one-second intervals using the UrbanScanner platform,\nwhich integrates a UFP DiscMini, wind anemometer, and GPS, along routes covering major roads,\nresidential, industrial, and highways as described by Ganji et al. [196]. In addition, 12 DiscMini\nand 5 UFP Partector instruments were deployed in residential backyards from July 9 to 30, 2021,\nproviding continuous UFP measurements over a 3-week period (Figure 4.1). Backyard DiscMini\nsensors recorded UFP concentrations at 10-second intervals, while Partector sensors collected data\nevery second. To align with the mobile monitoring campaign, we used data recorded by these sensors\nbetween 8 AM and 7 PM.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "69-1", + "page": 65, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "concentrations at 10-second intervals, while Partector sensors collected data\nevery second. To align with the mobile monitoring campaign, we used data recorded by these sensors\nbetween 8 AM and 7 PM. The DiSCmini instruments were factory-calibrated by the manufacturer\nprior to deployment, following procedures consistent with previous Mobile sensing platform for\nair pollution and imagery (UrbanScanner) campaigns [114, 19].\nThe performance of this class\nof diffusion-charger–based UFP sensors has been independently validated in laboratory studies;\nfor example, Mills et al. [231] compared a DiSCmini-type instrument with reference condensation\nparticle counter (CPC) and scanning mobility particle sizer (SMPS) instruments, reporting strong\nagreement (R2 > 0.9) across particle sizes from 10 to 300 nm and number concentrations up to\n106 particles cm−3.\nThe instrument measures particles between 10 and 300 nm, with a lower", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "69-2", + "page": 65, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "reporting strong\nagreement (R2 > 0.9) across particle sizes from 10 to 300 nm and number concentrations up to\n106 particles cm−3.\nThe instrument measures particles between 10 and 300 nm, with a lower\ndetection cut near 10 nm that excludes nucleation-mode particles but does not affect relative spatial\nvariability across road segments.\n4.3.2\nPredictor variables\nTo ensure spatial consistency, we defined 100-meter distance points along the mobile monitoring\nroutes. At each of these points, air pollution measurements [232] were aggregated and a median\nvalue reported [233].\nSpatial buffers were then created around these 100-meter interval points\nto extract land use and traffic-related variables within predefined buffer zones.\nTo model UFP\nconcentrations, we incorporated a comprehensive set of 194 predictor variables, categorized into\nland use, traffic-related, distance-based, and temporally changing variables (Table A-1).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "69-3", + "page": 65, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 4.1: Study area in Toronto with mobile monitoring routes and stationary backyard sensor\nlocations.\n4.3.3\nModel development\nWe used XGBoost, a gradient boosting decision tree algorithm, to develop the LUR models for\nUFP. Outlier removal by road type was applied using the Interquartile Range (IQR) method, where\noutliers were identified separately for each road type and removed if they fell outside Q1-1.5 ×\nIQR and Q3 + 1.5 × IQR. Finally, normalization was applied to standardize continuous predictors.\nFrom this dataset, we created seven models to evaluate different CV strategies for hyperparameter\ntuning and feature selection. In addition, we explored two alternative outlier-handling techniques:\nremoving values above the 95th percentile or retaining all spikes, by building four additional models\nfollowing two modeling frameworks: random CV alone, and spatiotemporal CV combined with\nfeature selection.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "70-0", + "page": 66, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "values above the 95th percentile or retaining all spikes, by building four additional models\nfollowing two modeling frameworks: random CV alone, and spatiotemporal CV combined with\nfeature selection.\nWe used 80% of the data points as training set for hyperparameter tuning and feature selec-\ntion.\nAll model development and validation procedures were implemented in Python using the\nscikit-learn and XGBoost libraries.\n4.3.4\nCross-validation frameworks and model configurations\nWe defined four CV schemes for model tuning, each using K-fold CV on the training dataset:\n• Random CV (RCV): Data were randomly split into training and validation (10-folds),\nignoring spatial or temporal structure (step-by-step details are provided in the Appendix A).\n• Spatial CV (SCV): We generated 400 geographic clusters using the 100 m-aggregated train-\ning records (see Figure A-1a). These clusters were randomly divided into 10 folds for SCV,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "70-1", + "page": 66, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "the Appendix A).\n• Spatial CV (SCV): We generated 400 geographic clusters using the 100 m-aggregated train-\ning records (see Figure A-1a). These clusters were randomly divided into 10 folds for SCV,\nensuring each cluster served as the validation region once ( Figure A-1a). A 300 m dead buffer\nwas imposed around validation clusters to limit spatial leakage [15] (step-by-step details are", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "70-2", + "page": 66, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "provided in the Appendix A).\n• Temporal CV (TCV): We partitioned the data by sampling day. We sorted the mobile\ndata chronologically and created 8 folds, each corresponding to a subset of days (Figure A-\n2). In each fold, models were trained on 7/8th of the days and validated on the remaining\n1/8th (validation involved predicting data from dates entirely unseen in training) (step-by-step\ndetails are provided in the Appendix A).\n• Spatiotemporal CV (STCV): We combined spatial and temporal blocking by holding out\ndata that were independent in both space and time. In our implementation, we ensured that\nno location from a certain set of clusters on certain days appeared in training if it was in the\nvalidation set. Essentially, we withheld entire day-of-cluster combinations. This STCV ap-\nproach is very stringent: for example, the model might be trained on data from most locations\nand days but tested on data from a particular group of locations on specific days (neither", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "71-0", + "page": 67, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "This STCV ap-\nproach is very stringent: for example, the model might be trained on data from most locations\nand days but tested on data from a particular group of locations on specific days (neither\nthose locations nor dates appear in training). Because our dataset is large, we achieved STCV\nby first clustering locations (as in SCV) and then in each fold selecting a subset of clusters\nand specific days to hold out, such that validation data are separated in both dimensions.\nThis approach aligns with the “target-oriented” validation concept introduced by Meyer et\nal. (2018) [14] (step-by-step details are provided in the Appendix A).\nFor each CV scheme, we conducted hyperparameter tuning (HP) by Bayesian optimization using\nonly the training folds and evaluated performance on the validation fold in each iteration. Hyper-\nparameters such as tree depth, learning rate, and regularization parameters were optimized based\non the average validation performance across the folds.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "71-1", + "page": 67, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "the validation fold in each iteration. Hyper-\nparameters such as tree depth, learning rate, and regularization parameters were optimized based\non the average validation performance across the folds.\nIn addition to the full-feature models, we built models with FFS under the three structured CV\nframeworks (See Appendix A\nfor details). We employed a forward stepwise selection approach:\nwe added predictors one at a time, in each step including the variable that most improved the CV\nmean squared error (MSE) under the given CV scheme. This process was repeated until adding\nany remaining variable did not increase the MSE. The feature selection was performed within each\nCV strategy (See Appendix A\nfor details). This ensured that selected predictors contributed to\ngeneralizable skill rather than fitting coincidences of the training data [15, 14]. Figure 4.2 presents\na conceptual diagram of the STCV coupled with FFS modeling framework, illustrating the iterative", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "71-2", + "page": 67, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "eneralizable skill rather than fitting coincidences of the training data [15, 14]. Figure 4.2 presents\na conceptual diagram of the STCV coupled with FFS modeling framework, illustrating the iterative\nworkflow from spatiotemporal fold definition and hyperparameter tuning to feature selection and\nfinal model evaluation. While this schematic highlights the methodological innovation specific to\nthe STCV–FFS framework, the same forward feature selection workflow was also applied under the\nspatial and temporal CV schemes, with fold definitions differing across frameworks.\nUltimately, we trained and evaluated 11 models (Table 4.1) to examine how modeling frameworks\n(including hyperparameter tuning and feature selection) and outlier treatment influence model per-\nformance and robustness.\nFirst, we evaluated seven modeling frameworks: RCV HP, SCV HP,\nTCV HP, STCV HP, SCV FFS, TCV FFS, and STCV FFS which differ in their hyperparame-\nter tuning and feature selection processes.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "71-3", + "page": 67, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "obustness.\nFirst, we evaluated seven modeling frameworks: RCV HP, SCV HP,\nTCV HP, STCV HP, SCV FFS, TCV FFS, and STCV FFS which differ in their hyperparame-\nter tuning and feature selection processes. Next, to examine how outlier handling and sampling\nchanges affect the RCV HP and STCV FFS frameworks, four additional models were developed:\nRCV HP 95, STCV FFS 95, RCV HP noOut, STCV FFS noOut. Notably, RCV HP, RCV HP 95,\nand RCV HP noOut all follow the same random CV approach but vary in how they address outliers,\nwhile STCV FFS, STCV FFS 95, and STCV FFS noOut share the same spatiotemporal CV and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "71-4", + "page": 67, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 4.2: Conceptual diagram of the framework integrating FFS and STCV, representing the\niterative workflow from data partitioning and hyperparameter tuning to FFS combined with STCV\nand final model evaluation.\nfeature selection strategy, they differ by their outlier-removal methods. The subscript 95 represents\nthe removal of values in the top 5% or 95th percentile and noOut refers to no outlier-removal.\nModels without either of these suffixes follow the baseline outlier-handling approach, which applies\nIQR-based removal by road type.\n4.3.5\nModel implementation and evaluation\nAppropriate selection of evaluation metrics is essential to accurately reflect model performance and\ngeneralizability in air pollution exposure assessment studies [172, 13, 15]. Given the spatial and\ntemporal complexities involved in mobile monitoring datasets, we employed two distinct evaluation\nframeworks to check generalizability and overfitting of the models. First, to assess overfitting and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "72-0", + "page": 68, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "temporal complexities involved in mobile monitoring datasets, we employed two distinct evaluation\nframeworks to check generalizability and overfitting of the models. First, to assess overfitting and\nperformance on unseen data, we used a 20% random holdout sample excluded from training, tun-\ning, and feature selection. We computed median concentrations to mitigate transient spikes. This\ndataset, consistently used across all models, provided a standardized comparison basis, reporting R2,\nMAE, and RMSE as evaluation metrics (See Appendix A for details). Second, to assess whether a\nmodeling framework improved generalization, we reallocated all data into random, spatial, temporal,\nand spatiotemporal folds. The tuning and feature selection process relied on a specific data split,\nwhereas final split for checking the models were developed using the entire dataset with different\ncluster allocations. Figure A-1b illustrates an example of a single fold from the 10-fold spatial CV", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "72-1", + "page": 68, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "whereas final split for checking the models were developed using the entire dataset with different\ncluster allocations. Figure A-1b illustrates an example of a single fold from the 10-fold spatial CV\nused during model tuning and feature selection, while Figure A-3 shows an example of a single fold\nfrom the 10-fold spatial CV used during the evaluation phase. Once each model had undergone", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "72-2", + "page": 68, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Table 4.1: Summary of model configurations\nModel Name\nCross-validation\nFeature\nOutlier\nHan-\nNote\nfor\nhyperparame-\nSelection\ndling\nter tuning\nRCV HP\nRandom CV\nNo\nBy road-type IQR-\nHyperparameter tuning with random CV. All\nbased removal\nfeatures are used as inputs.\nSCV HP\nSpatial CV\nNo\nBy road-type IQR-\nHyperparameter tuning with Spatial CV. All\nbased removal\nfeatures are used as inputs.\nTCV HP\nTemporal CV\nNo\nBy road-type IQR-\nUses day-wise splits to handle temporal vari-\nbased removal\nability.\nHyperparameter tuning done under\nTemporal CV while using all features.\nSTCV HP\nSpatiotemporal CV\nNo\nBy road-type IQR-\nHyperparameter\ntuning\ndone\nunder\nSpa-\nbased removal\ntiotemporal CV using all features.\nSCV FFS\nSpatial CV\nYes\nBy road-type IQR-\nForward feature selection integrated with Spa-\nbased removal\ntial CV. Hyperparameter tuning with Spatial\nCV.\nTCV FFS\nTemporal CV\nYes\nBy road-type IQR-\nFFS integrated with Temporal CV. Hyperpa-\nbased removal\nrameter tuning with Temporal CV while using", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "73-0", + "page": 69, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "d removal\ntial CV. Hyperparameter tuning with Spatial\nCV.\nTCV FFS\nTemporal CV\nYes\nBy road-type IQR-\nFFS integrated with Temporal CV. Hyperpa-\nbased removal\nrameter tuning with Temporal CV while using\nselected features.\nSTCV FFS\nSpatiotemporal CV\nYes\nBy road-type IQR-\nFFS integrated with Spatiotemporal CV. Hy-\nbased removal\nperparameter tuning with Temporal CV while\nusing selected features.\nRCV HP 95\nRandom CV\nNo\nRemove\ntop\n5%\nThis model is similar to RCV HP. Tests the\noutliers\neffect of removing extreme spikes (above the\n95th percentile) while using RCV HP frame-\nwork.\nSTCV FFS 95\nSpatiotemporal CV\nYes\nRemove\ntop\n5%\nThis model is similar to STCV FFS. Tests\noutliers\nthe effect of removing extreme spikes (above\nthe 95th percentile) while using STCV FFS\nframework.\nRCV HP noOut\nRandom CV\nNo\nNo outlier removal\nThis model is similar to RCV HP. Tests the\neffect of not removing extreme spikes while us-\ning RCV HP framework.\nSTCV FFS noOut\nSpatiotemporal CV\nYes\nNo outlier removal", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "73-1", + "page": 69, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "andom CV\nNo\nNo outlier removal\nThis model is similar to RCV HP. Tests the\neffect of not removing extreme spikes while us-\ning RCV HP framework.\nSTCV FFS noOut\nSpatiotemporal CV\nYes\nNo outlier removal\nThis model is similar to STCV FFS. Tests the\neffect of not removing extreme spikes while us-\ning STCV FFS framework.\nNote: Model names indicate the cross-validation (CV) strategy used for hyperparameter tuning (RCV =\nRandom CV, STCV = Spatiotemporal CV), followed by whether forward feature selection (FFS) was\napplied, and then followed by outlier-handling approach. Models with suffixes like “ 95” or “ noOut” test\nthe effect of removing or retaining outliers.\ntuning and feature selection with its designated strategy, we tested it using all four newly defined\nCV methods in a comparable manner (e.g., consistent cluster assignments for SCV and STCV,\nday-based splits for TCV). This ensures a direct comparison of performance while minimizing con-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "73-2", + "page": 69, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ewly defined\nCV methods in a comparable manner (e.g., consistent cluster assignments for SCV and STCV,\nday-based splits for TCV). This ensures a direct comparison of performance while minimizing con-\nfounding effects from varying fold assignments (See Appendix A for details). Evaluating each model\nunder these four CV schemes enables a nuanced assessment of performance across varying levels of\ngeographic and temporal separation between training and validation subsets.\nGiven our modeling context, the interpretation of CV metrics, especially R2, requires caution.\nMobile monitoring data aggregated from limited visits per location is susceptible to episodic pollution\nspikes and temporal variability, potentially inflating prediction errors and lowering R2, especially\nunder rigorous CV schemes (SCV, TCV, STCV). Consequently, while R2 is essential for seeing the\nimprovement in generalization, it does not necessarily represent the model’s capability to accurately", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "73-3", + "page": 69, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ly\nunder rigorous CV schemes (SCV, TCV, STCV). Consequently, while R2 is essential for seeing the\nimprovement in generalization, it does not necessarily represent the model’s capability to accurately\nestimate average pollutant levels over a given period. A model evaluated under TCV, for example,\nmight exhibit low R2 due to temporal variability but still accurately capture average spatial pollution\npatterns when evaluated on continuous data representative of the given period [8, 207, 219, 232].\nIn this study, CV performance metrics (e.g., R2, Mean Squared Error (MSE)) were primarily used", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "73-4", + "page": 69, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "for hyperparameter tuning and feature selection and then each model was tested across four CV.\nOverall, by integrating rigorous CV for evaluation (to inform robust model development and prevent\noverfitting) and an internal holdout sample evaluation, we aim to provide a balanced, transparent,\nand realistic assessment of each model’s capability to generalize.\nSensitivity analysis was conducted using SHAP-based feature importance plots, which reveal\nhow each predictor influences UFP. By averaging the absolute SHAP values, we identified the most\ncritical predictors. Finally, we generated a 50 m resolution UFP map for Toronto by applying the\nfinal XGBoost model to a regular grid, incorporating land-use, traffic, and distance-based predictors.\nThis approach captures spatial distributions and allows visual comparisons across different model\npredictions.\nTo independently explore model performance, we performed external validation against a sta-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "74-0", + "page": 70, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s approach captures spatial distributions and allows visual comparisons across different model\npredictions.\nTo independently explore model performance, we performed external validation against a sta-\ntionary network of 17 UFP fixed monitors installed in backyards of private homes throughout the\ncity of Toronto (12 DiscMini and 5 Partector) and which continuously measured UFP from July 9\nto July 30, 2021, while our mobile campaign concluded on June 23, 2021. Despite non-overlapping\nperiods, meteorological and emission conditions were sufficiently similar, enabling the backyard data\nto serve as a valuable continuous external holdout sample. Model performance was evaluated by\ncomputing the APE between predicted and observed mean UFP concentrations at each site (APE;\nSee Appendix A for the equation).\n4.4\nResults\n4.4.1\nMeasured concentrations\nThe mobile monitoring campaign yielded an extensive dataset of UFP measurements (over 440,000\nvalid observations after quality control).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "74-1", + "page": 70, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "A for the equation).\n4.4\nResults\n4.4.1\nMeasured concentrations\nThe mobile monitoring campaign yielded an extensive dataset of UFP measurements (over 440,000\nvalid observations after quality control). Concentrations varied greatly from day to day, reflecting\nchanging traffic, meteorology, and local emissions (Figure A-4). To further characterize the dataset,\nwe first examined the distribution of UFP concentrations after applying IQR-based filtering by road\ntype.\nThis method excluded road-specific statistical outliers and resulted in a mean concentra-\ntion of 16,817 particles/cm³ and a median of 13,856 particles/cm³. We also assessed the impact\nof two alternative outlier strategies: retaining all values without removal and excluding the top\n5% of concentrations (95th percentile removal). The unfiltered dataset yielded the highest mean\n(23,906 particles/cm³) and a median of 14,777 particles/cm³, indicating a right-skewed distribution\ninfluenced by extreme values.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "74-2", + "page": 70, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "h percentile removal). The unfiltered dataset yielded the highest mean\n(23,906 particles/cm³) and a median of 14,777 particles/cm³, indicating a right-skewed distribution\ninfluenced by extreme values. The 95th percentile removal approach produced similar summary\nstatistics to the IQR method, with a mean of 16,816 particles/cm³ and a median of 14,017 parti-\ncles/cm³. These results, summarized in Table A-2, highlight how different filtering strategies can\nshape the overall distribution of mobile UFP data.\nMeasurements from the backyard sensors over the three-week monitoring period (Figure A-5)\nreveal substantial variability in UFP concentrations across locations, likely reflecting differences in\nproximity to emission sources such as roads, buildings, or vegetation. Summary statistics for each\nmonitor (Table A-3) further highlight this heterogeneity: median UFP concentrations ranged from\nas low as 2,915 particles/cm³ to as high as 9,018 particles/cm³. Compared to mobile monitoring", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "74-3", + "page": 70, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "for each\nmonitor (Table A-3) further highlight this heterogeneity: median UFP concentrations ranged from\nas low as 2,915 particles/cm³ to as high as 9,018 particles/cm³. Compared to mobile monitoring\ndata, with a median concentration of approximately 13,856 particles/cm³ (after IQR-based filtering),\nstationary measurements exhibited lower medians and reduced day-to-day variability.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "74-4", + "page": 70, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "4.4.2\nImpact of modeling framework\nSeven modeling frameworks were evaluated to assess the impact of hyperparameter tuning and\nfeature selection combined with CV strategies on performance, generalization, and overfitting. Table\nA-4 compares 80% training and 20% holdout performance, highlighting fitting, overfitting, and how\nCV strategies and FFS affect robustness. The RCV HP model exhibits a near-perfect training R2 of\n0.999, with low RMSE (400) and MAE (261). However, its test R2 declines to 0.715 (RMSE = 6,692,\nMAE = 3,869) which is a sign of overfitting. Models tuned using structured CV approaches—such\nas SCV HP, TCV HP, and STCV HP—show more moderate training R2 values (0.544–0.825) and\nachieve test set R2 values between 0.510 and 0.660.\nIn addition to the models without feature\nselection, Table A-4 shows how incorporating FFS combined with different CV strategies impacts\nmodel performance. For instance, the SCV FFS model yields a training R2 of 0.957 and a test R2\nof 0.717.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "75-0", + "page": 71, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "selection, Table A-4 shows how incorporating FFS combined with different CV strategies impacts\nmodel performance. For instance, the SCV FFS model yields a training R2 of 0.957 and a test R2\nof 0.717. Meanwhile, models with feature selection merged and tuned under Temporal (TCV FFS)\nand Spatiotemporal (STCV FFS) CV schemes exhibit lower overall R2 values on both training and\ntest sets: 0.413 and 0.441 in training, and 0.398 and 0.450 in testing, indicating that overfitting was\neffectively avoided.\nFigure 4.3 presents scatter plots of the predicted versus actual UFP concentrations (parti-\ncles/cm³) for all the seven modeling frameworks using the 20% random holdout internal test set.\nThe red dashed line represents perfect prediction. Models using Random CV tend to have points\nclustered near this line, indicating high accuracy in a similar data split; however, this may hide\noverfitting to local patterns. In contrast, models using spatial, temporal, or spatiotemporal CV and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "75-1", + "page": 71, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "clustered near this line, indicating high accuracy in a similar data split; however, this may hide\noverfitting to local patterns. In contrast, models using spatial, temporal, or spatiotemporal CV and\nespecially those with FFS, show a wider spread in the data. This spread reflects the challenge of\npredicting pollution in different locations and times. Lower R2 values in these cases do not automat-\nically mean the models perform poorly, they reflect the limited visits per location and the episodic\nnature of mobile monitoring.\nTable A-5 presents the performance of seven XGBoost-based models under four CV approaches:\nRandom, Spatial, Temporal, and Spatiotemporal.\nEach model’s performance is summarized by\nthree metrics: R2, RMSE, and MAE. When using Random CV, RCV HP attains an R2 of 0.735,\nwith an RMSE of 6245 and an MAE of 3656, while SCV FFS records R2 = 0.717, RMSE = 6406,\nMAE = 3964. Models such as TCV FFS and STCV FFS show lower R2 values (0.389 and 0.413,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "75-2", + "page": 71, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "attains an R2 of 0.735,\nwith an RMSE of 6245 and an MAE of 3656, while SCV FFS records R2 = 0.717, RMSE = 6406,\nMAE = 3964. Models such as TCV FFS and STCV FFS show lower R2 values (0.389 and 0.413,\nrespectively) yet vary in their RMSE and MAE. Under Spatial CV, RCV HP R2 is 0.355, while\nSCV FFS reaches 0.418. Similarly, TCV FFS and STCV FFS have R2 values of 0.265 and 0.278,\nrespectively. Under Temporal CV, the R2 values typically drop further, reflecting the challenge of\nday-to-day variability in meteorology, traffic flows, and emission sources that can dominate mobile\nmonitoring data. For instance, RCV HP experiences a stark decrease from 0.355 in Spatial CV to\na mere 0.035 in Temporal CV, indicating its inability to generalize across distinct sampling days.\nModels specifically incorporating feature selection and/or tuned for temporal splits like TCV FFS,\nachieve a higher R2 of 0.173. Some models, such as SCV FFS, score as low as 0.054 in R2. RMSE", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "75-3", + "page": 71, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "days.\nModels specifically incorporating feature selection and/or tuned for temporal splits like TCV FFS,\nachieve a higher R2 of 0.173. Some models, such as SCV FFS, score as low as 0.054 in R2. RMSE\nand MAE similarly show larger absolute values for most models under this partitioning. Under\nspatiotemporal CV evaluation, RCV HP has an R2 of 0.105, and STCV FFS achieves 0.186.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "75-4", + "page": 71, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 4.3: Predicted versus actual UFP concentrations on the 20% random holdout test set for\neach model. The dashed line denotes perfect 1:1 agreement.\n4.4.3\nImpact of outliers on predictions\nTable A-6 presents the training and test set performance of four additional models differing by outlier\nhandling (none or 95th-percentile removal) and modeling framework (random vs. spatiotemporal\nCV with feature selection). Comparing the metrics for this table with the results from Table A-\n4 shows that the models tuned under Random CV often exhibit near-perfect R2 on the training\nset yet experience marked declines on the test set, suggesting overfitting when extreme values are\neither retained or filtered. Models that have feature selection combined with spatiotemporal CV\ntend to show lower training R2 but more balanced performance on the test set, indicating better\ngeneralization. The four additional models have also been tested under four different CV strategies.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "76-0", + "page": 72, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "CV\ntend to show lower training R2 but more balanced performance on the test set, indicating better\ngeneralization. The four additional models have also been tested under four different CV strategies.\nThe results of this evaluation (Table A-7) reveal that modeling frameworks with feature selection\nin conjunction with proper CV can improve generalization across time and space. Although models", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "76-1", + "page": 72, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tuned under random CV might yield deceptively high R2, imposing stricter splits along spatial and\ntemporal dimensions forces the models to learn predictors that hold up under varied conditions.\nBy comparing scatter plots in Figure 4.3 and Figure A-6, we observe two distinct groups of models\nwith similar behavior. The models STCV FFS, STCV FFS 95, and STCV FFS noOut exhibit sim-\nilar scatter patterns, suggesting consistent predictive structure even under different outlier-handling\nstrategies. In contrast, the RCV HP, RCV HP 95, and RCV HP noOut models display another\nshared pattern reinforcing the idea that model structure and CV strategy have a more pronounced\ninfluence on predictions than minor variations in outlier removal.\n4.4.4\nComparing with stationary measurements\nFigure 4.4 summarizes mean APE values by comparing average of stationary measurements with the\npredictions of the models. For instance, RCV HP shows notably high mean APE: 134%. Meanwhile,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "77-0", + "page": 73, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "asurements\nFigure 4.4 summarizes mean APE values by comparing average of stationary measurements with the\npredictions of the models. For instance, RCV HP shows notably high mean APE: 134%. Meanwhile,\nTCV FFS and STCV FFS markedly reduce these errors. TCV FFS achieves 65%, while STCV FFS\n78%. Models incorporating temporal or spatiotemporal CV, particularly when combined with FFS,\nappear to track average of continuous fixed-site sensor data more reliably than those relying on\nrandom or purely spatial CV for tuning and feature selection.\nAdditionally, comparing models\nwith different outlier-handling approaches reveals that RCV HP and its variations (RCV HP 95\nand RCV HP noOut) can produce substantially higher APEs (greater than 130%) indicating that\nmodels tuned with random CV do not align well with three-week average concentrations from sta-\ntionary sensors. In contrast, models using spatiotemporal CV with feature selection (STCV FFS,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "77-1", + "page": 73, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ting that\nmodels tuned with random CV do not align well with three-week average concentrations from sta-\ntionary sensors. In contrast, models using spatiotemporal CV with feature selection (STCV FFS,\nSTCV FFS 95, or STCV FFS noOut) consistently achieve lower APE, suggesting they may better\ncapture average pollution trends across the city, rather than being overly influenced by episodic\nmobile measurements.\n4.4.5\nSelected predictors and development of exposure surfaces\nSHAP analysis was applied to each of the XGBoost models to quantify the influence of individual\npredictors on UFP concentrations. In models without feature selection (e.g., RCV HP, SCV HP,\nTCV HP, STCV HP), the SHAP plots (Figure A-7) highlight a wide range of predictors, prominently\nfeaturing traffic-related variables (such as distance to highways, road area across multiple buffers,\nand AADT) and several meteorological variables. In contrast, models incorporating forward feature", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "77-2", + "page": 73, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "featuring traffic-related variables (such as distance to highways, road area across multiple buffers,\nand AADT) and several meteorological variables. In contrast, models incorporating forward feature\nselection (SCV FFS, TCV FFS, STCV FFS, STCV FFS 95, and STCV FFS noOut) retain fewer\npredictors, emphasizing more robust and interpretable features such as proximity to major roads,\nroad density, and specific land-use categories. Among the spatiotemporal CV models, STCV FFS\nretains 17 predictors, STCV FFS 95 includes 20, and STCV FFS noOut reduces the count to just\n9, demonstrating the effect of both CV strategy and outlier treatment on model parsimony and\ninterpretability.\nFigure 4.5 presents the exposure surfaces. Models relying on Random CV (RCV HP), Spatial CV\nfor hyperparameter tuning (SCV HP), Spatial CV with feature selection (SCV FFS), and Temporal\nCV without feature selection (TCV HP) exhibit abrupt “lines” where UFP levels shift dramatically,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "77-3", + "page": 73, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tial CV\nfor hyperparameter tuning (SCV HP), Spatial CV with feature selection (SCV FFS), and Temporal\nCV without feature selection (TCV HP) exhibit abrupt “lines” where UFP levels shift dramatically,\nresulting in concentration surfaces that appear less representative of known spatial trends in Toronto.\nThese models also show abrupt pollutant changes that deviate from expected averages, resulting\nin exposure surfaces with artifacts and indicating overfitting and a failure to reproduce realistic", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "77-4", + "page": 73, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 4.4: Mean APE(%) of UFP predictions compared with stationary measurements across all\nmodels.\nspatial gradients. In contrast, the model using spatiotemporal CV, even without feature selection\n(STCV HP), produces spatial patterns that better follow traffic-related variations, with elevated\nUFP levels along major roadways and lower concentrations in parks and open areas. The TCV FFS\nand STCV FFS models refine these trends further, assigning lower UFP values to vegetated areas\nand higher concentrations to major corridors. Their predicted gradients capture the expected UFP\nbehavior near roads more accurately and appear more physically consistent and better aligned with\naverage observations over the study period.\nFigure 4.5 also includes variations of these models under different outlier treatments. For Ran-\ndom CV (RCV HP, RCV HP 95, RCV HP noOut), altering the outlier threshold leads to noticeably", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "78-0", + "page": 74, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "y period.\nFigure 4.5 also includes variations of these models under different outlier treatments. For Ran-\ndom CV (RCV HP, RCV HP 95, RCV HP noOut), altering the outlier threshold leads to noticeably\ndifferent spatial surfaces, further illustrating the instability of random CV-based models. In contrast,\nSTCV FFS models (STCV FFS, STCV FFS 95, STCV FFS noOut) produce consistently smooth\nand interpretable concentration maps regardless of the outlier threshold, demonstrating that combin-\ning spatiotemporal CV with feature selection improves robustness to data variability and enhances\nspatial reliability.Residual spatial autocorrelation was evaluated using global Moran’s I on model\nresiduals, with results reported in the Supporting Information (Table A-8).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "78-1", + "page": 74, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 4.5: Predicted UFP exposure surfaces based on all models, highlighting the spatial vari-\nability associated with different cross-validation, feature selection, and outlier treatment strategies.\n4.5\nDiscussion\nIn this study, two internal evaluation approaches were used: a 20% random holdout and four CV\nschemes, which were applied to assess each LUR model using mobile monitoring data. Because\nmobile data are collected over short periods and can exhibit high temporal and spatial variability,\nthese internal evaluations may not fully capture true long-term pollution levels. In addition, struc-\ntured CV can impose overly stringent spatial and temporal constraints, creating test scenarios that\nare more challenging than typical real-world conditions and may not accurately reflect a model’s\ntrue predictive power. Therefore, the metrics derived from mobile data under different CV schemes\nshould not be interpreted as indicators of each model’s ability to predict longer-term concentrations,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "79-0", + "page": 75, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "which is the ultimate goal of modeling. Instead, they should be used to compare robustness, over-\nfitting, and improvement in generalization among the models. Accordingly, the reported R2 values\nshould be interpreted in a comparative sense, highlighting relative model stability across validation\nstrategies rather than as absolute measures of predictive accuracy.\n4.5.1\nImpact of cross validation for hyperparameter tuning\nThe choice of CV for tuning and feature selection had a large impact on model performance and\ngeneralizability.\nModels tuned with conventional random CV appeared to fit the training data\nextremely well but failed to generalize to unseen space or time. For instance, RCV HP achieved\nan almost perfect fit to its training data (R2 ≈0.999) but saw its R2 drop to 0.71 on a hold-out\ntest set. This sharp decline, alongside a jump in error measures, illustrates severe overfitting when\nspatial and temporal autocorrelations are ignored during hyperparameter tuning.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "80-0", + "page": 76, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "0.71 on a hold-out\ntest set. This sharp decline, alongside a jump in error measures, illustrates severe overfitting when\nspatial and temporal autocorrelations are ignored during hyperparameter tuning. In contrast, models\ntuned under structured CV scheme yielded more moderate training fits (training R2 0.54–0.83) and\ncomparable test performance (test R2 0.51–0.66).\nComparing the performance of models without feature selection under different CV methods\nindicate that using conventional random CV for tuning inflates model performance because it mixes\nspatially correlated data within both training and test sets. For instance, models like RCV HP\nexhibit very high R2 values of 0.735 under random CV, yet their performance sharply declines\nwhen evaluated with structured CV approaches, dropping to as low as 0.035 under Temporal CV.\nOn the other hand, using more strict splits for hyperparameter tuning (e.g., STCV HP) improve\ngeneralization across sampling days and locations.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "80-1", + "page": 76, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "roaches, dropping to as low as 0.035 under Temporal CV.\nOn the other hand, using more strict splits for hyperparameter tuning (e.g., STCV HP) improve\ngeneralization across sampling days and locations. Models tuned with temporal or spatiotemporal\nCV align closely with the central tendency of aggregated mobile measurements at locations with\nsimilar traffic patterns, rather than reproducing short-term peaks from transient events. Moreover,\nthe prediction power of models tuned with spatially and temporally blocked CV was evident when\nevaluating on truly independent continuous\n4.5.2\nFeature selection combined with CV methods and predictor impor-\ntance\nFFS further enhanced model generalizability and interpretability, especially when coupled with struc-\ntured CV. Without feature selection, even a spatially blocked model could retain spurious predictors\nand overfit certain conditions. For example, the Spatial CV model with all features (SCV HP) and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "80-2", + "page": 76, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "-\ntured CV. Without feature selection, even a spatially blocked model could retain spurious predictors\nand overfit certain conditions. For example, the Spatial CV model with all features (SCV HP) and\neven with FFS (SCV FFS) showed relatively high training R2 ( 0.95) alongside a lower test R2\n( 0.72), suggesting some overfitting persisted. This outcome indicates that spatial blocking alone, or\nspatial blocking plus FFS, may not fully prevent overfitting if temporal variability is unaccounted\nfor. In contrast, introducing temporal separation during feature selection and tuning had a clear\nregularizing effect: the temporally blocked (TCV FFS) and spatiotemporal (STCV FFS) models\nended up with much lower training R2 (≈0.41–0.44) that were nearly matched by their test R2\n(≈0.40–0.45). The very small training/testing performance gap for STCV FFS implies that over-\nfitting was reduced, yielding a model that prioritizes generalizable predictors. Such conservative", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "80-3", + "page": 76, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "est R2\n(≈0.40–0.45). The very small training/testing performance gap for STCV FFS implies that over-\nfitting was reduced, yielding a model that prioritizes generalizable predictors. Such conservative\nmodels might initially appear to perform worse (because they do not chase every idiosyncrasy in the\ntraining data), but they in fact proved more robust on independent data and produced more credible\nspatial predictions. Combining feature selection with temporal or spatiotemporal CV narrows the", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "80-4", + "page": 76, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "models to predictors that reflect stable patterns rather than short-lived spikes. Our study focuses on\ngeneralizable prediction, aiming to estimate exposure surfaces (long-term average concentrations)\nacross new locations and periods rather than interpolate within the sampled domain. Unlike spa-\ntiotemporal interpolation methods such as kriging, which deliberately exploit spatial dependence\nas the predictive signal[234], our objective is to learn transferable relationships between predictors\nand pollution levels. In this context, spatial autocorrelation can inflate CV performance, nudging\nmodels to reproduce patterns instead of capturing generalizable processes[13, 221, 14].\nThe robustness of models with feature selection combined with spatiotemporal CV was evident\nwhen evaluating on truly independent data. The random CV model showed poor agreement, whereas\nthe spatiotemporal CV with feature selection had an error nearly half as large. In fact, a model", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "81-0", + "page": 77, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "evident\nwhen evaluating on truly independent data. The random CV model showed poor agreement, whereas\nthe spatiotemporal CV with feature selection had an error nearly half as large. In fact, a model\ntrained with random CV and no outlier filtering completely failed to capture average pollution levels\n(overall APE ≈217%), whereas the STCV FFS, STCV FFS 95, and STCV FFS noOut maintained\nerror around 78–79% even with different outlier treatment.\nThis contrast in external validation\nperformance confirms that models tuned with naive random resampling were overfit to peculiarities\nof the training campaign, whereas incorporating feature selection and spatiotemporal CV yielded\nmodels that better generalize to independent data.\nThe model interpretation using SHAP plots provides insight into how feature selection using\ndifferent CV methods narrowed the predictors down to the most important ones. The overfit random-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "81-1", + "page": 77, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ent data.\nThe model interpretation using SHAP plots provides insight into how feature selection using\ndifferent CV methods narrowed the predictors down to the most important ones. The overfit random-\nCV model that used the full feature set (RCV HP) effectively relied on all 194 candidate features,\nmaking it difficult to interpret and indicating the model was fitting many spurious relationships.\nIn contrast, the models with FFS combined spatiotemporal CV retained a much smaller subset of\nvariables. The SHAP analysis consistently identified a few dominant drivers of UFP across these\nmodels including traffic and road infrastructure such as road area within 50 m, distance to the\nnearest highway, and traffic volume (AADT). This further demonstrates that the STCV framework\nsuccessfully identified these variables as sources of non-generalizable, spurious correlation. Even\nmeteorological or site-specific variables, which appeared in some less constrained models, were largely", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "81-2", + "page": 77, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "essfully identified these variables as sources of non-generalizable, spurious correlation. Even\nmeteorological or site-specific variables, which appeared in some less constrained models, were largely\nabsent or deemphasized in the STCV FFS models, likely because those factors do not generalize\nspatially if they were tied to specific days or locations. The fact that the same core features emerged\nunder different outlier-handling scenarios suggests that these predictors represent enduring signals\nrather than artifacts of preprocessing choice.\n4.5.3\nImpact of outlier treatment\nHandling of extreme UFP readings (“outliers”) had a notable effect on models tuned with random\nCV but had minimal impact on those built with the spatiotemporal CV + feature selection frame-\nwork. When using random CV for model tuning, the inclusion or exclusion of transient extreme\nspikes significantly altered both error metrics and predicted spatial patterns, indicating a tendency\nto “chase” these peaks.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "81-3", + "page": 77, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "andom CV for model tuning, the inclusion or exclusion of transient extreme\nspikes significantly altered both error metrics and predicted spatial patterns, indicating a tendency\nto “chase” these peaks. In contrast, the models developed with spatiotemporal CV and feature selec-\ntion demonstrated a remarkable robustness to outlier handling. The STCV FFS models performed\nnearly identically regardless of whether we kept all spikes or filtered them out. For these models,\nwhen predictions were compared with stationary measurements, the errors remained consistently\nlow across all outlier treatments, indicating that the model had learned the underlying signal in the\ndata. Despite this improved performance, the STCV FFS, STCV FFS 95, and STCV FFS noOut\nmodels still tended to overpredict relative to stationary backyard monitors, resulting in a mean", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "81-4", + "page": 77, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "absolute percentage error of about 78–79%. This discrepancy is expected because mobile monitor-\ning measurements were collected directly on roadways, whereas stationary monitors are typically\nlocated in residential backyards farther from emission sources. UFP concentrations are known to\ndecrease steeply with distance from traffic, often dropping by 50–70% within the first 100–300 m\nand approaching background levels beyond 300 m[128, 235]. Therefore, higher predicted values com-\npared to backyard observations are physically consistent with the well-established near-road decay\nbehavior of UFP. This difference in measurement environments represents a limitation of the valida-\ntion approach, as direct comparison between on-road and residential measurements may introduce\nsystematic differences in observed concentrations. The robustness of the STCV FFS framework to\noutliers is driven by its ability to remove predictors that do not meaningfully contribute to persis-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "82-0", + "page": 78, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ce\nsystematic differences in observed concentrations. The robustness of the STCV FFS framework to\noutliers is driven by its ability to remove predictors that do not meaningfully contribute to persis-\ntent patterns, while retaining those linked to stable sources such as traffic and road networks. This\nproperty is especially important for exposure mapping: it means the model can incorporate genuine\nhigh-pollution events (e.g., data recorded while passing a truck or during high traffic hours) into\nthe predictions without being misled by them. This guarantees that true extreme pollution events\nare kept and correctly shown as hotspots if they are part of consistent patterns, such as persistently\nhigh concentrations along highway interchanges.\n4.5.4\nModeling recommendations\nBased on our findings and previous studies, we propose a set of best practices for developing gener-\nalizable models using autocorrelated air quality data collected via mobile monitoring:\nA.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "82-1", + "page": 78, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "endations\nBased on our findings and previous studies, we propose a set of best practices for developing gener-\nalizable models using autocorrelated air quality data collected via mobile monitoring:\nA. Define the modeling objective: Clearly distinguish between interpolation (within-sample\nprediction) and generalization (prediction at new locations or times), as this determines how\ntraining and validation data are split.\nB. Align cross-validation with the modeling goal and data structure: The structure of\nCV and the definition of folds should reflect both the study objective and the data’s spatiotem-\nporal correlation structure.\nC. Design folds based on autocorrelation range: Spatial folds and cluster sizes should\nreflect the data’s autocorrelation range, with buffer zones applied around validation clusters\nto prevent spatial leakage.\nD. Integrate CV into feature selection and tuning: Perform feature selection and hyperpa-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "82-2", + "page": 78, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ata’s autocorrelation range, with buffer zones applied around validation clusters\nto prevent spatial leakage.\nD. Integrate CV into feature selection and tuning: Perform feature selection and hyperpa-\nrameter tuning within the structured CV to prevent overfitting and the selection of spurious\npredictors tied to specific sampling patterns.\nE. Prioritize external and visual validation: For mobile data, when the goal is to estimate\naverage surfaces over time, rely on independent continuous measurements and visual inspection\nto confirm expected spatial trends and detect artifacts. Internal CVs using mobile data should\nnot be interpreted as performance indicators. When comparing several models, internal CV\nmetrics should be interpreted comparatively to assess robustness and overfitting. Random CV\nreflects the model’s fit within the sampled data, while structured CV (spatial, temporal, or\nspatiotemporal) provides a more realistic measure of generalization. Models that maintain", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "82-3", + "page": 78, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tting. Random CV\nreflects the model’s fit within the sampled data, while structured CV (spatial, temporal, or\nspatiotemporal) provides a more realistic measure of generalization. Models that maintain\nmore balanced metrics across these CV schemes generally demonstrate stronger real-world\nreliability.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "82-4", + "page": 78, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "4.5.5\nFuture directions and implications\nFuture work should investigate the transferability of the feature selection approach combined with\nspatiotemporal CV by applying the model to different time periods and additional urban areas.\nTesting the model’s performance on multi-year mobile monitoring datasets in diverse cities or time\nperiods would help determine whether the identified predictors remain robust over time and across\nvarious urban contexts. Future work could also test the applicability of this framework to other\npollutants such as NO2 or PM2.5, which spatial and temporal characteristics differ from UFP. For\nsuch pollutants, the spatial folds in blocked cross-validation should be defined according to the\nrange and structure of their respective autocorrelation patterns. Another promising direction is to\nintegrate supplementary datasets, such as satellite-derived imagery or detailed meteorological data,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "83-0", + "page": 79, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ange and structure of their respective autocorrelation patterns. Another promising direction is to\nintegrate supplementary datasets, such as satellite-derived imagery or detailed meteorological data,\nto further refine the model’s predictive capabilities and capture finer-scale variations. Moreover,\nexploring alternative feature selection techniques, or hybrid methods that combine statistical and\ndomain-expertise-driven approaches, may yield further improvements in model performance and\ninterpretability, ultimately leading to more reliable air pollution exposure assessments.\nThis study demonstrates that aligning model development with the spatiotemporal structure of\nmobile air pollution data and the objective of modeling are critical for reliable predictions. Among\nthe tested frameworks, models using spatiotemporal CV with forward feature selection performed\nbest, avoiding overfitting and producing stable, interpretable UFP maps. In contrast, conventional", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "83-1", + "page": 79, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Among\nthe tested frameworks, models using spatiotemporal CV with forward feature selection performed\nbest, avoiding overfitting and producing stable, interpretable UFP maps. In contrast, conventional\nrandom CV overlooked spatial–temporal autocorrelation, encouraging data reproduction and un-\nstable predictions. By focusing on a small set of meaningful predictors and using rigorous CV, we\nimproved both robustness and external validity. Feature selection and objective-oriented CV are es-\nsential for building exposure models from mobile data, offering a reliable path toward high-resolution\npollution mapping in urban environments.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "83-2", + "page": 79, + "document_type": "thesis", + "page_header": "CHAPTER 4. BRIDGING THE GAP BETWEEN DATA REPRODUCTION AND PREDICTION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Chapter 5\nPost-COVID traffic composition: Why are reduc-\ntions in ultrafine particles and black carbon stalling?\n5.1\nChapter overview\nUnderstanding how traffic composition shapes urban air pollution is critical for effective transporta-\ntion and public health planning, yet vehicle-specific traffic patterns are rarely captured at high spatial\nresolution. This study integrates mobile air-quality monitoring and image-based traffic sensing to\nquantify the influence of fleet composition on air pollutant concentrations across Toronto, Canada.\nA 360◦camera, along with UFP and BC monitors, was mounted on a mobile platform and deployed\nduring two summer monitoring campaigns in 2021 and 2023. A deep-learning object-detection model\nwas used to classify vehicles in street-level imagery, and land-use regression models were developed\nto generate high-resolution spatial surfaces for both pollutants and vehicle categories.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "84-0", + "page": 80, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ion model\nwas used to classify vehicles in street-level imagery, and land-use regression models were developed\nto generate high-resolution spatial surfaces for both pollutants and vehicle categories.\nResults indicate that UFP and BC hotspots aligned with freight corridors and industrial zones\ndominated by heavy-duty trucks. Pollutant concentrations increased city-wide from 2021 to 2023,\nmirroring growth in freight and delivery activity. Increases in single-unit and combination-truck\nactivity along highways and major arterial roads between 2021 and 2023 drove corresponding rises\nin pollutant concentrations, reflecting post-pandemic growth in freight traffic. Heavy-duty vehicles,\nrather than overall traffic volume, were the dominant contributors to near-road UFP and BC pat-\nterns. These findings demonstrate the value of vision-based traffic characterization for identifying\nemission hotspots and for informing targeted urban air-quality interventions.\n5.2\nIntroduction", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "84-1", + "page": 80, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "at-\nterns. These findings demonstrate the value of vision-based traffic characterization for identifying\nemission hotspots and for informing targeted urban air-quality interventions.\n5.2\nIntroduction\nUrban traffic is a primary source of air pollution, but not all vehicles contribute equally [236, 237].\nHeavy-duty vehicles such as diesel trucks and buses emit disproportionately high levels of BC and\nUFP compared to light-duty vehicles, leading to elevated concentrations along corridors with frequent\ntruck and bus activity [40, 82, 18, 8]. As a result, the composition of traffic plays a critical role in\nshaping the spatial distribution of traffic-related air pollution [83, 238, 84]. While this relationship\nis captured in emission inventories and source apportionment studies [40, 239, 240, 241, 242, 243],\nit remains underexplored in empirical air quality modeling due to the lack of high-resolution traffic\ncomposition data as a model input.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "84-2", + "page": 80, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and source apportionment studies [40, 239, 240, 241, 242, 243],\nit remains underexplored in empirical air quality modeling due to the lack of high-resolution traffic\ncomposition data as a model input. Understanding how vehicle-specific traffic patterns influence\npollutant concentrations is essential for targeted mitigation and exposure reduction in urban areas.\nLUR is widely used for modeling the spatial distribution of air pollutants in urban environ-\nments [11, 165, 109]. By statistically linking observed concentrations to spatially referenced pre-\ndictors, such as road network features, land use characteristics, proximity to emission sources, and\npopulation density, LUR models can generate high-resolution surfaces suitable for exposure [244].\nInitially developed using data from stationary monitoring networks [245], LUR models have increas-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "84-3", + "page": 80, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ingly incorporated mobile monitoring campaigns to overcome the limitations of sparse fixed-site ob-\nservations [165, 170, 206, 246]. Mobile-based LUR enables the capture of fine-scale spatial gradients.\nThis evolution has significantly improved the capacity of LUR to represent the spatial complexity of\nurban air quality [232, 4, 174, 247]. Dense mobile monitoring campaigns equipped with low-cost sen-\nsors have been instrumental in capturing hyperlocal variations of BC, UFP, NO2, and PM2.5 across\ndiverse urban settings. For example, research in cities such as Amsterdam, Copenhagen, Milan,\nSeoul, and Auckland has demonstrated that these models can reliably characterize neighborhood-\nlevel pollution gradients and exposure disparities [29, 248, 249, 250]. Recent methodological ad-\nvancements have further enhanced the capabilities of mobile-based LUR models, enabling more\naccurate and granular representation of urban air quality. Several studies have employed advanced", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "85-0", + "page": 81, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "gical ad-\nvancements have further enhanced the capabilities of mobile-based LUR models, enabling more\naccurate and granular representation of urban air quality. Several studies have employed advanced\nfeature selection, integrated meteorological data, or combined stationary and mobile measurements\nto improve model robustness and spatial precision [165, 251, 252, 127, 167, 205].\nThese efforts\nunderscore the growing utility of LUR frameworks enhanced by mobile monitoring and advanced\nstatistical techniques for high-resolution air pollution mapping in complex urban environments.\nConcentrations of UFP and BC often peak near highways and freight corridors, closely track-\ning heavy-duty vehicle activity [233, 163].\nBeyond these spatial patterns, long-term temporal\nchanges—such as shifts in traffic volume, fleet composition, and policy interventions—can signif-\nicantly reshape pollutant levels [253, 229, 254]. A notable example was the COVID-19 pandemic,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "85-1", + "page": 81, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "mporal\nchanges—such as shifts in traffic volume, fleet composition, and policy interventions—can signif-\nicantly reshape pollutant levels [253, 229, 254]. A notable example was the COVID-19 pandemic,\nwhen widespread reductions in commuter and freight across many cities led to marked declines in\ntraffic-related pollutants [255, 256, 257]. To better capture these evolving dynamics, LUR–based\nmodels can be applied to data collected across multiple years. Comparing these model outputs across\ntime enables researchers to quantify the impact of long-term traffic changes on pollution patterns.\nThis approach can reveal persistent trends such as decreasing emissions from heavy-duty vehicles\nor spatial shifts in air pollution hotspots driven by modified traffic routes. Achieving such insights\nrequires accounting for both pollutant concentrations and the concurrent changes in traffic activity.\nTraditional methods for capturing traffic data such as loop detectors, inductive sensors, or man-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "85-2", + "page": 81, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "requires accounting for both pollutant concentrations and the concurrent changes in traffic activity.\nTraditional methods for capturing traffic data such as loop detectors, inductive sensors, or man-\nual counts often suffer from limited spatial coverage, infrequent updates, and lack of vehicle-type\nresolution [258]. These limitations present challenges when attempting to characterize traffic compo-\nsition at the level of detail required for high-resolution analysis [259]. To address these limitations,\nresearchers have turned to computer vision and street-level imagery as emerging tools for urban\nanalytics [260]. These technologies allow information to be extracted directly from images captured\nin real time, enabling broad spatial coverage and flexible deployment [261]. When combined with\nmobile platforms, such as vehicles equipped with 360-degree cameras, this approach enables the col-\nlection of continuous visual data along road networks [262]. Object detection models trained using", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "85-3", + "page": 81, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "mobile platforms, such as vehicles equipped with 360-degree cameras, this approach enables the col-\nlection of continuous visual data along road networks [262]. Object detection models trained using\ndeep learning techniques can process these images to detect and classify vehicles by type with high\naccuracy [263]. Algorithms such as the YOLO (You Only Look Once) family have demonstrated\nstrong performance in identifying and classifying vehicles including passenger cars and buses and\ntrucks under a wide range of urban conditions. These models can generate spatially explicit traffic\ncomposition datasets across large areas [19].\nIn this study, we investigate how traffic composition and its changes influence urban air pollution\nin Toronto by combining mobile air quality monitoring with image-based traffic sensing. UFP and BC\nconcentrations were measured during the summers of 2021 and 2023 using low-cost sensors mounted\non a mobile platform.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "85-4", + "page": 81, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "bining mobile air quality monitoring with image-based traffic sensing. UFP and BC\nconcentrations were measured during the summers of 2021 and 2023 using low-cost sensors mounted\non a mobile platform.\nSimultaneously, 360-degree images were captured at one-second intervals", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "85-5", + "page": 81, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and processed using a deep learning–based object detection model to quantify vehicle types such\nas passenger cars, single-unit trucks, and combination trucks along road segments. LUR models\nwere then developed to predict the spatial distribution of both air pollutants and vehicle types for\neach year. Comparing generated surfaces for these models allows us to examine temporal shifts\nin traffic-related air pollution, detect persistent or emerging pollution hotspots, and evaluate the\ninfluence of changing fleet composition over time. The framework and findings provide a practical\nbasis for cities to integrate vision-based traffic characterization into air-quality management and to\ntrack the impacts of transportation policies over time.\n5.3\nMethods\n5.3.1\nStudy area and campaign design\nThis study was conducted in Toronto, the most populous city in Canada, characterized by diverse\nland uses, complex traffic patterns, and high vehicle volumes. As a major metropolitan area, Toronto", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "86-0", + "page": 82, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "gn\nThis study was conducted in Toronto, the most populous city in Canada, characterized by diverse\nland uses, complex traffic patterns, and high vehicle volumes. As a major metropolitan area, Toronto\nexperiences high volumes of vehicular traffic, making it a suitable setting for investigating traffic-\nrelated air pollution dynamics. For this study, two mobile monitoring campaigns were carried out\nusing the UrbanScanner platform [196], a vehicle equipped with air pollution sensors and a 360°\ncamera, during spring and summer 2021 (34 days between April and June) and 2023 (30 days\nbetween June to early August). Sampling routes covered highways, arterial roads, and residential\nstreets, with some segments visited multiple times to capture temporal variability and others visited\nonly once. To ensure consistent spatial modeling, the road network was divided into standardized\n100-m segments, and monitoring data were assigned to these segments. In 2023, 14 sampling days", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "86-1", + "page": 82, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "visited\nonly once. To ensure consistent spatial modeling, the road network was divided into standardized\n100-m segments, and monitoring data were assigned to these segments. In 2023, 14 sampling days\ncoincided with wildfire smoke events from Northern Ontario and Quebec; these days were identified\nusing satellite data and reference air quality measurements and excluded from model development to\navoid confounding [244]. After exclusion, 8,419 unique road segments were available in 2021 and 7,584\nin 2023. We aggregated measurements to these segments and computed median concentrations for\nUFP and BC for each segment. An overview of the spatial coverage of the 2021 and 2023 (excluding\nwildfire days) monitoring campaigns, is shown in Figure 5.1 .\n5.3.2\n360-Degree imagery and traffic composition extraction\nTo capture the surrounding traffic environment, a high-resolution 360° camera (V.360°, VSN Mobil)\nwas mounted on the rooftop of the UrbanScanner vehicle (Figure B-1).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "86-2", + "page": 82, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "nd traffic composition extraction\nTo capture the surrounding traffic environment, a high-resolution 360° camera (V.360°, VSN Mobil)\nwas mounted on the rooftop of the UrbanScanner vehicle (Figure B-1). Images were recorded at\n1 Hz and synchronized with pollutant and GPS data, producing 433,088 images in 2021 and 476,623\nin 2023 (234,351 during wildfire days; 242,272 during non-wildfire days). Only images with valid\nGPS coordinates and concurrent UFP measurements were retained for analysis.\nIn this study, object detection was performed using the YOLOv8 architecture, initialized with\npretrained weights and further fine-tuned on a custom-labeled dataset to enhance detection accuracy\nand enable finer classification of vehicle types. The generic ”truck” class from the pretrained model\nwas expanded into four distinct categories. These categories were pickup, light commercial vehicle,\nsingle-unit truck and combination truck (Figure B-2) to better reflect their differing emission profiles.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "86-3", + "page": 82, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "panded into four distinct categories. These categories were pickup, light commercial vehicle,\nsingle-unit truck and combination truck (Figure B-2) to better reflect their differing emission profiles.\nWhile the base YOLOv8 model includes general categories such as person, bicycle, car, motorcycle,\nbus, train, and truck, it lacks the resolution needed to distinguish between truck subtypes.\nTo", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "86-4", + "page": 82, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 5.1:\nSpatial coverage of the UrbanScanner mobile monitoring campaigns conducted in\nToronto. Points represent mobile measurements collected during 34 sampling days in summer 2021\n(red) and during the summer 2023 campaign excluding wildfire days (blue). Star symbols indicate\nthe locations of government reference stations.\naddress this limitation, we manually annotated 1,543 images (Figure B-3) for training and 126\nimages for testing, expanding the label set to include: car, bus, person, motorbike, bicycle, pickup,\nlight commercial vehicle, single-unit truck, combination truck, and tram [19]. This enhancement was\ncritical for differentiating between heavy-duty and light-duty vehicles, which have distinct emissions\nprofiles. The image detection model’s performance was evaluated using mean Average Precision\n(mAP) on the test set.\nAfter training, the number of vehicles in each category was extracted\nfrom all images. Detected counts were aggregated to 100-m road segments.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "87-0", + "page": 83, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "uated using mean Average Precision\n(mAP) on the test set.\nAfter training, the number of vehicles in each category was extracted\nfrom all images. Detected counts were aggregated to 100-m road segments. For each segment, the\nmean number of detected vehicles per class was calculated across all segments within a given year,\nproducing standardized traffic composition indicators compatible with the air pollution data.\n5.3.3\nLand use regression modeling\nWe developed LUR models using XGBoost.\nA comprehensive set of spatial and environmental\npredictors was used to model BC, UFP, and vehicle-type activity derived from 360-degree street-level\nimagery. Predictors were generated using a GIS-based framework and grouped into four primary\ncategories: land-use characteristics, road network features, meteorological conditions, and a year\nindicator. Among the traffic-related predictors, AADT was included to represent traffic intensity.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "87-1", + "page": 83, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "y\ncategories: land-use characteristics, road network features, meteorological conditions, and a year\nindicator. Among the traffic-related predictors, AADT was included to represent traffic intensity.\nAADT data for 2020 were obtained from previously developed city-wide traffic volume models for\nToronto, including an image-enhanced product that integrates municipal traffic counts with vehicle\ndetection from aerial imagery to improve spatial coverage [182, 162]. These values were spatially\nmatched to the 100 m road segments used in the LUR framework. AADT was used exclusively as a\npredictor in the pollutant LUR models and served as a benchmark for comparison with image-derived\ntraffic prediction models. An overview of all predictors is provided in Table B-1.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "87-2", + "page": 83, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "To enhance model robustness, we combined XGBoost with forward feature selection and STCV,\nfollowing the methodology established in our previous work [264]. This strategy was designed to\nreduce overfitting and improve generalizability by explicitly accounting for both spatial and temporal\nautocorrelation inherent in mobile monitoring data.\nThe full dataset was randomly partitioned\ninto an 80% training set and a 20% holdout test set.\nAll model development steps, including\nhyperparameter tuning and feature selection, were performed exclusively on the training data. Final\nmodel performance was then evaluated using the holdout test set to assess generalization.\nWe developed LUR models using XGBoost. A comprehensive set of spatial and environmental\npredictors was used to develop LUR models for BC, UFP concentrations, and vehicle-type-specific\ntraffic counts derived from 360-degree imagery. These predictors were generated using a GIS-based", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "88-0", + "page": 84, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ental\npredictors was used to develop LUR models for BC, UFP concentrations, and vehicle-type-specific\ntraffic counts derived from 360-degree imagery. These predictors were generated using a GIS-based\nframework and grouped into four primary categories: land use characteristics, road network features,\nmeteorological conditions, and a temporal year indicator. An overview of all predictors is provided\nin Table B-1. To enhance model robustness, we combined XGBoost with FFS and STCV, following\nthe methodology established in our previous work. This strategy was selected to reduce overfitting\nand improve generalizability by accounting for both spatial and temporal autocorrelation inherent in\nmobile monitoring data. The full dataset was randomly partitioned into an 80% training set and a\n20% holdout test set. All model development including hyperparameter tuning and feature selection\nwas performed exclusively on the training data. Final model performance was then evaluated on", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "88-1", + "page": 84, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "t and a\n20% holdout test set. All model development including hyperparameter tuning and feature selection\nwas performed exclusively on the training data. Final model performance was then evaluated on\nthe hold-out test set to assess generalization.\nThe modeling process included the following steps:\n• Hyperparameter Tuning: Model hyperparameters (e.g., learning rate, tree depth, regu-\nlarization parameters) were optimized using Bayesian optimization within a STCV scheme.\nSampling days were used to define temporal folds, and spatial clustering was applied within\neach fold. A 300 m dead buffer zone was implemented between training and validation data\nto minimize spatial leakage.\n• Forward Feature Selection: After tuning, FFS was performed under the same STCV\nsetup.\nPredictors were incrementally added, and only those improving the cross-validated\nmean squared error were retained. The process continued until no further improvement in\nmodel performance was observed.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "88-2", + "page": 84, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "redictors were incrementally added, and only those improving the cross-validated\nmean squared error were retained. The process continued until no further improvement in\nmodel performance was observed. After choosing features, hyperparameters again were opti-\nmized under STCV.\n• Final Model Training: The final model was retrained on the full training set using the\nselected predictors and re-optimized hyperparameters, and its performance was evaluated on\nthe test set.\nAll models were developed using FFS combined with STCV, ensuring that predictor choice and\nhyperparameter tuning were guided by data splits across both space and time.\nThis approach\nemphasizes predictive generalizability rather than maximizing fit to localized episodic events.\nTo accommodate interannual variability, a dummy variable representing the year of sampling\n(2021 or 2023) was included in all models. This allowed the model to distinguish between structural", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "88-3", + "page": 84, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "events.\nTo accommodate interannual variability, a dummy variable representing the year of sampling\n(2021 or 2023) was included in all models. This allowed the model to distinguish between structural\ndifferences in pollution and traffic patterns across the two campaigns while enabling prediction over\nthe pooled dataset. Separate LUR models were developed for each of the following target variables:\nUFP, BC, passenger cars, pickups, light commercial vehicles, single-unit trucks, and combination", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "88-4", + "page": 84, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "trucks. This multi-target modeling approach enabled a comprehensive analysis of both pollutant\ndistributions and the spatial patterns of traffic composition across the urban landscape.\n5.3.4\nModel evaluation\nTo assess the performance and generalizability of each LUR model, we reserved a random 20%\nholdout set from the aggregated dataset for internal validation.\nThis subset was not used dur-\ning training, hyperparameter tuning, or feature selection, ensuring a truly independent evaluation.\nModel robustness were evaluated using three standard metrics: coefficient of determination (R2),\nroot mean square error (RMSE) and mean absolute error (MAE). These metrics were consistently\napplied across all pollutant and traffic-related models. In addition, for truly evaluating the UFP\nmodel, we used continuous measurements from 17 stationary backyard sensors (12 DiscMini, 5 UFP\nPartector) deployed between July 9 and 30, 2021 (Figure 5.1). DiscMini units recorded at 10-second", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "89-0", + "page": 85, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "e UFP\nmodel, we used continuous measurements from 17 stationary backyard sensors (12 DiscMini, 5 UFP\nPartector) deployed between July 9 and 30, 2021 (Figure 5.1). DiscMini units recorded at 10-second\nintervals and Partector units at 1-second intervals; only data collected between 08:00 and 19:00 were\nretained to match the mobile campaign period. This allowed us to evaluate the spatial accuracy\nof predicted concentrations against independent continuous observations. The absolute percentage\nerror (APE) was calculated for each sensor site to quantify discrepancies between observed and\npredicted average UFP concentrations over the campaign period (see SI for details).\n5.3.5\nSpatial surface generation and comparison\nThe final LUR models were trained for UFP, BC and vehicle types such as cars, pickup trucks, light\ncommercial vehicles, single-unit trucks and combination trucks. From these models we generated", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "89-1", + "page": 85, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "rison\nThe final LUR models were trained for UFP, BC and vehicle types such as cars, pickup trucks, light\ncommercial vehicles, single-unit trucks and combination trucks. From these models we generated\nhigh-resolution spatial surfaces to visualize and compare spatial patterns across years. For traffic-\nrelated models, predictions were made at 50-meter spaced points along all roads across Toronto\nby applying the final XGBoost models with the optimized predictor set. These predictions were\ngenerated for both 2021 and 2023, enabling detailed comparisons of vehicle-type-specific traffic\nintensity across these years.\nFor air pollution models, we generated a 50-meter resolution UFP and BC prediction surface\nby applying the trained models to a regularly spaced grid. To interpret the role of each predictor\nin driving the model outputs, SHAP plot were created for all final models. SHAP summary plots", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "89-2", + "page": 85, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "urface\nby applying the trained models to a regularly spaced grid. To interpret the role of each predictor\nin driving the model outputs, SHAP plot were created for all final models. SHAP summary plots\nhighlight the direction and magnitude of each feature’s influence, offering insight into the factors\ncontributing to pollution and traffic patterns across space. To evaluate temporal change, surfaces\nfrom 2021 and 2023 were compared at the road segment level (for traffic) and grid cell level (for\npollutants).\nDifference maps were generated to visualize interannual trends and highlight areas\nof significant change in vehicle activity or pollution hotspots. In fact, we integrated SHAP-based\ninterpretation with surface generation and interannual comparison. Together these elements offer\na robust framework for understanding how traffic composition and pollution levels vary over space\nand time.\nFor traffic-related models including those developed for cars, pickup trucks, light commercial", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "89-3", + "page": 85, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "obust framework for understanding how traffic composition and pollution levels vary over space\nand time.\nFor traffic-related models including those developed for cars, pickup trucks, light commercial\nvehicles, single-unit trucks, and combination trucks, we aggregated the predicted number of vehicles\nacross all classes for each road segment.\nWe then examined how prediction from image-derived\ntraffic presence relates to—and differs from—flow-based AADT. We compared model predictions\nwith AADT using three groupings of vehicles: (i) all vehicles, (ii) all vehicles except personal cars,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "89-4", + "page": 85, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and (iii) freight and service vehicles only (light commercial, single-unit, and combination trucks). For\neach grouping, we assessed the relationship with AADT using Spearman and Pearson correlations,\nand calculated R2 values after min–max normalizing both the LUR-based counts and AADT. To\ncharacterize trend differences, we produced segment-level difference maps (normalized LUR-based\nfreight and service vehicles minus normalized AADT). Because AADT measures throughput over a\nfull day, whereas image-based detections are sensitive to dwell time and parked/slow-moving vehicles,\nthese comparisons were designed to assess complementarity and divergence rather than to validate\none against the other.\n5.4\nResults\n5.4.1\nData summary\nMean of all records for UFP concentration increased from ∼23,900 particles/cm3 in 2021 to ∼27,300\nparticles/cm3 on non-wildfire days in 2023, while mean of records for BC rose from ∼1,429 to\n∼1,543 µg/m3.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "90-0", + "page": 86, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "of all records for UFP concentration increased from ∼23,900 particles/cm3 in 2021 to ∼27,300\nparticles/cm3 on non-wildfire days in 2023, while mean of records for BC rose from ∼1,429 to\n∼1,543 µg/m3. Wildfire-affected days in 2023 showed substantially lower UFP (∼14,200 particles/\ncm3) and slightly higher BC (∼1,580 µg/m3) concentrations. Table B-2 summarizes UFP and BC\nconcentrations by year and wildfire status, reporting the number of segments for median based ag-\ngregation, number of records, median, mean, and standard deviation values for all records. To ensure\nmodels reflected typical urban conditions, all wildfire days were excluded from further analyses. We\nused UFP concentrations from backyard sensors to truly evaluate model performance (summary in\nTable B-3), and these measurements were generally lower than those from mobile monitoring.\nThe YOLOv8 model, fine-tuned on 1,543 annotated training images and evaluated on 126 test", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "90-1", + "page": 86, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "rmance (summary in\nTable B-3), and these measurements were generally lower than those from mobile monitoring.\nThe YOLOv8 model, fine-tuned on 1,543 annotated training images and evaluated on 126 test\nimages, achieved an overall precision of 0.63, recall of 0.51, mean average precision at IoU=0.5\n(mAP@0.5) of 0.55, and mAP@[0.5:0.95] of 0.37. These results indicate moderate detection accuracy,\nwhich is sufficient for our application since the counts are intended to represent the surrounding\ntraffic composition rather than provide exact number of vehicle in images. The normalized confusion\nmatrix is provided in the Supporting Information (Figure B-4). Vehicle counts were extracted from\nall 360° imagery recorded in 2021 and 2023 using the trained YOLO model. Mean counts of detected\npassenger cars and pickups per images were slightly lower in 2023 non-wildfire days compared to\n2021 (cars: 5.75 vs. 5.95; pickups: 0.25 vs. 0.29), while light commercial vehicles, single-unit trucks,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "90-2", + "page": 86, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "assenger cars and pickups per images were slightly lower in 2023 non-wildfire days compared to\n2021 (cars: 5.75 vs. 5.95; pickups: 0.25 vs. 0.29), while light commercial vehicles, single-unit trucks,\nand combination trucks all showed modest increases (light commercial: 0.23 vs. 0.17; single-unit\ntrucks: 0.27 vs. 0.20; combination trucks: 0.10 vs. 0.08) (Details in Table B-4).\n5.4.2\nModel performance\nTable A-4 reports the R2, RMSE, and MAE for both training and 20% holdout test sets across pol-\nlutants and vehicle classes. The results show similar performance between the training and holdout\nsets for both pollutants and vehicle categories, suggesting limited overfitting. For pollutants, both\nBC and UFP models achieved modest R2 values, reflecting the high variability in mobile monitor-\ning data. These metrics mainly serve as an internal consistency check and should be interpreted\ncautiously as indicators of potential overfitting rather than true predictive performance.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "90-3", + "page": 86, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ile monitor-\ning data. These metrics mainly serve as an internal consistency check and should be interpreted\ncautiously as indicators of potential overfitting rather than true predictive performance.\nFigure 5.2 shows scatter plots of observed versus predicted values for the holdout set. The models\ncapture overall patterns but smooth out extreme highs and lows common in mobile monitoring.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "90-4", + "page": 86, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Table 5.1: Performance of Models on Training and 20% Random Holdout Test Data.\nR2 (Train)\nR2 (Test)\nModel Name\\Metric\nRMSE (Train)\nMAE (Train)\nRMSE (Test)\nMAE (Test)\nBC\n0.340\n795.839\n559.452\n0.343\n812.149\n565.597\nUFP\n0.284\n31888.747\n14461.235\n0.292\n30611.937\n13972.312\nCar\n0.533\n2.090\n1.572\n0.521\n2.104\n1.571\nCombination trucks\n0.678\n0.238\n0.115\n0.667\n0.246\n0.118\nSingle-unit trucks\n0.429\n0.309\n0.201\n0.424\n0.316\n0.202\nLight commercial\n0.128\n0.216\n0.146\n0.124\n0.218\n0.146\nPickups\n0.180\n0.159\n0.116\n0.172\n0.162\n0.117\nPredictions cluster around average ranges, reflecting the focus on stable, generalizable patterns\nrather than short-lived fluctuations. As a result, R2 values appear modest, but this aligns with the\ngoal of producing robust spatial surfaces rather than reproducing transient noise.\nFigure 5.2: Scatter plots of observed versus predicted values for the holdout test set: (a) UFP, (b)", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "91-0", + "page": 87, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "aligns with the\ngoal of producing robust spatial surfaces rather than reproducing transient noise.\nFigure 5.2: Scatter plots of observed versus predicted values for the holdout test set: (a) UFP, (b)\nBC, (c) personal cars, (d) combination trucks, (e) single-unit trucks, (f) light commercial vehicles,\nand (g) pickups.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "91-1", + "page": 87, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "To independently evaluate model performance, continuous UFP measurements from backyard\nstationary sensors were used as an external reference. Unlike the mobile monitoring data, which\nprovides episodic coverage, these stationary sensors record continuous concentrations, allowing for a\nrobust comparison of modeled versus observed values. APE across the monitoring network ranged\nfrom ∼25% to ∼257%, with most units falling between ∼60–140%.\nThe mean APE across all\nstationary\n5.4.3\nFeature importance and model interpretation\nWe began with a large number of predictors, but after applying FFS combined with STCV, each\nLUR model retained only a limited subset (with the number of features and their importance ranked\nby SHAP plots, Figure B-5, to aid interpretation). For both pollutants, traffic-correlated predictors\nsuch as road area, highway length, and AADT consistently emerged as dominant drivers, while open\nand water areas showed negative associations consistent with dispersion effects.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "92-0", + "page": 88, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "-correlated predictors\nsuch as road area, highway length, and AADT consistently emerged as dominant drivers, while open\nand water areas showed negative associations consistent with dispersion effects. Passenger cars were\nprimarily linked to residential and commercial features, including restaurants and accommodations,\nwhereas pickup and light commercial vehicles were associated with local roads, parking areas, and\nbus stops, reflecting their role in mixed-use and delivery traffic. Heavy-duty vehicles, particularly\nsingle-unit and combination trucks, were more strongly tied to major corridors, highways, and\nindustrial-related predictors.\n5.4.4\nSpatial prediction surfaces\nThe predicted concentration surfaces for pollutants (Figure 5.3) highlight distinct spatial patterns\nin both 2021 and 2023. For UFP, concentrations were highest along highways in both years, with\nelevated values extending to major arterials. Levels decreased sharply in proximity to parks and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "92-1", + "page": 88, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "atterns\nin both 2021 and 2023. For UFP, concentrations were highest along highways in both years, with\nelevated values extending to major arterials. Levels decreased sharply in proximity to parks and\ngreen areas, reflecting localized dilution and lower traffic intensity. BC displayed a similar highway-\ndominated pattern across both years, with hotspots concentrated along major transportation corri-\ndors, consistent with diesel-intensive traffic activity.\nVehicle-related spatial surfaces revealed distinct patterns across categories (Figure 5.4 for 2021;\nFigure ?? for 2023). Personal cars were most concentrated in downtown, residential neighborhoods,\nand parking areas, indicating their role in localized traffic and short-distance trips. Combination\ntrucks displayed their strongest presence along highways, with moderate activity in industrial zones,\nbut near-zero values across most other areas. Single-unit trucks followed a similar pattern, with", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "92-2", + "page": 88, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "trucks displayed their strongest presence along highways, with moderate activity in industrial zones,\nbut near-zero values across most other areas. Single-unit trucks followed a similar pattern, with\nconcentrations highest on highways and major roads. Light commercial vehicles were more prominent\nin downtown, reflecting service-related traffic linked to commercial activity. Finally, pickups were\nmore evenly distributed across the city, showing less spatial variation compared to other vehicle\ntypes.\n5.4.5\nTemporal comparison of spatial surfaces\nWe compared 2023 and 2021 prediction surfaces on non-wildfire days (difference = 2023 minus 2021)\nto assess how pollution and traffic changed over time (Figure 5.5). UFP and BC both increased\nacross much of Toronto, with the largest positive changes tracking transportation corridors. The\nhighest UFP increases were concentrated on highways, whereas the largest BC increases occurred", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "92-3", + "page": 88, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 5.3: Predicted spatial concentration surfaces for UFP and BC across Toronto for 2021 and\n2023. (a) UFP (2021), (b) UFP (2023), (c) BC (2021), (d) BC (2023).\nFigure 5.4: Predicted spatial surfaces of vehicle categories across Toronto for 2021: (a) personal\ncars, (b) combination trucks, (c) single-unit trucks, (d) light commercial vehicles, and (e) pickups.\nCorresponding surfaces for 2023 are shown in Figure S6 (SI).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "93-0", + "page": 89, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "along major arterial roads. While concentrations rose across most of the city, some spots such as\nthe downtown core decreased.\nFigure 5.5: Temporal comparison of predicted concentrations between 2021 and 2023. Difference\nmaps (2023 minus 2021) show spatial shifts in pollutant (a) UFP, (b) BC\nVehicle activity also displayed distinct temporal shifts (Figure 5.6). Personal cars increased along\nhighways and within the downtown core, but declined across many residential corridors. Combina-\ntion trucks showed the growth on highways and near industrial corridors. Single-unit trucks rose\nalong major arterials, with modest downtown decreases. Light commercial vehicles increased notably\nin the downtown and around highway access points. Pickups show slight increases near highways\nand downtown, with decreases across many other segments.\n5.4.6\nTraffic trends alignment with AADT\nWe compared image-derived traffic predictions with AADT using three groupings: (i) all vehicles,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "94-0", + "page": 90, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ways\nand downtown, with decreases across many other segments.\n5.4.6\nTraffic trends alignment with AADT\nWe compared image-derived traffic predictions with AADT using three groupings: (i) all vehicles,\n(ii) all except personal cars, and (iii) freight and service vehicles only (Table 5.2).\nAgreement\nwith AADT was weak when all vehicles were included, particularly in 2021. Excluding personal\ncars improved alignment substantially (Pearson r ≈0.79–0.81), with further gains when focusing\nonly on freight and service vehicles (R2 ≈0.48–0.52). These results indicate that heavy-duty traffic\ncaptured in imagery provides a stronger proxy for actual vehicle throughput and is especially relevant\nfor emissions analysis.\nAlignment of image-detected traffic with AADT (correlations and R2; AADT and\nTable 5.2:\ndetections normalized for R2).\nR2\nCondition\nYear\nSpearman\nPearson\n−4.745\nAll traffic\n2021\n0.144\n0.325\n−1.648\nAll traffic\n2023\n0.397\n0.565\nNo personal cars\n2021\n0.608\n0.791\n0.375\nNo personal cars", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "94-1", + "page": 90, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "5.2:\ndetections normalized for R2).\nR2\nCondition\nYear\nSpearman\nPearson\n−4.745\nAll traffic\n2021\n0.144\n0.325\n−1.648\nAll traffic\n2023\n0.397\n0.565\nNo personal cars\n2021\n0.608\n0.791\n0.375\nNo personal cars\n2023\n0.643\n0.809\n0.397\nNo personal cars + no pickups\n2021\n0.642\n0.794\n0.479\nNo personal cars + no pickups\n2023\n0.660\n0.811\n0.516\nIn addition to statistical comparisons, we generated a 2023 difference map (Figure 5.7) comparing\nthe spatial patterns of min–max–normalized image-based traffic predictions (freight and service\nvehicles) and min–max–normalized AADT; the map visualizes pattern divergence between the two\nnormalized surfaces.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "94-2", + "page": 90, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 5.6: Temporal comparison of predicted vehicle activity between 2021 and 2023. Difference\nmaps (2023 minus2021) show spatial shifts in traffic patterns: (a) personal cars, (b) combination\ntrucks, (c) single-unit trucks, (d) light commercial vehicles, and (e) pickups.\n5.5\nDiscussion\n5.5.1\nModel validation and interpretation of traffic representation\nIn this study, we developed LUR models for both pollutants (UFP, BC) and vehicle types derived\nfrom street-level imagery, pooling 2021 and 2023 (non-wildfire days). The results of the models were\nused to quantify how traffic composition relates to pollutants’ concentrations and to describe the spa-\ntial patterns and between-year changes across Toronto. We applied FFS within STCV framework and\ntuned hyperparameters accordingly to mitigate overfitting and enhance generalizability [14, 15, 228,\n13]. This method removes spurious features and chooses generalizable features and hyperparameters.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "95-0", + "page": 91, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tuned hyperparameters accordingly to mitigate overfitting and enhance generalizability [14, 15, 228,\n13]. This method removes spurious features and chooses generalizable features and hyperparameters.\nWe took several steps to validate the results, including internal holdout testing, comparisons with\nindependent stationary measurements, and visual inspection of the predicted concentration maps", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "95-1", + "page": 91, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure 5.7: Spatial difference map showing the mismatch between normalized image-based model\npredictions (freight and service vehicles only) and normalized AADT for 2023. Because these two\nmeasures represent different traffic characteristics, the map should be interpreted as a difference in\nspatial patterns rather than a direct comparison. Positive values (red segments) indicate segments\nwhere the image-based surface exceeds the AADT pattern, while negative values (blue segments)\nindicate the opposite.\nto ensure realistic spatial patterns. The modeling framework helped ensure stable performance, as\nreflected in the similar metrics (R2, RMSE, and MAE) across training and test sets for our LUR\nmodels. Previous studies have reported internal R2 values for BC and UFP ranging from 0.15 to\n0.7, reflecting variability across models and settings [8, 4, 174, 29, 86]. However, internal evaluation\nmetrics should not be viewed as definitive indicators of real-world performance. Mobile monitoring", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "96-0", + "page": 92, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "reflecting variability across models and settings [8, 4, 174, 29, 86]. However, internal evaluation\nmetrics should not be viewed as definitive indicators of real-world performance. Mobile monitoring\ndata, by nature, is collected over short periods and can be highly variable in time and space. These\ntransient conditions, along with autocorrelation, mean that internal splits often reuse similar routes\nand contexts, potentially overstating model performance [228, 149, 265]. In our study, R2 values\nfor BC and UFP were moderate, consistent with the goal of modeling broader spatial trends under\ntransient conditions inherent to mobile monitoring. Our moderate R2 values are similar to earlier\nstudies, which found that mobile models can still capture longer-term trends even if internal perfor-\nmance is low [166, 266]. To truly validate the power of LUR models using mobile data, independent\ncontinuous data is needed [165, 128]. For external validation, we compared predicted UFP values", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "96-1", + "page": 92, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "rfor-\nmance is low [166, 266]. To truly validate the power of LUR models using mobile data, independent\ncontinuous data is needed [165, 128]. For external validation, we compared predicted UFP values\nto 2021 stationary backyard sensor data. Despite partial alignment, most predictions exceeded ob-\nserved values, leading to a 96% mean APE. This mismatch is expected given that mobile monitoring\nwas conducted on roadways, whereas the stationary monitors were located in residential backyards.\nPrior studies have consistently shown that UFP concentrations decline steeply with distance from\nthe source, often dropping by 50–70% within the first 100–300 m and approaching background levels\nbeyond 300 m [267, 101, 5]. Therefore, differences between predictions and backyard observations are\nconsistent with established near-road decay patterns of UFP. We also conducted visual inspections", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "96-2", + "page": 92, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "of the predicted pollution maps for each year. This step is particularly important when working\nwith autocorrelated data, where misleading predictors or overfitting can produce unrealistic spatial\npatterns or artifacts. As emphasized by Meyer et al. (2019) [15], visual assessment complements\nstatistical validation by revealing issues such as linear artifacts or overfitting to geolocation vari-\nables that numerical metrics alone may overlook. The use of STCV for both feature selection and\nmodel tuning allowed us to filter out spurious and overly localized predictors, ultimately improving\nmodel robustness. This approach helped ensure that only stable and relevant features were retained,\nreducing the likelihood of artifacts or unrealistic spatial patterns in the predicted maps. Our sur-\nface maps (Figure 5.3) revealed consistent and interpretable spatial gradients, further supporting\nthe generalizability of the trained models. By comparison, several prior LUR and hybrid modeling", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "97-0", + "page": 93, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ur-\nface maps (Figure 5.3) revealed consistent and interpretable spatial gradients, further supporting\nthe generalizability of the trained models. By comparison, several prior LUR and hybrid modeling\nstudies—despite producing generally strong statistical performance—show visual artifacts in their\npublished output surfaces, including unrealistic spatial discontinuities or grid patterns that likely\nstem from overfitting to localized variables. Examples include Xu et al. (2021) [184], Lloyd et al.\n(2023) [86], Shairsingh et al. (2020) [268], and Yuan et al. (2023) [165], among others.\nWe summed predicted vehicle classes and compared them with AADT to explore where these\ntwo align and where they represent distinct traffic dynamics. First, we added up all vehicle types.\nHowever, the correlation with AADT was weaker than expected. This was likely due to two main\nissues: parked vehicles being counted by the camera, and traffic speed differences. For example,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "97-1", + "page": 93, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "le types.\nHowever, the correlation with AADT was weaker than expected. This was likely due to two main\nissues: parked vehicles being counted by the camera, and traffic speed differences. For example,\ntwo roads might show the same number of vehicles in images, but one may have parked or slow-\nmoving cars while the other has fast-moving traffic that contributes more to AADT. These issues\nmade the comparison with AADT less accurate when all vehicles were included. However, when we\nfocused on heavier vehicle classes – excluding personal cars, and especially excluding both personal\ncars and pickups – the alignment improved markedly. For heavy-duty/service vehicles only, Pearson\ncorrelations rose to 0.80 and R2 increased to 0.5 (2021–2023). This shows that our model’s truck\nand commercial vehicle predictions matched the spatial distribution of actual traffic counts very\nwell (Table 5.2). This robust agreement for freight-oriented traffic builds confidence that the model", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "97-2", + "page": 93, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and commercial vehicle predictions matched the spatial distribution of actual traffic counts very\nwell (Table 5.2). This robust agreement for freight-oriented traffic builds confidence that the model\naccurately captures “where the trucks are,” which is critical since those are the vehicles most linked\nto emissions.\nUsing vehicle counts from street-level imagery may offer a more meaningful representation of\npollution exposure than traditional AADT metrics.\nThis is because slower vehicles—especially\ntrucks—spend more time within each road segment, emitting pollutants over a longer duration.\nAADT, by contrast, only reflects the average number of vehicles passing a point per day and does\nnot account for how long those vehicles remain in the area. As a result, segments with slow-moving\nor idling trucks may have lower AADT but still produce more localized pollution. This difference\nbecomes especially apparent when comparing the spatial patterns of normalized model predictions", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "97-3", + "page": 93, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "moving\nor idling trucks may have lower AADT but still produce more localized pollution. This difference\nbecomes especially apparent when comparing the spatial patterns of normalized model predictions\nfor trucks and commercial vehicles with normalized AADT values (Figure 5.7). In our difference\nmaps, roads with high truck presence but low AADT often corresponded to slower traffic conditions,\nsuch as industrial zones or downtown corridors with congestion.\nThese findings underscore the\nlimitations of relying solely on AADT for characterizing traffic–pollution relationships and highlight\nthe value of computer vision–based traffic indicators as a complementary metric for urban air quality\nstudies.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "97-4", + "page": 93, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "5.5.2\nPrincipal findings\nThis study shows that LUR modeling for pollutants and traffic composition using mobile monitoring\nand image-based traffic data can effectively capture the link between traffic composition and air\npollution.\nWe observed clear spatial patterns, with UFP and BC levels consistently higher on\nroads dominated by heavy-duty vehicles, and lower in areas with mostly light-duty cars. In this\ncontext, traffic composition refers to the relative presence and distribution of different vehicle types\nacross road segments, rather than simply the total number of vehicles. This indicates that pollution\nintensity is more strongly influenced by the presence of heavy-duty vehicles than by total traffic\nvolume alone. These findings align with well-established emission patterns: heavy-duty diesel trucks\nemit significantly more BC and UFP than gasoline-powered cars [267, 235].\nPrior studies also\nsupport this relationship; Peters et al.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "98-0", + "page": 94, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "with well-established emission patterns: heavy-duty diesel trucks\nemit significantly more BC and UFP than gasoline-powered cars [267, 235].\nPrior studies also\nsupport this relationship; Peters et al. (2014) [101], found elevated pollutant levels on roads with\nhigh truck shares, and Pattinson et al. (2014) [233] observed spikes in UFP and BC near highways\nduring truck-heavy periods. Our results confirm that truck traffic is a key contributor to urban air\npollution.\nUFP and BC levels were higher in 2023 compared to 2021 (Figure 5.5), with increases most\npronounced on highways and major arterials, in line with increased traffic activity. Between 2021\nand 2023, traffic patterns in Toronto changed significantly (Figure 5.6). Passenger car counts declined\non many residential streets but rose downtown and along highways, reflecting shifts in commuting\nbehavior during the post-pandemic recovery. Freight and service vehicles, particularly combination", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "98-1", + "page": 94, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ned\non many residential streets but rose downtown and along highways, reflecting shifts in commuting\nbehavior during the post-pandemic recovery. Freight and service vehicles, particularly combination\nand single-unit trucks, increased noticeably on highways, major arterials, and near industrial areas\nin 2023. Notably, more combination trucks were observed on highways, while single-unit trucks\nsurged on urban major arterials. Meanwhile, downtown saw a decline in single-unit truck activity\ndespite an increase in personal and light commercial vehicles. These traffic changes directly aligned\nwith observed pollution trends. The difference in where UFP and BC increased the most reflects the\nvehicle types driving each pollutant. UFP rose mainly on highways due to increased overall traffic,\nas it’s emitted by both light- and heavy-duty vehicles. In contrast, BC increased most on major\narterials, where single-unit truck activity grew sharply. The slight drop in UFP and BC downtown,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "98-2", + "page": 94, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ffic,\nas it’s emitted by both light- and heavy-duty vehicles. In contrast, BC increased most on major\narterials, where single-unit truck activity grew sharply. The slight drop in UFP and BC downtown,\ndespite more gasoline cars in 2023, can be attributed to fewer single-unit trucks, which are major\ncontributors to both pollutants. Similar observations were made by Xu et al. (2022) [19], where the\nspatial distribution of elevated UFP short-term spikes followed the distribution of single-unit trucks.\nThe results of our study confirm that the presence and spatial distribution of heavy-duty vehicles\nplays a more dominant role than total traffic volume in shaping urban air quality and highlight a\ntemporal transition from pandemic-suppressed travel in 2021 to freight-driven activity in 2023, likely\nreflecting economic rebound and increased demand for delivery services.\nThis study shows that LUR modeling for pollutants and traffic composition using mobile moni-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "98-3", + "page": 94, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ht-driven activity in 2023, likely\nreflecting economic rebound and increased demand for delivery services.\nThis study shows that LUR modeling for pollutants and traffic composition using mobile moni-\ntoring and image-based traffic data can effectively capture the link between traffic and air pollution.\nWe observed clear spatial patterns, with UFP and BC levels consistently higher on roads dominated\nby heavy-duty vehicles, and lower in areas with mostly light-duty cars. This indicates that pollution\nintensity is more strongly influenced by the presence of heavy-duty vehicles than by total traffic\nvolume alone. These findings align with well-established emission patterns: heavy-duty diesel trucks\nemit significantly more BC and UFP than gasoline-powered cars [267, 235].", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "98-4", + "page": 94, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "5.5.3\nImplications & future work\nOur findings have both methodological and policy implications. The disproportionate impact of\nheavy-duty vehicles on BC and UFP concentrations suggests that targeted traffic policies, such as\nrestricting trucks from residential areas, promoting cleaner freight technologies, or implementing\nlow-emission zones, could significantly improve urban air quality. From a modeling perspective,\nthis study illustrates the value of combining mobile air quality monitoring with image-based traffic\ndetection. This fusion enabled us to produce high-resolution maps of both traffic composition and\npollution, capturing fine-grained spatial variability that conventional traffic counts often miss. The\nstrong agreement between image-derived truck counts and AADT data further supports the utility\nof computer vision as a low-cost alternative to fixed traffic sensors.\nBuilding on this framework, future work could expand the temporal scope by deploying year-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "99-0", + "page": 95, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ADT data further supports the utility\nof computer vision as a low-cost alternative to fixed traffic sensors.\nBuilding on this framework, future work could expand the temporal scope by deploying year-\nround mobile campaigns or repeating annually to capture seasonal trends and long-term shifts (such\nas the effects of rising electric vehicle adoption). Additionally, incorporating pollutants like PM2.5\nor NOx would offer a broader view of traffic-related impacts. Improvements to the image-processing\nalgorithm, including motion-based filtering or video-based tracking, could help distinguish moving\nvehicles from parked ones and improve light-duty flow estimates while including impact of traffic\nspeed. Integrating stationary and mobile monitoring, whether through calibration or data fusion,\ncan improve model accuracy while preserving spatial detail. Beyond these technical refinements,\nthis approach can be extended to other cities to test generalizability, especially in the context of", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "99-1", + "page": 95, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "n,\ncan improve model accuracy while preserving spatial detail. Beyond these technical refinements,\nthis approach can be extended to other cities to test generalizability, especially in the context of\nnew infrastructure or policy changes. We also see strong potential for using image-extracted traffic\nfeatures as predictors in LUR models, developing short-term prediction frameworks, and ultimately\nreplacing costly traffic sensors. This approach could ultimately replace expensive traffic and pollution\nsensors by leveraging street-level imagery to estimate traffic flow and pollutant concentrations based\non observed traffic composition and background pollution levels.\n5.5.4\nLimitations\nDespite the strengths, several limitations should be noted. Seasonal mismatches across campaigns\nmay introduce biases in cross-year comparisons. Our results reflect summer, daytime campaigns\nin two years only, so they do not represent winter/school-year or nighttime conditions. The traffic", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "99-2", + "page": 95, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "aigns\nmay introduce biases in cross-year comparisons. Our results reflect summer, daytime campaigns\nin two years only, so they do not represent winter/school-year or nighttime conditions. The traffic\ndetection via imagery, while novel and informative, has its shortcomings. The most notable is the\ndifficulty in distinguishing truly moving traffic volume from static or incidental vehicles in the images.\nAs discussed, our method likely over-counts vehicles in areas with parking, since a car lingering in\nview gets counted similarly to one in motion.\n5.6\nConclusion\nThis study demonstrates the power of combining mobile monitoring and image-based traffic detection\nto better understand the role of traffic composition in shaping urban air pollution. By integrating\ndetailed vehicle classification into LUR models, we captured strong spatial patterns in UFP and\nBC, with clear links to the presence of heavy-duty and single-unit trucks—key contributors to emis-\nsions.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "99-3", + "page": 95, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tailed vehicle classification into LUR models, we captured strong spatial patterns in UFP and\nBC, with clear links to the presence of heavy-duty and single-unit trucks—key contributors to emis-\nsions. Temporal comparisons revealed how shifts in traffic, particularly increases in truck activity\non highways and major arterials between 2021 and 2023, drove corresponding increases in pollutant", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "99-4", + "page": 95, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "levels. Meanwhile, the slight downtown improvements underscore the air quality benefits of reduced\nsingle-unit truck presence, despite rising light-duty car counts. Overall, this work highlights the im-\nportance of monitoring and managing freight movement in urban environments and supports future\nefforts to integrate image-derived traffic features into scalable and cost-effective pollution prediction\nframeworks. Beyond pollution prediction, the results also offer practical insights for traffic man-\nagement. Since image-derived metrics reveal where and when freight activity intensifies, they can\ninform strategies for regulating truck volumes, rerouting freight, or optimizing delivery schedules\nto reduce exposure in vulnerable areas. Overall, this work supports the use of image-based traffic\ndata not only for pollution modeling but also for designing smarter, more targeted traffic and urban\nplanning interventions.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "100-0", + "page": 96, + "document_type": "thesis", + "page_header": "CHAPTER 5. POST-COVID TRAFFIC COMPOSITION: WHY ARE ...", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Chapter 6\nSummary and conclusion\n6.1\nSummary of chapters\nThis dissertation presents three integrated studies aimed at advancing high-resolution mapping of\nTRAP in complex urban environments by developing, testing, and interpreting spatial modeling\nworkflows that translate mobile monitoring measurements into coherent, city-wide concentration\nsurfaces. Leveraging mobile monitoring platforms, emerging sensing technologies, and rigorous model\ndevelopments, this work addresses the challenge of capturing fine-scale spatial variability while\nensuring that predictive models generalize beyond observed data. Particular emphasis is placed on\nPM2.5 exposure mapping using mobile sensors installed on courier trucks (Chapter 3), alongside\ncomplementary analyses of traffic-related pollutants with stronger near-road gradients, including\nUFP and BC (Chapters 4 and 5). Across these studies, a recurring methodological challenge is", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "101-0", + "page": 97, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "longside\ncomplementary analyses of traffic-related pollutants with stronger near-road gradients, including\nUFP and BC (Chapters 4 and 5). Across these studies, a recurring methodological challenge is\nthat mobile monitoring datasets are spatially and temporally autocorrelated, as measurements are\ncollected along road networks and often repeated across the same routes, making model development\nand evaluation especially sensitive to validation design, feature selection, and preprocessing choices,\nparticularly when applying ML-based spatial modeling frameworks. Together, these studies address\nthe following three main research questions:\n1. Mobile monitoring to exposure surfaces: How can route-constrained mobile measurements\ncollected along urban road networks be aggregated, modeled, and validated to produce stable,\nhigh-resolution, city-wide exposure surfaces, and how do different predictor choices influence the\nstability and generalizability of predictions beyond the sampled routes?\n2.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "101-1", + "page": 97, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "idated to produce stable,\nhigh-resolution, city-wide exposure surfaces, and how do different predictor choices influence the\nstability and generalizability of predictions beyond the sampled routes?\n2. Bridging the gap between data reproduction and prediction under dependence:\nHow do validation designs and feature-selection workflows that account for spatial/temporal\ndependence affect estimated model performance, robustness, and transferability (i.e., prediction\nbeyond data reproduction)?\n3. Traffic composition impact and interpretation: How can street-level or 360-degree imagery\nbe used as an interpretability layer, by extracting vehicle presence and fleet composition, to\nexplain and validate spatial patterns in traffic-related pollutant surfaces, particularly the shifts\nobserved during and after the COVID-19 period?\nIn Chapter 3, we developed a practical framework for generating city-wide PM2.5 exposure sur-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "101-2", + "page": 97, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "raffic-related pollutant surfaces, particularly the shifts\nobserved during and after the COVID-19 period?\nIn Chapter 3, we developed a practical framework for generating city-wide PM2.5 exposure sur-\nfaces using low-cost sensors installed on a courier truck operating across downtown Toronto. A\ncentral challenge in this setting is that mobile monitoring data are inherently route-constrained\nand temporally uneven, as measurements are collected along specific delivery corridors and under\natmospheric conditions that can vary substantially from day to day. To address this challenge, we\ndesigned a modeling workflow that explicitly tests how key methodological choices influence both", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "101-3", + "page": 97, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "predictive performance and the plausibility of the resulting exposure maps. In particular, we evalu-\nated the impact of different spatial aggregation strategies and found that aggregating observations\nat fine road-based intervals (on the order of ∼100 m) offers a strong balance between capturing\nnear-road gradients and maintaining statistical stability by increasing the number of observations\navailable per spatial unit. We then assessed how different categories of predictors contribute to\nmodel skill. Base models relying primarily on spatial covariates, representing road-network struc-\nture, land use, and local built-environment characteristics, were able to recover some spatial patterns\nbut showed limited explanatory power overall, reflecting the fact that PM2.5 variability in Toronto\nis strongly influenced by temporally varying factors such as meteorology and regional background\nconditions. In contrast, model performance improved markedly once temporally varying covariates", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "102-0", + "page": 98, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "oronto\nis strongly influenced by temporally varying factors such as meteorology and regional background\nconditions. In contrast, model performance improved markedly once temporally varying covariates\nwere incorporated. Adding meteorological predictors (e.g., temperature, humidity, and wind speed)\nand regional concentration signals derived from stationary monitoring stations helped separate local\nspatial gradients from broader background conditions and increased predictive skill substantially,\nwith test R2 values reaching approximately 0.82 in the best-performing framework. Importantly,\nthis improvement was not achieved through aggressive preprocessing alone, but rather reflected the\nneed to condition mobile observations on the atmospheric state at the time of sampling.Finally, we\nevaluated the credibility of the resulting fleet-based exposure surfaces through comparisons with in-\ndependent regulatory monitoring stations. These comparisons showed that, when temporal drivers", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "102-1", + "page": 98, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "we\nevaluated the credibility of the resulting fleet-based exposure surfaces through comparisons with in-\ndependent regulatory monitoring stations. These comparisons showed that, when temporal drivers\nare properly accounted for, courier-truck monitoring can reproduce regional PM2.5 patterns and\nsupport coherent city-wide mapping, even when direct sampling does not uniformly cover all neigh-\nborhoods. Overall, Chapter 3 demonstrates that mobile fleet sensing can provide a scalable pathway\nfor high-resolution PM2.5 exposure assessment, provided that spatial aggregation is carefully se-\nlected and that both meteorological variability and stationary-station background information are\nintegrated into the modeling framework. This chapter answered the first research question of the\ndissertation.\nIn Chapter 4, we addressed the gap between data reproduction and true prediction by testing\nhow spatiotemporal autocorrelation and validation design shape the credibility of UFP exposure", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "102-2", + "page": 98, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "issertation.\nIn Chapter 4, we addressed the gap between data reproduction and true prediction by testing\nhow spatiotemporal autocorrelation and validation design shape the credibility of UFP exposure\nmaps derived from mobile monitoring. Because mobile datasets are collected along road networks\nwith repeated coverage, nearby observations share similar predictors and pollutant levels; as a re-\nsult, standard random cross-validation can leak information across folds, producing inflated perfor-\nmance estimates and encouraging models to “memorize” campaign-specific structure rather than\nlearn transferable relationships. To address this issue, we implemented dependence-aware modeling\nframeworks that enforce geographic and temporal separation during both hyperparameter tuning\nand forward feature selection, aligning the full workflow with the intended objective: predicting\naverage concentrations over a defined time period at unobserved locations and conditions, rather", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "102-3", + "page": 98, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ning\nand forward feature selection, aligning the full workflow with the intended objective: predicting\naverage concentrations over a defined time period at unobserved locations and conditions, rather\nthan reproducing nearby sampled points. Across the evaluated frameworks, random cross-validation\nmodels exhibited clear overfitting behavior, characterized by near-perfect training performance but\nsubstantially lower test performance. In contrast, structured cross-validation approaches (spatial,\ntemporal, and especially spatiotemporal) produced more conservative yet more defensible estimates\nof predictive skill. Importantly, the spatiotemporal cross-validation combined with feature selection\nsubstantially improved external validity. When evaluated against independent stationary measure-\nments, prediction errors were markedly lower than those from random cross-validation baselines, in\nsome cases nearly halved. These autocorrelation-aware models were also robust to outlier handling,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "102-4", + "page": 98, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "producing stable exposure surfaces and relying on a smaller, more interpretable set of traffic-relevant\npredictors. Overall, Chapter 4 reframed land-use regression as an objective-driven, dependence-\naware prediction problem, showing that validation and feature-selection strategies must be aligned\nwith the modeling objective and explicitly account for spatiotemporal autocorrelation to support\nreliable generalization and physically plausible exposure surfaces. This chapter answered the second\nresearch question of the dissertation.\nIn Chapter 5, we integrated mobile monitoring with computer-vision-based traffic sensing to in-\nvestigate why reductions in UFP and BC have reversed in the post-COVID era. Using fine-tuned\nobject-detection models applied to 360-degree street-level imagery, we quantified the spatial activity\nof multiple vehicle classes and then used a LUR framework to translate these image-derived detec-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "103-0", + "page": 99, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "bject-detection models applied to 360-degree street-level imagery, we quantified the spatial activity\nof multiple vehicle classes and then used a LUR framework to translate these image-derived detec-\ntions into city-wide vehicle activity surfaces, representing high-resolution maps of fleet composition\nfor two summer monitoring campaigns in 2021 and 2023. This integration enabled pollutant maps\nto be interpreted not only as statistical outputs, but as spatial patterns that can be linked directly\nto measurable shifts in fleet activity captured from imagery. The results showed that city-wide in-\ncreases in UFP and BC from 2021 to 2023 were most consistent with the rebound and redistribution\nof freight activity along key transportation corridors. The spatial signatures differed by pollutant.\nIncreases in UFP were strongest along highways, whereas increases in BC were most pronounced on\nmajor arterial roads. These differences were explained by image-derived changes in fleet composition.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "103-1", + "page": 99, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ".\nIncreases in UFP were strongest along highways, whereas increases in BC were most pronounced on\nmajor arterial roads. These differences were explained by image-derived changes in fleet composition.\nCombination trucks increased primarily on highways and near industrial corridors, while single-unit\ntrucks increased along arterial roads, with some decreases observed in the downtown core, aligning\nwith locations where BC concentrations rose most strongly.\nOverall, the findings reinforce that\ntraffic composition, rather than traffic volume alone, drives localized pollution patterns, with roads\ndominated by heavy-duty vehicles exhibiting consistently higher pollutant levels. A key contribu-\ntion of this chapter is demonstrating how imagery can serve as an interpretability layer for exposure\nmodeling, where vehicle-type-specific indicators provide a direct and observable explanation for why\npollution hotspots emerge and how they shift over time. This chapter answered the third research", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "103-2", + "page": 99, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "e\nmodeling, where vehicle-type-specific indicators provide a direct and observable explanation for why\npollution hotspots emerge and how they shift over time. This chapter answered the third research\nquestion of the dissertation.\n6.2\nResearch contributions\nThis dissertation addressed fundamental methodological gaps in the field of high-resolution urban\nexposure assessment, contributing both to sensing technology application and modeling rigor.\n6.2.1\nContributions to air pollution sensing and modeling\nIn Chapter 3, we demonstrated that routine courier-truck operations can be transformed into a\npractical, fleet-based mobile monitoring system for generating city-wide PM2.5 exposure surfaces.\nThis work provided applied guidance on spatial aggregation and showed that mobile monitoring\ncan move beyond route-level snapshots when models explicitly account for the temporal state of\nthe atmosphere, including meteorological conditions and regional background signals. The result is", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "103-3", + "page": 99, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "oring\ncan move beyond route-level snapshots when models explicitly account for the temporal state of\nthe atmosphere, including meteorological conditions and regional background signals. The result is\na scalable blueprint for long-term, city-wide monitoring using low-cost sensors, without requiring\ndense stationary monitoring networks or dedicated research vehicles.\nIn Chapter 5, we advanced source attribution and interpretability by demonstrating how com-\nputer vision applied to street-level and 360-degree imagery can extract vehicle-type-specific activity", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "103-4", + "page": 99, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "and support city-wide traffic composition mapping. By linking these imagery-derived fleet indicators\nto UFP and BC patterns across monitoring campaigns, the dissertation shows how exposure surfaces\ncan be interpreted through measurable shifts in freight and service activity, improving explanatory\npower beyond what is achievable using aggregate traffic indicators alone.\n6.2.2\nContributions to applied machine learning: validation and general-\nization\nA major contribution of this dissertation is demonstrating that in autocorrelated mobile datasets,\nstrong headline performance can mask models that primarily reproduce training structure rather\nthan generalize to new settings. Chapter 4 quantified the risk of information leakage under random\ncross-validation and showed that aligning validation, feature selection, and hyperparameter tuning\nwith the intended deployment objective, namely estimating time-averaged concentrations at unob-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "104-0", + "page": 100, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ndom\ncross-validation and showed that aligning validation, feature selection, and hyperparameter tuning\nwith the intended deployment objective, namely estimating time-averaged concentrations at unob-\nserved locations and conditions, yields models that are more robust, interpretable, and transferable.\nThis reframes land-use regression as an objective-driven prediction problem in which transferability\nis prioritized over apparent goodness of fit.\nAlthough this contribution is developed using environmental exposure data, its implications\nextend more broadly to applied machine learning. Across domains involving autocorrelated datasets,\nincluding geospatial analysis, sensor and mobility data, time-series monitoring, stacked or multi-\nstage machine-learning pipelines, and panel or repeated-measures studies, the same failure mode\ncan arise. Models may appear accurate because they exploit dependence in the training data rather\nthan learning relationships that hold in genuinely new settings.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "104-1", + "page": 100, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "asures studies, the same failure mode\ncan arise. Models may appear accurate because they exploit dependence in the training data rather\nthan learning relationships that hold in genuinely new settings. The key takeaway is that modeling\nframeworks, particularly how feature selection and hyperparameter tuning are embedded within\nan appropriate cross-validation design, shape outcomes as much as the choice of algorithm itself,\nbecause they determine what is learned, what remains interpretable, and what can be trusted beyond\nthe training data.\n6.3\nRecommendations for future research\nTo further advance high-resolution urban air pollution modeling, future research should focus on\nimproving the physical plausibility, robustness, interpretability, and scalability of exposure mapping\nframeworks. Several promising directions are outlined below.\n1. Hybrid physics–machine learning for physically consistent exposure surfaces. Future", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "104-2", + "page": 100, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "terpretability, and scalability of exposure mapping\nframeworks. Several promising directions are outlined below.\n1. Hybrid physics–machine learning for physically consistent exposure surfaces. Future\nexposure-mapping frameworks should move beyond purely statistical prediction by embedding\nphysical constraints directly into the learning process through physics-informed loss functions\nor constrained optimization. In practice, this could involve decomposing pollutant concentra-\ntion into a regional background component and a local increment, where the local increment\nis constrained to follow near-road dispersion behavior. Examples include enforcing monotonic\ndecay with distance from major roads under comparable conditions, smooth variation along\nprevailing wind directions, and bounded gradients away from emission sources. Hybrid models\ncould also incorporate meteorology-dependent structure, such as stronger sensitivity to verti-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "104-3", + "page": 100, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "iation along\nprevailing wind directions, and bounded gradients away from emission sources. Hybrid models\ncould also incorporate meteorology-dependent structure, such as stronger sensitivity to verti-\ncal mixing during stable nighttime conditions and limits on unrealistic short-range variability", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "104-4", + "page": 100, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "during well-mixed daytime periods. A particularly useful direction is to treat dispersion con-\nsistency as a regularization mechanism. In this setting, the model remains data-driven but is\npenalized when it violates physically expected behavior, such as predicting local peaks far from\nplausible sources, generating inconsistent gradients across parallel corridors, or creating disconti-\nnuities when meteorological conditions change smoothly. This approach is especially important\nduring winter inversions and sharp diurnal transitions, when sparse sampling can cause purely\ndata-driven models to extrapolate in ways that appear smooth but are physically implausible.\nThe outcome would be exposure surfaces that are not only accurate but also mechanistically\ndefensible, thereby increasing trust for health and planning applications.\n2. Road-network-aware spatiotemporal learning using graph-based modeling. Most cur-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "105-0", + "page": 101, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "t only accurate but also mechanistically\ndefensible, thereby increasing trust for health and planning applications.\n2. Road-network-aware spatiotemporal learning using graph-based modeling. Most cur-\nrent land-use regression predictors rely on Euclidean buffers, which implicitly assume that nearby\nspace influences concentrations equally in all directions. However, road emissions and human ex-\nposure propagate through a non-Euclidean domain, where connected corridors, ramps, merges,\nand intersections create dependencies that are more naturally described by network topology\nthan by circular distance. Future work should explicitly represent the road network using graph-\nbased spatiotemporal models, in which road segments and intersections are represented as nodes\nand edges. Graph-based representations can encode real mobility structure, including upstream\nand downstream connectivity, corridor continuity, and intersection complexity. Such models", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "105-1", + "page": 101, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ented as nodes\nand edges. Graph-based representations can encode real mobility structure, including upstream\nand downstream connectivity, corridor continuity, and intersection complexity. Such models\ncould better capture micro-hotspots that traditional regression approaches tend to smooth over,\nsuch as near highway on-ramps, major merges, bus hubs, and freight terminals. At the same\ntime, they can reduce unrealistic spreading of predicted hotspots into nearby residential streets\nthat are close in Euclidean distance but disconnected in traffic flow. A promising direction is to\ncombine graph structure with time-aware inputs, such as meteorology and hour-of-day effects,\nto learn when and where network-driven hotspots intensify.\n3. Probabilistic exposure mapping and uncertainty quantification. To make high-resolution\nexposure surfaces more actionable, future models should provide calibrated uncertainty estimates\nrather than point predictions alone.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "105-2", + "page": 101, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "re mapping and uncertainty quantification. To make high-resolution\nexposure surfaces more actionable, future models should provide calibrated uncertainty estimates\nrather than point predictions alone. Two distinct sources of uncertainty should be addressed:\ndata uncertainty, arising from sensor noise, calibration drift, and sampling representativeness;\nand model uncertainty, arising from limited spatial coverage and extrapolation into under-\nsampled neighborhoods or road types. Future frameworks should produce uncertainty maps\nthat clearly distinguish regions with strong data support from those where predictions rely\nheavily on extrapolation. Potential methods include Bayesian regression or boosting variants,\nensemble approaches with spatial or spatiotemporal blocking, and conformal prediction tech-\nniques that yield valid prediction intervals under realistic assumptions. Beyond visualization,\nuncertainty should be made decision-relevant by distinguishing high-exposure, high-certainty", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "105-3", + "page": 101, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tion tech-\nniques that yield valid prediction intervals under realistic assumptions. Beyond visualization,\nuncertainty should be made decision-relevant by distinguishing high-exposure, high-certainty\nhotspots from high-exposure, low-certainty areas that warrant targeted follow-up monitoring.\nThis would also strengthen epidemiological applications by enabling sensitivity analyses that\nexplicitly propagate exposure uncertainty into health-effect estimates.\n4. Multi-city generalization through transfer learning and domain adaptation. Because\nmobile and imagery-based monitoring campaigns are resource intensive, future research should\nassess whether partially universal predictor sets exist for traffic-related pollutants, such as pre-\ndictors that consistently capture heavy-duty or diesel-dominated corridor effects. Transferability", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "105-4", + "page": 101, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "challenges arise from both covariate shift, including differences in road geometry, land use, and\nfleet composition, and conditional shift, where the same predictor implies different emissions or\ndispersion behavior across cities. Future studies should explicitly test cross-city generalization\nthrough controlled experiments, such as training models in one city and evaluating performance\nin another without retraining, followed by fine-tuning using a limited number of local sampling\ndays. Transfer-learning approaches could reuse learned representations of road-network struc-\nture or imagery-derived features while recalibrating only the final layers. A practical objective is\nto establish a minimum-data protocol that identifies the smallest amount of local data required\nto adapt a model with acceptable uncertainty, making deployment feasible for resource-limited\nmunicipalities.\n5. Causal and counterfactual modeling for intervention-relevant inference. Predictive", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "106-0", + "page": 102, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "red\nto adapt a model with acceptable uncertainty, making deployment feasible for resource-limited\nmunicipalities.\n5. Causal and counterfactual modeling for intervention-relevant inference. Predictive\nexposure maps are not automatically suitable for answering policy-relevant questions about in-\nterventions. Future work should complement mapping with causal frameworks that estimate\nthe effects of transportation interventions while controlling for confounding factors such as me-\nteorology, temporal trends, and regional background transport. This can be achieved through\nquasi-experimental designs centered on real-world interventions, including low-emission zones,\ntruck restrictions, delivery re-timing, signal-timing changes, or street redesign. A key direction is\nto formalize the causal question explicitly, such as whether reducing heavy-duty vehicle activity\non a corridor lowers nearby ultrafine particle or black carbon exposure relative to what would\nhave occurred otherwise.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "106-1", + "page": 102, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "causal question explicitly, such as whether reducing heavy-duty vehicle activity\non a corridor lowers nearby ultrafine particle or black carbon exposure relative to what would\nhave occurred otherwise. Methods may include difference-in-differences with matched control\ncorridors, synthetic control approaches, or structural causal models that treat traffic composi-\ntion as a mediator. These tools would enable credible counterfactual simulations and reduce\nthe risk of misattributing observed pollution changes to policy actions when they are driven by\nweather or background variability.\n6. Richer imagery-derived features for interpretation and urban morphology. Beyond\nvehicle counts, street-level imagery can support interpretation by enabling the extraction of\nmicro-scale urban form features that influence pollutant dispersion and retention.\nFuture\npipelines should expand computer vision methods to estimate variables such as street-canyon", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "106-2", + "page": 102, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "the extraction of\nmicro-scale urban form features that influence pollutant dispersion and retention.\nFuture\npipelines should expand computer vision methods to estimate variables such as street-canyon\nenclosure, building-height proxies, sky-view factor approximations, tree-canopy density, curbside\nparking intensity, and intersection complexity. These features can help explain why roads with\nsimilar traffic volumes exhibit markedly different pollution patterns, particularly for ultrafine\nparticles and black carbon. An especially valuable direction is mechanism-aware interpretation,\nin which morphology features are linked to pollutant behavior under different meteorological\nregimes. For example, stronger pollutant retention may occur along canyon-like arterials during\nstable evening conditions. This would elevate imagery from a traffic-sensing tool to a broader\nurban microenvironment sensor, improving attribution and supporting actionable urban design\ninsights.\n7.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "106-3", + "page": 102, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "stable evening conditions. This would elevate imagery from a traffic-sensing tool to a broader\nurban microenvironment sensor, improving attribution and supporting actionable urban design\ninsights.\n7. Motion-aware traffic sensing and engine-load proxies. Static imagery cannot reliably dis-\ntinguish moving traffic from parked vehicles and cannot capture driving dynamics that strongly\ninfluence emissions, particularly for ultrafine particles and black carbon. Future work should\nincorporate video-based tracking, multi-frame sampling, or motion-aware filtering to estimate\nspeed, queuing, stop-and-go behavior, and acceleration proxies. These variables more closely", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "106-4", + "page": 102, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "reflect engine load, which is often a stronger determinant of particle emissions than vehicle\ncounts alone. Operationally, this could involve producing corridor-level indicators such as con-\ngestion intensity, idling likelihood near intersections or loading zones, and acceleration events\nnear ramps and merges. Integrating these motion-aware indicators into exposure models could\nimprove emissions relevance, sharpen identification of short-range hotspots, and clarify interven-\ntion levers, such as reducing queuing or smoothing traffic flow, even when total traffic volume\nremains unchanged.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "107-0", + "page": 103, + "document_type": "thesis", + "page_header": "CHAPTER 6. SUMMARY AND CONCLUSION", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Table A-1: Summary of Predictor Variables Used for Land Use Regression Modeling\nAttribute\nType of Measurement\nLand use related variables\nParking lots areas\nArea within buffer per area of buffer\nCommercial areas\nArea within buffer per area of buffer\nGovernmental areas\nArea within buffer per area of buffer\nIndustrial areas\nArea within buffer per area of buffer\nOpen area\nArea within buffer per area of buffer\nResidential area\nArea within buffer per area of buffer\nWaterbody area\nArea within buffer per area of buffer\nParks\nArea within buffer per area of buffer\nDistance to the lake\nClosest distance from Lake Ontario in meters\nPopulation\nPopulation density in each buffer\nDistance to Billy Bishop Toronto City Air-\nClosest distance from Toronto Island Airport\nport\nAccommodation points\nNumber of points in buffers\nChimney\nNumber of points in buffers\nGas station\nNumber of points in buffers\nRestaurant\nNumber of points in buffers\nTraffic related variables\nRoad area", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "108-0", + "page": 124, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Accommodation points\nNumber of points in buffers\nChimney\nNumber of points in buffers\nGas station\nNumber of points in buffers\nRestaurant\nNumber of points in buffers\nTraffic related variables\nRoad area\nArea within buffer per area of buffer\nLength of highway in a buffer\nLength in buffer in meters\nLength of railway in a buffer\nLength in buffer in meters\nLength of major roads in a buffer\nLength in buffer in meters\nLength of bus lines\nLength in buffer in meters\nLength of all-road in a buffer\nLength in buffer in meters\nClosest distance to highway\nClosest distance to nearest highway\nClosest distance to major road\nClosest distance to nearest major road\nClosest distance to rail line\nClosest distance to nearest rail line\nNumber of intersections\nNumber of points in buffers\nNumber of traffic signals\nNumber of points in buffers\nNumber of bus stops\nNumber of points in buffers\nAADT traffic count for four different years\nTraffic count per area of buffer\nTemporally changing variables\nAir pollution", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "108-1", + "page": 124, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ls\nNumber of points in buffers\nNumber of bus stops\nNumber of points in buffers\nAADT traffic count for four different years\nTraffic count per area of buffer\nTemporally changing variables\nAir pollution\nHourly air pollutant (PM2.5 and NO2) concentrations\nfrom a reference station\nWeather\nWind speed, relative humidity, and tem-\nData obtained from a stationary monitoring station\nperature\nNote: Predictor variables were grouped into land use, traffic, temporally changing, and weather-related\ncategories for model development.\nA.1.3\nModel development\nHyperparameter tuning under each CV\nTo optimize model performance, we fine-tuned key XGBoost hyperparameters using Bayesian opti-\nmization, which efficiently explores the hyperparameter space by iteratively refining the search based\non past evaluations. This approach allows for a more targeted search than traditional grid search,\nleading to improved model performance. The following XGBoost hyperparameters were optimized:", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "108-2", + "page": 124, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "h based\non past evaluations. This approach allows for a more targeted search than traditional grid search,\nleading to improved model performance. The following XGBoost hyperparameters were optimized:\n• Learning Rate (learning rate): Controls the step size in updating model weights to balance\nconvergence speed and accuracy.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "108-3", + "page": 124, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "• Number of Estimators (n estimators): Defines the number of boosting iterations, im-\npacting model complexity.\n• Max Depth (max depth): Limits the depth of each decision tree to prevent overfitting\nwhile capturing important feature interactions.\n• Subsample Ratio (subsample): Specifies the fraction of training data used per boosting\niteration to introduce randomness and reduce overfitting.\n• L1 Regularization (reg alpha): Applies an L1 penalty to shrink coefficients and enhance\nfeature selection.\n• L2 Regularization (reg lambda): Applies an L2 penalty to prevent large weight values\nand improve generalization.\nBayesian optimization dynamically adjusts the search space, refining hyperparameter selection\nbased on prior evaluations.\nHyperparameters were sampled from appropriate distributions, and training data was split into\nseveral folds based on chosen cross validation for each model. An XGBoost regressor was trained on", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "109-0", + "page": 125, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tions.\nHyperparameters were sampled from appropriate distributions, and training data was split into\nseveral folds based on chosen cross validation for each model. An XGBoost regressor was trained on\nthe scaled feature using the MinMaxScaler with performance assessed using cross-validation R2. The\nfinal models were trained using the best hyperparameters identified through optimization, ensuring\nstrong generalizability across different validation sets.\nTo reduce overfitting, we used four different CV strategies for tuning—Random, Spatial, Tem-\nporal, and Spatiotemporal CV. We built four separate models, each tuned using one of these CV\napproaches. Each method partitions the dataset in a unique way to address biases from spatial\nautocorrelation, temporal trends, or both.\na) Random Cross-Validation (RCV).\nWe employed 10-fold RCV on the 80% training set—derived from the 100 m-aggregated dataset—for\nhyperparameter tuning, randomly splitting these records into ten folds of similar size.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "109-1", + "page": 125, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "-Validation (RCV).\nWe employed 10-fold RCV on the 80% training set—derived from the 100 m-aggregated dataset—for\nhyperparameter tuning, randomly splitting these records into ten folds of similar size. At each it-\neration, one-fold acted as the validation set, while the remaining nine served as the training set.\nAlthough this conventional approach is fast and straightforward, it can overestimate model perfor-\nmance when the data exhibit strong spatial or temporal correlations.\nb) Spatial Cross-Validation (SCV).\nWe generated 400 geographic clusters across the study area (Figure A-1a) using the 100 m–\naggregated UFP measurements from the training set. The number and size of clusters were selected\nto reflect the short-range spatial autocorrelation of UFP, which exhibit steep concentration gradients\nand rapid decay with distance from emission sources. Within the 80% training set, the clusters were", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "109-2", + "page": 125, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "reflect the short-range spatial autocorrelation of UFP, which exhibit steep concentration gradients\nand rapid decay with distance from emission sources. Within the 80% training set, the clusters were\nrandomly divided into 10 folds for SCV, ensuring that each spatial region was used for validation\nonly once. To further minimize spatial dependence between training and validation data, a 300 m\nexclusion buffer (“dead buffer”) was applied around validation clusters. This buffer distance was\nchosen to exceed the effective spatial correlation range of UFP, thereby preventing information\nleakage from adjacent road segments influenced by the same traffic emissions.\nc) Temporal Cross-Validation (TCV).\nIn air pollution research, pollutant levels often vary significantly by season, day of the week, and\ntime of day, due to changes in meteorology, traffic, and emission sources. Consequently, random or", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "109-3", + "page": 125, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "(a)\n(b)\nFigure A-1: (a) Overview of the 400 geographic clusters used for SCV. (b) Illustration of a single\nfold from the 10-fold spatial cross-validation. Red points represent the validation fold, and blue\npoints represent the training set. Any points within the specified buffer zone around the validation\nfold were excluded and are not shown.\nspatial cross-validation alone may not fully capture these temporal fluctuations. To address this,\nwe performed TCV before aggregating the data at 100 m intervals. Specifically, we divided the 33\ndistinct sampling days into k folds, ensuring each day was exclusively used for training or validation\nin any given fold. Afterward, each fold’s data was aggregated at 100 m intervals along the sampling\nroutes (Figure A-2). This method maintains strict temporal boundaries and helps with choosing\nhyperparameters that can enhance the model’s ability to generalize across new sampling days.\nd) Spatiotemporal Cross-Validation (STCV).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "110-0", + "page": 126, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "od maintains strict temporal boundaries and helps with choosing\nhyperparameters that can enhance the model’s ability to generalize across new sampling days.\nd) Spatiotemporal Cross-Validation (STCV).\nWhile TCV handles temporal variability and SCV addresses geographic autocorrelation, STCV\nintegrates both methods to ensure folds are disjoint in time and space. In this study, we first split\nthe dataset by sampling day (temporal axis) and then further partitioned the training sets into\ngeographic clusters (spatial axis), with a 300 m buffer zone around validation clusters. Specifically:\n(a) Temporal Partition:\nWe grouped the data by sampling day, assigning each distinct day to one of several time-\nbased folds. This enforces strict chronological separation—no single day contributes to both\nthe training and validation subsets in a given fold.\n(b) Spatial Partition:\nNext, the training data within each temporal fold was divided into spatial clusters (e.g., K-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "110-1", + "page": 126, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ingle day contributes to both\nthe training and validation subsets in a given fold.\n(b) Spatial Partition:\nNext, the training data within each temporal fold was divided into spatial clusters (e.g., K-\nmeans), forming region-based folds. By designating one of the spatial folds as “validation fold”\nwe created spatially separate areas. Moreover, we excluded training points within 300 m of\nthose validation regions to reduce spatial leakage.\n(c) Fold Combination:\nEach temporal fold was paired with each spatial fold to yield fully disjoint subsets, enforcing both\ntime and space separation. A custom cross-validator systematically yielded training/validation\nsplits, tagging records based on the combined day–cluster assignment and buffer exclusion.\nBy integrating temporal and spatial blocks together with a dead buffer zone, our STCV frame-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "110-2", + "page": 126, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure A-2:\nExample of aggregated dataset after temporal cross-validation by sampling days.\nTraining days are shown in blue, while validation days appear in red.\nwork strictly isolates training and validation sets in both time and space. We employed this setup\nfor hyperparameter tuning to better equip the model for truly unseen conditions, reducing the risk\nof learning artifacts unique to specific regions or sampling periods and thereby simulating more\nrealistic, real-world scenarios.\nFeature selection method\nIn addition to hyperparameter tuning, we integrated a feature selection procedure with each of the\nthree advanced cross-validation strategies—SCV, TCV, and STCV. The general workflow for feature\nselection follows three main steps:\na) Hyperparameter Tuning to define a base model\nb) Forward Feature Selection combined with CV strategies\nc) Hyperparameter Re-Tuning on Selected Features\nThe specific details of each step vary slightly depending on whether we are dealing with SCV,", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "111-0", + "page": 127, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "b) Forward Feature Selection combined with CV strategies\nc) Hyperparameter Re-Tuning on Selected Features\nThe specific details of each step vary slightly depending on whether we are dealing with SCV,\nTCV, or STCV, as described below.\na) Feature selection integrated with SCV:\nFeature selection was integrated with SCV through a three-step process:\nStep 1: Hyperparameter Tuning with SCV\nWe optimized the model’s hyperparameters using a SCV that ensures training and validation sets", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "111-1", + "page": 127, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "are geographically distinct. In each trial, an XGBoost model was trained on the spatially isolated\nfolds, and the mean R2 score across folds was used to iteratively refine the hyperparameters to\nchoose these parameters for a base model.\nStep 2: Forward Feature Selection with SCV\nWe implemented a forward selection strategy under the same SCV framework. Starting with an\nempty set of selected features, each candidate predictor was added one at a time to the current\nfeature set. For each candidate, model performance was evaluated using cross validation score for\nthe base model. The evaluation metric was the negative mean squared error. The feature that\nprovided the greatest performance improvement was permanently added to the selected subset and\nremoved from the candidate list. This process was repeated iteratively until no further improvement\nwas achieved. Parallel processing was used to accelerate the evaluation.\nStep 3: Hyperparameter Tuning on Selected Features", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "112-0", + "page": 128, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "te list. This process was repeated iteratively until no further improvement\nwas achieved. Parallel processing was used to accelerate the evaluation.\nStep 3: Hyperparameter Tuning on Selected Features\nAfter determining the optimal feature subset, we re-tuned the hyperparameters using Bayesian\noptimization. In this step, the model was evaluated via SCV using only the selected features, with\nthe objective of maximizing the mean R2 score for spatial folds. The best hyperparameters obtained\nwere then used to train the final model, ensuring that both the feature set and tuning parameters\nare optimized for robust performance on unseen spatial data.\nb) Feature selection integrated with TCV\nFeature selection was integrated with TCV through a three-step process:\nStep 1: Hyperparameter Tuning with TCV\nFor the temporal approach, we began by tuning the model’s hyperparameters using TCV. The\ndataset—comprising measurements from 33 distinct sampling days—was partitioned into folds us-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "112-1", + "page": 128, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Tuning with TCV\nFor the temporal approach, we began by tuning the model’s hyperparameters using TCV. The\ndataset—comprising measurements from 33 distinct sampling days—was partitioned into folds us-\ning KFold, with data aggregated at 100 m intervals. Each fold was labeled as training or validation\nbased on its sampling day, and hyperparameters were optimized to maximize the mean R2 score.\nThese hyperparameters were used to define a base model.\nStep 2: Forward Feature Selection with TCV\nWe then applied a forward selection strategy under TCV. Starting with an empty set of features,\neach remaining candidate was added one at a time and the model’s performance was evaluated via\ncross validation score for the base model using custom temporal splits. The evaluation metric was\nthe negative mean squared error. Importantly, after each feature was added, we regenerated the\ntemporal folds by randomly selecting from the 33 sampling days—ensuring that each iteration of", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "112-2", + "page": 128, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "etric was\nthe negative mean squared error. Importantly, after each feature was added, we regenerated the\ntemporal folds by randomly selecting from the 33 sampling days—ensuring that each iteration of\nthe forward selection was evaluated using a fresh temporal cross-validation split. The candidate\nfeature that yielded the highest performance improvement was permanently added to the selected\nset and removed from the candidate pool.\nStep 3: Hyperparameter Tuning on Selected Features with TCV\nAfter identifying the optimal feature subset, we re-tuned the hyperparameters using Bayesian opti-\nmization under TCV, using only the selected features. The objective was to maximize the mean R2\nscore across the newly formed temporal folds. The best hyperparameters obtained were then used\nto train the final model, ensuring that both the feature set and tuning parameters are optimized for\nrobust performance on unseen temporal data.\nc) Feature selection integrated with STCV", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "112-3", + "page": 128, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "re then used\nto train the final model, ensuring that both the feature set and tuning parameters are optimized for\nrobust performance on unseen temporal data.\nc) Feature selection integrated with STCV\nFeature selection under STCV follows the same three-step framework used for TCV, with the key\ndistinction that the dataset is split along both spatial and temporal dimensions.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "112-4", + "page": 128, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "A.1.4\nModel implementation and evaluation\nTo check overfitting, we employed an external evaluation using a random 20% holdout sample from\nthe aggregated dataset. This subset was not used in model training, hyperparameter tuning, or\nfeature selection.\nTo mitigate the effects of short-term transient pollution spikes, we computed median concen-\ntrations. The external holdout evaluation employed metrics including the R2, RMSE, and MAE.\nThis external validation dataset was used consistently across all models, providing a standardized\nbasis for objective comparisons of model performance and generalizability. For these evaluations, we\nreported metrics including R2, RMSE, and MAE. Each metric provides complementary information,\nand their combined use has been recommended by numerous air pollution modeling studies.\nA.1.5\nModel evaluation metrics\nR2 indicates the proportion of variability in pollutant concentrations explained by the model, guiding\nhyperparameter tuning and feature selection.\nPn", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "113-0", + "page": 129, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "modeling studies.\nA.1.5\nModel evaluation metrics\nR2 indicates the proportion of variability in pollutant concentrations explained by the model, guiding\nhyperparameter tuning and feature selection.\nPn\ni=1(yi −ˆyi)2\nR2 = 1 −\nPn\ni=1(yi −¯y)2\nMAE and RMSE provide absolute measures of prediction error in the original units of pollutant\nconcentrations. RMSE emphasizes larger errors more significantly, whereas MAE is less sensitive to\noutliers and extreme spike.\nv\nn\nu\nt 1\nX\nu\n(yi −ˆyi)2\nRMSE =\nn\ni=1\nn\nMAE = 1\nX\n|yi −ˆyi|\nn\ni=1\nDifferent CV evaluation metrics were employed exclusively during the hyperparameter tuning and\nfeature selection phases to optimize model structure and complexity. Four CV strategies—RCV, SCV\n(Figure A-1b), TCV (Figure A-2), and STCV—were used independently for tuning and selecting\noptimal feature subsets.\nTo ensure consistency in evaluating each final model’s performance, we applied four CV strate-\ngies—Random, Spatial, Temporal, and Spatiotemporal.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "113-1", + "page": 129, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "tly for tuning and selecting\noptimal feature subsets.\nTo ensure consistency in evaluating each final model’s performance, we applied four CV strate-\ngies—Random, Spatial, Temporal, and Spatiotemporal. Specifically, once a model had undergone\nhyperparameter tuning and feature selection under its designated CV strategy (the final models),\nwe tested the final chosen models with all the four CV strategies by reallocating all the data points\nto folds in a comparable manner (e.g., identical cluster assignments for SCV and STCV, consistent\nday-based splits for TCV). This approach guarantees a fair, direct comparison of how each model\nperforms under different CV, while maintaining the integrity of the data splits and minimizing any\nconfounding factors introduced by varying fold assignments. By comparing Figure A-1b, and Fig-\nure A-3, we can see CV used for evaluation of generalization is different from the designated CV for\ntunning and feature selection.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "113-2", + "page": 129, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure A-3: Illustration of a single fold from the 10-fold SCV used for model evaluation.\nA.1.6\nExternal validation with stationary sensors\nTo independently assess model performance mobile-based LUR models were externally validated\nagainst data from a stationary network of backyard sensors.\nTo quantify the discrepancy between predicted and observed backyard sensor concentrations, we\nemployed the APE. For each backyard sensor, let yi represent the average measured concentration\nduring the one-month sampling period, and let ˆyi be the corresponding model prediction. The APE\nis then calculated as:\nAPEi(%) = |yi −ˆyi|\n× 100\nyi\nBy analyzing APE values across all backyard sensor sites, we can evaluate the relative prediction\nerror of each model and further assess its robustness for estimating average exposure levels over a\nperiod.\nA.2\nResults\nA.2.1\nData summary\nEnsuring data quality in UFP concentration measurements was critical for reliable model develop-\nment.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "114-0", + "page": 130, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "s robustness for estimating average exposure levels over a\nperiod.\nA.2\nResults\nA.2.1\nData summary\nEnsuring data quality in UFP concentration measurements was critical for reliable model develop-\nment. The raw dataset initially contained 448,309 UFP measurements collected over 34 sampling\ndays. After removing erroneous records (i.e., values below 1,000 or above 1,000,000 UFP counts),", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "114-1", + "page": 130, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "data points outside Toronto, and measurements recorded outside the 8 AM to 7 PM window, the\ndataset was reduced to 440,323 valid observations.\nFigure A-4 presents a violin and boxplot of all recorded UFP concentrations per sampling day,\nillustrating both the distribution and variability of recorded values. The violin plot represents the\nprobability density of UFP concentrations, while the overlaid boxplot highlights quartiles, medians,\nand whiskers extending to 1.5 times the IQR. This visualization provides insight into the spread of\nthe data and the presence of high variation in UFP level. The variability across days highlights the\ninfluence of traffic patterns, meteorology, and local emission sources on measured concentrations.\nFigure A-4: Violin and boxplot of UFP concentrations per sampling day, showing distribution and\nsummary statistics while minimizing extreme outliers. Y-axis constrained for visual clarity.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "115-0", + "page": 131, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ntrations.\nFigure A-4: Violin and boxplot of UFP concentrations per sampling day, showing distribution and\nsummary statistics while minimizing extreme outliers. Y-axis constrained for visual clarity.\nFollowing data cleaning, we developed seven models using IQR-based filtering by road type to\nevaluate the impact of different modeling frameworks on predictive performance. To assess the effect\nof outlier handling on different modelling approaches, four additional models were created using two\nalternative approaches: no outlier removal, which retained all data points to preserve peak pollution\nvalues, and 95th percentile removal, which excluded the top 5% of UFP concentrations. A summary\nof key statistics for each of these outlier handling methods is provided in Table A-2.\nTable A-2: Summary statistics of UFP measurements under different outlier handling methods.\nOutlier Handling Method\nCount\nMean\nStd\nMedian\n25%\n75%\nMax\nNo Outlier Removal\n440,323\n23,906\n45,900\n14,777\n8,175\n24,571\n999,226", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "115-1", + "page": 131, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "statistics of UFP measurements under different outlier handling methods.\nOutlier Handling Method\nCount\nMean\nStd\nMedian\n25%\n75%\nMax\nNo Outlier Removal\n440,323\n23,906\n45,900\n14,777\n8,175\n24,571\n999,226\nIQR-Based Filtering by Road Type\n413,407\n16,817\n12,744\n13,856\n7,815\n22,206\n107,009\n95th Percentile Removal\n418,122\n16,816\n11,677\n14,017\n7,879\n22,564\n59,828", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "115-2", + "page": 131, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Overall, these preprocessing choices significantly affect the mean, standard deviation, and max-\nimum UFP concentrations, as shown in Table A-2. Notably, the median remains relatively stable\nacross the three outlier-handling methods, but large spikes inflate the mean and standard devia-\ntion when no outlier removal is applied. No outlier removal retains the most data but includes\nextreme events; IQR-based filtering—applied separately by road type—offers a more conservative\napproach to removing outliers; and 95th-percentile removal provides an intermediate balance. Each\noutlier-handling strategy was carried forward into subsequent model development and CV exper-\niments, forming the basis for evaluating how each modeling method can get influenced by data\npreprocessing.\nIn addition to the mobile data, UFP concentrations were continuously recorded at 17 residential\nbackyard locations using DiscMini and Partector sensors during the same three-week campaign.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "116-0", + "page": 132, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "rocessing.\nIn addition to the mobile data, UFP concentrations were continuously recorded at 17 residential\nbackyard locations using DiscMini and Partector sensors during the same three-week campaign.\nMeasurements were restricted to the 8 AM to 7 PM window to align with the mobile campaign.\nFigure A-5 illustrates the variation in UFP concentrations across backyard sites, reflecting localized\ndifferences in environmental conditions and proximity to emission sources such as roads, buildings,\nor vegetation. Summary statistics for each backyard sensor are provided in Table A-3.\nFigure A-5: UFP concentrations recorded at 17 residential backyard sites during the three-week\nmonitoring period (8 AM–7 PM), highlighting variability across locations due to local environmental\nconditions and proximity to emission sources.\nA.2.2\nImpact of modeling framework\nTable A-4 presents the performance metrics (R2, RMSE, and MAE) for all model configurations\nevaluated using an 80/20 random train-test split.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "116-1", + "page": 132, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "y to emission sources.\nA.2.2\nImpact of modeling framework\nTable A-4 presents the performance metrics (R2, RMSE, and MAE) for all model configurations\nevaluated using an 80/20 random train-test split.\nThese results provide insight into in-sample", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "116-2", + "page": 132, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Table A-3: Summary statistics of UFP concentrations measured by each backyard DiscMini (DM)\nand Partector (PRT) sensor, including count, mean, median, and range of observations.\nUnit\nCount\nMean\nMedian\nStd\nMin\nMax\nDM10\n63,481\n6,210.31\n3,492\n14,526.14\n1,002\n1,104,573\nDM14\n63,292\n8,982.06\n6,238\n41,160.28\n1,398\n4,167,965\nDM15\n63,063\n7,756.52\n4,929\n8,867.13\n1,001\n704,203\nDM16\n63,140\n7,186.12\n5,093\n12,482.03\n1,026\n715,998\nDM17\n62,724\n8,344.16\n5,853\n13,516.68\n1,001\n1,166,068\nDM18\n62,781\n6,889.49\n4,399\n33,584.85\n1,001\n4,398,565\nDM19\n63,143\n5,250.45\n3,201\n22,174.91\n1,001\n2,997,100\nDM22\n61,134\n3,627.11\n2,915\n2,744.29\n1,001\n212,064\nDM3\n63,628\n7,571.47\n3,970\n29,409.06\n1,002\n2,176,635\nDM5\n63,299\n6,795.16\n5,197\n6,170.19\n1,001\n227,475\nDM6\n60,447\n19,545.02\n9,018\n113,663.32\n1,134\n3,706,109\nDM7\n63,306\n10,845.56\n7,777\n21,069.02\n1,002\n2,184,976\nDM8\n63,212\n10,512.09\n7,803\n8,045.62\n1,225\n184,156\nPRT4\n636,700\n7,764.10\n5,485\n11,787.02\n1,001\n1,615,623\nPRT5\n651,316\n8,320.63\n5,989\n8,737.13\n1,001\n611,245\nPRT6", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "117-0", + "page": 133, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ",845.56\n7,777\n21,069.02\n1,002\n2,184,976\nDM8\n63,212\n10,512.09\n7,803\n8,045.62\n1,225\n184,156\nPRT4\n636,700\n7,764.10\n5,485\n11,787.02\n1,001\n1,615,623\nPRT5\n651,316\n8,320.63\n5,989\n8,737.13\n1,001\n611,245\nPRT6\n643,395\n7,746.33\n5,043\n7,639.47\n1,001\n80,509\nPRT8\n632,728\n7,184.91\n5,363\n7,097.68\n1,001\n784,929\nPRT9\n610,515\n9,082.05\n6,072\n14,374.65\n1,001\n3,386,064\nperformance and help establish a baseline for comparison across different modeling frameworks.\nThe table reflects how hyperparameter tuning, feature selection, and outlier-handling strategies\naffect model fit when evaluated on randomly held-out data.\nTable A-4: Performance of Models on Training and 20% Random Holdout Test Data.\nTraining Set\nTest Set\nModel Name / Metric\nR2\nR2\nRMSE\nMAE\nRMSE\nMAE\nRCV HP\n0.999\n400\n261\n0.715\n6,692\n3,869\nSCV HP\n0.825\n5,070\n3,441\n0.659\n7,319\n4,638\nTCV HP\n0.544\n8,189\n5,472\n0.510\n8,770\n5,856\nSTCV HP\n0.579\n7,867\n5,171\n0.550\n8,401\n5,537\nSCV FFS\n0.957\n2,529\n1,736\n0.717\n6,662\n4,013\nTCV FFS\n0.413\n9,290\n6,550\n0.398\n9,722", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "117-1", + "page": 133, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "5,070\n3,441\n0.659\n7,319\n4,638\nTCV HP\n0.544\n8,189\n5,472\n0.510\n8,770\n5,856\nSTCV HP\n0.579\n7,867\n5,171\n0.550\n8,401\n5,537\nSCV FFS\n0.957\n2,529\n1,736\n0.717\n6,662\n4,013\nTCV FFS\n0.413\n9,290\n6,550\n0.398\n9,722\n6,816\nSTCV FFS\n0.441\n9,127\n6,441\n0.450\n9,288\n6,521\nNote: Model names indicate the CV strategy used for hyperparameter tuning and whether FFS was\napplied. For example, RCV HP refers to a model using random CV without feature selection. SCV HP,\nTCV HP, and STCV HP use spatial, temporal, and spatiotemporal CV, respectively, also without feature\nselection. SCV FFS, TCV FFS, and STCV FFS apply forward feature selection with spatial, temporal,\nand spatiotemporal CV, respectively.\nAbbreviations: CV = Cross-validation; FFS = Forward feature selection; R2 = Coefficient of\ndetermination; RMSE = Root mean squared error; MAE = Mean absolute error.\nTable A-5 reports the performance of each model under four cross-validation strategies: random,\nspatial, temporal, and spatiotemporal.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "117-2", + "page": 133, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ation; RMSE = Root mean squared error; MAE = Mean absolute error.\nTable A-5 reports the performance of each model under four cross-validation strategies: random,\nspatial, temporal, and spatiotemporal. This comparison highlights the variability in model general-\nizability depending on the type of validation used. It further demonstrates the limitations of random\nCV and emphasizes the value of using spatially and temporally structured CV to evaluate model\nrobustness in the context of autocorrelated environmental data.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "117-3", + "page": 133, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Table A-5: Model performance (R2, RMSE, MAE) under Random, Spatial, Temporal, and Spa-\ntiotemporal CV.\nRandom Cross Validation\nSpatial Cross Validation\nTemporal Cross Validation\nSpatiotemporal Cross Validation\nModel Name / Metric\nR2\nR2\nR2\nR2\nRMSE\nMAE\nRMSE\nMAE\nRMSE\nMAE\nRMSE\nMAE\nRCV HP\n0.735\n6,245\n3,656\n0.355\n9,175\n6,127\n0.035\n15,928\n10,999\n0.105\n14,533\n10,111\nSTCV HP\n0.682\n6,845\n4,401\n0.378\n9,026\n6,222\n0.088\n15,508\n10,704\n0.136\n14,264\n9,950\nTCV HP\n0.520\n8,428\n5,618\n0.345\n9,324\n6,400\n0.130\n15,159\n10,471\n0.154\n14,139\n9,933\nSTCV HP\n0.547\n8,188\n5,393\n0.357\n9,265\n6,334\n0.110\n15,305\n10,696\n0.138\n14,245\n10,085\nSCV FFS\n0.717\n6,406\n3,964\n0.418\n8,701\n5,978\n0.054\n15,768\n10,866\n0.107\n14,488\n10,168\nTCV FFS\n0.389\n9,505\n6,723\n0.265\n9,811\n7,057\n0.173\n14,733\n10,391\n0.173\n13,863\n9,959\nSTCV FFS\n0.413\n9,310\n6,575\n0.278\n9,710\n6,980\n0.168\n14,776\n10,410\n0.186\n13,760\n9,868\nNote: Model names indicate the CV strategy used for hyperparameter tuning and whether FFS was\napplied.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "118-0", + "page": 134, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "13,863\n9,959\nSTCV FFS\n0.413\n9,310\n6,575\n0.278\n9,710\n6,980\n0.168\n14,776\n10,410\n0.186\n13,760\n9,868\nNote: Model names indicate the CV strategy used for hyperparameter tuning and whether FFS was\napplied. For example, RCV HP refers to a model using random CV without feature selection. SCV HP,\nTCV HP, and STCV HP use spatial, temporal, and spatiotemporal CV, respectively, also without feature\nselection. SCV FFS, TCV FFS, and STCV FFS apply forward feature selection with spatial, temporal,\nand spatiotemporal CV, respectively.\nAbbreviations: CV = Cross-validation; FFS = Forward feature selection; R2 = Coefficient of\ndetermination; RMSE = Root mean squared error; MAE = Mean absolute error.\nA.2.3\nImpact of outliers on predictions\nTable A-6 compares the performance of four representative models trained under two outlier-handling\nstrategies: capping extreme values at the 95th percentile ( 95) and keeping all values, including spikes\n( noOut).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "118-1", + "page": 134, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ares the performance of four representative models trained under two outlier-handling\nstrategies: capping extreme values at the 95th percentile ( 95) and keeping all values, including spikes\n( noOut). Models are evaluated using R2, RMSE, and MAE on both the training and test sets.\ntest performance (R2, RMSE, MAE) for four models under different\nTable A-6: Training vs.\noutlier-handling strategies and modeling methods.\nTraining Set\nTest Set\nModel Name / Metric\nR2\nR2\nRMSE\nMAE\nRMSE\nMAE\nRCV HP 95\n0.999\n274\n224\n0.728\n5,056\n3,297\nSTCV FFS 95\n0.325\n7,853\n5,939\n0.335\n7,902\n5,968\nRCV HP noOut\n1.000\n466\n379\n0.627\n15,520\n6,462\nSTCV FFS noOut\n0.240\n25,531\n10,599\n0.302\n21,223\n10,273\nNote: Model names reflect the type of CV used during hyperparameter tuning—RCV = Random CV,\nSTCV = Spatiotemporal CV—whether FFS was applied, and the outlier-handling strategy. Models with\nthe 95 suffix (e.g., RCV HP 95, STCV FFS 95) remove the top 5% of extreme UFP values to limit the\ninfluence of spikes.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "118-2", + "page": 134, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "mporal CV—whether FFS was applied, and the outlier-handling strategy. Models with\nthe 95 suffix (e.g., RCV HP 95, STCV FFS 95) remove the top 5% of extreme UFP values to limit the\ninfluence of spikes. Models with noOut (e.g., RCV HP noOut, STCV FFS noOut) retain all UFP values,\nincluding spikes, to evaluate robustness under unfiltered conditions. Specifically, RCV HP 95 and\nRCV HP noOut represent models tuned under random CV without feature selection, while STCV FFS 95\nand STCV FFS noOut use spatiotemporal CV with forward feature selection.\nAbbreviations: CV = Cross-validation; FFS = Forward feature selection; UFP = Ultrafine particles; R2 =\nCoefficient of determination; RMSE = Root mean squared error; MAE = Mean absolute error.\nTable A-7 presents model performance (R2, RMSE, MAE) under four cross-validation strate-\ngies—random, spatial, temporal, and spatiotemporal—for four model configurations with different\noutlier-handling strategies.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "118-3", + "page": 134, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "sents model performance (R2, RMSE, MAE) under four cross-validation strate-\ngies—random, spatial, temporal, and spatiotemporal—for four model configurations with different\noutlier-handling strategies. The table compares RCV HP 95 and STCV FFS 95 (which exclude the\ntop 5% of UFP values) to RCV HP noOut and STCV FFS noOut (which retain all data, including\nextreme spikes).\nFigure A-6 displays scatter plots of predicted versus measured UFP concentrations for four\nmodel frameworks that differ by outlier treatment and the use of cross-validation/feature selection\nstrategies: (a) RCV HP 95, (b) STCV FFS 95, (c) RCV HP noOut, and (d) STCV FFS noOut.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "118-4", + "page": 134, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Performance metrics (R2, RMSE, MAE) across Random, Spatial, Temporal, and\nTable A-7:\nSpatiotemporal CV splits for four models under different outlier-handling strategies and modeling\nmethods.\nRandom Cross Validation\nSpatial Cross Validation\nTemporal Cross Validation\nSpatiotemporal Cross Validation\nModel Name / Metric\nR2\nR2\nR2\nR2\nRMSE\nMAE\nRMSE\nMAE\nRMSE\nMAE\nRMSE\nMAE\n−0.011\n−0.055\nRCV HP 95\n0.726\n4,996\n3,262\n0.303\n7,625\n5,736\n11,888\n9,188\n11,773\n9,305\nSTCV FFS 95\n0.292\n8,033\n6,083\n0.172\n8,288\n6,359\n0.104\n11,220\n8,892\n0.056\n11,167\n8,712\n−0.960\n−0.034\n−0.183\nRCV HP noOut\n0.607\n18,053\n6,375\n27,958\n13,135\n58,600\n24,148\n52,453\n24,493\nSTCV FFS noOut\n0.282\n8,090\n6,059\n0.113\n8,471\n6,383\n0.043\n56,872\n21,723\n0.028\n49,697\n20,893\nNote: Model names reflect the type of CV used during hyperparameter tuning—RCV = Random CV,\nSTCV = Spatiotemporal CV—whether FFS was applied, and the outlier-handling strategy. Models with", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "119-0", + "page": 135, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "697\n20,893\nNote: Model names reflect the type of CV used during hyperparameter tuning—RCV = Random CV,\nSTCV = Spatiotemporal CV—whether FFS was applied, and the outlier-handling strategy. Models with\nthe 95 suffix (e.g., RCV HP 95, STCV FFS 95) remove the top 5% of extreme UFP values to limit the\ninfluence of spikes. Models with noOut (e.g., RCV HP noOut, STCV FFS noOut) retain all UFP values,\nincluding spikes, to evaluate robustness under unfiltered conditions. Specifically, RCV HP 95 and\nRCV HP noOut represent models tuned under random CV without feature selection, while STCV FFS 95\nand STCV FFS noOut use spatiotemporal CV with forward feature selection.\nAbbreviations: CV = Cross-validation; FFS = Forward feature selection; UFP = Ultrafine particles; R2 =\nCoefficient of determination; RMSE = Root mean squared error; MAE = Mean absolute error.\nAs in Figure A-6, each panel shows the predicted UFP values on the x-axis and the corresponding", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "119-1", + "page": 135, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "articles; R2 =\nCoefficient of determination; RMSE = Root mean squared error; MAE = Mean absolute error.\nAs in Figure A-6, each panel shows the predicted UFP values on the x-axis and the corresponding\nmeasured UFP values on the y-axis, with the red dashed line representing perfect 1:1 agreement.\nFigure A-6: Predicted versus actual UFP concentrations on the 20% random holdout test set for the\nfour models—(a) RCV HP 95, (b) STCV FFS 95, (c) RCV HP noOut, and (d) STCV FFS noOut.\nA.2.4\nSelected predictors and development of exposure surfaces\nTo better understand how individual predictors influence model predictions, we conducted SHAP\nanalysis across all seven modeling frameworks. SHAP assigns additive importance values to each", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "119-2", + "page": 135, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "predictor, indicating their influence—positive or negative—on UFP concentration estimates relative\nto a baseline. Figure A-7 presents the SHAP summary plots for all models, including RCV HP 95,\nSTCV FFS 95, RCV HP noOut, and STCV FFS noOut. Models such as RCV HP, SCV HP, TCV HP,\nSTCV HP, RCV HP 95, and RCV HP noOut, which used all 194 features without feature selection,\nshow contributions from many predictors—some of which may be spurious—suggesting potential\noverfitting. In contrast, models using spatiotemporal cross-validation combined with forward feature\nselection such as STCV FFS, STCV FFS 95, and STCV FFS noOut retain fewer, more informative\npredictors, with 17, 20, and 9 features respectively.\nFigure A-7: SHAP summary plots for each of the seven XGBoost models—revealing which features\nmost strongly increase (pink) or decrease (blue) predicted UFP concentrations.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "120-0", + "page": 136, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "A.2.5\nResidual spatial autocorrelation of model errors\nResidual spatial autocorrelation was evaluated using global Moran’s I calculated on model residuals\n(observed −predicted) for both the training set and the hold-out random test set across all modeling\nframeworks. This analysis provides a quantitative diagnostic of remaining spatial dependence.\nFor UFP data collected through mobile monitoring, some residual spatial dependence is expected\nbecause point-level measurements are influenced by short-term transient conditions, such as passing\nhigh-emitting vehicles and localized congestion, while the modeling objective is to estimate average\nexposure surfaces over a period. In this context, moderate residual spatial dependence indicates\nthat transient sampling biases have been averaged rather than reproduced.\nTable A-8 reports global Moran’s I values for RCV, SCV, TCV, and STCV cross-validation\nframeworks, with and without FFS and under different outlier-handling strategies. Moran’s I is", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "121-0", + "page": 137, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "than reproduced.\nTable A-8 reports global Moran’s I values for RCV, SCV, TCV, and STCV cross-validation\nframeworks, with and without FFS and under different outlier-handling strategies. Moran’s I is\ninterpreted based on its magnitude rather than statistical significance, as large sample sizes yield\nsignificant p-values even for weak residual dependence.\nModels trained using RCV generally exhibit very low or near-zero residual Moran’s I values on the\ntest set (e.g., RCV HP and RCV HP 95). This pattern reflects spatial leakage, whereby spatially\nproximate training and test observations allow models to reproduce local patterns and transient\nfluctuations, leading to artificially decorrelated residuals rather than improved spatial realism.\nIn contrast, models trained using spatiotemporally structured cross-validation, particularly when\ncombined with forward feature selection (e.g., STCV FFS), show moderate residual Moran’s I values\non the test set.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "121-1", + "page": 137, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "models trained using spatiotemporally structured cross-validation, particularly when\ncombined with forward feature selection (e.g., STCV FFS), show moderate residual Moran’s I values\non the test set. These values indicate that short-term transient variability has been averaged rather\nthan memorized, while large-scale spatial artifacts associated with overfitting are reduced. The re-\nmaining residual spatial dependence reflects unresolved local variability that is physically consistent\nwith long-term average exposure estimation from mobile monitoring data.\nConsistent with the interpretation proposed by Deppner et al.(2024) [213], Moran’s I is used here\nas a comparative diagnostic of remaining spatial dependence in model residuals, rather than as an\nabsolute indicator of model quality or a strict requirement for spatial independence. In the context\nof mobile monitoring and average exposure estimation, non-zero residual Moran’s I is expected and", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "121-2", + "page": 137, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "absolute indicator of model quality or a strict requirement for spatial independence. In the context\nof mobile monitoring and average exposure estimation, non-zero residual Moran’s I is expected and\nappropriate, as it reflects the removal of transient and sampling-related noise rather than overfitting\nto episodic measurements.\nTable A-8: Global Moran’s I of model residuals (observed −predicted) for training and test sets.\nTest set\nTrain set\nModel\nMoran’s I\np-value\nMoran’s I\np-value\n−0.0866\nRCV HP\n0.0339\n0.007\n0.001\nSCV HP\n0.0633\n0.001\n0.0949\n0.001\nTCV HP\n0.1376\n0.001\n0.2623\n0.001\nSTCV HP\n0.1047\n0.001\n0.2158\n0.001\nSCV FFS\n0.0401\n0.003\n0.0163\n0.003\nTCV FFS\n0.2190\n0.001\n0.3636\n0.001\nSTCV FFS\n0.1996\n0.001\n0.3560\n0.001\n−0.0960\nRCV HP 95\n0.0365\n0.003\n0.001\nSTCV FFS 95\n0.2436\n0.001\n0.3782\n0.001\n−0.0832\nRCV HP noOut\n0.0661\n0.001\n0.001", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "121-3", + "page": 137, + "document_type": "thesis", + "page_header": "APPENDIX A. SUPPLEMENTARY MATERIAL FOR CHAPTER 4", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Appendix B\nSupplementary material for chapter 5\nB.1\nMaterials and methods\nB.1.1\nMobile monitoring and pollutant measurement\nMobile monitoring campaigns were conducted with the UrbanScanner platform, which is a vehicle\nequipped with air pollution sensors and 360-degree imaging equipment. A Global Positioning System\n(GPS) unit was used to geotag all measurements, enabling high-resolution spatial mapping of pollu-\ntant concentrations [196]. UFP concentrations were measured using two nanoparticle counters: the\nDiSCmini (TSI Inc., USA) and the Partector2 (Naneos, Switzerland). Both instruments provided\nparticle number concentrations and were factory-calibrated prior to deployment.\nBC concentra-\ntions were recorded using the AE51 microaethalometer, which operates at a 10-second resolution.\nBC concentrations were recorded using two different microaethalometers: the MA350 (MicroAeth)\nin 2021 and the AE51 (AethLabs) in 2023, both operating at a 10-second resolution. Pollutant", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "122-0", + "page": 138, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "resolution.\nBC concentrations were recorded using two different microaethalometers: the MA350 (MicroAeth)\nin 2021 and the AE51 (AethLabs) in 2023, both operating at a 10-second resolution. Pollutant\nmeasurements were temporally synchronized with GPS data using shared timestamps. The data\nwere spatially joined to 100-meter road segments using a GIS-based approach, and pollutant values\nwere aggregated using the median across all valid passes within each year. The 2023 dataset was\nstratified into wildfire-affected and non-wildfire days. This resulted in segment-level UFP and BC\nconcentrations that form the basis for land use regression modeling.\n360° Imagery and traffic composition extraction\nB.1.2\nDuring both campaigns, the 360° camera was mounted on a Nissan Micra (Figure B-1) and oper-\nated continuously throughout each monitoring run [196]. The camera was configured to capture\nstreet-level panoramic images one image per second throughout the mobile monitoring campaigns.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "122-1", + "page": 138, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": ") and oper-\nated continuously throughout each monitoring run [196]. The camera was configured to capture\nstreet-level panoramic images one image per second throughout the mobile monitoring campaigns.\nAll instruments, including the camera, GPS, and pollutant sensors, were time-synchronized before\ndeployment to ensure accurate alignment across datasets.\nIn order to improve the YOLOv8 model’s ability to capture the diversity of vehicles observed\nin Toronto, the generic “truck” category was subdivided into four distinct classes: pickup, light\ncommercial vehicle, single-unit truck, and combination truck. Figure B-2 provides representative\nexamples of these subcategories, highlighting the differences in vehicle size and form that justify\ntheir separation in the classification scheme.\nTo fine-tune the model, a set of 360° images collected by the UrbanScanner platform was manually\nannotated to create training and validation data. Figure B-3 illustrates this process by showing", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "122-2", + "page": 138, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "heme.\nTo fine-tune the model, a set of 360° images collected by the UrbanScanner platform was manually\nannotated to create training and validation data. Figure B-3 illustrates this process by showing\na raw 360° image alongside its annotated counterpart with bounding boxes for multiple vehicle\nclasses. These annotations were essential for expanding the label set and ensuring that heavy-duty\nvehicles—key contributors to BC and UFP emissions—were accurately distinguished from lighter\nvehicle categories.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "122-3", + "page": 138, + "document_type": "thesis", + "page_header": null, + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure B-1: UrbanScanner vehicle equipped with a rooftop-mounted V.360° camera and air pol-\nlution sensors for synchronized image and pollutant data collection.\nB.1.3\nPredictor variables\nWe used a wide range of predictors to develop the LUR models, capturing land use, road network,\nmeteorological, and temporal characteristics of the study area. A complete set of all predictors is\nsummarized in Table B-1. Most spatial predictors were calculated as continuous or categorical vari-\nables within defined circular buffers surrounding each 100-meter road segment. Land use predictors\nincluded the area, length or count of various urban land use classes such as residential, commercial,\nindustrial and recreational spaces. These predictors captured the spatial distribution of built envi-\nronment features that may influence pollutant concentrations and traffic patterns. These metrics\nwere computed at multiple buffer radii (50, 100, 200, 300, 500, 1000, and 2000 meters) to represent", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "123-0", + "page": 139, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "i-\nronment features that may influence pollutant concentrations and traffic patterns. These metrics\nwere computed at multiple buffer radii (50, 100, 200, 300, 500, 1000, and 2000 meters) to represent\nboth localized and neighborhood-scale effects.\nRoad network predictors accounted for physical infrastructure and traffic-related activity. These\nincluded the total length of highways, arterial and local roads, railways, and bus lines within each\nbuffer. Additional variables captured road surface area, the number of intersections, traffic signals,\nand bus stops. To represent traffic volume, annual average daily traffic (AADT) counts were included\nfor pollutant models. However, AADT was excluded from models predicting vehicle-type-specific\ntraffic counts, as it was reserved for comparing model predictions against an independent traffic\ndataset.\nMeteorological and background pollution variables were obtained from a nearby regulatory sta-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "123-1", + "page": 139, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "raffic counts, as it was reserved for comparing model predictions against an independent traffic\ndataset.\nMeteorological and background pollution variables were obtained from a nearby regulatory sta-\ntion and temporally aligned with each mobile monitoring day. These included hourly measurements\nof PM2.5, NO2 concentrations, wind speed, temperature, and relative humidity. As these variables", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "123-2", + "page": 139, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "(a) Pickups\n(b) Light commercial\n(c) Single-unit truck\n(d) Combination truck\nFigure B-2: Representative examples of truck subtypes used in the expanded YOLOv8 classifica-\ntion: pickup, light commercial vehicle, single-unit truck, and combination truck.\nwere spatially invariant across the study area on each day, they were incorporated without spatial\nbuffering.\nTo account for interannual variability, a dummy variable was introduced to distinguish between\nthe 2021 and 2023 sampling campaigns. This allowed for a pooled modeling framework without the\nneed to train separate models for each year.\nB.1.4\nModel evaluation\nTo assess potential overfitting, we conducted an external evaluation using a randomly selected 20%\nholdout sample from the aggregated dataset. This subset was excluded from model training, hyper-\nparameter tuning, and feature selection. Model performance on the holdout set was evaluated using", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "124-0", + "page": 140, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "20%\nholdout sample from the aggregated dataset. This subset was excluded from model training, hyper-\nparameter tuning, and feature selection. Model performance on the holdout set was evaluated using\nthe coefficient of determination (R2), root mean squared error (RMSE), and mean absolute error\n(MAE) [11].\nThe R2 measures how much of the variation in pollutant concentrations is explained by the\nmodel:", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "124-1", + "page": 140, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "(a)\n(b)\nFigure B-3: Example of training data used for vehicle detection model development. (a) Raw\n360-degree image captured using the V.360° camera mounted on the UrbanScanner vehicle. (b)\nAnnotated version of the same image with ground truth bounding boxes manually labeled for vehicle\nclasses including cars, buses, light commercial vehicles, single-unit trucks, and combination trucks.\nPn\ni=1(yi −ˆyi)2\nR2 = 1 −\n(B.1)\nPn\ni=1(yi −¯y)2\nMAE and RMSE quantify prediction errors in the same units as pollutant concentrations. RMSE\nplaces greater weight on large errors, while MAE provides a more balanced view by being less\ninfluenced by outliers and extreme spikes [140].\nv\nn\nu\nt 1\nX\nu\n(yi −ˆyi)2\nRMSE =\n(B.2)\nn\ni=1\nn\nMAE = 1\nX\n|yi −ˆyi|\n(B.3)\nn\ni=1\nIn 2021, UFP concentrations were continuously recorded at 17 residential backyard locations\nusing DiscMini and Partector sensors during the three-week campaign (From July 9 to 30, 2021).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "125-0", + "page": 141, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "(B.3)\nn\ni=1\nIn 2021, UFP concentrations were continuously recorded at 17 residential backyard locations\nusing DiscMini and Partector sensors during the three-week campaign (From July 9 to 30, 2021).\nMeasurements were restricted to the 8 AM to 7 PM window to align with the mobile campaign. To\ntruly assess model performance, mobile-based LUR model for UFP were externally validated against\nthe data from the stationary backyard sensors.\nWe used Absolute Percentage Error (APE) to measure the difference between predicted and\nobserved backyard sensor concentrations. For each backyard sensor, let yi represent the average\nmeasured concentration during the three-week sampling period, and let ˆyi be the corresponding\nmodel prediction. The APE is then calculated as:\n× 100\nyi −ˆyi\n(B.4)\nAPEi(%) =\nyi\nAnalyzing APE values across backyard sensor sites allowed us to assess the model’s relative\nprediction error and evaluate its robustness in estimating average exposure levels over time.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "125-1", + "page": 141, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Table B-1: Summary of Predictor Variables Used for Land Use Regression Modeling\nAttribute\nType of Measurement\nLand use related variables\nParking lots areas\nArea within buffer per area of buffer\nCommercial areas\nArea within buffer per area of buffer\nGovernmental areas\nArea within buffer per area of buffer\nIndustrial areas\nArea within buffer per area of buffer\nOpen area\nArea within buffer per area of buffer\nResidential area\nArea within buffer per area of buffer\nWaterbody area\nArea within buffer per area of buffer\nParks\nArea within buffer per area of buffer\nDistance to the lake\nClosest distance from Lake Ontario in meters\nPopulation\nPopulation density in each buffer\nDistance to Billy Bishop Toronto City Air-\nClosest distance from Toronto Island Airport\nport\nAccommodation points\nNumber of points in buffers\nChimney\nNumber of points in buffers\nGas station\nNumber of points in buffers\nRestaurant\nNumber of points in buffers\nTraffic related variables\nRoad area", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "126-0", + "page": 142, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Accommodation points\nNumber of points in buffers\nChimney\nNumber of points in buffers\nGas station\nNumber of points in buffers\nRestaurant\nNumber of points in buffers\nTraffic related variables\nRoad area\nArea within buffer per area of buffer\nLength of highway in a buffer\nLength in buffer in meters\nLength of railway in a buffer\nLength in buffer in meters\nLength of major roads in a buffer\nLength in buffer in meters\nLength of bus lines\nLength in buffer in meters\nLength of all roads in a buffer\nLength in buffer in meters\nClosest distance to highway\nClosest distance to nearest highway\nClosest distance to major road\nClosest distance to nearest major road\nClosest distance to rail line\nClosest distance to nearest rail line\nNumber of intersections\nNumber of points in buffers\nNumber of traffic signals\nNumber of points in buffers\nNumber of bus stops\nNumber of points in buffers\nAADT traffic count\nTraffic count per area of buffer\nTemporally changing variables\nPM2.5 and NO2 concentrations", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "126-1", + "page": 142, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ffic signals\nNumber of points in buffers\nNumber of bus stops\nNumber of points in buffers\nAADT traffic count\nTraffic count per area of buffer\nTemporally changing variables\nPM2.5 and NO2 concentrations\nPollutant concentrations from a stationary monitoring\nstation\nWeather-related variables\nWind speed, relative humidity, and tem-\nMeteorological variables from a stationary monitoring\nperature\nstation\nTemporal indicator\nSampling year (2021 vs. 2023)\nBinary variable (1 = 2021, 0 = 2023)\nNote: Predictor variables were grouped into land use, traffic, temporally changing, weather-related, and\ntemporal indicator categories for model development.\nB.2\nResults\nB.2.1\nData summary\nTable B-2 provides detailed descriptive statistics for UFP and BC concentrations by campaign year\nand wildfire status, including the number of segments, number of samples, median, mean, and\nstandard deviation. These values highlight differences between 2021, 2023 non-wildfire days, and\n2023 wildfire-affected days.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "126-2", + "page": 142, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Table B-2: Summary statistics for UFP and BC concentrations by year and wildfire status.\nPollutant\nYear / Status\nN (segments)\nNumber of samples\nMedian\nMean\nStd\nUFP\n2021\n8,419\n440,323\n14,777\n23,906.729\n45,900\nUFP\n2023 wildfire-affected\n7,420\n230,611\n7,355\n14,171.680\n31,328.200\nUFP\n2023 none wildfire\n7,584\n232,881\n16,120\n27,330.539\n48,552.254\nBC\n2021\n7,559\n44,410\n899\n1,428.960\n3,937.358\nBC\n2023 wildfire-affected\n6,094\n23,251\n1,231\n1,579.951\n2,003.630\nBC\n2023 none wildfire\n5,942\n23,002\n956\n1,543.159\n3,469.490\nNote: UFP and BC summary statistics were calculated across road segments for each year and wildfire\ncondition. Concentrations are in particles/cm3 for UFP and ng/m3 for BC.\nUFP concentrations were also recorded at 17 residential backyards using DiscMini and Partector\nsensors during the three-week campaign, restricted to 8 AM–7 PM to align with mobile sampling\n(see Table B-3 for summary statistics).", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "127-0", + "page": 143, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "recorded at 17 residential backyards using DiscMini and Partector\nsensors during the three-week campaign, restricted to 8 AM–7 PM to align with mobile sampling\n(see Table B-3 for summary statistics).\nTable B-3: Summary statistics of UFP concentrations measured by each backyard DiscMini (DM)\nand Partector (PRT) sensor, including count, mean, median, and range of observations.\nUnit\nCount\nMean\nMedian\nStd\nMin\nMax\nDM10\n63,481\n6,210.31\n3,492\n14,526.14\n1,002\n1,104,573\nDM14\n63,292\n8,982.06\n6,238\n41,160.28\n1,398\n4,167,965\nDM15\n63,063\n7,756.52\n4,929\n8,867.13\n1,001\n704,203\nDM16\n63,140\n7,186.12\n5,093\n12,482.03\n1,026\n715,998\nDM17\n62,724\n8,344.16\n5,853\n13,516.68\n1,001\n1,166,068\nDM18\n62,781\n6,889.49\n4,399\n33,584.85\n1,001\n4,398,565\nDM19\n63,143\n5,250.45\n3,201\n22,174.91\n1,001\n2,997,100\nDM22\n61,134\n3,627.11\n2,915\n2,744.29\n1,001\n212,064\nDM3\n63,628\n7,571.47\n3,970\n29,409.06\n1,002\n2,176,635\nDM5\n63,299\n6,795.16\n5,197\n6,170.19\n1,001\n227,475\nDM6\n60,447\n19,545.02\n9,018\n113,663.32\n1,134\n3,706,109\nDM7\n63,306", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "127-1", + "page": 143, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "7.11\n2,915\n2,744.29\n1,001\n212,064\nDM3\n63,628\n7,571.47\n3,970\n29,409.06\n1,002\n2,176,635\nDM5\n63,299\n6,795.16\n5,197\n6,170.19\n1,001\n227,475\nDM6\n60,447\n19,545.02\n9,018\n113,663.32\n1,134\n3,706,109\nDM7\n63,306\n10,845.56\n7,777\n21,069.02\n1,002\n2,184,976\nDM8\n63,212\n10,512.09\n7,803\n8,045.62\n1,225\n184,156\nPRT4\n636,700\n7,764.10\n5,485\n11,787.02\n1,001\n1,615,623\nPRT5\n651,316\n8,320.63\n5,989\n8,737.13\n1,001\n611,245\nPRT6\n643,395\n7,746.33\n5,043\n7,639.47\n1,001\n80,509\nPRT8\n632,728\n7,184.91\n5,363\n7,097.68\n1,001\n784,929\nPRT9\n610,515\n9,082.05\n6,072\n14,374.65\n1,001\n3,386,064\nNote: UFP concentrations are expressed in particles/cm3. Each row represents an individual DiscMini\n(DM) or Partector (PRT) sensor deployed in backyard monitoring sites.\nFigure B-4 shows the normalized confusion matrix for YOLOv8 detections on the test dataset.\nThe matrix highlights per-class accuracy and misclassification tendencies: passenger cars and buses", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "127-2", + "page": 143, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "sites.\nFigure B-4 shows the normalized confusion matrix for YOLOv8 detections on the test dataset.\nThe matrix highlights per-class accuracy and misclassification tendencies: passenger cars and buses\nachieved relatively higher accuracy, while pickups, light commercial vehicles, and single-unit trucks\nshowed greater confusion due to visual similarity. Pedestrians and bicycles were detected less con-\nsistently, reflecting their smaller size and limited presence in the training data. Overall, the matrix\nconfirms that the model provides adequate resolution to differentiate heavy-duty vehicle types, which\nare the most relevant for analyzing traffic-related emissions.\nThe YOLOv8 model was applied to 360° images to detect and count vehicles by class. Table B-\n4 summarizes median, mean, and variability in detected counts for passenger cars, pickups, light\ncommercial vehicles, single-unit trucks, and combination trucks per image for 2021 and 2023 non-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "127-3", + "page": 143, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure B-4: Normalized confusion matrix for YOLOv8 detections on the test set, showing per-class\naccuracy and misclassification patterns.\nwildfire days.\nTable B-4: Summary statistics of detected vehicles by type and year.\nVehicle Type\nYear / Status\nMedian\nMean\nMax\nStd\nPassenger cars\n2021\n5\n5.946\n33\n4.030\nPassenger cars\n2023 none wildfire\n5\n5.748\n34\n4.519\nPickups\n2021\n0\n0.287\n6\n0.599\nPickups\n2023 none wildfire\n0\n0.250\n6\n0.546\nLight commercial vehicles\n2021\n0\n0.167\n7\n0.443\nLight commercial vehicles\n2023 none wildfire\n0\n0.228\n8\n0.513\nSingle-unit trucks\n2021\n0\n0.205\n8\n0.520\nSingle-unit trucks\n2023 none wildfire\n0\n0.267\n9\n0.587\nCombination trucks\n2021\n0\n0.084\n8\n0.386\nCombination trucks\n2023 none wildfire\n0\n0.102\n7\n0.422\nNote: Vehicle counts were extracted from 360° street-level imagery using object-detection models. Values\nrepresent median, mean, and variability (standard deviation) of detected vehicles per road segment.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "128-0", + "page": 144, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "B.2.2\nFeature importance and model interpretation\nTo interpret the models, SHAP values were used to rank predictors by their contribution (Figure B-\n5). A year dummy variable (1 = 2021, 0 = 2023) was included to allow generation of concentration\nsurfaces for each year, though its influence varied across models. After feature selection, only the\nstrongest predictors were retained, ensuring that model interpretation focused on variables with\nconsistent and meaningful contributions. For UFP, the most influential predictors were road area\n(100 m) and AADT (100 m), both directly reflecting local traffic activity. Notably, road area within\n100 m reaches its highest values on highways, highlighting the central role of major corridors in\ndriving UFP levels. Other predictors included traffic lights (100–2000 m), parking areas (2000 m),\nand commercial activity (500 m), with finer contributions from restaurants, bus stops, chimneys,\nrail lines, and gas stations.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "129-0", + "page": 145, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "predictors included traffic lights (100–2000 m), parking areas (2000 m),\nand commercial activity (500 m), with finer contributions from restaurants, bus stops, chimneys,\nrail lines, and gas stations. BC showed a different profile, with highway length, AADT, and road\narea as dominant drivers, consistent with emissions from diesel-heavy fleets.\nFor passenger cars, the most important predictors were road area (50 m), restaurants (100–300\nm), residential area (250 m), and accommodations (2000 m), reflecting their strong association with\nresidential and commercial activity in dense urban areas. Pickups were primarily linked to road area\n(50 m), restaurants (100 m), bus-way length (1000 m), residential area (50 m), and parking area\n(50 m), highlighting their presence in dense road networks, commercial zones, and neighborhood\nactivity. Light commercial vehicles were most influenced by highway length (50 m, 300 m), major-", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "129-1", + "page": 145, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "ng area\n(50 m), highlighting their presence in dense road networks, commercial zones, and neighborhood\nactivity. Light commercial vehicles were most influenced by highway length (50 m, 300 m), major-\nway length (2000 m), total road length (100 m), and traffic lights (100 m), underscoring their reliance\non major corridors. Restaurants (100 m) and bus stops (200 m) also contributed, reflecting their\nactivity in mixed-use urban areas. Single-unit trucks were mainly driven by road area (50–100 m)\nand distance to highways or major roads, confirming their dependence on major corridors. Parking\nareas (50 m), intersections (50–500 m), and industrial land uses (resource area within 2000 m) also\nplayed important roles, reflecting their activity in dense traffic networks and freight-related zones.\nCombination trucks were similarly concentrated along freight corridors, with strong contributions\nfrom road area (50–100 m), major-way length (100 m), bus-way length (50 m), and proximity to\nhighways.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "129-2", + "page": 145, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Combination trucks were similarly concentrated along freight corridors, with strong contributions\nfrom road area (50–100 m), major-way length (100 m), bus-way length (50 m), and proximity to\nhighways. Additional predictors such as intersections (50–2000 m) and industrial land uses (resource\nareas at 1000–2000 m) emphasized their strong links to logistics and industrial activity.\nThese\npatterns highlight a clear stratification in vehicle-type predictors: passenger cars are more strongly\nlinked with residential and commercial density, light- and medium-duty vehicles with local road\nand service infrastructure, and heavy-duty vehicles with major corridors and industrial land uses.\nOverall, these results underscore the central role of traffic corridors and land-use patterns in shaping\nboth pollutant concentrations and vehicle-type-specific traffic distributions, while also highlighting\nthe distinct urban versus freight-related activity captured across different vehicle categories.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "129-3", + "page": 145, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "Figure B-5: SHAP summary plots ranking predictor contributions across pollutants and vehicle\ncategories. a) UFP, b) BC, c) Personal cars, d) Combination trucks, e) Single-unit trucks, f) Light\ncommercial vehicles, g) Pickups.", + "source": "main.pdf", + "file_type": "pdf", + "chunk_id": "130-0", + "page": 146, + "document_type": "thesis", + "page_header": "APPENDIX B. SUPPLEMENTARY MATERIAL FOR CHAPTER 5", + "repository": null, + "relative_path": "main.pdf", + "section": null + }, + { + "text": "# Artificial Neural Networks with PyTorch\n\nWelcome to this repository on **Artificial Neural Networks (ANNs)** using **PyTorch**. This project is a hands-on introduction to neural networks, progressing from fundamental concepts to practical implementations and real-world applications.\n\nThe repository is organized as a series of notebooks, each focusing on a different aspect of building and training neural networks.\n\n---\n\n# Repository Roadmap\n\n## Notebook 1 — ANN Fundamentals with PyTorch\n\n

\n \n

\n\n### Overview\n\nThis notebook introduces the core concepts behind artificial neural networks and demonstrates how to implement them using PyTorch. Starting from the mathematical intuition of neural networks, it covers forward propagation, gradient descent, automatic differentiation, and binary classification.\n\n**Topics Covered**\n\n* Artificial neurons\n* Perceptrons\n* Activation functions\n* Feedforward neural networks", + "source": "README.md", + "file_type": "md", + "chunk_id": "131-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "README.md", + "section": null + }, + { + "text": "rward propagation, gradient descent, automatic differentiation, and binary classification.\n\n**Topics Covered**\n\n* Artificial neurons\n* Perceptrons\n* Activation functions\n* Feedforward neural networks\n* Forward propagation\n* Binary Cross-Entropy Loss\n* Gradient Descent\n* Backpropagation\n* PyTorch tensors\n* Automatic differentiation (Autograd)\n* Binary classification\n* Iris dataset\n* GPU acceleration (CUDA)\n\n**Notebook**\n\n```text\n01_ann_fundamentals_pytorch.ipynb\n```\n\n---\n\n## Notebook 2 — Multi-Class Neural Networks\n\n### Overview\n\nThis notebook demonstrates how to build, train, and evaluate a fully connected neural network for multi-class image classification using the MNIST handwritten digit dataset. It introduces softmax-based classification, multiclass loss functions, model evaluation, and inference on both MNIST samples and external handwritten images.\n\n### Topics Covered\n\n- MNIST dataset preparation\n- Fully connected neural network architecture\n- Softmax activation", + "source": "README.md", + "file_type": "md", + "chunk_id": "131-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "README.md", + "section": null + }, + { + "text": "el evaluation, and inference on both MNIST samples and external handwritten images.\n\n### Topics Covered\n\n- MNIST dataset preparation\n- Fully connected neural network architecture\n- Softmax activation\n- Cross-Entropy Loss\n- Multi-class classification\n- Model training and validation\n- Performance evaluation\n- Inference on external images\n- Image preprocessing for prediction\n\n### Notebook\n\n- `02_multiclass_neural_networks.ipynb`\n\n---\n## Notebook 3 — Neural Network Applications\n\n### Overview\n\nThis notebook explores both the theoretical foundations and practical applications of neural networks. It begins by implementing a two-layer neural network from scratch for multiclass classification using the Iris dataset, and then applies PyTorch to perform binary image classification on the Cars vs. Trucks dataset.\n\n### Classification Tasks\n\n| Task | Dataset |\n|------|---------|\n| 🌸 **Tabular Multiclass Classification** | Iris Dataset |\n| 🚗🚚 **Binary Image Classification** | Cars vs.", + "source": "README.md", + "file_type": "md", + "chunk_id": "131-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "README.md", + "section": null + }, + { + "text": "the Cars vs. Trucks dataset.\n\n### Classification Tasks\n\n| Task | Dataset |\n|------|---------|\n| 🌸 **Tabular Multiclass Classification** | Iris Dataset |\n| 🚗🚚 **Binary Image Classification** | Cars vs. Trucks Dataset |\n\n#### 🌸 Tabular Multiclass Classification (Iris)\n\n

\n \n

\n\n### Topics Covered\n\n- Building a neural network from scratch\n- Forward and backward propagation\n- Gradient verification\n- Iris flower classification (3 classes)\n- Cars vs. Trucks image classification\n- Dataset preprocessing\n- Model training and evaluation\n- Hyperparameter tuning\n- Performance visualization\n- Comparing ANN and CNN models\n\n### Notebook\n\n- `03_neural_network_applications.ipynb`\n\n---\n\n# Technologies\n\n* Python\n* PyTorch\n* NumPy\n* Matplotlib\n* Scikit-learn\n* Jupyter Notebook\n\n---\n\n# Skills Demonstrated\n\n* Deep Learning\n* Artificial Neural Networks\n* PyTorch\n* Machine Learning\n* Binary Classification\n* Multiclass Classification", + "source": "README.md", + "file_type": "md", + "chunk_id": "131-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "README.md", + "section": null + }, + { + "text": "Matplotlib\n* Scikit-learn\n* Jupyter Notebook\n\n---\n\n# Skills Demonstrated\n\n* Deep Learning\n* Artificial Neural Networks\n* PyTorch\n* Machine Learning\n* Binary Classification\n* Multiclass Classification\n* Gradient Descent\n* Backpropagation\n* Automatic Differentiation\n* Model Evaluation\n\n---\n\n## Getting Started\n\nClone the repository:\n\n```bash\ngit clone https://github.com/Miladsaeedi70/artificial-neural-networks-pytorch.git\n```\n\nInstall the required packages manually:\n\n```bash\npip install torch torchvision numpy pandas matplotlib scikit-learn jupyter\n```\n\nLaunch Jupyter Notebook:\n\n```bash\njupyter notebook\n```\n\n---\n\n# Repository Structure\n\n```text\nartificial-neural-networks-pytorch/\n│\n├── notebooks/\n│ ├── 01_ann_fundamentals_pytorch.ipynb\n│ ├── 02_multiclass_neural_networks.ipynb\n│ └── 03_neural_network_applications.ipynb\n│\n├── images/\n│ ├── ann_fundamentals.png\n│ ├── multiclass_classification.png\n��� └── neural_network_applications.png\n│\n├── README.md\n├── requirements.txt\n├── LICENSE", + "source": "README.md", + "file_type": "md", + "chunk_id": "131-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "README.md", + "section": null + }, + { + "text": "3_neural_network_applications.ipynb\n│\n├── images/\n│ ├── ann_fundamentals.png\n│ ├── multiclass_classification.png\n│ └── neural_network_applications.png\n│\n├── README.md\n├── requirements.txt\n├── LICENSE\n└── .gitignore\n```\n\n---\n\n# Future Enhancements\n\n* Convolutional Neural Networks (CNNs)\n* Hyperparameter optimization\n* Regularization techniques\n* Model checkpointing\n* Additional benchmark datasets\n\n---\n\n# References\n\n* PyTorch Documentation\n* Scikit-learn Documentation\n* Fisher's Iris Dataset\n\n---\n\n# License\n\nThis project is licensed under the MIT License.", + "source": "README.md", + "file_type": "md", + "chunk_id": "131-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "README.md", + "section": null + }, + { + "text": "Explanation:\n# Artificial Neural Networks from Scratch and with PyTorch\n\n## Overview\n\nThis notebook demonstrates the implementation of Artificial Neural Networks (ANNs), beginning with manual forward propagation and gradient descent before introducing PyTorch for automatic differentiation and model training.\n\nThe notebook covers:\n\n- Forward propagation\n- Gradient descent\n- Cross-entropy loss\n- Binary classification\n- Multiclass classification\n- The Iris dataset\n- Building neural networks using PyTorch", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "132-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Artificial Neural Networks from Scratch and with PyTorch" + }, + { + "text": "Explanation:\n## Learning Objectives\n\nAfter completing this notebook you will be able to:\n\n- Understand ANN architecture\n- Perform forward propagation\n- Compute gradients using gradient descent\n- Understand cross-entropy loss\n- Build neural networks using PyTorch\n- Train and evaluate a classifier on the Iris dataset\n\nPython implementation:\nimport math\n\ndef sample_forward_pass(x, w):\n y=[]\n #forward pass\n for n in range(len(x)):\n v = 0\n # compute w.x\n for p in range(len(x[0])):\n v = v + x[n][p]*w[p]\n\n #sigmoidal activation\n y.append(1 / (1 + math.e**(-v)))\n\n #output model prediction\n return y\n\nExplanation:\nMake predictions on sample x. Note to keep things simple and not have to track the bias and weights separately, we make the first column as 1's. The goal here is to build a network that can make the correct predictions which are given as: 0, 0, 0, and 1 for the four samples provided.\n\nPython implementation:\n# initial weights\nw = [1, -1, 1]\n\n# data (first column is for the bias term)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "133-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Learning Objectives" + }, + { + "text": "can make the correct predictions which are given as: 0, 0, 0, and 1 for the four samples provided.\n\nPython implementation:\n# initial weights\nw = [1, -1, 1]\n\n# data (first column is for the bias term)\nx = [[1, 0.1,-0.2],\n [1,-0.1, 0.9],\n [1, 1.2, 0.1],\n [1, 1.1, 1.5]]\n\ny = sample_forward_pass(x, w)\n\nExplanation:\nIf we were lucky to select working weights then the forward pass is all that would be required in order to use our model to make predictions.\n\nPython implementation:\n# Magic weights\n\n# w = [ __ , __ , __ ]\n\ny = sample_forward_pass(x, w)\n\nExplanation:\nThe problem is that finding working weights is not trivial. To solve this we use gradient descent (i.e. continuous optimization) to learn a good set of weights. The following examples show implementatoins of gradient descent using a MSE and CrossEntropy loss function.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "133-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Learning Objectives" + }, + { + "text": "Explanation:\n#### Example: 1-layer ANN with MSE and Gradient Descent\n\nPython implementation:\nimport math\n\n# data (first column is the bias term)\nx = [[1, 0.1,-0.2],\n [1,-0.1, 0.9],\n [1, 1.2, 0.1],\n [1, 1.1, 1.5]]\n\n# labels (desired output)\nt = [0, 0, 0, 1]\n\n# initial weights\nw = [1, -1, 1]\n\niterations = 50\nlearning = 10\n\ndef simple_ann(x, w, t, iterations, learning):\n\n E = []\n\n #iterate over epochs\n for ii in range(iterations):\n err = []\n y = []\n #iterate over all the samples x\n for n in range(len(x)):\n v = 0\n # compute w.x\n for p in range(len(x[0])):\n v = v + x[n][p]*w[p]\n\n #sigmoidal activation\n y.append(1 / (1 + math.e**(-v)))\n\n #MSE classification error\n err.append((y[n]-t[n])**2)\n\n #gradient descent to compute new weights\n for p in range(len(w)):\n d = x[n][p]*(y[n]-t[n])*(1-y[n])*(y[n])\n w[p] = w[p] - learning*d\n\n #sum up classification error\n E.append(sum(err)/len(x))\n\n return (y, w, E)\n\n(y, w, E) = simple_ann(x, w, t, iterations, learning)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "134-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Example: 1-layer ANN with MSE and Gradient Descent" + }, + { + "text": "Explanation:\n#### Example: 1-layer ANN with Cross-Entropy and Gradient Descent\n\nPython implementation:\nimport math\n\n# data (first column is the bias term)\nx = [[1, 0.1,-0.2],\n [1,-0.1, 0.9],\n [1, 1.2, 0.1],\n [1, 1.1, 1.5]]\n\n# labels (desired output)\nt = [0, 0, 0, 1]\n\n# initial weights\nw = [1, -1, 1]\n\niterations = 50\nlearning = 10\n\ndef simple_ann(x, w, t, iterations, learning):\n\n E = []\n\n #iterate over epochs\n for ii in range(iterations):\n err = []\n y = []\n\n #iterate over all the samples x\n for n in range(len(x)):\n v = 0\n\n #compute w.x\n for p in range(len(x[0])):\n v = v + x[n][p]*w[p]\n\n #sigmoidal activation\n y.append(1 / (1 + math.e**(-v)))\n\n #cross-entropy classification error\n err.append(-t[n]*math.log(y[n]) - (1-t[n])*math.log(1-y[n]))\n\n #gradient descent to compute new weights\n for p in range(len(w)):\n d = x[n][p]*(y[n]-t[n]) #cross_entropy\n w[p] = w[p] - learning*d\n\n #sum up classification error\n E.append(sum(err))\n\n return (y, w, E)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "135-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Example: 1-layer ANN with Cross-Entropy and Gradient Descent" + }, + { + "text": "adient descent to compute new weights\n for p in range(len(w)):\n d = x[n][p]*(y[n]-t[n]) #cross_entropy\n w[p] = w[p] - learning*d\n\n #sum up classification error\n E.append(sum(err))\n\n return (y, w, E)\n\n(y, w, E) = simple_ann(x, w, t, iterations, learning)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "135-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Example: 1-layer ANN with Cross-Entropy and Gradient Descent" + }, + { + "text": "Explanation:\n### Exploring the Iris dataset\n\nLet us modify the above code to work with the iris data set.\n\nTo begin, load the iris data into Google Colab. If you have difficulty with loading the data it is suggested that you use Chrome.\n\nPython implementation:\n# use sklearn.datasets to load iris data\n\nfrom sklearn.datasets import load_iris\nfeatures, labels = load_iris(return_X_y=True)\n\nExplanation:\nThe iris data has 150 samples spread across three classes:\n1. Iris-setosa,\n2. Iris-versicolor,\n3. Iris-virginica.\n\nThere are three features used:\n1. sepal length in cm\n2. sepal width in cm\n3. petal length in cm\n4. petal width in cm\n\nTo keep things simple, let us pick two of the classes and perform binary classification with our sample code. We will select **Iris-setosa** and **Iris-versicolor** to start:\n\nPython implementation:\n# classification iris-setosa and iris-versicolor\nimport numpy as np\n\nindices = np.array(range(0,100))\n\n#setup x matrix\nx = np.zeros((len(indices), 4 + 1))", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "136-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Exploring the Iris dataset" + }, + { + "text": "rsicolor** to start:\n\nPython implementation:\n# classification iris-setosa and iris-versicolor\nimport numpy as np\n\nindices = np.array(range(0,100))\n\n#setup x matrix\nx = np.zeros((len(indices), 4 + 1))\nx[:,1:5] = features[indices,:]\n\n#add bias column\nx[:,0] = np.ones(len(indices))\n\n#labels\nt = labels[indices]\n\n# initial weights\nw = np.random.rand(5)\n\niterations = 100\nlearning = 0.0001\n\nPython implementation:\nimport math\n\ndef simple_ann(x, w, t, iterations, learning):\n\n E = []\n\n #iterate over epochs\n for ii in range(iterations):\n err = []\n y = []\n\n #iterate over all the samples x\n for n in range(len(x)):\n\n v = 0\n #compute w.x\n for p in range(len(x[0])):\n v = v + x[n,p]*w[p]\n\n #sigmoidal activation\n y.append(1 / (1 + math.e**(-v)))\n\n #cross-entropy classification error\n err.append(-t[n]*math.log(y[n]+ 0.000001) - (1-t[n])*math.log(1-y[n]+ 0.000001))\n\n #gradient descent to compute new weights\n for p in range(len(w)):\n d = x[n][p]*(y[n]-t[n]) #cross_entropy\n w[p] = w[p] - learning*d", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "136-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Exploring the Iris dataset" + }, + { + "text": "*math.log(y[n]+ 0.000001) - (1-t[n])*math.log(1-y[n]+ 0.000001))\n\n #gradient descent to compute new weights\n for p in range(len(w)):\n d = x[n][p]*(y[n]-t[n]) #cross_entropy\n w[p] = w[p] - learning*d\n\n #sum up classification error\n E.append(sum(err))\n\n return (y, w, E)\n\n(y, w, E) = simple_ann(x, w, t, iterations, learning)\n\nExplanation:\nWe're able to successfully classify iris-setosa and iris-versicolor.\n\nPython implementation:\n# classification iris-versicolor and iris virginica\nimport numpy as np\n\nindices = np.array(range(50,150))\n\n#setup x matrix\nx = np.zeros((len(indices), 4 + 1))\nx[:,1:5] = features[indices,:]\n\n#add bias column\nx[:,0] = np.ones(len(indices))\n\n#labels\nt = labels[indices]-1\n\n# initial weights\nw = np.random.rand(5)\n\niterations = 100\nlearning = 0.0001\n\nimport math\n\ndef simple_ann(x, w, t, iterations, learning):\n\n E = []\n\n #iterate over epochs\n for ii in range(iterations):\n err = []\n y = []\n\n #iterate over all the samples x\n for n in range(len(x)):\n\n v = 0\n #compute w.x", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "136-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Exploring the Iris dataset" + }, + { + "text": "le_ann(x, w, t, iterations, learning):\n\n E = []\n\n #iterate over epochs\n for ii in range(iterations):\n err = []\n y = []\n\n #iterate over all the samples x\n for n in range(len(x)):\n\n v = 0\n #compute w.x\n for p in range(len(x[0])):\n v = v + x[n,p]*w[p]\n\n #sigmoidal activation\n y.append(1 / (1 + math.e**(-v)))\n\n #cross-entropy classification error\n err.append(-t[n]*math.log(y[n]+ 0.000001) - (1-t[n])*math.log(1-y[n]+ 0.000001))\n\n #gradient descent to compute new weights\n for p in range(len(w)):\n d = x[n][p]*(y[n]-t[n]) #cross_entropy\n w[p] = w[p] - learning*d\n\n #sum up classification error\n E.append(sum(err))\n\n return (y, w, E)\n\n(y, w, E) = simple_ann(x, w, t, iterations, learning)\n\nExplanation:\nThe performance on the iris-versicolor and iris virginica is not as good. To find out why, we will try to visualize the data.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "136-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Exploring the Iris dataset" + }, + { + "text": "Explanation:\n### Visualize Iris Dataset\nSince the Iris dataset has only 4 inputs we can try to visualize it on a 2-dimensional plane to get a better idea of what is happening.\n\nPython implementation:\n#scatter plot of iris-setosa and iris-versicolor\nimport numpy as np\nfrom matplotlib import pyplot as plt\n\nindices = np.array(range(0,100))\n\nselected_features = features[indices,:]\nselected_labels = labels[indices]\n\nfeature_name = ['sepal length in cm', 'sepal width in cm', 'petal length in cm', 'petal width in cm']\n\nx_index = 0\ny_index = 1\n\nplt.figure(figsize=(5, 4))\nplt.scatter(selected_features[:,x_index], selected_features[:,y_index], c= selected_labels)\nplt.colorbar(ticks=[0, 1, 2])\nplt.xlabel(feature_name[x_index])\nplt.ylabel(feature_name[y_index])\n\nplt.tight_layout()\nplt.show()\n\nPython implementation:\n# scatter plot of iris-versicolor and iris virginica\nimport numpy as np\nfrom matplotlib import pyplot as plt\n\nindices = np.array(range(50,150))\n\nselected_features = features[indices,:]", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "137-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Visualize Iris Dataset" + }, + { + "text": "mplementation:\n# scatter plot of iris-versicolor and iris virginica\nimport numpy as np\nfrom matplotlib import pyplot as plt\n\nindices = np.array(range(50,150))\n\nselected_features = features[indices,:]\nselected_labels = labels[indices]\n\nfeature_name = ['sepal length in cm', 'sepal width in cm', 'petal length in cm', 'petal width in cm']\n\nx_index = 0\ny_index = 1\n\nplt.figure(figsize=(5, 4))\nplt.scatter(selected_features[:,x_index], selected_features[:,y_index], c= selected_labels)\nplt.colorbar(ticks=[0, 1, 2])\nplt.xlabel(feature_name[x_index])\nplt.ylabel(feature_name[y_index])\n\nplt.tight_layout()\nplt.show()", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "137-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Visualize Iris Dataset" + }, + { + "text": "Explanation:\n### Nonlinear Separation\nOur 1-layer ANN was only successful on one of the binary classification combinations. A 1-layer ANN is unable to handle nonlinear separations (or decisions boundaries). To address this we can introduce a second layer known as a hidden layer. How could we do this?\n\nWe can just include an additional 1-layer networks as shown in the image below.\n\n![alt text](https://miro.medium.com/max/1000/1*sX6T0Y4aa3ARh7IBS_sdqw.png)\n\nWe would follow the same process as with the 1-layer network:\n\n1. write out the equations for the forward pass\n2. error term can stay the same MSE or Cross-Entropy\n3. Gradient descent would be applied now to two layers of weights\n\nFirst we would consider the forward pass for a 2-layer ANN. We could use a sigmoidal (logistic) activation function to keep things consistent with our earlier example. Note that the activation function will be applied once on the hidden layer, and also on the output layer.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "138-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Nonlinear Separation" + }, + { + "text": "sigmoidal (logistic) activation function to keep things consistent with our earlier example. Note that the activation function will be applied once on the hidden layer, and also on the output layer.\n\nComputing the gradient with respect to the different layers of weights will become more difficult, but still manageable. There are just some additional terms in the chain rule. The second layer weights will be almost identical to what we computed for a 1-layer network, except the input will be the hidden layer activation.\n\nInstead of spending the time to compute gradients with each change to the network, which can be a fun mathematical exercise, we will instead focus on using the PyTorch libraries which handle all of this internally.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "138-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Nonlinear Separation" + }, + { + "text": "Explanation:\n# Building a Two-Layer Artificial Neural Network\nBuild a 2-layer network using cross-entropy. Determine the gradients with resepect to the layer 1 and layer 2 weights. How could you validate if the gradients were computed correctly?", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "139-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Building a Two-Layer Artificial Neural Network" + }, + { + "text": "Explanation:\n### What is PyTorch?\n\nPyTorch is a scientific computing package that builds on the NumPy library to numerically computer the gradients and allow for the use of GPUs. It incorporates deep learning capabilities while maximizing flexibility and speed.\n\n### PyTorch Basics\n\nTo use PyTorch you must first import the library", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "140-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "What is PyTorch?" + }, + { + "text": "Explanation:\n#### Tensors\nTensors are n-dimensional arrays that allow that can be used with a GPU to accelerate computing. There are several ways to work with Tensors:\n\nPython implementation:\n# initialize a random tensor\nx = torch.rand(4, 3)\nprint(x)\n\nPython implementation:\n# initialize a tensor loaded with zeros\nx = torch.zeros(4, 3)\nprint(x)\n\nPython implementation:\n# initialized with data entered manually\nx = torch.tensor([2.1, 4.0, -5.2])\nprint(x)\n\nPython implementation:\n# initialized with data from numPy\nimport numpy as np\ndata = np.array([2.1, 4.0, -5.2])\nprint(data)\nx = torch.tensor(data)\nprint(x)\n\n# note you can easily convert fron tensor to numpy\nx_np = x.numpy()\nprint(x_np)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "141-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Tensors" + }, + { + "text": "Explanation:\n### Tensor size and shape\n\nPython implementation:\n# obtain size of tensor data structure\nx = torch.zeros(4, 3)\nprint(x.size())", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "142-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Tensor size and shape" + }, + { + "text": "Explanation:\n### Operations\n\nPython implementation:\n# tensor addition\nx = torch.rand(4, 3)\ny = torch.ones(4, 3)\nprint(x + y)\n\nPython implementation:\n# tensor multiplication\nx = torch.rand(4, 3)\ny = torch.ones(4, 3)\nprint(x * y)\n\nPython implementation:\n# Provide output tensor as argument\nresult = torch.ones(4,3)\ntorch.add(x, y, out = result)\nprint(result)\n\nPython implementation:\n# resize and reshape tensors\nx = torch.randn(4, 3)\ny = x.view(12)\nz = x.view(-1, 4) # the size -1 is inferred from other dimensions\nprint(x.size(), y.size(), z.size())\n\nPython implementation:\n# convert one element tensor to a Python number\nx = torch.randn(1)\nprint(x)\nprint(x.item())", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "143-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Operations" + }, + { + "text": "Explanation:\n### Automatic Differentiation\n\nThe PyTorch autograd package allows for easy computation of derivative. This is handled automatically using a define-by-run framework, which works as you write your code.\n\nTo enable this feature you need to set the Tensor attribute .required_grat to True, at which point it begins to track all operations performed on it. After your computation have been completed you can call .backward() and all the gradients will be computed for you. The gradient for each tensor will be stored in the tensor attribute .grad.\n\nEach tensor also has a .grad_fn attribute which references Function that has created it.\n\nPython implementation:\n# Example computation of gradients\nx = torch.rand(4, 3, requires_grad=True)\nprint(x)\n\n#perform some operations\ny = x + 10\nz = y*y\nout = z.mean()\n\nprint(z, out)\n\nPython implementation:\n# compute the gradient with respect to the output\nout.backward()\nprint(x.grad)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "144-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Automatic Differentiation" + }, + { + "text": "Explanation:\n## PyTorch Simple Neural Networks\n\nThe following is an example of a 1-layer and 2-layer neural network using PyTorch.\n\nExplanation:\nPyTorch - 1-layer neural network\n\nPython implementation:\nimport torch\n\nx = torch.ones(5) # input tensor\ny = torch.zeros(3) # expected output\n\nw = torch.randn(5, 3, requires_grad=True)\nb = torch.randn(3, requires_grad=True)\nz = torch.matmul(x, w)+b\n\n# assumes binary output\nloss = torch.nn.functional.binary_cross_entropy_with_logits(z, y)\n\nExplanation:\nobtain gradients for 1-layer neural network\n\nExplanation:\nPyTorch - 2-layer neural network\n\nPython implementation:\n#2-layer neural network\nimport torch\n\nnum_hidden = 3\nx = torch.ones(5) # input tensor\ny = torch.zeros(3) # expected output\n\n# layer 1\nw = torch.randn(5, num_hidden, requires_grad=True)\nb = torch.randn(num_hidden, requires_grad=True)\nz = torch.matmul(x, w)+b\nz = torch.sigmoid(z)\n\n# layer 2\nw2 = torch.randn(num_hidden, 3, requires_grad=True)\nb2 = torch.randn(3, requires_grad=True)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "145-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "PyTorch Simple Neural Networks" + }, + { + "text": "e)\nb = torch.randn(num_hidden, requires_grad=True)\nz = torch.matmul(x, w)+b\nz = torch.sigmoid(z)\n\n# layer 2\nw2 = torch.randn(num_hidden, 3, requires_grad=True)\nb2 = torch.randn(3, requires_grad=True)\nz2 = torch.matmul(z, w2)+b2\n\n# assumes binary output\nloss = torch.nn.functional.binary_cross_entropy_with_logits(z2, y)\n\nExplanation:\nobtain gradients for 2-layer neural network\n\nExplanation:\nPyTorch computational graphs make gradient calculations stright forward. Much of this is hidden away allowing you to focus more on developing and testing your model architectures.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "145-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "PyTorch Simple Neural Networks" + }, + { + "text": "Explanation:\n## Artificial Neural Networks in PyTorch\nIn this example we will train an \"artificial pigeon\" to perform a digit recognition\ntask. That is, we will use the MNIST dataset of hand-written digits, and train\nthe pigeon to **recognize a small digit, namely a digit that is less than 3**.\nThis problem is a **binary classification problem** we want to predict\nwhich of two classes an input image is a part of.\n\n### 1) Load MNIST Data\nThe MNIST dataset contains hand-written digits that are 28x28 pixels large.\nHere are a few digits in the dataset:\n\nPython implementation:\nfrom torchvision import datasets, transforms\n\n# load the data\nmnist_train = datasets.MNIST('data', train=True, download=True)\nmnist_train = list(mnist_train)[:2000]\n\nPython implementation:\n# plot the first 18 images in the training data\nfor k, (image, label) in enumerate(mnist_train[:18]):\n plt.subplot(3, 6, k+1)\n plt.imshow(image)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "146-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Artificial Neural Networks in PyTorch" + }, + { + "text": "Explanation:\n### 2) Defining the ANN Forward Pass\nHere is an implementation of the artificial pigeon brain in PyTorch.\nDon't worry if this code or the explanations don't make sense yet.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torchvision import datasets, transforms\nimport matplotlib.pyplot as plt # for plotting\n\nimport torch.optim as optim\n\ntorch.manual_seed(1) # set the random seed\n\nclass Pigeon(nn.Module):\n def __init__(self):\n super(Pigeon, self).__init__()\n self.layer1 = nn.Linear(28 * 28, 30)\n self.layer2 = nn.Linear(30, 1)\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = self.layer1(flattened)\n activation1 = F.relu(activation1)\n activation2 = self.layer2(activation1)\n return activation2\n\npigeon = Pigeon()\n\nExplanation:\nIn this network, there are 28x28 = 784 input neurons, to work with our 28x28 pixel images. We have a single output neuron and a hidden layer of 30 neurons.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "147-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "2) Defining the ANN Forward Pass" + }, + { + "text": "tivation2\n\npigeon = Pigeon()\n\nExplanation:\nIn this network, there are 28x28 = 784 input neurons, to work with our 28x28 pixel images. We have a single output neuron and a hidden layer of 30 neurons.\n\nThe variable `pigeon.layer1` contains information about the connectivity\nbetween the input layer and the hidden layer (stored as a matrix), and the\nbiases (stored as a vector).\n\nSimilarly, the variable `pigeon.layer2` contains information about the weights\nbetween the hidden layer and the output layer, and the bias.\n\nThe weights and biases adjust during training, so they are called the model's\n**parameters**.\n\nPython implementation:\n# view parameters\nfor w in pigeon.layer1.parameters():\n print(w)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "147-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "2) Defining the ANN Forward Pass" + }, + { + "text": "Explanation:\n### 3) Test Network Forward Pass\nHere is an example of using the network to classify whether the\nimage contains a small digit.\n\nPython implementation:\n# make predictions for the first 10 images in mnist_train\nimg_to_tensor = transforms.ToTensor() #transform the image data into a 28x28 matrix of numbers\n\nfor k, (image, label) in enumerate(mnist_train[:10]):\n inval = img_to_tensor(image)\n outval = pigeon(inval) # find the output activation given input\n prob = torch.sigmoid(outval) # turn the activation into a probability\n print(prob)\n\nExplanation:\nSince we haven't trained the network\nyet, the predicted probability of images containing a small digit\nis close to half. The \"pigeon\" is unsure.\n\nIn order for the network to be useful, we need to actually train it, so\nthat the weights are actually meaningful, non-random values.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "148-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "3) Test Network Forward Pass" + }, + { + "text": "Explanation:\n### 4) Updating Network Parameters with Gradient Descent and Cross-Entropy Loss\n\nTo train the neural network, we first perform a forward pass to generate predictions for the input data. These predictions are then compared with the true labels using a loss function.\n\nThe loss measures how closely the predicted outputs match the ground truth. In this example, we use **binary cross-entropy**, a standard loss function for binary classification problems.\n\nTraining a neural network can be formulated as an optimization problem: the objective is to find the set of model parameters (weights and biases) that minimizes the loss over the training data.\n\nTo accomplish this, we use **stochastic gradient descent (SGD)**, an optimization algorithm that iteratively updates the model parameters in the direction that reduces the loss. During each iteration, gradients are computed through backpropagation and used to improve the network's predictions.\n\nPython implementation:", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "149-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "4) Updating Network Parameters with Gradient Descent and Cross-Entropy Loss" + }, + { + "text": "e model parameters in the direction that reduces the loss. During each iteration, gradients are computed through backpropagation and used to improve the network's predictions.\n\nPython implementation:\n# simplified training code to train `pigeon` on the \"small digit recognition\" task\nimport torch.optim as optim\ncriterion = nn.BCEWithLogitsLoss()\noptimizer = optim.SGD(pigeon.parameters(), lr=0.005, momentum=0.9)\n\nExplanation:\nNow, we can start to train the pigeon network, similar to the way we would train\na real pigeon:\n\n1. We'll show the network pictures of digits, one by one\n2. We'll see what the network predicts\n3. We'll check the loss function for that example digit, comparing the network prediction against the ground truth\n4. We'll make a small update to the parameters to try and improve the loss for that digit\n5. We'll continue doing this many times -- let's say 1000 times\n\nFor simplicity, we'll use 1000 images, and show the network each image only once.\n\nPython implementation:", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "149-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "4) Updating Network Parameters with Gradient Descent and Cross-Entropy Loss" + }, + { + "text": "rove the loss for that digit\n5. We'll continue doing this many times -- let's say 1000 times\n\nFor simplicity, we'll use 1000 images, and show the network each image only once.\n\nPython implementation:\nfor (image, label) in mnist_train[:1000]:\n # actual ground truth: is the digit less than 3?\n actual = torch.tensor(label < 3).reshape([1,1]).type(torch.FloatTensor)\n # pigeon prediction\n out = pigeon(img_to_tensor(image)) # step 1-2\n # update the parameters based on the loss\n loss = criterion(out, actual) # step 3\n loss.backward() # step 4 (compute the updates for each parameter)\n optimizer.step() # step 4 (make the updates for each parameter)\n optimizer.zero_grad() # a clean up step for PyTorch\n\nExplanation:\nIt is very common to run into errors with changing different data types to tensors with the correct shape.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "149-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "4) Updating Network Parameters with Gradient Descent and Cross-Entropy Loss" + }, + { + "text": "Explanation:\n### 5) Test Updated Network Forward Pass\n\nPython implementation:\n# make predictions for the first 10 images in mnist_train\nfor k, (image, label) in enumerate(mnist_train[:10]):\n print(label, torch.sigmoid(pigeon(img_to_tensor(image))))\n\nExplanation:\nNot bad! We'll use the probability 50% as the cutoff for making a\ndiscrete prediction. Then, we can compute the accuracy on the 1000\nimages we used to train the network.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "150-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "5) Test Updated Network Forward Pass" + }, + { + "text": "Explanation:\n### 6)a) Validation of Network on Training Data\n\nPython implementation:\n# computing the error and accuracy on the training set\nerror = 0\nfor (image, label) in mnist_train[:1000]:\n prob = torch.sigmoid(pigeon(img_to_tensor(image)))\n if (prob < 0.5 and label < 3) or (prob >= 0.5 and label >= 3):\n error += 1\nprint(\"Training Error Rate:\", error/1000)\nprint(\"Training Accuracy:\", 1 - error/1000)\n\nExplanation:\nThe accuracy on those 1000 images is 96%, which is really good considering\nthat we only showed the network each image only once.\n\nHowever, this accuracy is not representative of how well the network is doing,\nbecause the network was *trained* on the data. The network had a chance to\nsee the actual answer, and learn from that answer. To get a better sense of\nthe network's predictive accuracy, we should compute accuracy numbers on\na **test set**: a set of images that were not seen in training.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "151-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "6)a) Validation of Network on Training Data" + }, + { + "text": "Explanation:\n### 6)b) Validation of Network on Testing Data\n\nPython implementation:\n# computing the error and accuracy on a test set\nerror = 0\nfor (image, label) in mnist_train[1000:2000]:\n prob = torch.sigmoid(pigeon(img_to_tensor(image)))\n if (prob < 0.5 and label < 3) or (prob >= 0.5 and label >= 3):\n error += 1\nprint(\"Test Error Rate:\", error/1000)\nprint(\"Test Accuracy:\", 1 - error/1000)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "152-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "6)b) Validation of Network on Testing Data" + }, + { + "text": "Explanation:\n### Putting it all together\n\nPython implementation:\n# putting the above code together gives us...\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nimport matplotlib.pyplot as plt # for plotting\nimport torch.optim as optim #for gradient descent\n\ntorch.manual_seed(1) # set the random seed\n\n# obtain data\nfrom torchvision import datasets, transforms\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)\nmnist_train = mnist_data[:4096]\nmnist_val = mnist_data[4096:5120]\n\nPython implementation:\n#Artificial Neural Network Architecture (aka MLP)\nclass MNISTClassifier(nn.Module):\n def __init__(self):\n super(MNISTClassifier, self).__init__()\n self.fc1 = nn.Linear(28 * 28, 50)\n self.fc2 = nn.Linear(50, 20)\n self.fc3 = nn.Linear(20, 10)\n\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = F.relu(self.fc1(flattened))\n activation2 = F.relu(self.fc2(activation1))", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "153-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Putting it all together" + }, + { + "text": "= nn.Linear(50, 20)\n self.fc3 = nn.Linear(20, 10)\n\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = F.relu(self.fc1(flattened))\n activation2 = F.relu(self.fc2(activation1))\n output = self.fc3(activation2)\n return output\n\nPython implementation:\ndef get_accuracy(model, train=False):\n if train:\n data = mnist_train\n else:\n data = mnist_val\n\n correct = 0\n total = 0\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n\n output = model(imgs)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total\n\nPython implementation:\ndef train(model, data, batch_size=64, num_epochs=1):\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "153-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Putting it all together" + }, + { + "text": "der(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n for epoch in range(num_epochs):\n for imgs, labels in iter(train_loader):\n\n out = model(imgs) # forward pass\n\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model, train=True)) # compute training accuracy\n val_acc.append(get_accuracy(model, train=False)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n plt.show()", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "153-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Putting it all together" + }, + { + "text": "(model, train=False)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\nPython implementation:\nmodel = MNISTClassifier()\n\n#proper model\ntrain(model, mnist_train, num_epochs=5)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "153-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Putting it all together" + }, + { + "text": "Explanation:\n## Enable GPU\nPyTorch allows you to run the computations on a GPU to speed up the processing. In order to enable GPUs you will need to:\n1. select GPUs in \"Notebook Settings\" found under the \"Edit\" menu option.\n2. setup model to work with the cuda\n3. make sure image and labels data are stored placed on the GPU\n\nAn example of this is provided below.\n\nPython implementation:\ndef get_accuracy(model, train=False):\n if train:\n data = mnist_train\n else:\n data = mnist_val\n\n correct = 0\n total = 0\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n output = model(imgs)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "154-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Enable GPU" + }, + { + "text": "odel(imgs)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total\n\nPython implementation:\ndef train(model, data, batch_size=64, num_epochs=1):\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n for epoch in range(num_epochs):\n for imgs, labels in iter(train_loader):\n\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "154-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Enable GPU" + }, + { + "text": "#############################################\n\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model, train=True)) # compute training accuracy\n val_acc.append(get_accuracy(model, train=False)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "154-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Enable GPU" + }, + { + "text": "aining Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\nPython implementation:\nuse_cuda = True\n\nmodel = MNISTClassifier()\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\n#proper model\ntrain(model, mnist_train, num_epochs=5)", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "154-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Enable GPU" + }, + { + "text": "Explanation:\n# Conclusions\n\nIn this notebook we implemented artificial neural networks both manually and using PyTorch.\n\nKey concepts covered include:\n\n- Forward propagation\n- Gradient descent\n- Cross-entropy loss\n- Automatic differentiation\n- Binary and multiclass classification\n- Neural network implementation in PyTorch\n\nThese concepts form the foundation for training deeper neural networks on more complex datasets.", + "source": "01_ann_fundamentals_pytorch.ipynb", + "file_type": "ipynb", + "chunk_id": "155-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/01_ann_fundamentals_pytorch.ipynb", + "section": "Conclusions" + }, + { + "text": "Explanation:\n## PreLab 1B - Multi-Class ANNs", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "156-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "PreLab 1B - Multi-Class ANNs" + }, + { + "text": "Explanation:\n# Multi-Class Image Classification with Artificial Neural Networks (PyTorch)\n\nThis notebook demonstrates how to build, train, and evaluate a fully connected Artificial Neural Network (ANN) for multi-class image classification using the MNIST handwritten digit dataset. The workflow covers model design, training, performance evaluation, and inference on external images.\n\n### Multi-Class Classification Pipeline\n\n```text\nMNIST Images\n │\n ▼\nData Preprocessing\n │\n ▼\nFlatten Images\n(28 × 28 → 784)\n │\n ▼\nFully Connected ANN\n │\n ▼\nOutput Layer\n(10 Classes)\n │\n ▼\nSoftmax Probabilities\n │\n ▼\nPredicted Digit\n```", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "157-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Multi-Class Image Classification with Artificial Neural Networks (PyTorch)" + }, + { + "text": "Explanation:\n## Dataset Preparation\n\nThe MNIST dataset contains 70,000 grayscale images of handwritten digits (0–9), each with a resolution of 28 × 28 pixels. In this section, the dataset is downloaded, transformed into PyTorch tensors, and prepared for model training and evaluation.\n\nPython implementation:\n# obtain data\nfrom torchvision import datasets, transforms\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)\nmnist_train = mnist_data[:4096]\nmnist_val = mnist_data[4096:5120]", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "158-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Dataset Preparation" + }, + { + "text": "Explanation:\n### Multi-Class ANN Architecture\nIn this example we will be using a 3-layer ANN with ReLU activation functions applied on the first and second hidden layers. The softmax activation will be used for outputting class probabilities and is not included in the architecture setup.\n\nA fully connected neural network is implemented to classify handwritten digits into one of ten classes. The model receives a flattened image as input and produces class probabilities through the final output layer.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nimport matplotlib.pyplot as plt # for plotting\nimport torch.optim as optim #for gradient descent\n\ntorch.manual_seed(1) # set the random seed\n\nclass MNISTClassifier(nn.Module):\n def __init__(self):\n super(MNISTClassifier, self).__init__()\n self.layer1 = nn.Linear(28 * 28, 50)\n self.layer2 = nn.Linear(50, 20)\n self.layer3 = nn.Linear(20, 10)\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "159-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Multi-Class ANN Architecture" + }, + { + "text": "r(MNISTClassifier, self).__init__()\n self.layer1 = nn.Linear(28 * 28, 50)\n self.layer2 = nn.Linear(50, 20)\n self.layer3 = nn.Linear(20, 10)\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = F.relu(self.layer1(flattened))\n activation2 = F.relu(self.layer2(activation1))\n output = self.layer3(activation2)\n return output\n\nmodel = MNISTClassifier()\n\nprint('done')", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "159-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Multi-Class ANN Architecture" + }, + { + "text": "Explanation:\n## Model Training\nThis section defines the training and validation workflow, including the accuracy metric, optimization procedure, and loss computation. The model parameters are updated using backpropagation and gradient descent while performance is monitored on both the training and validation datasets.", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "160-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Model Training" + }, + { + "text": "Explanation:\n### Function to Obtain Accuracy\nThe get_accuracy function is used to compute the accuracy on training or validation data.\n\nPython implementation:\ndef get_accuracy(model, train=False):\n if train:\n data = mnist_train\n else:\n data = mnist_val\n\n correct = 0\n total = 0\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n output = model(imgs)\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total\n\nprint ('done')", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "161-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Function to Obtain Accuracy" + }, + { + "text": "Explanation:\n### Function to perform Training and Validation\nThe train function puts everything together. You can provide arguments to adjust the batch size and number of training epochs.\n\nPython implementation:\ndef train(model, data, batch_size=64, num_epochs=1):\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n for epoch in range(num_epochs):\n for imgs, labels in iter(train_loader):\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "162-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Function to perform Training and Validation" + }, + { + "text": "ates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model, train=True)) # compute training accuracy\n val_acc.append(get_accuracy(model, train=False)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\nprint('done')", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "162-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Function to perform Training and Validation" + }, + { + "text": "Explanation:\n### Sanity Check\nVerify that the model is able to overfit on a single batch of data.\n\nPython implementation:\n#overfitting the model (sanity check)\ndebug_data = mnist_train[:64]\nmodel = MNISTClassifier()\ntrain(model, debug_data, num_epochs=500)\n\n#obtain accuracy on 64 samples\ncorrect = 0\ntotal = 0\nfor imgs, labels in torch.utils.data.DataLoader(debug_data, batch_size=64):\n output = model(imgs)\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\nprint('Accuracy on batch of 64: ', correct / total)\n\nExplanation:\nNote the Final Training and Validation accuracy is obtained on the full training data and validation data. It does not reflect the performance on the 64 samples that were overfit.", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "163-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Sanity Check" + }, + { + "text": "Explanation:\n### Run Training and Validation\nNow that we've validated that our model can overfit a relatively small amount of training data (i.e. 64 samples), we can proceed to train our model on all of the training data.\n\nWe will be training our model over 5 epochs (how many training iterations is that?) to ensure that we can complete this tutorial in a reasonable time. In your free time you are welcome to explore the model accuracy as you increase the number of epochs.\n\nPython implementation:\n#proper model\nmodel = MNISTClassifier()\ntrain(model, mnist_train, num_epochs=5)", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "164-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Run Training and Validation" + }, + { + "text": "Explanation:\n### Continuing Training\nAt this stage we can consider adjusting our model architecture:\n- number of hidden layers,\n- hidden units,\n- activation functions,\n- optimizers,\n- learning rate,\n- momentum,\n- batch size,\n- training iterations,\n\nto evaluate the performance of our ANN model. Can we do better?\n\n**Tip:** Once you have searched through the hyperparameters and obtained model parameters that work reasonably well. You may want to save your model so that you don't have to retrain the model next time you open the Google Colab file.\n\nPython implementation:\n# save the model for next time\ntorch.save(model.state_dict(), \"saved_model\")", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "165-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Continuing Training" + }, + { + "text": "Explanation:\n### Test one image\nAt this point we have trained our model and observed accuracy scores on the training data and validation data. We haven't really looked at the data. For the next stage of the tutorial we will try to understand what the data looks like and consider what is required to classify new images obtained from the internet, or even our cell phone camera.\n\nPython implementation:\n#laod new image for testing\nmnist_sample = mnist_data[19120] #samples with indices > 5120 can be used for testing\nimg, label = list(mnist_sample) #obtain a single image and label\n\n#plot sample image\nprint('image dimensions: ', img.shape)\nplt.imshow(img.view(-1,28)) #make image 28 x 28 (not 1 x 28 x 28 as required by model)\n\n#test new image\nout = model(img)\nprob = F.softmax(out, dim=1)\nprint('output dimensions: ', out.shape)\nprint('output probabilities: ', prob, 'sum: ', torch.sum(prob))\n\n#print max index and compare with label", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "166-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Test one image" + }, + { + "text": "ew image\nout = model(img)\nprob = F.softmax(out, dim=1)\nprint('output dimensions: ', out.shape)\nprint('output probabilities: ', prob, 'sum: ', torch.sum(prob))\n\n#print max index and compare with label\nprint('output: ', prob.max(1, keepdim=True)[1].item(), 'with a probability of', prob.max(1, keepdim=True)[0].item())\nprint('label: ', label)", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "166-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Test one image" + }, + { + "text": "Explanation:\n### Exploring the MNIST data\nBefore we can load new data for testing we should understand what preprocessing went into making the training data. We will explore:\n- Data Type\n- Data Dimensions\n- Data Normalization\n- Orientation\n\nPython implementation:\n# max and min values\nprint('min val:', torch.min(img).item(), ' max val:', torch.max(img).item())\n\n#histogram of values\nplt.hist(img.view(-1,28*28), bins=30)\nplt.show()\n\nExplanation:\nNow that we know a little bit about our data we should be able to generate new samples of the data.", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "167-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Exploring the MNIST data" + }, + { + "text": "Explanation:\n### Load New Image\nIn the Google Colab environment there are a number of ways to load data samples. If you are using Chrome or Chromium you should be able to just load the data into the workspace using the following code. Alternatively, you can always read an image posted online.\n\nPython implementation:\n#optional for laoding data from a file\nfrom google.colab import files\nimg_new = files.upload()\n\nPython implementation:\n#loading data from the internet\nimg_new = plt.imread('https://www.researchgate.net/profile/Hariton_Costin/publication/311806756/figure/fig1/AS:542753920229376@1506414026147/Sample-of-the-MNIST-dataset-of-handwritten-digits.png')", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "168-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Load New Image" + }, + { + "text": "Explanation:\n### Examine the image\nHow does the image we loaded differ from the one in the MNIST dataset?\n\nPython implementation:\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n# convert from colour to grayscale\ndef rgb2gray(rgb):\n return np.dot(rgb[...,:3], [0.299, 0.587, 0.144])\n\nimg_gray = rgb2gray(img_new)\n\nplt.title(\"New Image\")\nplt.imshow(img_gray)\nplt.show()\n\n# compare to original MNIST image\nplt.title(\"Original MNIST\")\nplt.imshow(img.view(-1,28))\n\nExplanation:\nNotice that the image colours are inverated? How will this affect our classification on new data?", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "169-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Examine the image" + }, + { + "text": "Explanation:\n### Inverting Colours\nOne option is to just load an image with the colours inverted.\n\nPython implementation:\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n#load an image with black and white matching the MNIST data\nimg_new = plt.imread('https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcT5ZR8ImWkYVd2FRMZgUvCdNkHx0uKjjSAtTEJ0U-x0SPWQFxqnbg')\nprint('Image Dimensions', img_new.shape)\n\ndef rgb2gray(rgb):\n return np.dot(rgb[...,:3], [0.299, 0.587, 0.144])\n\nimg_gray = rgb2gray(img_new)\n\nplt.title(\"New Image\")\nplt.imshow(img_gray)\nplt.show()\n\n# compare to original MNIST image\nplt.title(\"Original MNIST\")\nplt.imshow(img.view(-1,28))", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "170-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Inverting Colours" + }, + { + "text": "Explanation:\n### Cropping the Image\nThe images used to train our model were centered on the handwritten digit and resized to 28 x 28 pixesl. We will need to do the same to our new data in order for it work with our Multi-Class ANN model.\n\nPython implementation:\n#cropping\nimg_cropped = img_gray[95:120,5:33]\nplt.title(\"Image Cropped\")\nplt.imshow(img_cropped)\nplt.show()\n\n#resize image\nfrom skimage.transform import rescale, resize, downscale_local_mean\nimg_resized = resize(img_cropped, (28,28), anti_aliasing=True)\n\n#plot resized image\nplt.title(\"Image Resized\")\nplt.imshow(img_resized)\nplt.show()\n\n#image max and min values\nprint(np.amax(img_resized))\nprint(np.amin(img_resized))\n\n#normalize to range 0 to 1\nimg_resized = img_resized / np.amax(img_resized)\nplt.title(\"Image Normalized\")\nplt.imshow(img_resized)\nplt.show()\n\n#verify max and min values\nprint(np.amax(img_resized))\nprint(np.amin(img_resized))\n\nExplanation:\nIf required, how could you invert the colours?", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "171-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Cropping the Image" + }, + { + "text": "Explanation:\n### Testing a New External Image\n\nPython implementation:\n#test new external image\n\n#plot resized image\nplt.title(\"New Image\")\nplt.imshow(img_resized)\nplt.show()\n\n#convert image to torch tensor\nimg_new = torch.tensor(img_resized)\nprint('Initial Dimensions: ', img_new.shape)\n\n#make our image match the model dimensions 1 x 28 x 28 and tensor type\nimg_new = img_new.unsqueeze(0).type(torch.FloatTensor)\nprint('Updated Dimensions: ', img_new.shape)\n\n#perform forward pass on ANN model and generate an output\nout = model(img_new)\nprob = F.softmax(out, dim=1)\n\n#examine output properties\nprint('output dimensions: ', out.shape)\nprint('output probabilities: ', prob, 'sum: ', torch.sum(prob))\n\n#print max index\nprint('Predicted Output: ', prob.max(1, keepdim=True)[1].item(), 'with a probability of', prob.max(1, keepdim=True)[0].item())", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "172-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Testing a New External Image" + }, + { + "text": "Explanation:\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Build a fully connected neural network in PyTorch.\n- Train a model for multi-class image classification.\n- Evaluate model performance using accuracy.\n- Perform inference on unseen images.\n- Preprocess external images for deployment.\n- Analyze prediction errors using a confusion matrix.", + "source": "02_multiclass_neural_networks.ipynb", + "file_type": "ipynb", + "chunk_id": "173-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/02_multiclass_neural_networks.ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Neural Network Applications with PyTorch\n\nThis notebook explores practical neural network applications using PyTorch. It begins by implementing a two-layer neural network from scratch to understand the fundamentals of forward and backward propagation, and then applies deep learning to a real-world image classification task using the Cars vs. Trucks dataset.\n\n$\\color{blue}{\\text{Name: Milad Saeedi}}$", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "174-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Neural Network Applications with PyTorch" + }, + { + "text": "Explanation:\n## Overview\n\nThis notebook demonstrates both the theoretical foundations and practical applications of Artificial Neural Networks (ANNs). The first part focuses on implementing a neural network from scratch for multiclass classification, while the second part develops, trains, and evaluates image classification models using PyTorch.\n\nExplanation:\n#PART A: Building a Neural Network from Scratch", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "175-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Overview" + }, + { + "text": "Explanation:\nBefore we get into using PyTorch to train our classifier we will go through the process of creating our neural network from scratch.\n\n## Part 0. Helper Functions\n\nTo help guide the construction we will use the iris dataset. Provided are some helper code to get us started:\n\nPython implementation:\nimport pandas as pd\n\ndata = pd.read_csv(\"Iris_3class.csv\")\n\nPython implementation:\nimport pandas as pd\nimport numpy as np\n\nraw_data = pd.read_csv(\"Iris_3class.csv\", header = None)\nraw_data = raw_data.values\nnp.random.shuffle(raw_data)\n\nPython implementation:\nimport numpy as np\n#raw_data = raw_data.values\n\n# split your data into training and validation\nX_train = raw_data[0:100,:4]\ny_train = raw_data[0:100,4:5].astype(int)\nX_val = raw_data[100:,:4]\ny_val = raw_data[100:,4:5].astype(int)\n\nprint(X_train.shape, y_train.shape)\nprint(X_train.dtype, y_train.dtype)\nprint(X_val.shape, y_val.shape)\nprint(X_val.dtype, y_val.dtype)\n\nExplanation:", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "176-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "[100:,:4]\ny_val = raw_data[100:,4:5].astype(int)\n\nprint(X_train.shape, y_train.shape)\nprint(X_train.dtype, y_train.dtype)\nprint(X_val.shape, y_val.shape)\nprint(X_val.dtype, y_val.dtype)\n\nExplanation:\nSince the labels are provided as integers we will need to convert them into one-hot vectors to match the neural network output format.\n\nPython implementation:\n#Convert array to one-hot encoding\ndef to_one_hot(Y):\n n_col = np.amax(Y) + 1\n binarized = np.zeros((len(Y), n_col))\n for i in range(len(Y)):\n binarized[i, Y[i]] = 1.\n return binarized\n\nPython implementation:\ny_train = to_one_hot(y_train)\nprint(X_train.shape, y_train.shape)\nprint(X_train.dtype, y_train.dtype)\n\ny_val = to_one_hot(y_val)\nprint(X_val.shape, y_val.shape)\nprint(X_val.dtype, y_val.dtype)\n\nPython implementation:\n#verify one-hot encoding\ny_train[0:5,:]", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "176-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "Explanation:\n## Part 1. Building the Network\nAt its core a 2-layer neural network is just a few lines of code. Most of the complexity comes from setting up the training of the network.\n\nUsing vectorized form, set up the neural network training to use a cross-entropy loss function and determine the gradients with resepect to the layer 1 and layer 2 weights.\n\nPython implementation:\n# write code to create a 2-layer ANN in vectorized form\n\n#define sigmoid\ndef sigmoid(x):\n return 1/(1+np.exp(-x))\n\n#define softmax\ndef softmax(x):\n e = np.exp(x)\n return e/e.sum(axis=1, keepdims = True)\n\ndef ann(W, X_train, y_train):\n\n num_hidden = 5\n num_features = 4\n num_outputs = 3\n\n #Weights\n w0 = W[:20].reshape(num_features, num_hidden)\n w1 = W[20:].reshape(num_hidden, num_outputs)\n\n #Feed forward\n layer0 = X_train\n layer1 = sigmoid(np.dot(layer0, w0))\n layer2 = np.dot(layer1, w1)\n\n # softmax\n output = softmax(layer2)\n\n #Back propagation using gradient descent\n\n #cross-entropy loss", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "177-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 1. Building the Network" + }, + { + "text": "eed forward\n layer0 = X_train\n layer1 = sigmoid(np.dot(layer0, w0))\n layer2 = np.dot(layer1, w1)\n\n # softmax\n output = softmax(layer2)\n\n #Back propagation using gradient descent\n\n #cross-entropy loss\n error =np.sum(-y_train*np.log(output))# TO BE COMPLETED\n\n #initialize gradients to zero\n dw0 =np.zeros((4,5)) # TO BE COMPLETED\n dw1 =np.zeros((5,3)) # TO BE COMPLETED\n\n #calculate gradients\n # TO BE COMPLETED\n dL_dz = output-y_train\n du_dv_hat = w1.T\n dv_hat_dv = layer1*(1-layer1)\n dv_dw0 = X_train\n dz_dw1 = layer1\n\n #determine gradients\n dw1 +=dz_dw1.T.dot(dL_dz) # TO BE COMPLETED\n dw0 +=dv_dw0.T.dot((dL_dz).dot(du_dv_hat)*(dv_hat_dv)) # TO BE COMPLETED\n\n #combine gradients into one vector\n dW =np.array(list(dw0.flatten()) + list(dw1.flatten())) # TO BE COMPLETED\n\n return (error, dW, output)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "177-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 1. Building the Network" + }, + { + "text": "Explanation:\n## Part 2. Training the Network\n\nPython implementation:\nnum_hidden = 5\nnum_features = 4\nnum_outputs = 3\n\n#initialize weights\nw0 = 2*np.random.random((num_features, num_hidden)) - 1\nw1 = 2*np.random.random((num_hidden, num_outputs)) - 1\n\n#combine weights into a single vector\nW = np.array(list(w0.flatten()) + list(w1.flatten()))\n\n#train network\nn = 0.001\niterations = 100000\nerrors = []\nfor i in range(iterations):\n (error, dW, y_pred) = ann(W, X_train, y_train)\n W += -dW * n\n errors.append(error)\n\nPython implementation:\n#examine predictions on training data\n(_, _, y_pred) = ann(W, X_train, y_train)\npred = np.round(y_pred, 0)\npred[:5]\n\nPython implementation:\n#examine ground truth training data\ntrain = np.round(y_train, 0)\ntrain[:5]\n\nExplanation:\nIn the next line, all rows in training set have been compared :\n\nBased on the results, we can observe that 100 of the 100 data points in the training set were accurately calculated.", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "178-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 2. Training the Network" + }, + { + "text": "line, all rows in training set have been compared :\n\nBased on the results, we can observe that 100 of the 100 data points in the training set were accurately calculated.\n\nPython implementation:\naccuracy=0\nfor i in range(len(train)):\n if (train[i:i+1]==pred[i:i+1]).all():\n accuracy +=1\n\nprint(\"Number of correct predictions in training set=\", accuracy)\n\nExplanation:\nIn the following line, the model's performance for the validation data set has been assessed:\n\nBased on the results, we can observe that 49 of the 50 data points in the validation data set were accurately calculated.\n\nPython implementation:\n(_, _, y_pred_val) = ann(W, X_val, y_val)\npred_val = np.round(y_pred_val, 0)\n\nval = np.round(y_val, 0)\nN_accuracy=0\nfor i in range(len(val)):\n if (val[i:i+1]==pred_val[i:i+1]).all():\n N_accuracy +=1\n\nprint('Number of correct predictions in validation set=', N_accuracy)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "178-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 2. Training the Network" + }, + { + "text": "Explanation:\n## Part 3. Gradient Verification\nTo verify the correctness of the implementation, the analytical gradients computed during backpropagation are compared against numerical approximations. This validation step confirms that the gradient calculations are implemented correctly.\n\n$\\color{blue}{\\text{ }}$\nThe numerical values of gradients are the same as our gradients. As a result, we can conclude that they are correct.\n\nPython implementation:\n#write code to numerical verify the gradients you calculated\n\nnum_hidden = 5\nnum_features = 4\nnum_outputs = 3\n\n#initialize weights\nw0 = 2*np.random.random((num_features, num_hidden)) - 1\nw1 = 2*np.random.random((num_hidden, num_outputs)) - 1\n\n#combine weights\nW =np.array(list(w0.flatten()) + list(w1.flatten())) # TO BE COMPLETED\n\n#compute gradients analytically\n(error, dW, y_pred) = ann(W, X_train, y_train)\n\n#compute gradients numerically\ndW_num = np.zeros((len(W),1))\n\nfor ind in range(len(W)):\n #reset gradients", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "179-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 3. Gradient Verification" + }, + { + "text": "BE COMPLETED\n\n#compute gradients analytically\n(error, dW, y_pred) = ann(W, X_train, y_train)\n\n#compute gradients numerically\ndW_num = np.zeros((len(W),1))\n\nfor ind in range(len(W)):\n #reset gradients\n We1 = np.array(list(w0.flatten()) + list(w1.flatten()))\n We2 = np.array(list(w0.flatten()) + list(w1.flatten()))\n\n #increment slightly\n We1[ind] =We1[ind] + 0.000001 # TO BE COMPLETED\n We2[ind] =We2[ind] - 0.000001 # TO BE COMPLETED\n\n #compute errors\n (error_e1, dW_e1, y_pred1) = ann(We1, X_train, y_train)\n (error_e2, dW_e2, y_pred2) = ann(We2, X_train, y_train)\n\n #obtain numerical gradients\n grad_num = (error_e1-error_e2)/0.000002# TO BE COMPLETED\n\n #display difference between numerical and analytic gradients\n print(round(abs(grad_num - dW[ind]), 4), grad_num, dW[ind])", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "179-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 3. Gradient Verification" + }, + { + "text": "Explanation:\n# PART B: Image Classification with PyTorch\n\nIn the second part, we will see how we can use PyTorch to train a neural network to identify Cars and Trucks.\n\nPython implementation:\nimport numpy as np\nimport time\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport torchvision\nfrom torch.utils.data.sampler import SubsetRandomSampler\nimport torchvision.transforms as transforms", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "180-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "PART B: Image Classification with PyTorch" + }, + { + "text": "Explanation:\n## Part 0. Helper Functions\n\nWe will be making use of the following helper functions. You will be asked to look\nat and possibly modify some of these, but you are not expected to understand all of them.\n\nPython implementation:\n###############################################################################\n# Data Loading\n\ndef get_relevant_indices(dataset, classes, target_classes):\n \"\"\" Return the indices for datapoints in the dataset that belongs to the\n desired target classes, a subset of all possible classes.\n\n Args:\n dataset: Dataset object\n classes: A list of strings denoting the name of each class\n target_classes: A list of strings denoting the name of desired classes\n Should be a subset of the 'classes'\n Returns:\n indices: list of indices that have labels corresponding to one of the\n target classes\n \"\"\"\n indices = []\n for i in range(len(dataset)):\n # Check if the label is in the target classes\n label_index = dataset[i][1] # ex: 9", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "ices that have labels corresponding to one of the\n target classes\n \"\"\"\n indices = []\n for i in range(len(dataset)):\n # Check if the label is in the target classes\n label_index = dataset[i][1] # ex: 9\n label_class = classes[label_index] # ex: 'truck'\n if label_class in target_classes:\n indices.append(i)\n return indices\n\ndef get_data_loader(target_classes, batch_size):\n \"\"\" Loads images of cars and trucks, splits the data into training, validation\n and testing datasets. Returns data loaders for the three preprocessed datasets.\n\n Args:\n target_classes: A list of strings denoting the name of the desired\n classes. Should be a subset of the argument 'classes'\n batch_size: A int representing the number of samples per batch\n\n Returns:\n train_loader: iterable training dataset organized according to batch size\n val_loader: iterable validation dataset organized according to batch size\n test_loader: iterable testing dataset organized according to batch size", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "training dataset organized according to batch size\n val_loader: iterable validation dataset organized according to batch size\n test_loader: iterable testing dataset organized according to batch size\n classes: A list of strings denoting the name of each class\n \"\"\"\n\n classes = ('plane', 'car', 'bird', 'cat',\n 'deer', 'dog', 'frog', 'horse', 'ship', 'truck')\n ########################################################################\n # The output of torchvision datasets are PILImage images of range [0, 1].\n # We transform them to Tensors of normalized range [-1, 1].\n transform = transforms.Compose(\n [transforms.ToTensor(),\n transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5))])\n # Load CIFAR10 training data\n trainset = torchvision.datasets.CIFAR10(root='./data', train=True,\n download=True, transform=transform)\n # Get the list of indices to sample from\n relevant_indices = get_relevant_indices(trainset, classes, target_classes)\n\n # Split into train and validation", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "in=True,\n download=True, transform=transform)\n # Get the list of indices to sample from\n relevant_indices = get_relevant_indices(trainset, classes, target_classes)\n\n # Split into train and validation\n np.random.seed(1000) # Fixed numpy random seed for reproducible shuffling\n np.random.shuffle(relevant_indices)\n split = int(len(relevant_indices) * 0.8) #split at 80%\n\n # split into training and validation indices\n relevant_train_indices, relevant_val_indices = relevant_indices[:split], relevant_indices[split:]\n train_sampler = SubsetRandomSampler(relevant_train_indices)\n train_loader = torch.utils.data.DataLoader(trainset, batch_size=batch_size,\n num_workers=1, sampler=train_sampler)\n val_sampler = SubsetRandomSampler(relevant_val_indices)\n val_loader = torch.utils.data.DataLoader(trainset, batch_size=batch_size,\n num_workers=1, sampler=val_sampler)\n # Load CIFAR10 testing data\n testset = torchvision.datasets.CIFAR10(root='./data', train=False,\n download=True, transform=transform)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "nset, batch_size=batch_size,\n num_workers=1, sampler=val_sampler)\n # Load CIFAR10 testing data\n testset = torchvision.datasets.CIFAR10(root='./data', train=False,\n download=True, transform=transform)\n # Get the list of indices to sample from\n relevant_test_indices = get_relevant_indices(testset, classes, target_classes)\n test_sampler = SubsetRandomSampler(relevant_test_indices)\n test_loader = torch.utils.data.DataLoader(testset, batch_size=batch_size,\n num_workers=1, sampler=test_sampler)\n return train_loader, val_loader, test_loader, classes\n\n###############################################################################\n# Training\ndef get_model_name(name, batch_size, learning_rate, epoch):\n \"\"\" Generate a name for the model consisting of all the hyperparameter values\n\n Args:\n config: Configuration object containing the hyperparameters\n Returns:\n path: A string with the hyperparameter name and value concatenated\n \"\"\"\n path = \"model_{0}_bs{1}_lr{2}_epoch{3}\".format(name,\n batch_size,", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "nfiguration object containing the hyperparameters\n Returns:\n path: A string with the hyperparameter name and value concatenated\n \"\"\"\n path = \"model_{0}_bs{1}_lr{2}_epoch{3}\".format(name,\n batch_size,\n learning_rate,\n epoch)\n return path\n\ndef normalize_label(labels):\n \"\"\"\n Given a tensor containing 2 possible values, normalize this to 0/1\n\n Args:\n labels: a 1D tensor containing two possible scalar values\n Returns:\n A tensor normalize to 0/1 value\n \"\"\"\n max_val = torch.max(labels)\n min_val = torch.min(labels)\n norm_labels = (labels - min_val)/(max_val - min_val)\n return norm_labels\n\ndef evaluate(net, loader, criterion):\n \"\"\" Evaluate the network on the validation set.\n\n Args:\n net: PyTorch neural network object\n loader: PyTorch data loader for the validation set\n criterion: The loss function\n Returns:\n err: A scalar for the avg classification error over the validation set\n loss: A scalar for the average loss function over the validation set\n \"\"\"\n total_loss = 0.0\n total_err = 0.0", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "unction\n Returns:\n err: A scalar for the avg classification error over the validation set\n loss: A scalar for the average loss function over the validation set\n \"\"\"\n total_loss = 0.0\n total_err = 0.0\n total_epoch = 0\n for i, data in enumerate(loader, 0):\n inputs, labels = data\n labels = normalize_label(labels) # Convert labels to 0/1\n outputs = net(inputs)\n loss = criterion(outputs, labels.float())\n corr = (outputs > 0.0).squeeze().long() != labels\n total_err += int(corr.sum())\n total_loss += loss.item()\n total_epoch += len(labels)\n err = float(total_err) / total_epoch\n loss = float(total_loss) / (i + 1)\n return err, loss\n\n###############################################################################\n# Training Curve\ndef plot_training_curve(path):\n \"\"\" Plots the training curve for a model run, given the csv files\n containing the train/validation error/loss.\n\n Args:\n path: The base path of the csv files produced during training\n \"\"\"\n import matplotlib.pyplot as plt", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "ng curve for a model run, given the csv files\n containing the train/validation error/loss.\n\n Args:\n path: The base path of the csv files produced during training\n \"\"\"\n import matplotlib.pyplot as plt\n train_err = np.loadtxt(\"{}_train_err.csv\".format(path))\n val_err = np.loadtxt(\"{}_val_err.csv\".format(path))\n train_loss = np.loadtxt(\"{}_train_loss.csv\".format(path))\n val_loss = np.loadtxt(\"{}_val_loss.csv\".format(path))\n plt.title(\"Train vs Validation Error\")\n n = len(train_err) # number of epochs\n plt.plot(range(1,n+1), train_err, label=\"Train\")\n plt.plot(range(1,n+1), val_err, label=\"Validation\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"Error\")\n plt.legend(loc='best')\n plt.show()\n plt.title(\"Train vs Validation Loss\")\n plt.plot(range(1,n+1), train_loss, label=\"Train\")\n plt.plot(range(1,n+1), val_loss, label=\"Validation\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"Loss\")\n plt.legend(loc='best')\n plt.show()", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "181-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 0. Helper Functions" + }, + { + "text": "Explanation:\n## Part 1. Exploring the Dataset\nWe will make use of some of the CIFAR-10 data set, which consists of\ncolour images of size 32x32 pixels belonging to 10 categories. You can\nfind out more about the dataset at https://www.cs.toronto.edu/~kriz/cifar.html\n\nFor this assignment, we will only be using the car and truck categories.\nWe have included code that automatically downloads the dataset the\nfirst time that the main script is run.\n\nPython implementation:\n# This will download the CIFAR-10 dataset to a folder called \"data\"\n# the first time you run this code.\ntrain_loader, val_loader, test_loader, classes = get_data_loader(\n target_classes=[\"car\", \"truck\"],\n batch_size=1) # One image per batch", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "182-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 1. Exploring the Dataset" + }, + { + "text": "Explanation:\n### Part (a)\n\nVisualize some of the data by running the code below.\nInclude the visualization in your writeup.\n\n(You don't need to submit anything else.)\n\n$\\color{blue}{\\text{ }}$ - The image of some of data has been added to this markdown: \n\n![Pic1.png](data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAV0AAADhCAYAAABr92YNAAAABHNCSVQICAgIfAhkiAAAAAlwSFlzAAALEgAACxIB0t1+/AAAADh0RVh0U29mdHdhcmUAbWF0cGxvdGxpYiB2ZXJzaW9uMy4yLjIsIGh0dHA6Ly9tYXRwbG90bGliLm9yZy+WH4yJAAAgAElEQVR4nOy9Wa9kWZbn9Vt7OufYdCd3D4/IiMzKGrK6qlF10w8FQjy0hPgkvPEBeOhXJF75ODwgBIimKTUIJKDVXdmVlWR2ZAwePtx7bTjDnnjY+5jZ9Yjw8BspJQ/4kszNr9mxY8f2sNZ//ddwJOfMB/kgH+SDfJA/jKj/ry/gg3yQD/JB/v8kH5TuB/kgH+SD/AHlg9L9IB/kg3yQP6B8ULof5IN8kA/yB5QPSveDfJAP8kH+gPJB6X6QD/JBPsgfUMy73vzH//Gf5OVywx//8T/kMOx5c/+K++037A63bJYLnDUIQoyR7XZLjJGYAilCjrBeXdA2Hcv1Cm0cyqzo2g03N5/y/PnHfPqTz/hf/uZ/4O9+9a+JeYuoyHLVEkLg9vYOyRkNhDARY+Ty6oZFt+Lj53/Ek5vn/Ad//U+5uX7KJx//lH/7d/+Wv/vV3/Gv/tX/ztcvPud+94KUE01zibUtXXdJCCPTNDCMAyFMhODJOfE", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "3bJYLnDUIQoyR7XZLjJGYAilCjrBeXdA2Hcv1Cm0cyqzo2g03N5/y/PnHfPqTz/hf/uZ/4O9+9a+JeYuoyHLVEkLg9vYOyRkNhDARY+Ty6oZFt+Lj53/Ek5vn/Ad//U+5uX7KJx//lH/7d/+Wv/vV3/Gv/tX/ztcvPud+94KUE01zibUtXXdJCCPTNDCMAyFMhODJOfE3/83fyPsO2D/9L/9lBsgpQ/3UnHYn5R2Y0/AyzC9BJnP8yPEYmY958Ho++/Dp/OUPeesD9dz1c/n0oeNhuX7rt7IDj8fkb733L//r//S9xwTgn/1X/yyHuOMw/ZrFOnL1JHJ5dcVqtWaYIjFB2zmM0Sw6CzkRw4Q1BmMMzy4/pmuWvPjqltevb/mf/vm/YPQDWXm0zmiTmMZACBGjEm2r+ct/7wmLpWGz1mjVYdWKxj3BmEt+/esXvHmz58WLHj8lfND0w8Dt3R3LVcdq1fEf/fV/wh999qf89Pk/oGvWNHpFjuDHgNIaMZohBno/8d/9z/8jL159w3/xn/3n7z0u//2L+GBYRcrj+DcZoawjobwp5Y2zY8pxdcKOb4rIW+cqb8vZx+fPqgffKySOK+y4bmaZL1gqHjutX45r5Xhs/ec/vDbvPSZXm5tMhlTX7LxqH565/D6UQhBE9PHCUopkMnI2SEopjLP88T/4c/7i3//HXD59wmK15HDYkmKibVtiCGzvXhOTJ8UR2zRY51he3mBdg9Ia4xounzwn54T3I7/527/lt7/8W/7N//Z/8c0XX5dryHXvn12x1Iktc6IB4fbNl987Ju9UuilOxDgyTQe0ylxsloS4w4cDxmq00SgRkIwooF5MzomUMjF6fNSMfsBKZr3YsN50fPbTj7m8uGK5atBGSERAyFkxjp4UQx3SRMqRmCZC8gzDlpwDd3cv0Vrx4pvfgWQur67ZH7Zst/fc77Zs93tCTIgCqxWSI8P+Fh8mpmlgnAZC8MQYyDm973qpMyyc1Fj", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Nl987Ju9UuilOxDgyTQe0ylxsloS4w4cDxmq00SgRkIwooF5MzomUMjF6fNSMfsBKZr3YsN50fPbTj7m8uGK5atBGSERAyFkxjp4UQx3SRMqRmCZC8gzDlpwDd3cv0Vrx4pvfgWQur67ZH7Zst/fc77Zs93tCTIgCqxWSI8P+Fh8mpmlgnAZC8MQYyDm973qpMyyc1FjdBsJxBUrODzZNmZX6dFxV9Zi6h+RMQR7fPzvPvGFyBpFcFO/85VXRitTn+sbx47lsvEzZpPnsGrLUy8i89d7j5c3uS2Ia6Ic9U85kk+mnl7jbWwYvZDQXF5c4Z1mPDa0zLLuOrrV0rSX4Hdtxy939a/aHPauNZpFbbLdE6YRSkeEw4QePn3qcg+XKsVhq2kWG7MlxSzYOMYJt9tjmwN3d5+x2nsMhoZTGWkfXWFaLBOmWcXjBdn9JCAdSk8lZEWNGK4PFsRsO3O0OfP67L/ni668eNSZanxSXvDWXJ8WYz9576836YTVP5DxRPF7pnp9/Xnon3VHWTs7z+hGUUvNbD+Q77PbjRUDOFp9Us1IMz/zbBNEGpRRaa3LKpJQIoTxzPL7+ypCYhoHd/R3rywuUUjS2IelIDJHgPSlGyKmOVS7KtT+Qgse6Bq001lrG8cDh/jXj4Y5p2JNJKK0hVIWv5O0tepqD95AfULqe4CcOhy2utSy6lt3OYLUmxciUIkoK0k0pMtvP2aKnFIlhwntFJjOMB8bxwDDu2fcau9WEMBXFrQxIJARPiqkMbE5ITlUxZmIKhDCxP2yx1vHm9iXWOl6/ecnt3Rvut3d4P5FypliBjA8TkhUpCiF6vB9JMZBTLDZfHrdqylosiyNnjsulTKIcldxxIo7/rRN1dvxJ8db/HIHM22j2BDXk7M/5+bRoy4sZOUPdD5HuOeI5nV7KhpPM+y+dh9JPr0nJ42NATZnDoY4Rgo+KLIpp0giW3CwgGbQ0kCBOkFImx4yiwZrMxeYJPnmSTKA", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "/ecnt3Rvut3d4P5FypliBjA8TkhUpCiF6vB9JMZBTLDZfHrdqylosiyNnjsulTKIcldxxIo7/rRN1dvxJ8db/HIHM22j2BDXk7M/5+bRoy4sZOUPdD5HuOeI5nV7KhpPM+y+dh9JPr0nJ42NATZnDoY4Rgo+KLIpp0giW3CwgGbQ0kCBOkFImx4yiwZrMxeYJPnmSTKACqIBrG4wFM40Yo0BflXVjNGSNKAOyItFi3ArXCChNZiLEESKE1NNNkRAzo79lGFfc775inPaEkMhJGIeE0gZtHEPIjD4w+gOTPzxqTI5jfZyn0/M8zae1MhtzecuzmY/jbPLlgXKdj/v2zOUH3/mt6zu7uBnwHr87nxmBB5/5Nip+rDzEHvWM9QeKCMYalFYY4xAliNLEECAEJM5m4fRcdHcmp0xOiRgDMRQlW/TSyVDNylIpjTa2IF7bsFhf4NoOYwzTkBn7PX4cyMEjOaNEyEf0A1lyMQzztTxi27xT6cY4MU0H7u9fcmWuWa+v2O4adjuN9wMxBgBS/aH5TOEqLcTkwSfyEFFhAtEYq7nfviLEgWnqGacepQStHZnI/jASYzhTuvGodEOYyDlxd/+KmCJffv05IUWaxYKvv/mCl69fMExDVS4KcmYcB3KEFAryjimQq4OlfgSjLeq0jCU/dMvl6Kl9B2w8V5A5PzgmV2v/bdrgdL4ZxZ6+4/itb503P9jIx8v4jnPPKJfjhvtxChdg139ZvK8kJBQ56/oNmoAGMQy9Q1JL7i4Qa9E0ZD8yTSMpJHICJSu6ZsHTp45hOnC3f0UST1Ye2yxQ4gi+oMisN0TtELeArCEZskAEms6wCC3K/IasemIeCWHC9wNNN7DwE4fxa7Z98bSsXdKPW0KA/TYABqUcpt2AsozTPZPf/oiRyQ8MnZwZPjkztKf/52rAzxXvA36Ao4H+3p1e3d75W2dbeg7N8tn8U/6TU6p/54Iusny/Lvl9FG7+ro9XF10pXNtirKF", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "3f0UST1Ye2yxQ4gi+oMisN0TtELeArCEZskAEms6wCC3K/IasemIeCWHC9wNNN7DwE4fxa7Z98bSsXdKPW0KA/TYABqUcpt2AsozTPZPf/oiRyQ8MnZwZPjkztKf/52rAzxXvA36Ao4H+3p1e3d75W2dbeg7N8tn8U/6TU6p/54Iusny/Lvl9FG7+ro9XF10pXNtirKFpGqovwDiOFUydzM25Uci5eNg5RaKf8ONI9L4gaa2L4lRCqsBIaYN1LV23ou0WXD39BOMcxjkOKdHf3zEd9kQ/UnSFkGZsJCelr848jvfdPz9ILwQP+90tIhGlE/d3t4zjgA8DKcXjZo8pnAaijkuMnpQCYjJZMqM/sNsbvn7xOcvFmtVqAyTW6w1ta0k5Mk4DOSdEVFUeCqX00R2IMZJ0JISJ+/s3iAjaOF6+fkk/HAghkKoyzIBKVTEKiFKoyrlk0o9CdlorjlzojCLrF8zKVPLZed/WvUcFe0IYks92wvHpbEGR6zG5Itizz5+dW3jrc9VdnOnnNPPInBb/zCPmowH5cYrX2Y5cKSVFJvgJo1uWS4MyLaIskjVGCTlm+v3AdLjHGI/VnhTLuGw2n+CaBjHPGSaPmB3aKnSjsW6FNi1at2htWK4LR2wbRwoQPATfk+NATAeUHvjJpx0Xl1tevfqS/f6ON7dfYlwDusE0C+yiJetAlIEhvMFPiX0/kHMxHEt1jW06njzTmGbxqDFRMm/EOu5HN/80X2+DSQXEnEgx4X2JOSzaDtGqgsF5u+dvOWkiJyU5qyZ12o71UycKKoTqpks5t1KKlIrSVxyZtLMvOJ4E4MgNP0rOUH+eV+LZdVMBm1aCUZBzJMSEIqK1oLQu++y4lk8DGGPEj2OlETKzBkkpEmMgTL6e37G5eMLFk4+4uHpC2y1YrNcA9MOu0hGRFNNxDxW9phBVFLAoQWt9vOgUYzVaP7x/3q10U4AAfb8jE0ESu8OOcRyJcaqKqyiblGJFuYp", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "qpks5t1KKlIrSVxyZtLMvOJ4E4MgNP0rOUH+eV+LZdVMBm1aCUZBzJMSEIqK1oLQu++y4lk8DGGPEj2OlETKzBkkpEmMgTL6e37G5eMLFk4+4uHpC2y1YrNcA9MOu0hGRFNNxDxW9phBVFLAoQWt9vOgUYzVaP7x/3q10U4AAfb8jE0ESu8OOcRyJcaqKqyiblGJFuYp5dlIqF22SgpTxfqDvFa/fvGAce7yfgMxyuWC5WpBS4uWrr8vGVUW5kTSiIpJUVQzFZfBhYre/J5FJotjeHxinnhBjcVUBciZmKrooA6XQZ3P/SD4XUEpOzMBbaOGIQuVstX4neq3mshrubwe43t4kMwqWtxDKDBlOCODBPnmgqEHlt67r6BWVzfad1MN7ijNtNcKRTCTFgNHQtRrXNihlmQaNQkGEKXiG/o7GBZwLBeUqw9VVR9NcYN0Vjc8kGXBtS7NYYLs1xpWgqNb2ODoJhSeRfSTEN6RpS0oHlJp4+uyS9WaHNpfc3r3gMEZMI4gRVNOim6YYcyamuGMKgX7akaKQomCXoJslV9eapuseNSZyXAdn5vAMtB6V4bm3JJBTIsWAn0ZijHSuQZQ6IuKjb3OMws6vvIVwOf+eM9qAYmRjDIhIDbRVHlfKmCoR1LcWr1Q64BQUPie2fpzMwUM5rkelVAmOqUxMCZJHJBelq4SsNDnF0ylmI1ANVQoRUjrtoZRIMRJDqAEzx2J1xdXNx1zdPKPpFlhnCWGqwbdICvGov8rl1ZgIgjKqnMea4/eH6Ywj/wF5p9IlhYI+x4T3A/v9lpCK0hOdEVW0TM5FwarZNGZVlVwZVKQQ0CEOpDHg3wT6fk9/OHB5+YTlYs1isSTFhDUNMZQFl1Iq1ECuiq4S7NPkSann5asX2N09r+/vyEmTkyakUONMGkShy0UUpJcSOYeKggVBw7cW1rulGAM4xl2r0jt30+aN9HZ47Py1o7V+6PGdfb4cV5jpM3+sHq/FoyRgpRy", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "9lpCK0hOdEVW0TM5FwarZNGZVlVwZVKQQ0CEOpDHg3wT6fk9/OHB5+YTlYs1isSTFhDUNMZQFl1Iq1ECuiq4S7NPkSann5asX2N09r+/vyEmTkyakUONMGkShy0UUpJcSOYeKggVBw7cW1rulGAM4xl2r0jt30+aN9HZ47Py1o7V+6PGdfb4cV5jpM3+sHq/FoyRgpRyTcjF2mTL55ZFJGSJCzkKmXHuatXBV/sVOyO9FLzSmbJSmddXwKjpnCYNmYR2N6Xh68wlGdRhWhDAy2Au0GdF6ZAoBUGhzg22uuVr/KSlbLiZdAirGkJUBUSjlEIq7F2Ng1/fsdgNvXu9IwZMCSHYIhqZZ0rnAx+0TroZ7bp7/AucSziWadoOPjhx6FBnjGqzNXF05wEOesIsd2h749FOIqXnUmJwjzO97/fx9UaCVEMikGBj7Hu8Dy65DKTDGnBnVM6t57lkd5aRkj8efcZJ+Crz85lWZs6ZhsViwWHSEnElHdMd3hDxmxT4H3x41JG+Nw4PoB4KgRKGVwmhF4ww5CUYSY4Qpgk2C0pEwTSWLIM/YIuOnkd39Hbv7O6yztG2HUoVamMFJ0SGKplmyWl2zWt/guq6ASqUwrsE1HW27xJoGJboYAVFHTleJeuBVHO3eW57A98k7lW7OEbIQE4TsibM7KqCPbmmqx2ZSmi8KQJ3NcYacyMkTU8L7TKFphc36Eq0VRhsi6QFSLqNZuKWc5QggU0qEHOiHnikEBh+wpkOrtlALdQSK4tfH31OUWVE8D5He+4tIQQMzyJDqQsoR+soRAp+7Pvk8cHY62UOYezrhg0l9EHSjcLZaEkYCK6uwSohVoc8KN9ZHSpmpzmFMNRgwQ/Kq96Uq3izyQ+vle8UajbWKzbpDKYNIg7EJLRmnWhrdcbm8xugFKi0JYcLpBtEDogf05EkIxq4xdkPbXZOlQTeOJCUQNxuS4vqX3xsjjEPicJi42x7IaYIU0VI2h3Ut2mRaYzB", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "RAp+7Pvk8cHY62UOYezrhg0l9EHSjcLZaEkYCK6uwSohVoc8KN9ZHSpmpzmFMNRgwQ/Kq96Uq3izyQ+vle8UajbWKzbpDKYNIg7EJLRmnWhrdcbm8xugFKi0JYcLpBtEDogf05EkIxq4xdkPbXZOlQTeOJCUQNxuS4vqX3xsjjEPicJi42x7IaYIU0VI2h3Ut2mRaYzBugWk6tEponVBaEWt6Y5aESINS4DoHDAgJsQF0YuUA0T80DA9Ejv98x+uc7dNv0Q2FSgsx4oMnpXiM2M+OzbfPe/7i2wG0MyNbJaVE3/dorRER2rZ5QEUcH9+5Qc4yZR6/gU574MEJqgtw9NgEhSJLxihFyIUP16acI/oAko5ABSn0wjgMjMPANI441x4907c2HkppjLFY12JdS4wTMQW0cWjjMK5BlS8r1yOz58L36w0B9R5K5d1KN9WFnUL9phl3CTGmyvmn41wWt+D0+2ZkmlMi5YyPqW6YiRAi0+jZrK5omyXWtCXf937LOB5I2WO0pnELJi/4ACVEkmsUMuPHnhA8KiWCzmidyfEUwIHiisyjkFDkJCBzIE2+M0L7LpnTgDgaG3i7PebRJXnr7/Pjvuu1GQlLzg/m7YR0FTkFiJ42b1my4x89v+bJukUpQwaGEEgZQoKYMiELv9sqbgfFi61ijEJWp2tQFXHPYCA/ehcVub68ZL1a84s/+TOsWWDNCh8CIUYaa3DG8tGTn9O4JZ27BsnEHMhqIqmJKSRCAmUuyNIQWBCjMMRMzMVYpwg5ZUIq1x4IDOOBF998wfb+ntevXtbNkTFKo0TRD6YiyMpAZovRDqMtIjAZaN0KY4Tl+hJNQOIB4R64Q5kRUZ5+3B3jFu8rc8jgqA7P9OK5YjvOc85MPjH6yDgFUl3rQzVISWmUqq5/FvSZ8nqoyCvdUF8sfkc6KuyIkHJinDzWZrrz6znTfb8PafBOmR3i42DMX6jIaCavSdkgNCAF+JWRzxhr0SYTp5k", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "vXtbNkTFKo0TRD6YiyMpAZovRDqMtIjAZaN0KY4Tl+hJNQOIB4R64Q5kRUZ5+3B3jFu8rc8jgqA7P9OK5YjvOc85MPjH6yDgFUl3rQzVISWmUqq5/FvSZ8nqoyCvdUF8sfkc6KuyIkHJinDzWZrrz6znTfb8PafBOmR3i42DMX6jIaCavSdkgNCAF+JWRzxhr0SYTp5kCmONKCT9O7O7u6Hc7xsWSzcU1xlrG3YEYA6oG170fGfo9+909V88+xphCVVlgs7khjhPrzTVN8wXkEnDXRshZHxV4SonkU0XnoAW00Vinz4Lt3y0/gHQ5RrdnCD/DTTk75nwsT2kZJ3fmmM4RM5lEShCzh6QJPpSiihCJMRJCIIaAqFRQKfnB9QA1UJXIMRVrl3JhE1AVjcgpyFn/nhfmKSMwVzrkcUvrbZ6uvPZwHM7H4/vP8xAY1x8GWY55lUfjRR1/ysEpRyQNaPZcuA1Puow2RXGOoaDdkDOxItydh5Dh1aFs1iQPL0KKdX2Pq/5+WXRr1stLnlx9jDUdVi+ZQsSHgjqNMrRmiTMLWteVhamFKJ7IRBo8KST6QfAxMIQtIZXfE1MqtFYs68jHyipqmKae3XbLYb9l7PdYqzFalTUFTFP5ScbMrqsm6UTSZQ17k+s608ToENHonEhpJCeDkBGlyX4gx8f6ATN99NCJflv5UsmukEp6Wj9M7Pux/F6EYQxMEYZUKBytFBopm3dWrDNIrB7m7CkpAUNCS0Lr4haXMFkJBGlV3OejX5XLMpScq1d39mvOFnk5/+P9omNZw/G8UjIlauZCAWkQY2YKJa1TBJJQrrvyHceCBGbPrQb0QySnMrfGGLQxdc+egnApBIKf8NNAPuN+BcFYh7ENxrgSwCefxlPNm3XWJ4VjNlowWqO1ol00J2D2PfJupRvfUi+z9RNOCddneVeF+5Bj+kZKlXyO9bV4ohoCEyGXaKOfPNM4lajhNJZKJRsgJkIKpQqlDmjKCYc", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "8UjIlauZCAWkQY2YKJa1TBJJQrrvyHceCBGbPrQb0QySnMrfGGLQxdc+egnApBIKf8NNAPuN+BcFYh7ENxrgSwCefxlPNm3XWJ4VjNlowWqO1ol00J2D2PfJupRvfUi+z9RNOCddneVeF+5Bj+kZKlXyO9bV4ohoCEyGXaKOfPNM4lajhNJZKJRsgJkIKpQqlDmjKCYc6uSHKoDA43dG2a2JUpCTkOjn5HFLkiEqaYyXQ7F8/QrSuqvsM6c4TXkdt1pBnA1mfzrIUzv9f3jstHHWCApxmoFADKUVS3JOHV2T/gmu15ONO41xxhUJN+ym4rgaZksLqzFdbmHLhDROFslE5n+iF/OByHyWfPPtjri+f8mef/TVa6eLuzZyyLwqTkJEA2Xu0c9imY4iGKVje3L3h7r7n7z9/ye4w8no7EmIiRAgxFOOcStGNzxFRwuXVBiERxx1+ODDs77m8WNGsOoZxwIfAm7t7Uoam7bDG0bQtRjRGKbIqcYnF0tA1Diue1mpWbsl0GBl2htViQes0ftKkND1+YN5GJWe0lgCqAhgNDJPnxZsdr27veH17x7PLC7qm4XffHJhCZp/uECkJ/A5oM2SdQWW0lHUjqtAFYnRBXwpaiTiVeHq1pG0cWVuUUlxsNhhtWHYtVhuIuWb7vJ+/8+PMc1Xkko9UYgkSqqPnOWc9eT+W4LcStDZooyEW6mWObTAXKiAFbISINpqma+mWK1zTMvT7ch4tTOPA2PcM+3sO2zekaURiIseAkOnaBYemw1hXdVtEaYW2ugToM1irS8aH0TRW0zWGq8sVq1XL9dNrnHPvHIF3B9LO18vZCJfSvBIoy0etn6sV1qScSl5APrnfD4JBR2VXCx9yOi5Ooy2QWHSlQsRpVzIUyPjoySlhMsdUMe0WuNUlzi5LrmXvCSESa7lgrlQCUjRRVtWKU/ISH6tilKpY+Rjcq7/oHCl+i6flwbEPj5/fKzsyZ6lG4cyqzqeRYkyEBDl", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-9", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "W0zWGq8sVq1XL9dNrnHPvHIF3B9LO18vZCJfSvBIoy0etn6sV1qScSl5APrnfD4JBR2VXCx9yOi5Ooy2QWHSlQsRpVzIUyPjoySlhMsdUMe0WuNUlzi5LrmXvCSESa7lgrlQCUjRRVtWKU/ISH6tilKpY+Rjcq7/oHCl+i6flwbEPj5/fKzsyZ6lG4cyqzqeRYkyEBDlACmhjsE0JXs3G7jjmAkLCSMIKaA0qCnlG+DUPU2V1TCt7iOHfX5yzOGvR4oobB3Vj5bLPcmIMIyVwN2LIBGvofeDgPXe7iTfbkZdv9txve75+fU9KglKWEDx+mog5knImSt0IOpOjZ3/7DdNhz2F7i/zkOVbfsNvvGYaRr776mpgy7WKJs46u7QrfCyUKbYTN5Ypp0bJZLPHOojrD1I9MhwFJE8EKWU0UeusRcrRis0tTBve7MhWPHmJK+MnTHwbGboFWmm0fGabMdsogCWMyRjKWDLoU+Biq0hVdvkvrgnJVxhFwKqGNYRVh0QmSMramZWmZ09u+qzLyrZ9Ukd8cZnisnOi8eTyK6VFKoyoSn5GpcMalKko8QilIlAoxqN7AQzHG4poGY0wp5+YUpJt11zzWOaUC6uoGUMoUhKvma1E15qRRruz/1XqBsYa2a2icoWsN61VH1zrW6w5r361W3/0up414GrQ6p0qh9EP3Yy7ZIws5xbL5a9rFMT+1WvZEoRogknOsk6homwXQcHHhaGxD1ywRoxGl8DGQckKFWF2BlmaxYnn1ETkbcta8fv2Gw6Fn8j0pR8AzlzhCqnkdRdGnXAo6HiOiqnuU1be4XDgh3+OCnR/y0Oac/j5HyfP4nCvch98hKpMlonKJAJl2gV1umKZC0UwhkFJBmFLRvs4F6RjdYEwxkiXjYVa+xeH8sRsJoOt0iTZnW931VB9lriTDOAzElGGKKN/ggN3o2Q4TX73e8erNnt9+ccubux2/+fxrtDKsFxdM08jY90QVS8DLRIxRqDz", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-10", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "nXAo6HiOiqnuU1be4XDgh3+OCnR/y0Oac/j5HyfP4nCvch98hKpMlonKJAJl2gV1umKZC0UwhkFJBmFLRvs4F6RjdYEwxkiXjYVa+xeH8sRsJoOt0iTZnW931VB9lriTDOAzElGGKKN/ggN3o2Q4TX73e8erNnt9+ccubux2/+fxrtDKsFxdM08jY90QVS8DLRIxRqDzihx2/+9UvOWxvuX/1kjz9Fc5oXr1+zXa341//7S/xPrJcb2icY7noIAYkRWzXYBvH808/ZrVaYcWw7BriuiP2W8Lunn4/oLVnfdFg3eMCaQ+U1oMgqzw45twZk5zxw8h+u2PftuQkvLzP7IbM3a7wssr4AiZ0QiQiknCpzCHZlJiLKJRktAKTRyyBiPDkcsFPtcIItJLQSmFVzZKpObqz1/PttfBtBPzo9fLgt5dNIKJRymKsObnwcwZF5U1Eq2oZCoVobCSpWFIF58hxHUXXtHSrNdY2aGWQLEgWtCi0aLToYvhSJlY6M6WIKI0yDqVtTaVTKDTOWFIT0W2mcYaPP33GctVyc7OhcZa2tWhdvHwtD+mk75J3Kl1dOZCj7shz1kFGYnWFj/mDcESUx6jMQw4IarqFUjWdqbh3IonFYoGzLf/wL/8K6zRPnmywymJ1w+g9UwxMobiYclS6DbZdsri4IadCKxjdsNttef3mK6aplA2TIzl6SJEci5InparsH0svnAo13l5y3wqUnW+4B8ec0OvZP2cnqvnD37q0EhgsQTMhp1zKGbVF28qLTSV/U9XzCiW/0VT3SqJCy3nLkIJ43zfH8Ptks2lZdg3GWOZqwJJ2VedYF240pUg/DkTvScPIlAWP0CyXXOoFT+4D4u753Yu7QtOIQmtL02Q8niiRrmuwVnCqBHnTdCCOA3Ea8GOJXKdUghxahKSKG5+jZzhEkp9IfoSdQrRCa5guL3n+9Ck5WURp2rbD6muEAyKeptHo9+/rAlRFdtbRSHiLCz0HLAJWKZb", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-11", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "43zfH8Ptks2lZdg3GWOZqwJJ2VedYF240pUg/DkTvScPIlAWP0CyXXOoFT+4D4u753Yu7QtOIQmtL02Q8niiRrmuwVnCqBHnTdCCOA3Ea8GOJXKdUghxahKSKG5+jZzhEkp9IfoSdQrRCa5guL3n+9Ck5WURp2rbD6muEAyKeptHo9+/rAlRFdtbRSHiLCz0HLAJWKZbWsjCaTgkSRuKYaWKGlGnqJEksXLSXjJ1RbnW5I5qElKCkCBqFShMqe8bdlj2eexUxWkjR45oW13Zw5C6pirBe3/kePrvmR8afTyIKJbMeKEhSaYNUL/lEV54jFtBKo22Dc23p+ZICKUb2+21NMfUYqzFGF6MTSnwo5xpQDpFxOJDJuGbBcn3FxfUzusbidELn6hXP6DdEYkiEmGmtpls5/uyzK9arlouby1Ix17oS1NRS8oJzOuOQv1/eqXSVUkelm6rbmioMTyEVv2R2jYXqzteg2RwRf0BRyBENJxIp16owySy6jtXqkp///OcslwueP79BYZBouNtv2R4OjLPS9RFE0LrBNh3t6qKgu5I3ReMc+/0rUuxLPl/05DAVhZsiKfpKaTy+OELPYzLvmzM24buyEx4s2gzp3Il/Syk/ZF/Sd8AIVTg7ZSCXgoa5hrww1RnRBeGc5wErlQu1UL0TIZEr2i2HqMeyLN+Szbph0TiMLUi3GOeE1PzHmEG0IoXIfhwZY2IfEsq1iGtolte4VcuTXSKrBut+SwyZTFG6eub/CCwXDmcFqyamHEjTQJp6kh/xU0kZmgMkWimSSDE00TOOgTD2hKln9J5M4ehS8KUpSgUSnVmwXmhiasl5xLhiOB4jZR44U1Sn6qZcec1ZhIzVwtJVpasVhJGQJ5oINmdMTgW1Ab2Gg2QWorBwBBAjhSqaUukvYJSuaXSBabtj70fuJGCNKspCTh7becZCjSLUF06e2EN98njNO9MH1pqSiaE1SOVIZ7e+Bozm35RTQimDNY7lal1", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-12", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "OgTD2hKln9J5M4ehS8KUpSgUSnVmwXmhiasl5xLhiOB4jZR44U1Sn6qZcec1ZhIzVwtJVpasVhJGQJ5oINmdMTgW1Ab2Gg2QWorBwBBAjhSqaUukvYJSuaXSBabtj70fuJGCNKspCTh7becZCjSLUF06e2EN98njNO9MH1pqSiaE1SOVIZ7e+Bozm35RTQimDNY7lal14eVVymVGK4CfGfijgwpR1EmqTm1Itm/Ehst8fcK6hWyxZrC+5uHpK21icJFJNP5tSItfiiBgTISRap1hoxz/58+fcXC2xi0UZ41DS+kKMJB/IMWLM75m9oLWu6iGTshBiPBq/nApSzRX1lxLCUswQa4rYEdEdJ1WOilepMuFN6+i6ho8++ojLi6e03QIlwv1djzMtnWvwXjNNmikWdsAZWy2lxUfFeDedccOCsxajS5pMmgbIxf5rK+iKno0SGmeP3ZTeV1Q1NOm76tKznNEVckK7KaFyRBNJqJL+I6Zyq/XonGrgq7pQqJqlUdO6AEQhWoPSRO/pd3s+/+1v0eGAaWyJpErJyjBGH1P6jDYYrdBa0EkR684pIOzbFXY/RlpraKzBGEVOikhG5bMeDLm6iFqIJKYwsd0fYDiQtaVZAqZj199yGO/wqSelTKguLzGBAWsUjVNYlbh79TWH+1vS2ENtTBJDPAZlU4xM/YFxmsgxFFdaqxJgsrY0RomRcb9jaBzRD4TRsts5lpsVm/UT1puWprXE1BcP6RGij+zN2fqHB/a20pVoMp3WuEXHftGybx2D3xKHiY+9YH2mOXimEHkzjPhWETpNaJco24CfvbbaNAaFwqBrel3OsN8NjP3I7jCitcI2mvXa02fN5WbFxXqJodCmRpVq02NshlkJv6Mfw3uIiKFtLB8/W9F2lsWyJYZEjJmQVM28UUVR+pKqqnUFFtbx5Ok13WJZAqXTVBpcpQibNaVaNeCHPbfffIXv94hSvLnfkVNEk0le4ceB3d0rXn/zOf0", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-13", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "zjPhWETpNaJco24CfvbbaNAaFwqBrel3OsN8NjP3I7jCitcI2mvXa02fN5WbFxXqJodCmRpVq02NshlkJv6Mfw3uIiKFtLB8/W9F2lsWyJYZEjJmQVM28UUVR+pKqqnUFFtbx5Ok13WJZAqXTVBpcpQibNaVaNeCHPbfffIXv94hSvLnfkVNEk0le4ceB3d0rXn/zOf03v0HlhB89WRvy+ppX33zNr3/5t3zx+a2mrdIAACAASURBVJe8ebNnIYFo4bDb05mMngL9GPjq9Z5pCgxjYKMOdCpgN88Q075zDN6pdGf+UqSS15USACDlmkaWSDIXBwLEM+rgTKHluvjUKQUDXZSDtZb1esNmc4FShuADu92O1GicFmJU5TFTsqYpVlEMMZW8RplVlhTS2+hS1aLrdxmtsVpjraa1BmcMi0WLMY/j6eZka5Vn5/xsCebj0gQqykkFyZsc0TkUrg1FEFOKEerHVS4MqGKuDMs1dees07wqVTFaKUJMxHHi1auXtCqwWC9wzrBeFiJfK8Mcr1TqZOhK445TUGd2JXOu3vx38HbvI1YbjNZzXKScN1WtUwtKcqVlUo745JmmoRY+aDALJGUmfyh8fJrqOtLHgIcRg9alf4Mi0e/uGfb35FCQnORMqmmHcyzBe48fRyQnrDFoZ9FSPQJKn4MwjvhxJNfuVNM4kvOGpllzsblmtVkyjrtSFv+YtVKR7tzj45hbL/mYVTMPl+SMFqHVhqXWLLRi6ifCdGAVhc5n2nHgMPpirLwhRIPPUtLOQ0k1ioUIBdGgczWqpWnU4Euee+o9SguuMYwRTNtgjWbRll4WJSXt5MqdVrUc4wQ/ml0Qg7WWi82C5aphs+nwPhJ8wpeEJYYgxJQZxnAsYigIWLHoWhaLro6v0LYtOZd0uBhGQijZT8PunhwnEMXhMKJEWLYNOQXiNNHv79nevWJ32IKf6A8TWRvk8p7bV6949fULtrf39P2IUgmTFP0w0fca5TP", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-14", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "U4Euee+o9SguuMYwRTNtgjWbRll4WJSXt5MqdVrUc4wQ/ml0Qg7WWi82C5aphs+nwPhJ8wpeEJYYgxJQZxnAsYigIWLHoWhaLro6v0LYtOZd0uBhGQijZT8PunhwnEMXhMKJEWLYNOQXiNNHv79nevWJ32IKf6A8TWRvk8p7bV6949fULtrf39P2IUgmTFP0w0fca5TP3h4kX39wyDIFhCGR3ILuAkTVi3z0G75e9MCvJo5UrO/SI5ATygwoqKhqMx6h7GaTC5YYQEF0iilobmqbl6uqazfqCX/7yV9zd3vL557/h+bOf8Gd/8pcM00SEkneahYAho9FiUEZoDRgtWA373cikPM+fPSVcrXA/+xitBWc01mhcRWRWa7q2KcHAR4gxtaT2HBXmh8/zspQcKcE7j04DOgwY5UhKo7Uj1dStOXg2p3rNwQxSNXq5nC8rjRiDWMchZva7iX/xv/6fLFzmyc0FN9cX/PU/+Ucslwu6tnLtuShpowRrFOasHPiIxqnU0O/BMawWG4xpCXks9AJCqrwbKeND4OXr1+yHntt+xxgCpLEYLVGYPCJZM+xuOezuCKEnJfCq1NTHEFi1C5xREA4E3/P66y8Y91v8MDANE9MYGIeJYRgRoXoDO/rDnkEJ1lq6RVuCHjV9KMWI7w+Evimd8MKAHxPDZNn5jivZoIzBREN6JBt1dNHrCBeUe+ofW2cASRl8IA6B6XbAv3yN3N2h/A7tD3C7Ixwmtl/csfeR3ZS4N5o31vJVe0ewDa6mgWEqjadticQbe6Q0vK+AqKb0oQLOCH////yKX/z8U37x80/59OkN7XJRDD4CokubkHziV38frNu2lywWlpurinRXLVLyLshiSVnoQ62CDbVBuVbc3e+42+3Z7bZM3nOxWrBsOpwuGTLawDjsGPqAMBDGgc0KrLMondFKsWpKI/v7u9f89u/u+fKLX7J0DSoLX/67N4SYsG3LNA3s7t8w7nY4KQUrdyHxf/9", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-15", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "OCH////yKX/z8U37x80/59OkN7XJRDD4CokubkHziV38frNu2lywWlpurinRXLVLyLshiSVnoQ62CDbVBuVbc3e+42+3Z7bZM3nOxWrBsOpwuGTLawDjsGPqAMBDGgc0KrLMondFKsWpKI/v7u9f89u/u+fKLX7J0DSoLX/67N4SYsG3LNA3s7t8w7nY4KQUrdyHxf/92z7KdiCnRDxPfvLwrsypCXBn2XYfRIPbdC+UHiiNmzVJL6U7jXuM/cqZgZnV83oeAhyax8kMnt7tMoBKNMQ5jLNM00fc92+09m/UVMcVa2nvqahRSVRBS+lwqo3G25MxNgyUZTde20Gg2ywajVSlTNRpXH6W22z2eXqhpW6cKA8588zO/cf7xtbZSkUswR2rKjXpYWlhKiSuVUBForr0tTn1NCp0j6hSQDLEsilL9lUontZruMuOoGeEqVVFvLkbyWIiR83Fsjz1DHylGO7QyJx4uS20yUhFYLGWn+8OefjwQc0apebBKubBQeLo0t+Sbr6mmFipV8k5T8MRpYBr62lUq1fai5ZFiOhrTVKPTuU5FirYAQUrwRuWyIU0N8AiJlDwpexK+VqIr8vnif085V7gcVa0ckS+5rCcq2h7v9+y+eMXu1Sum7ZakRsieMPbQD4y7O/qQi/vtFXHU7IdAry3WuJLnanKNvJvaE9geiw6mKRBTrj0sShaMMZm9yzy9XLK92TBtVsTGleo7USjneNAw/e09/Ug7bWxp2+icrbxuyZnWShMwxCyoVLxs68pvslpx6AdIpaFN6amyxBhN6loQ0CZDnohBEaZCG+XkyQlMXfciBQTlnBiGPb3v0es1RmkO/QE/BeSwJ0dPCiOKiLPCGAsC34+JRGljO02ReNyMiqAbJu2IlG567xyDd70Zar9clCo9E0I8jrHKczVIWU5CJeTnhE+hTjZvKTYpSrRWFscg5FRaLmplinuji0ukVMK6Uu4oueTk+ZwYhwNKFM5Z2saxbha", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-16", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Q0CZDnohBEaZCG+XkyQlMXfciBQTlnBiGPb3v0es1RmkO/QE/BeSwJ0dPCiOKiLPCGAsC34+JRGljO02ReNyMiqAbJu2IlG567xyDd70Zar9clCo9E0I8jrHKczVIWU5CJeTnhE+hTjZvKTYpSrRWFscg5FRaLmplinuji0ukVMK6Uu4oueTk+ZwYhwNKFM5Z2saxbhasVgsuVktiuIPUozcbnFX80c8+w2iDoigYRUbVDa5rMcdjxM1Id87/nSW/pW+p2CZrRDVolVA6gWlAG0QZsqiKLzUZjSKgc8AQUcSaVlea2ZQtn8tC0iOyblk8fcJf/NFTbi4WXF40rBYtz549LTmFUmghIeOswUWNNQqXZvwlBUlnkEoVZSAoXeiHR4qRBYIjp6ko3Az9MDD2E8Y4vA98/c033O3uuR92LJYdT5/doCjNRJy1ZDEon5CQabQlIVjriDWI1jWWVWvYvrxl2L4hjD0xTPWOAsV99iESQqSzhmx0yfFOCa3KDWi0CG3b0S6X4CckJy5XCy4vL1g4g1WQ8kCWABbEFFd9mjzBD2wu339MTjDkjGKoXqKq3o3VGj9NvPjd7/j83/w9/8d/+8+xttJgny2xK8vdYUva73l59xVZNGqxQcZI10e+GAde1ABgKY0u4xF9xFpH23QlV1UpDvtSMCKq3JHBdpZ20bC+WnLhEmsHF62DlDgcSkHBzbOnGOdomrOE/zNg8FjpVjd0S8G1xUuNGYwIWsOb7UA/Jvb7hDaOpx9d44zCWUGpHd4nYp4IPqOfZtrG0G1KS0aRiJKEn/aMo2eaAm/udlirabqSUTONE0pbNtcb7g8j/eAxHzna1rF8uuSw7Xn91WucUayXDV0LKQqvXo+MU8I2LbaxqJzpVornH3+Ez8KUFU3XYhpHFE3+gVtP/kDvhTK4uVYCPYzOn9ziY3SAfESlx2jBHLSpnxOZ3ZOqvFImxsTkJyZf3M3iUhiMLgn3rSjEJJIK6JRru0d", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-17", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "En/aMo2eaAm/udlirabqSUTONE0pbNtcb7g8j/eAxHzna1rF8uuSw7Xn91WucUayXDV0LKQqvXo+MU8I2LbaxqJzpVornH3+Ez8KUFU3XYhpHFE3+gVtP/kDvhTK4uVYCPYzOn9ziY3SAfESlx2jBHLSpnxOZ3ZOqvFImxsTkJyZf3M3iUhiMLgn3rSjEJJIK6JRru0ddavqdLW6zLjlyzmiCczizKNUi7RKjC79ZcN8crioo+bHqRSs5SyI/21ZHpM8Zt1vGQLKuWQe6dOxQpTN+mZxUUbygU8SmgbWLdDpVMrdQKkXpKrJOZFH0Vx1eLvjk+VNuLpcsF4bWGbQ2dchPczXnlGslJZiW5egWSUyQY0GTUpLDkceh/yKOnDUh+OPcIhmlC9sfc2KcJg79wJvbO/phJFOamkhW6EZINOzud4z9WAssyliLklLSqUvvgXEY6A893pdCGASUM3TrBe1yQbtY0HQNJgbWl1dFaWhN0zjWqxUXV5dcXF2SQym2sVpYLDqSaEKCMCWGMdD3E7vdHqfu8bWnxWPkYfRfavZCBSrzeiwJsvT7Pdu7W15984L1csF6tSDHrijSDCFmdpMHlXDRQ4ZWQWuEBlXLvhNxipAiKUzFq7K6oF5RZH8gTR5BgdGIaTFolm3huf00Fg7ce3a7LaKEzdVFSTXEneIycub5PlLaBhpX0qzm8uO5n0ROqbYDKOX9MSa8EoiCT4VeTCGRiWXtaMNitapIPBP8yH53h7W+ZlmVkhHElqbltkXbuRJyx+iFGCGGjLUG17iSvqYFMabcnqciflHgbEPTNqAU1lqWqw6fhTEpTOPQRhPztxtivi0/0E+3NskQeaBwiwY48zNycZNTLpkLc8KfiBxTXGblphCMMrUTlhCqwt1u71BSAjFWWzq3ZNkuWC8WLEQRRNHuB6YYMapFKYVzpiqUjFNA9Cxcg12uMLrBWcOyLfdLyml27+Ysh3TqcfAIMXMZcM1XnjuLHh3", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-18", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "cnqciflHgbEPTNqAU1lqWqw6fhTEpTOPQRhPztxtivi0/0E+3NskQeaBwiwY48zNycZNTLpkLc8KfiBxTXGblphCMMrUTlhCqwt1u71BSAjFWWzq3ZNkuWC8WLEQRRNHuB6YYMapFKYVzpiqUjFNA9Cxcg12uMLrBWcOyLfdLyml27+Ysh3TqcfAIMXMZcM1XnjuLHh3JWeFmKndaAn4qRyQFRDegbXlAzVqAANg00oRbfrZOfLQoidgzhVAS3nXNIXRMmyf4Q8Mv/uSnXFysECmdqIIfOKag1cBVQfWCM4LLqhrMMjc5TaQ0MEfssm5LStojJadFSTIY+zL3NTVQLRyjz4TaHPz2bsuvf/07RBSu/ZKchJyEGBti1ozJEVBomW96mo6Nyp0rxmp7f8/d6zfs9wM5R7QzuM2CZbfkySfPuXn2jKZtiCnxsz/7c6ZppLWOrmm4WC352Wc/4aefflLSfULgxatyJ5JJLMkH+tHTNANds0P7L+mXW1bXqx+sNHpbZjM8c7lw1lejBi2NEqYUuH/1ktcvvubFl5+jnz7lorEQS1PsmDRTUrwaJhCh04ZGN1w0DTdLBxqG2p/icJfx48QQA41TrBqFsxptDNM2kNOAThGNZiHC9WLDzz55yrpz+HFgHEaGfuDlNy9A4PrmCq1ALTqOxfP59OseK5dLz2KhMdahq9K1SuG0IseSglXqlyKHcUCCQbTlMEHImhQiOpZGTsq64tlZi6iyV7b7AxmDtX0J2iuFNkvapuHq6hrTNJi2w/MSH2+Z+kyaJtq2RdDYxRKjBKwpHexCJEpAdGK1XnOxXmK7FbZtWVxe4FH4OWtJStrqD6X+/0AgLZ+Vnr5FzxZIN0PXCmyL4p3fUpWPFMk1IbpSEbkEeCRJ0YGxmE2tFU+un7Dq1rS65ePnz7m5vEG0Aa25Wk/ElDGmKc1L7FxumrBKY7Rm7AwxTGiJaKVYLlaQKXeUSLFWytXKph+ldFVVtic++1s", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-19", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "jBKwpHexCJEpAdGK1XnOxXmK7FbZtWVxe4FH4OWtJStrqD6X+/0AgLZ+Vnr5FzxZIN0PXCmyL4p3fUpWPFMk1IbpSEbkEeCRJ0YGxmE2tFU+un7Dq1rS65ePnz7m5vEG0Aa25Wk/ElDGmKc1L7FxumrBKY7Rm7AwxTGiJaKVYLlaQKXeUSLFWytXKph+ldFVVtic++1sK90gLyPFGlRJrkYDKJVVBlV4KOhfcrQE1Tsi0xeUSvV4YSvtBXRlFiUVhWrg9BCITL158yd2txdjC5bauUDSNK52T5k0/N/jRiop0qdxuQuOPHVOSUiXQ8thxMQv6oefLL75iGEf2fV+rexybyydk0XSrC1ZT4uJyR4yl+XwMuVIAJXHq8uoKZRtysyDmkgtZ7pelyPHAOB3op4nRe4xzWGe5eXrDYr1mfXXN9fUz1psNoks5+tPnzwnBY0TTWstm0bG52LBaLjgMI4nSHMiHxKGfSCkwTp79fuD+fsel61gYi1YXWPvumvq3ZY70H9sgzq8f36+rJyZ83xOmgUQoHoISQizdtLSfGIPH1+VuUkarTBRBV240kZAYkK1GtEZph1IWrS3GupK5YR1iAzkocjXmMZXS4VwBkPe+UinhWB02o/K5lF+ql/QjgC7LxtO64nXNrSTnZ+csXVa4xqFNSSdLCD5kmsawWnVEH1BKsVouWS6WWOfQ2pJRGNPi2hV+CqV9rCr6KUTF5KUE2SaQPuGHAFmRQulaJ6rQnM41kBN+Skw+4wNY61CNQtySbFdkuyQoy96X5lGp7vm66X+w4OrdKWNyWiRJai3EfL7zu9LCMf2pAsAyRbHwhXP9fXGlFKSMZIVKGmq7Pshoo3n+0ceQhM+efMrFxYZn10+PddQzdWFN6TUguiyEVHs3lHskXTDfU23OTwwhMI4DwXtCnkqeak4PAlnvK3bmdNNpKxXrf1K4M9JNUIKFitLmS6XSPU1nsqlM+NETSKg8wHRHk1uWyrGygtG", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-20", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "dKWNyWiRJai3EfL7zu9LCMf2pAsAyRbHwhXP9fXGlFKSMZIVKGmq7Pshoo3n+0ceQhM+efMrFxYZn10+PddQzdWFN6TUguiyEVHs3lHskXTDfU23OTwwhMI4DwXtCnkqeak4PAlnvK3bmdNNpKxXrf1K4M9JNUIKFitLmS6XSPU1nsqlM+NETSKg8wHRHk1uWyrGygtGKdLzDRangAc1059nHns8/f0mIga5raZuGZ09uWHQtzhqqP8iseHWlGEp8rgbyVEKpUDkIXTIq9I9Qum5J2A385jf/jldv3vDl11/j2iW27fjFX/wV64srlptrojhutoFh9PSVe8s+IikhKJ48e0q3XGPXV4SU6ceihERlvvn6t9zuRg7DQO8n2qZjvVnzR3/yp1zd3PDsJz8hZ01OpfFPTImPP/20eF8x02jF0lkuLy9ZrRb4lJhixMfM4APxUO6W4ifPTvfcmS2HdsnaNRhlyoZ8hDxAuszm+SztSkoaZo6Bab8jjAeyBESVeIMPiTQG1DQy+qko3SxF6eYaILINTdcUBRM8GIOERGGybLn5onU0zqJdg/KRjCeJMMVMSALKEDP4EJkmzziO+CmUrmTpRIWcgNexO8ijZdMFrCt0oDAHE8ujbS0Yh22WaONo2gWTj+wOI4vWQO7wk0dE2KxXrFcrGtuWdZsF4zqadk3fT6U6pMSKmUIBg2Y3kfDE1DPFgKRK3UmZjJQ0TdPip4mx75nGRAhC07Q45xC3IdoVyi3wCGHIJTNvLsBMpWOf/MAdaX5Y6dZA2DF+OydbU5SNksLJWFPKY6cznlTXnFyjM6IKSlRSGnCTBcmaThkWuuXJxTM+uvmYZbdEiWAyNLahbRzGGow+tWnTc4/cGv2eO4rB6RZBc9rVfF81ZyzD2NP3mRwKX/RjXCStKeeWE6V1xJLHAFo66TvJSIqIeCSNIC0oYart7FxWKEpDct0IpnOYGgUdQ3VdsiKlyOTH2tJAePnylq9fvOb", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-21", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "JWFPKY6cznlTXnFyjM6IKSlRSGnCTBcmaThkWuuXJxTM+uvmYZbdEiWAyNLahbRzGGow+tWnTc4/cGv2eO4rB6RZBc9rVfF81ZyzD2NP3mRwKX/RjXCStKeeWE6V1xJLHAFo66TvJSIqIeCSNIC0oYart7FxWKEpDct0IpnOYGgUdQ3VdsiKlyOTH2tJAePnylq9fvObfffkV+75nvWi5uFhzuVnROFu40DMsrlTCqoTVmUBRyJqMJFChNMHJinKn5x+hdCcfGabIYYh88+qeX/7qt9w8+5jLG0vWC1x3xUefrLicJhbrj5h8pB893ke8D+V+XSI8/eRTXLtAdStCgn6cGMeBYTyw279if2h49sknXN9c8tFF2Xif/fSzcmeJJOz2e/p+LEUkIkyx8Hs5BLCGxiiGaWA/7Lnb3rHd94y1mMJqg6SS99q6hqvLS25urrm5vi551KN/9LjMIt/xLBRQQogMh55pmE4oSRRWSiGMdRrbGS4v12Q01i5R1hIbi3IGawSCxyTPYCGgaHWDc5Zm4YrSsI7N1RPswhNSVTRKMBdPCHbNKIYUDW/6RNaB3lssmn4ssd8435VRZjLs/Je8P+1inMbY85atUgJ/kokhEz2EuEdJjx+Hwuv6gKSAMwGni0d3f3/LMAx89eqepm25uLjm0HvQLYPP7A4TWudjFo8I9Ie+tI1MHNtJWl2yVnKlUJU2aJuxqcQKRAnrzZrFoqNbrbDOsd/vytg5U+4soXXh61WeFeS7x+CdC0XmIMzM3QpJculteRYnUzVhPx2zAYpKNlrqbUbKxRelWxS4SqWdXaMNrbasF2suVhc0jcNqRed0Sb1Ic0qJorG2KtyzOa83mph/5nnAb+Y0oypJ4TEFpkkTojpOxGNF1/VyUrpzMcN59kJ1CY4EWEaISPIFCwuUSxAsYMh0EtEGTFOKDCKaKRUknCkLcn8oN2KSmNnterbbPd+8fMP9bk+/bEvqVCyJ8Mcy4iO9U5q", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-22", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "qbUbKxRelWxS4SqWdXaMNrbasF2suVhc0jcNqRed0Sb1Ic0qJorG2KtyzOa83mph/5nnAb+Y0oypJ4TEFpkkTojpOxGNF1/VyUrpzMcN59kJ1CY4EWEaISPIFCwuUSxAsYMh0EtEGTFOKDCKaKRUknCkLcn8oN2KSmNnterbbPd+8fMP9bk+/bEvqVCyJ8Mcy4iO9U5qfaF3oBQQ0EaUySlKlFQp6Sj+C7B6nUB+R7W7gxctb3PKaxaWAatBuyWZhiSnTdFf4kBinUBVuCZ5kgcsnzzBNS3YdIWb6cWK/33K/FVzToJ1lc32FSms+++gpq9WCp0+fMU6e3aFnHEZ2ux2mKa35Ug3uphDQAjGVuzEM08hh6Dn0e0IoN8acb7shSXDGsuw6lssly+USSZk4PU7pPohnng3pfM++uSgip4wfi0vPzJeK1ECvwmpBWcVi0ZLFot2KZDXRlUbtBlApYlKk05mIImmDsZbGGRpX0O5yrTFd6fGWKK039eKCqFu8lN4luzGjdGCK9Q7OY6aZMjGWZt6ls2B6ePPVR0jJTnoYqP1/SXvTJjmO7Uzz8S22XCqrUAUQIO4m3TtqtVqtHklt1qa2mb8+n2dM1jJNa1ojae5C8pIACrXkEpvv88EjC6RMIlktN4MBIApEZWSkx/Fz3vd5S+oMy/2b8a6opqIvBpkQi45HktFGo6RkHHuGcWbyR7rVGqWaAntC4UNmdoHGiMX5KpZ5hyemkjZiqiJRlSzI1KJfRSqFUpBNfDopd6s16/WatusQSuJOh+L6NAvA/AyWF/woLff3broF/FxE9YVtkEoFocXywS42zfOHukiyygdXy8zb1xdcXDRstyukFEzWPTESVAaD4M3nLdfXDRUJN4x89cVvqYzi8zeXaNlgZIdAIYRit7tA1PXTsaSgOcsb75zDOYe1bunffnIkpRSJKTw9RM5W5DMT9zmrUYvM61t/9by/fvr9uYtaBm45RmLyRDug6yKgNqY", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-23", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "oLff3broF/FxE9YVtkEoFocXywS42zfOHukiyygdXy8zb1xdcXDRstyukFEzWPTESVAaD4M3nLdfXDRUJN4x89cVvqYzi8zeXaNlgZIdAIYRit7tA1PXTsaSgOcsb75zDOYe1bunffnIkpRSJKTw9RM5W5DMT9zmrUYvM61t/9by/fvr9uYtaBm45RmLyRDug6yKgNqY4kHZiphKBFTOYERrQovQYZ1um6yF4Zmu5u7ujrRt2qy1gaJsVf/EXf4kyhlYLurbh5volpjorEJa2jRRoBJWR1EIQfCisguEjYdiTTw+Y7SWqXYFuyXkC3jzruvyf/9dfM00T948nqnbDn/zHv+RXf/yn/PQXv+Tq5Vt0uyoazAzri3X5oGWe7ORpGUoGIYkIXAj4EJYK98hh/4AQkfWqoZGXGAGfff6GyhT4j7We4+FECBFtTBnIxEBY8J0igZeC2Tn6SaBUwnpXmL+LxtfPtkR1p0j2nugsx+MBlSLBFinWq7e/ePY9U24Klodg6TOKb/02xuUBlCDrimAMziiSc+TsyacR6RwvUotarWlff05UiSAj7v6R0PeYcSB7z2OM+AxeZmqRWUVQSIQyHN++wlYtvagIGVwIKLUoNkKAGPkoHxmqEZUzRit+/bvf8+LFBe2mom0MbWvKg4LEd0fkP27VbbNINcs1OZ+WtRQk73GT525/LAnAC9gpLUUClNaiVJqUB5CKqukQIjHPJ46D5eEw4Xyibjp+9fO3rFctKWemeeL9+2/wscgK16sNbbuircv3M7tpuReL3M7NDl8FfIj85M1brq93/PyXPyeLjPx/HZlMt2rKXKGpC/VOlrblD63v33SXUt5oRYolWVcupPmcy9NHhMV6KwVpsZtqmagUXG5qrq9WXF0VBUE/lONCiAm9cF53W81mJVGygITHYSBUCh9WIBVChoXtsDTyv3MHf/p1jAFrLX0/4Jz7zqZ7tp4aY6jrRTWwyF6eTRlbqsAn3nz+bns", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-24", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "85M1brq93/PyXPyeLjPx/HZlMt2rKXKGpC/VOlrblD63v33SXUt5oRYolWVcupPmcy9NHhMV6KwVpsZtqmagUXG5qrq9WXF0VBUE/lONCiAm9cF53W81mJVGygITHYSBUCh9WIBVChoXtsDTyv3MHf/p1jAFrLX0/4Jz7zqZ7tp4aY6jrRTWwyF6eTRlbqsAn3nz+bnsDyqZ7rm/PG29aZDzJTQX4ocAoqNVMLQINlpRdAQHFiPei5GPFxDzPTNPM/jAQukyjW3Iu3NDrFy9oVysqmalN6UlJtVTxOX8K0wO0XNBz0ZPdiO/vycOe1N8jK4lUkOWBLJ/XuwT48OEW5zzz7NCm5uXLS25eveb65Wt0sypqjVjetlLpiKdNp1zPTCIx+vD0gI8p4BfCv7NjyT6TAmk0Rkmqpi75Wc4TY+nF5sV+nmNRc5yzxRRFHx5SxHrPZEUxlKSiYAg+gvNoIag134ogT0XZ4FxJmX3GEmdp4Xmj/WcbL/n8ujM+5ULr1ZqsFFErQkoknzEuoHzC5BojK9p2RZaBJB0+J+I8084WQsDESAC8kDRZsE6KZBLJZGRdM686wOBSLrl0ORXiWwikEBilJleexlTklDieeupGY63DGAmUYeJ3qvhnrLPsb2lPL4O5848iX3TOLZtuXv7rpwDZnCRSlvcRqZBaLz3YkWm0jMOIAJqm4frFC3bbDYjMMI4ljXye6ceJ9WrFer1h1XZopZjm+im5xtuASAqlEiYm1qs169WG9WpFIlJpRSZRGYk2CmM02hQesHzqK/7r63s33c26RUlFa+rFXeSfnCRiAWHPrmjmsqxKZRYFrQl0OvCrtxf87O0Vl1c3SGUY5/hU0QgxIGRP02nq2tHUE0pZ1l23ALE7JJqcMqbWS19Ko9W3e6ef1jjO3N8/8OWXX3I4HD5ZjxdqkVKS7XbD1dWuOJIET9X6c1Zt5HIc+ud/Lz9Bo87pxUUrGBEEopvw/ZH5tCcLwXa3Qxh", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-25", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "6RUlFa+rFXeSfnCRiAWHPrmjmsqxKZRYFrQl0OvCrtxf87O0Vl1c3SGUY5/hU0QgxIGRP02nq2tHUE0pZ1l23ALE7JJqcMqbWS19Ko9W3e6ef1jjO3N8/8OWXX3I4HD5ZjxdqkVKS7XbD1dWuOJIET9X6c1Zt5HIc+ud/Lz9Bo87pxUUrGBEEopvw/ZH5tCcLwXa3QxhJqguu0IvA7ALj7NGmQuqzZCry4cMt0zTx+PDIum2YjwNGC2otudpt2F5ckKNHSbH0MllKelmieFKCGGgkWCJ3D++YHt9x+N1fo8KICT1xfoPuLhnzb4hZAP/1WdflN//0T6W6lobPPr/mV3/8p+xu3tA0WzyKEDJiCXYMsfS+cn4STZCIi1tv+ZEDwU/0x3tOh3tO+4+46Uh2EylYohLMdkAgsJNlmmdijsWOLsRT6onUBWKtZcnjcD5wHDzTPDJbj3eejx8/Mo8T0/7AbrPi3/3ip2zampvLS16+vGG33TIP07M33W+3l8qLlU/H8pxLm86SGXPiRGTSirReETYdbtMwhZngPL1NaAeXoqFLDVJW1CrTCMinA/n2A5eTxYTixhNJYLzAoGiE5u7Sst+OVFcXnBrNad6TvGc8nRaSVqFkJR+Q6zW+rlHX11BXDMrQDom+P2K0YNO1kHQpWNIPbzD/fJ2xuFp8ipUSubTR1nVGAae2widB3W44y1VTXOLYF6W9TOXnEAKn44F5LO/nMDt+/ouf89lnL/mz//jHvLx+wWa7JoTA48MDv/3iK/7mb/87r159xtXVNS8udtR1jYsZ7yOn05Hj4cT7d7dPVL71ZovShoeHAz5YhuOxDDtFICwn7Lo2i5Zc8EPD+e+njEmFVqUnJIVBiZq2rWmaGmWKy8PHUv6nrJldYL2eqcREJWbaunqCI2cBSlYosfQgRAAxEZPD+hJEKWVmu92WaavpEItw/rsbfelRfnqrz9rBxZETAs59GkjIxRartCT4pkTFU76Hb+f", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-26", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Z7yOn05Hj4cT7d7dPVL71ZovShoeHAz5YhuOxDDtFICwn7Lo2i5Zc8EPD+e+njEmFVqUnJIVBiZq2rWmaGmWKy8PHUv6nrJldYL2eqcREJWbaunqCI2cBSlYosfQgRAAxEZPD+hJEKWVmu92WaavpEItw/rsbfelRfnqrz9rBxZETAs59GkjIxRartCT4pkTFU76Hb+faP+em+Q53ARZZzSfTRTibSWIsE2U7k+1EcjNh6VN5JZCVxqZYrIUEZheZXKJKAh3L3/fec9gfsM4SQmAaRx69Y7VqyG2NnWfmuiIFjxSCmPzC+CztBUHGevBJQXLICMJP4Ebi1EMcIY6osYcsmd1C9nrmEsu1rtuW9WrNdntBXdd8SrQ4D2Hg/Lg8n2DOkqFSk6Zl+OkI3uLmETcPuHkojAQ3lQGJVAVAk8EF9wS4T7kAffLSsihH2NLHO4Pzc0oEmXCuVMeznYu6JVik6Nhu11xsyo9V11G3DWLpOT5nfQfveaYK8cmVlpfXHHOpdAOAMeTKkGpDSoU54inH8UpJshR0KSBFpJKhqF5kRsdEFRIyg8pQCdA5oaPDBIv0CtxMtma5Fx1hOD5FIJEyKmbiPOCSY+w1xhtCqmhbxTicaGuDs13R0yNJ8fmfn/PY+ZxUcb48ORdlUExyUewUydx5mH+en5TWzPJgXYqSFBPRl1ZAThFygJzKqVsrmtqQjCbtdmz3B9abDcYYck7L6beGABm/dI7LHEWqYsqZrSVEx2k8kVLAaEPTVLy6cK0YfQAAIABJREFUvnoyO5mqzGKk/OFZ0fduukZV1EZzse7YrBuuL1fsdhdsNmu6zQXaNCjdlBwuL7AuM0wBN37EjreYfOR4mhjnE1I3VN0VSmt0pYnZEhKM0yMhDbz+ac/6IvPLN78sPn7EopxYjiSq0I+egumebuzyc7nAZ9eSxC4g6/Jnhd/pvS+e7IXM5IN/dl+30kte24K5BEn5REdMdogcmOexoAXdTLS", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-27", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Z0fduukZV1EZzse7YrBuuL1fsdhdsNmu6zQXaNCjdlBwuL7AuM0wBN37EjreYfOR4mhjnE1I3VN0VSmt0pYnZEhKM0yMhDbz+ac/6IvPLN78sPn7EopxYjiSq0I+egumebuzyc7nAZ9eSxC4g6/Jnhd/pvS+e7IXM5IN/dl+30kte24K5BEn5REdMdogcmOexoAXdTLSeOEz4/S3+tMfOIzEEGIci9t9W5QgbHC4JXBLs1omuMvTHR6Zx5MsvvkAIuNztGA6PvPt4y83LF1xeXqBbQ/e4KrpKAXVdLbpWtehyC0YyCU3SGhkElT3g7BFhR0IYcWEk5Xt0P3I4HrD++VP668sLqrrh8vo1L1+/5dXNK4IsvcNPH8zlWovSz36C8GhRNp2UmJwnBY8dT0z9gdP+I6fHO04Pt7jhSLQTu8sNWjXFApwz0zxgnS9KhaLVIKRITKkkVeSiWBEUhGLpu5XBSvCB06knzDOtFlzu1vzxr/6An//sJ/z8Jz+hXXeYuiLW3SJtfMY67ybnX58bmcuwOQORjM+ZIQSsENC0pK4lrltScEQfeNSGKDJ9rVlXmex61spi1UxbQ7WpSb2DnGmlRilZUjy8I08zMQ94H5iOHxnTgJsnnLWMD/flFLBAp5qqwc17xiHRnz4itaLZdvj5wM3liuQtRkqaeo1ZrN05ZbjZ/OhLUhIayoBw2T5JGXwqeFCjJHXl8FPk8VDQjForzkkOWounNAkyTy0lUoLlRDtPPYf9PcfTqcCf2gqtNe16ze7qitefv2EaR+4f7nlxeUUjOnxMzC5wtz/Q9yeGeWQlG6pK8vU37+j7E/cP97RNzf/+V/+Zn7x9zV/8r39azF3fGbD+8D3y/ZuurgooxigqUzzQTS1pG8mqWwDi7QYhDJmKEAXOCx7u4OHesn/omUcPyiG1pPMZUwuaVUWKDTF0HA4z1jn2h5l2NfPyeonJyLKoDbxDV9XC+Vxyi84v7Vut3fNwTGtNVVV", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-28", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "MzC5wtz/Q9yeGeWQlG6pK8vU37+j7E/cP97RNzf/+V/+Zn7x9zV/8r39azF3fGbD+8D3y/ZuurgooxigqUzzQTS1pG8mqWwDi7QYhDJmKEAXOCx7u4OHesn/omUcPyiG1pPMZUwuaVUWKDTF0HA4z1jn2h5l2NfPyeonJyLKoDbxDV9XC+Vxyi84v7Vut3fNwTGtNVVVLCyAtf1b6iEqpggbUpeXgw/z0NT92NaIoIsITYBnm/oAbB9Y6YGSiciUafh6GQo4aPGE4EZ0l+cJ2DcMJYRVOdaUqXZKPpCjypkjGSEmuDLvtCjK0RoHRNFUZZnjv2e/3DFOhZUkpadoKpRRG6yfQjak7hKqYpz2zy9j+QJgtQpXK34mESSX0jxARzz1GA2/fvCraym5FZSRuHkkqklRxC32yXItFWVE2x5gTJIF1Fhc849gzO8fQHxn7E9M04n05uWilMXVNXdVUpiqVYkr4WKLeQwpFXpcLlCTm85G0KD5gEZakMmSz84y3lnnoycGzu1hhtCKmyDRNPOz3tN5iqoroi8pgdXn9o69JjLFs+IvJaBHa8JTRRWa0xQGmlKGuGlbrLbqqiSwF8tK7FzKTK0nUYAnUEpLW2LYjriP2skLaiFwGp8GAnAQyOLwU+JyYcmTKER/jU4UoWCLZl4GmD7FYniOIIBFaMPU9h8fHAgUCdhcv6BbT0XML3U+jt1LFKnEW2RRbr5SZi3WLlJ79OJZTTypMXSQFUqSKw5Rc+MrlOabRtcE0NZDp+xO//d3v2O/3vN9ti+y0bjgcj4zTyDRNOGd52D9grWN0MM22/Plwoh/60gOeDIfTgWmaCLHkLlZ1RdXU1E2FjqmkDC+qKaXlv629UOmayiiaWtPUkrrKtDV0jWC9qqjblouLK0zVUNVbCkiy5je/FUxu4tdfvePDR0fIE0rB9kWi6SQXtGQfSHPg9sMtYz/xqz8aqNuBn/8kL9ZhQfSRcR6pUyTX9fLEk09vHny", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-29", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "NOGd52D9grWN0MM22/Plwoh/60gOeDIfTgWmaCLHkLlZ1RdXU1E2FjqmkDC+qKaXlv629UOmayiiaWtPUkrrKtDV0jWC9qqjblouLK0zVUNVbCkiy5je/FUxu4tdfvePDR0fIE0rB9kWi6SQXtGQfSHPg9sMtYz/xqz8aqNuBn/8kL9ZhQfSRcR6pUyTX9fLEk09vHnyqokpiqKKua9q2LWaKZZBy/hpjKqqqQpsC1RmGY/lgPGOtlyOqW7CFIUSOd+/Y332gXoOpoMHivWN//4CfM/MgcJMlzq5ExcSAn49kKXHisvBMN6sCIM8ZgiOEAnjpKo3iptC3okfHCrFeoaTAOcuH21vyck20UqxW3ROjWKoyUd1uE6Zq6fuJcQ6MD/d4N6OqNTYrbJS0WSMjyJCQz7wmAH/8v/whKcMcICtJf3xA1StU3SJMg1Sas5nmjAqKS2K0T5l+HLB24ng4MNmZ/cOBfhjoTwecnRAiU9c1RtR07QpTGchLVLfzuBAIMZCSICUI8VzplteiltoqC0H0nuAd4+mEnSaG4x5JZvXZFXWlsc7ycNiDhLZtqUxFjGVj+Mkv/vBHX5MYwtKPLA9ElfPCGwAop6XT/sBw6ql0RddtkKJCtQ0+L0f+VJJQAkCrSQYmHLVShLrGb3cI2TBUmeQTIVgEEak81VHS2hknBZZMT+CYAzZF/EJjg1J9apMICeYQsa6cAIuNP3GSirv375n6E8f9A2/f/pSrF9cL2eyZlD5KV7bY04vU9Kz0kcqgpOTlpaGpLN/cDWXAGIsz85zKK6Uih2Is0abAe6TWdNs1q4stj3d37B/v+W9/87cYU9F2LaYytKtCJjOVwllL8IGv332D0TWzF1jneXi4Z5oGTsd9uUtzwrmZmOIT0bDpWpq2xtQGnUGZ4gVIKdE0zQ9ekx+wAZfo7HHMxGCZp55p8hyOIy9mRbca8bPFmApjGnKWxKj46osv+PLLL/n4MLDvI7OfECoTzEi", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-30", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "ih2Is0abAe6TWdNs1q4stj3d37B/v+W9/87cYU9F2LaYytKtCJjOVwllL8IGv332D0TWzF1jneXi4Z5oGTsd9uUtzwrmZmOIT0bDpWpq2xtQGnUGZ4gVIKdE0zQ9ekx+wAZfo7HHMxGCZp55p8hyOIy9mRbca8bPFmApjGnKWxKj46osv+PLLL/n4MLDvI7OfECoTzEiXFGaVSVYRp4bJtkwu89XXt1ivELmha1rWTUM/HnnY37LbXbFdX/Dzn/+MzXaNFAsuMOmSxJANUHipXdctEqClDXAOxiTRrTqquiKnEuURQ1gALT9+1XEi54wRkclNTIcjj998xbvf/465y7Q600hLSpF+DijV0tU7YqD0EH1hLVTduuQ5VQ0OyTQ4UvLEYOmqmlobthcbqkqThCqVWwrU7QWbixdUTY2pKtIS5VotWubCJ1BoY4rgG0HKBQDv5nvsOBOmqfS+PcSkQDS4GCBFhFBo+QMU5n/pumiBCxE3z9wfbvn6wwOoClTF9as3dOstV9c3SF0YaoUlsQjjU2J/+Mg4nvjm/TvmeWZ2JQaFnDGmKkdgEnoZfjoX8IALHms9PpQMvbOAOp5B5jEuw6tSVYcQmMeBeew5HPbM08h4OqKl4HQ6cVdV/ObLr9huN9wdDqyqphhrfCLmzH/5q7/60dfksD+Uk4S1KCnK9H+pFoJ3BO/5+sMtp9NAVRUoS7va4mTExfCku9ZSoEWBrCgtlraExFFxDKWtpzrDWkhuTINKiehnRBKoo8cSsaKAc1Ja+t6IZWBaTohnH1U6TyYWp2RhY0DykbCcCsbTI7UqWtdzEfRjl3MRYxSb1pRgAV3iehCKtiqntIutZj157obEaAPHKbC72HBxsSlfi8S6sMxNinOuaRrazZruYss0nBiHI8fjASEll7ygyZm27XDzzDg4QkiLa1GhtQNhSCmVeHXZoszycBRw3D8yzyN2GksF/cVvmV1PvZLLzEIUVywsOErJ65vX/+o", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-31", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "hY0DykbCcCsbTI7UqWtdzEfRjl3MRYxSb1pRgAV3iehCKtiqntIutZj157obEaAPHKbC72HBxsSlfi8S6sMxNinOuaRrazZruYss0nBiHI8fjASEll7ygyZm27XDzzDg4QkiLa1GhtQNhSCmVeHXZoszycBRw3D8yzyN2GksF/cVvmV1PvZLLzEIUVywsOErJ65vX/+o1+AErSakUrU04GzgdPc5FhtGRqVlPFhEtWhfuQUrgHNy+e8f79x/Yn2ZOcxH1l3CAmaRq1jYRrSDMhtnVWB95f/vIZCM5ClZty9V2w+F0z4e7r3h5/Ybrq5fc3OzoOg3CAYKYKlLWy80DWmvqui7VbTrriMvPMUfqtsEYg7OlKoqxsFafs7SfgHJEdH7Cn/YcP77n7pvf4+pIrROdnBECgmhYrTPbzRWzzUhZaF4pZ3TTlI1RV4SUGSZL8jPRj9BlcgPKVFRNQx1EEfWngFoVIHvbttRNQ5ZFH70yVZFTqeUIpk3RX6bMsZ8ZpxlnZ+w0kOxc4lFiJmZFFqYMokIEZEmdeObSqoT/eTdx//E9f////P3y/1b8wS//iMura2olME1DkpqIKLCQZbO8v7+j7/f8/ssvmO2MUBVKaUxdrpNWmioFZIrM01h4rynhQsD58JRVJZZNNy+bFku1mYUkBo+3M8PxwHDac3h8ZJ5G7DxhtGYcRw7G8PWHWw7DyHGcWOmaSml65589YOz7gRgC0zCgFFQVQDmG2rkEaL77/TdYG9BGY2qD0i0n2+PmEylGUk6oxXzUmqJlizkSkXgMxyg4hsx6ZYi15nK1gHJ6hbAZ3Th8tPjkifk8PGbZeM8b5rIJ8ynDOedPKSZ5GSIG53ES5uHEqMDU9bMrXefjUxuwqTRtXS2brqSpit2/aVvaOfDqGNj3FpdGdhcXvHp5jY8laHJyfhmgOYzWrFctzXpNu91yW1cICcPYkxK03bpEj6XyGobhRMjnUASF0p7KtEghqCqNQdN", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-32", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "4hsx6ZYi15nK1gHJ6hbAZ3Th8tPjkifk8PGbZeM8b5rIJ8ynDOedPKSZ5GSIG53ES5uHEqMDU9bMrXefjUxuwqTRtXS2brqSpit2/aVvaOfDqGNj3FpdGdhcXvHp5jY8laHJyfhmgOYzWrFctzXpNu91yW1cICcPYkxK03bpEj6XyGobhRMjnUASF0p7KtEghqCqNQdN0LWdol3OWGAPj6cg8Rd69/4aQLKtdjTyHaS4MmgLf+Te0F9qmOMy0UoRQnlIf7yfuHiwPB0/b1Lx5taY2Aq1iEagfRw6TJAXJyxcvuboyfPHVO0LKtDoh/MDDh68I3uOtx84zIQc+PnzN4XTLu69/TV1pLi/XeD8y2T33d99wsbrmxWVFcK8gHigwh0uEWiOr6xI417Y0TUMIkWkacd5z7AdCDIRU0ioudpsyZIkBHxw+2GfdNP/4N/8HytR022vGyXM4jHjnkFIz+Yh1iSGXo4hpJCpKotREoYhCMtiecXIc5glTVby4eY11nruHB0S0iFTYApUx7MeeDHz48EhIAUTk6uqSz9+8RioDSpUNXEv+wy9+xnbV8eJqSwqReZzZ397z8cM97+7uOfYDd/s7fAwYqQmUFm6KsYC+l0rXKIX8AR7ov7Tef/iIdYG7fc/+UHpgp3Fimh33Hz9gtOHFi2u69Zo3P/sFq82Wy5tXzD4we8dw3DOOA/3piLWWulshqhqjOnIset0vfvNP7G8/8Pj4QAZ++stfYpoG2dTF0ZTKZkvKRGuLzNFZ0tJft/NEfzgwTSPTNBBDXCbYFU1TU7crVN0R0Uw2kvY9JzEiBRzHCf/MtstxfyB4T388kKIlprFI5UhY6/A+cPewJyWBaS5QQpR6MwVkCIXRoCDHTFzg2Wc0ZMwKLzWiW1OhYSXJtYKbNcIYhG7RB4d+PcLte+LjPZFAsLbYbWMpPM65hSkFcgrl+i0PrCwEIeriFkwgkRipmYeR7B1a6+9Gcv2Y++R+4mIDb9+", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-33", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "rVN0R0Uw2kvY9JzEiBRzHCf/MtstxfyB4T388kKIlprFI5UhY6/A+cPewJyWBaS5QQpR6MwVkCIXRoCDHTFzg2Wc0ZMwKLzWiW1OhYSXJtYKbNcIYhG7RB4d+PcLte+LjPZFAsLbYbWMpPM65hSkFcgrl+i0PrCwEIeriFkwgkRipmYeR7B1a6+9Gcv2Y++R+4mIDb9+u0HWFbqqScCEVTdNhTMWLFzeEKAjNa97fPmDdl+zWG3abLYd+xLrAdrtFKI00DUYrVrWi7Tq69Yqvf/8VUtekdCLE9KRBjs4zDSOP9weyLBZrkRVNK/ns9SWbzYbXr16SBfh0roQjv/v/DA93t6hsIZdqOOfI4Xh62myXhIBPMfLfs7530111xf0lpcY5T4w82TZTHhiqmUp7Kg1aeaZp5uGxJ6kL0BesV2uEXvHh9gHvA22lySIzDccnmlFK5cYfx8wsBN6OVEYxuxU5O2IaSS5jR8/+8ZZNJyE+IoRCaJBVxuQtWhuULkdrqVRBR8bCcPXLBjs7i/O+bMIxElN86vn92HX//ktM3eJjxrrMPBVhvhSSGEplKVPpQSmhiShiZvkhcDFgvSV4h4mBTUrYEAqVK3t0dkAmRE/vLD4mvv5wR0wBqSJRwvbykogjIoi5QMrnkFgLSV03BBxznHDjzOnhwP7jA/tTz2k6Ekl0dUMUECgs4xRK/5OcqStZZH3PXNNssS4UcLlzhBiXFJCRfv8IOXN8vGez3dK0DSl61ps13rlCtbITwTviWWESI7nsogRfcsvuP97y4fdfcn9/B1KwfXFFt9mwMqYMglIix0QOoST7Bk+YSyL0cCpaztPhEWcLu1mp0g+s65q6btB1g9KGjCg9z9mfBYkchh7/zFORtUVna6cJ70ecO5YIq7zI1UJknEZAoao1QkZyXto8cak7xSd4aAlflaWFIgRZKFRdUyFRLchaktqaVNeIbkfWEdgUqaKdiOlUNtP0XY16AUWl5XqnRUN", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-34", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "wfXFFt9mwMqYMglIix0QOoST7Bk+YSyL0cCpaztPhEWcLu1mp0g+s65q6btB1g9KGjCg9z9mfBYkchh7/zFORtUVna6cJ70ecO5YIq7zI1UJknEZAoao1QkZyXto8cak7xSd4aAlflaWFIgRZKFRdUyFRLchaktqaVNeIbkfWEdgUqaKdiOlUNtP0XY16AUWl5XqnRUNd2CTnuchZWpYz2NkSnUP9iOTbf776yVM3DUJXSFMXCI8qHGjTdmWYeHFFRvGSNSFJLt49sFmv2aw6fEhI6TFdh9QGqjLD6CpJ27V0Xdm4pVRPg9a43A/eWay1zLNF6QqpBaRCtd6sOi4vNrz57BWJzBT8corybNYr3NgWG3YKVJVCyuKCfcIhlLhtYk7/tk33v/zln1N0hQrvi9tofzhx6kfePzwyzBN//9vbMhjyha5+HDxv3l7y+u0VN6/e0nUXHI4zMSX++E/+hMf9gf/2t3/LNJ2YpkOBaitwsywPi5xJqWQZSZFARPphzzw6fv2bf2Aa7rnaqYJ3NAbTKjrlUFWFdI6YIUuJTwmfEsLoBZgM+75n/s2vqSuFVmUDfK4R+O6bL5G64vGhJ6OJWZPGEzWesCT61qsddVPz4mVBGt4dRk695dQHZh/w2TPaGRNrcvZAICeHkAlVjPQFOWhqkhHUNzsEkdYE2nWNT54xZKZQnEcmZj72I0Iprrcr7DDycH9PfzrhnSuvUStsVszO8/HuS1Iu+uCza67VmlprrravStTRM5dpN/g8M/s9MZW03e3ugvVmw3g8LptrefClFIpEbu6xtli3U/DInNiu18W6DSQ383A78nh3z+2Hd3z44jcc7j9inUVpzcf3X7OeLtFVQ8gZGwJxtkRnGU/H0krYPxKcZe5P5FykfU1ds1nvWG8uqJuGbr3FVDXb3SVam9JuCYHZW3yOxBw5Did8fN6m6920tK88MTlmbwmxZEi6oPBRMIbCkNbzjFYBvCUMPWE8Eee5PNA", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-35", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Ebu6xtli3U/DInNiu18W6DSQ383A78nh3z+2Hd3z44jcc7j9inUVpzcf3X7OeLtFVQ8gZGwJxtkRnGU/H0krYPxKcZe5P5FykfU1ds1nvWG8uqJuGbr3FVDXb3SVam9JuCYHZW3yOxBw5Did8fN6m6920tK88MTlmbwmxZEi6oPBRMIbCkNbzjFYBvCUMPWE8Eee5PNAxRdOcIzJnJApDxgi4udyhleHViw11JRnnvpgOxkydFOtuxX3X8dC2TLPDZ0oqxuII5GnjXR5ai9N08T6RUmS2M3cPj4xtzTR3pSrOcYm7et6mO3mYkyToFanuYLXCVGU+sdttWbUdb97+AUpV3LyB128eeXF1zXq9ZrNeFaUFmVPfM86WLz48lGuiNmhdaGpGmxI9nxPBO24/vsdozX5/T1ykiVfrjovtht3FlvV6xecvL9hdbNmtKyZrmYaR0+MDD4+P9Id7vJ95cbXGaMnbz14gtYZY0l1izlSL9FEb+W/bdK+vdhSJjyT4gn0rR7EWmzN6MJyOHu8gzA4XwEWxZAdplKoWrFxNAi62O6z1pFAuhrUzxixA8FSAEVoKYjSk6EsemKD0ZYkcj6elD7TGVBIRLS5PRNUXudE00vc91tqSeRQC/qzXzGlJ8wy0jaEyxSzwXHF3gRQXnz5L5pKIHpXiMq1OGL2iripWXcfkImM/MluHdaG48aR4ki2F4InRF1G3+FR5pEzpKUlFu2oRRGpli0mExSO+2HxTTpyOPToE3smEm2eOjweO48AUHV5kgpREZQgiMNsytCOfkY6SLJeAwEUC+dw1zY7JOubZFfDIGZ2pymRZpiWAUUp8iMzWcup7rPfMzjP50ts9D5jckgphnWX/eM/h4R47T+QUFnNBGVC5ecY7u8DwA24aCXZmHHr8PJf+r/ekGFGyBBU2bUPTrVhv1tRNS9OtClTblMFQzHmhW3lcLveQ9QEfnzd0naaRGDzzNDDbIlOKGFLWBDRJCrIq96C", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-36", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "AwEUC+dw1zY7JOubZFfDIGZ2pymRZpiWAUUp8iMzWcup7rPfMzjP50ts9D5jckgphnWX/eM/h4R47T+QUFnNBGVC5ecY7u8DwA24aCXZmHHr8PJf+r/ekGFGyBBU2bUPTrVhv1tRNS9OtClTblMFQzHmhW3lcLveQ9QEfnzd0naaRGDzzNDDbIlOKGFLWBDRJCrIq96CPiZw9IgaCt+VInCIpxWJGyOUEIgFyIgWPmyeUT0QdQO4QpiIHvySx5PKZkgJnNLGpSF6R4zn1YdG7f+vNPhuMnv6MT5ux94FY12RZ0KJQGBln2NOPXufplFAIpVFLtVvYyDWmrksP3xQDVgacdbRtW4qBxXHZNBX9MHJ/6JeEsBLJVBnzKdF4UTycX2LKeclcM4XAVpkC5SJj54lBSe7uJNNsuX/Yc/dwz8PDI6fTEe8tu4sVXVtxuduCEEwhluFsKpwKpVVxhf5bNt2fvXm9GAEWnXcC6yPOR352ONBPE9988yX90PPu/QeOo0UcJqqmKRPzRZZiTAnIu7l+iXP+qaE9jwPJQFQl9VUKQVsZFInoaoTQIAxaVAhR8fXXd5yOA039c5pW4TkRsscyYK3FOff0xhpjCieiMovzJzL0R06HB7q2Kn3ji9Wz0wBUXXK9dNaFNYojhJkcRqbhREyZ1csbNqsVL6+2vL/bc3t7i3XlulWVoKtrxmEkx8hwOBTThp2RRpKELBtuTtQyU9WaFzcvIAfC/EClarQoABtFppISmTK//h9/j/COv3YDSiuqrl4m0ZJeGuamxkaDExVT+IJsLcr1SC1RlaJeVXStImWPC8/fdf/h119hneXu4YFT3zNN09NQJgoJVcOq26KriofTyGF2vN8fn4Azk7V4H+j7Hr+AZpy1nE4nxuHE6Xik04LtesU8W7KAME9Ypen3jzjn6YeRsT8VdcLQl+qZjJKKzWpF09TF2LPe0K7XNG33lDqQRUFopZzxdjmGjiM2BVyKTHZ6trz", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-37", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "sLcr1SC1RlaJeVXStImWPC8/fdf/h119hneXu4YFT3zNN09NQJgoJVcOq26KriofTyGF2vN8fn4Azk7V4H+j7Hr+AZpy1nE4nxuHE6Xik04LtesU8W7KAME9Ypen3jzjn6YeRsT8VdcLQl+qZjJKKzWpF09TF2LPe0K7XNG33lDqQRUFopZzxdjmGjiM2BVyKTHZ6trzwm69+h/eO/vSI9Y7JzdSra6pmh1nvELpBqQ0pWMbhEREsMkyk4JfTYwE3KRURStJVJfPNOstpHLhz7xkmR4yZZvVfefHyhrq7IPuA7494MqPIzOsamXcI15O9fQrwPPdzz/r37zjolp/T2fGVMqZu2V7dLHJExTz1pGdW/1qVFAuhNLqqaVZr6qqiMoaqbtG6JsSI0onNpmOz3fD6zedFX68UQpV7cxx6+r4np8CpH7h7ONLoFbvthq5pqCpD1zZoJVmv1hhjaJqSF6e1pmtq2qoiRc/Qe/7u//7vpJQY54lpthxPPf1xoO/78n1sOv7iz/43Xr16wauXF1jv+bg/4Hxg9h61hGyenZHfew2+/wLJxUgjFo2lQKlMZVLpDbYGFW8YpxWtNvSz4/I0U7cX1E2FnXqi97h5pKobcnTkWI6RImckErXAcuQ52ZfyVM9R0HQrttuXNOYCo1dI5iJlkS1CNkgqUuCpipjmqegqq4q2q9HagNaknHDeFog6iRAcQgSc1xQKy49fUtdopemarvTPyaToSd4uWWxpUXF47j/es388LJtxsUvLJFESamXIWeBmWyJIckEuqnN6r5JcmIq2brjZXSFEwrkKskAkha4Sbcx0dYsEHvaPOF/ix6umohINSyewTIVlxewdpMj19RUytBjfkEVJKO7aqoA8UpFVPXcdT0ORbzlPBpp2CURUkr6fcS4yzRbhyhRdKInQ+smuO1uH955h6PHOMfQnvHeMw4B3FkmiqTva2iy4v1jurWmif3zEOscwFiWCtzMsusqu7aiMYbt", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-38", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "2brjZXSFEwrkKskAkha4Sbcx0dYsEHvaPOF/ix6umohINSyewTIVlxewdpMj19RUytBjfkEVJKO7aqoA8UpFVPXcdT0ORbzlPBpp2CURUkr6fcS4yzRbhyhRdKInQ+smuO1uH955h6PHOMfQnvHeMw4B3FkmiqTva2iy4v1jurWmif3zEOscwFiWCtzMsusqu7aiMYbtaUzc1q/WKuu2omhZlqtL/p9zn0ftv9QAtIfpFcpgW7fjzhkb7/QPOztx/vMUFx+Qdm0vN6qJi1eSiO9flhJHUWCrKRJlXKAFaQCgDau885HKvhlQUG9Z5jscB5yMfP34EKdld7kg+YseRUIgWzH4mUaRabV1DymUIthxrlCq67rOj03v/pP6JMZXJ/2Kxn60lpsKdDee++zNWTvkpFzHG8t7HVPCNPiS0ijjvF133knZiqsXo8+n6V3VDl+H1Z5+x7vtyIuxackqLyUhilCIrVd61mPDWggCHxE8To5JPpoa8DM2sdfgQsb7MGjKZzWbF9fUln79+w6tX11zsWmZnCVIyW0dl7ULKK2EP4geuyffzdJdvXqliShBCceZPrTtNTCve7FZ4Hzj9ZGJygcPk2feOQ2/5ePeRfpgZZ8tqtcZNR6IbkNmhRcZIRb0YMMTC7TJCoLImOcHFzQ1/9Ks/Z7u+oWsv+N3v/gfz3KPNJabqkKLFJstsjwxjzzAOrNY13ari5csrqqoiCoH1jlOfcVahdCZlh3WJ2QpSem6l21E3DZfXr8pRQpRBSwqeZgw4EfAhE44jd7ffcBwnjo8HtKnQVY0MRXO6rroiRj8ORYuKpEJhhKZWBmMUb9sVu/WGP3r9FmU0I5Zxtuz7sSgPELzcbFAp89dffcUhRsZUqG9t0+BiyY7ebrfoes1s9wSteHH1h9Q4VmlkdpZpnli+dBkuPp+9cHu/LzQub9FGsbu6om0bTFXx+69vCYee+8eP+BBRWj1BtM/VlXWlndCfjoTgCc6W9ot", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-39", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "QVY0MRXO6rroiRj8ORYuKpEJhhKZWBmMUb9sVu/WGP3r9FmU0I5Zxtuz7sSgPELzcbFAp89dffcUhRsZUqG9t0+BiyY7ebrfoes1s9wSteHH1h9Q4VmlkdpZpnli+dBkuPp+9cHu/LzQub9FGsbu6om0bTFXx+69vCYee+8eP+BBRWj1BtM/VlXWlndCfjoTgCc6W9ot3JcXYaLYXG7brNTElprlI4IKz2HF8aimlBTDTtkWTfXPzkqbpuNjuMMZQNU2B4CwEPWA5yheaW0qREDw+eHxwJVkgi0XR8bwTwPt3XzH0A1/+5re4GLEx8vJt5spJPlu/oa0V0tQgA7FyhCSwacKoYpEWIiBk4PC4J8wz79xUOK6VLvI8H7i7e2ScLNdvPqM/9fz05z9BpMywPzDHQJ9csdIvD9ZGF01rWja+syzq2/b5EMLTQLMfihU358w8z+z3B5Qpcj6l4rNzBlPMS6uyDBJjzNgcCCFTaU/OinacyVmyWXmqJYYLzq7qcr9UVUNV1/y7f//v6fsT24uvGWdHPxY7tJaSWpuSSBIiMQemU084nyCW99yHsDwAzknnxQBh6pqq1jSrmp/89HN+9tO3/Ic/+VNubl4gVWJyE/VmxTBOnIaBYZqZrSM7uQQk/Ovr+zddqSkN9+WpWIgISxtQoykJD0knKlnhYma9Tmw6y3Y1k7xDERDZobLlq9/+A3cPD7jpQPIzilQCGxGlXyVL8KNe8pOkkEihqZuW9XIcTNkTc4GyqKpYeisDTS1JUZHizDzC7fsAUjD7gHWW4+lIf9pzPDxgtMBoQaNvSjzpM5Z1M0pLlC7gkhyXNFqlYRGZ22kmL+mqwTlkCkRXWAyzqYhS4UNhFAdXJuQyR5JRpFkTlj6r6Ht019GikXWNNwXe/MXjY8kQzoLq4oJWSMThgLYzrVasjeZSG4JQBCJXTY1pKo7SM0vHqlXUqmZjNJV1mKkm+EAMiSgqfihC+l9azpdjVYhpGS6", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-40", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "zPDxgtMBoQaNvSjzpM5Z1M0pLlC7gkhyXNFqlYRGZ22kmL+mqwTlkCkRXWAyzqYhS4UNhFAdXJuQyR5JRpFkTlj6r6Ht019GikXWNNwXe/MXjY8kQzoLq4oJWSMThgLYzrVasjeZSG4JQBCJXTY1pKo7SM0vHqlXUqmZjNJV1mKkm+EAMiSgqfihC+l9azpdjVYhpGS6kJ5KblgudLXqCm0tWlRTIMxENgfPF3RecI8ZQrmuKkNIiVVqE6KL8PSkV3pbqUFAMK5pUTCN1zWq1oapbNtsdVdVQd2vUkhN2JuykpTccY5EQBmeJKRKCI6ZISqFAmuRygnn2BuMgBbQoGX9SgRaF1ZHDDH5Ciq7c602H0aoUIDkic0TnBDEi3IydRoZTXu41ifWBcSohoFVVcXh4IKeIVplKa0wqqRmzn5cObEaFcsI8J/qaxYyg9adtQEpZrocvhhPrShjmmdZ3xl2mFIl+KsPJZywfFuVQTE/W46L1K0O7IuUMqKULgeFdAAAgAElEQVSSR2pMjkV7LuR3e8hZoKWirhouLi4Z51seHu44HfeMQ7GQO+uIoUjggi8oyyfrf2aReRX7vNaG7WbNxW7H528/5/Lygt3ugp/95HNeXF2yu7iiMi2IQKWha1b4kJGTI2e7pEWLJcn6X18/btMVanH6FEeNkKCkfsrrIsGqEhSIoWSzmtiME0N/gDiTvMAHy++/+EcOhxN+OpBDKpvuQqHXqsBZ9BKeWDbd8m/XVUPXrWnalpBmQiqDBy0lSgkqnYlGkGtFCjNTtJyO94QYOS5V0OF0YJ4GxuFI1xqa2rDb1kjRPeumsXbCmOIRJxXma/kgKPLSj7XzvBzHHNF5ZI74hfmqZUUSipRU2XStReSEIZC0JBlNiKmkGDwYVN3QZoNoWsZVx3Da88XHd7DwU64ursjawPGIzon2omOtFDtVwgoTis+airo1vFeeITpWjaapDeuVppodpmnwsyX6QKJ+shU/Zzk", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-41", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "V0OF0YJ4GxuFI1xqa2rDb1kjRPeumsXbCmOIRJxXma/kgKPLSj7XzvBzHHNF5ZI74hfmqZUUSipRU2XStReSEIZC0JBlNiKmkGDwYVN3QZoNoWsZVx3Da88XHd7DwU64ursjawPGIzon2omOtFDtVwgoTis+airo1vFeeITpWjaapDeuVppodpmnwsyX6QKJ+shU/ZzkfF5982XjxaWnnlMSKykhIjugmxnEoUB5TdJ5CCLwvZo3gHSlGyOXoKnKR8Ktz3xGxxNnLIlpf4NtKKarKsGpruvWGzcUVdd3RbbZoU1O1XWFAKE3mnKdXXIulVVG0wDEGQrCFSCaW2HqpSwrzc3k3qbQEtDBIKaiUwkhVXluwZD+jdUm4EE0L1LDqSH4me0srC6BHhJm5qQjRllaHFKRxZLIWISWm0hwfHrDjgBSRrm253l3ig8O5CRkTMkGlNEosbjNZHJxnylZceAx6afl471HOMY5jkUSJczqCKDyJFLF2Ij5zuBgWnva3WccldTw/STjLpuux3hcHYw6orJ82fxDLiabwvauqZrvd8e7DLY+P95xOh8LwmEoU0zxNxbLv3VMfWym9gKEKDGe96ui6ls/fvObt55/zn/7Tn/Hm9WvevP6MzWZNXddoXVGEBYGkBW2zYrIeIVVxWKZvwWC+Z31/T9e0lEp38Ucvb1ZJCA4LuCQVjKpQpVebEk1VIaTgD3/6ms9uLtgfjozTzNfvbpHJ8Obm4knSJdWZhrU4sHMRhUjpSHFgmvYcji2ZwMPhPYfjHbf3XyKlpm62+OCZ5uEpOSKEQIxxwbFFJucIMTC7ufTp3Mxm09K1NW2VWXXts26a9+/eUdflBo8hlKHLNOHsjJAVoq3RTVOgKYcSjT264vRCVqh2TWUqgl2876Kg7FZNxRNBJC3RNZ+/5nK1IlxeMsfE7/aP3E8DRccbiClySh4S3NmBk525/fgNr1/ecLXdLMD5IuQWOXLRVtQV1Kt2IdIVVrE", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-42", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "pOSKEQIxxwbFFJucIMTC7ufTp3Mxm09K1NW2VWXXts26a9+/eUdflBo8hlKHLNOHsjJAVoq3RTVOgKYcSjT264vRCVqh2TWUqgl2876Kg7FZNxRNBJC3RNZ+/5nK1IlxeMsfE7/aP3E8DRccbiClySh4S3NmBk525/fgNr1/ecLXdLMD5IuQWOXLRVtQV1Kt2IdIVVrEQNbWpi74VQ37mMRogxHI/pFgm7ilG3GjJOTNNQwHa2KI+0IX3SQ6+/FtCEN2SW+fdUvksff9czAHBF0DNuRI7w8q1MdSr1dLH71hvLmi7UuUqpfEuLL1DgdYGYwpu8gkRmAtcKQtJbTShNABLAbHMD1DVU7TOc1bX1hihiTv5BNxuLy+ptys65anzQJs1YIjCILVG122xyFtJpzKGRPPZK2L0vPzsJd47hrEnZsH9w57L3TVdtyZ7Dzlxf3/PuPAikAIjNXG2eOuY4kRMicE6hJRs15vS41423XPySoyhJCcEz+nUU7LDChFObTcoU6hfDx8PzPPwrGsiZJkzzNYRU0ErniE4mfIAnGZLTKD7HusDIbE8ICqMLgPy8hBeQDlC0DQVu+2a1y+v+epqx+HxESFEyb7TRY9dmeJYLeSxjrqpubi4oGlarl9c0XUtN9fXbDZrXt7c0LZd4biYAo+KxEVCNzK7iX2/px/G0q7IudxXP6Ly/36eri5MA/ITH2mxvC2xwymDjJBKSq3ICZE9RmuEEry43LLdNHSLvON0PCBywO3W+Jzxy5QvU1I0c06ksLASciBnTwgW6yYmOzDNJ/rhwP7wSEqZqtqQUsT7Ge893gecs/gQmKa5CPRDIOZYNoUUSvWROmJoOBw6vHvepns6HZhGTfSlMnSzXSzFiYvLF1RVhaoNKQh8zrilKj9vgLJqCkUqFjAJSiMrQ7PePDEjznKr9uUruosNtDV+mjk+3DPl+CQZCynipcBpiRWZIRRnW9vU+BgxUiFNEYnnHKm1LMc1bYq", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-43", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "hwP7wSEqZqtqQUsT7Ge893gecs/gQmKa5CPRDIOZYNoUUSvWROmJoOBw6vHvepns6HZhGTfSlMnSzXSzFiYvLF1RVhaoNKQh8zrilKj9vgLJqCkUqFjAJSiMrQ7PePDEjznKr9uUruosNtDV+mjk+3DPl+CQZCynipcBpiRWZIRRnW9vU+BgxUiFNEYnnHKm1LMc1bYqFnYLZU1oWNr9USNT/VHshLQL7lPIiCSzvRQieYSpDsRgWzqlY5EkxPW3wKURysciVnDKKXVVQ+r5x6WGqp1SQM3WuMHybpqFbrWm7DU27QiqDEGKpqjJCOsiFES1EeS/E0to4V3FKSTIlpTplWTL6pAFpFk3r865J09QoEqtOIqRGGUPdtlRNhSEi44yODZBIQiB1scfGoErwoQQjJcZ0IAqpzNoJHqCqaxCCdrVit7tkPp7wztKPB1JOTHYu0quq5NJFn5idw4fAaZqKFVfphSubn+49vwyQ7DQv2vxyapMS/GpFThGZS7vRzSPj0D/rmoilUg2LPVtp9ZQ2zGJT9iGQhWSyc5GDKU29KCjSQpvjW4NAKEP/tm2K1vZiy253Ua5lSqxXq8JlaYt54uJiy3q9pus6rq6u6Lpu2WRbri53S/VfqtozRL1U4sXZOruJyU5M04R1thQB3z4G/cDD+Qc2XQ2UvqoUEin1U4pmXtB0mOXDEUskTg4elSMmJ2rTklJk3e7wIXB9dYO1jtNwwoWAXbK/nHOFJuVmDo93zNby8HjCqJLEeX19xe7ymn/4x8TpeOL9h2+K6H9xIYVvy1aWnpXzJQU2hPJ9Fn1sRslESDOT9Xzz4Uv0MyVjfX+CDPcPDwt1v8QEeR/4ma7YKo3RikDi5GastVjvMfUyIV5v0W3L+4cvC+1+tuwur3j16i1SaYQ0RO8QwOWf/2euX13z+tUVp77n9u/+jvD119z+0z+QvUQmw2d/8kdcXV3hrq9obm/53fFIqhvSqqO9vOLi8pJ5KJb", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-44", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "WnpXzJQU2hPJ9Fn1sRslESDOT9Xzz4Uv0MyVjfX+CDPcPDwt1v8QEeR/4ma7YKo3RikDi5GastVjvMfUyIV5v0W3L+4cvC+1+tuwur3j16i1SaYQ0RO8QwOWf/2euX13z+tUVp77n9u/+jvD119z+0z+QvUQmw2d/8kdcXV3hrq9obm/53fFIqhvSqqO9vOLi8pJ5KJbcu37AxUCzXAuEf6pKRYrlwaeWB+1z16KmT6GIz30Ihf7lHM6WwVRcNNMhxE8OsrxojsNChAtLmuoS9qd0iWqJITBPU9GghsJxXW93aGPYbLdUdUXdtGhVg9KL7K78G0IUZcIZgMMiDzMLoQqWzUDq8sEnIXVFrVsQFWRFTu7pSPtj1y/+4BcMp5nfju/RuqLpVqimQhrNcLgnhshYPZKEYszF1l20tpYcLLuVoakUVVsobZXWDP3Ah3e3PD7smcaZtmt58eoG27XM08jjF3vmcebj7UeapmW9WUNIqLNDC0rQpiwDch8j+8OxbDRNu4SEFnKbnWfmfiQEj3MT2TkIfgkwEDw+7LHz9KxrYmqDVIJpnEuc+vLfz98bOT9Vtwzjkh2X0WosDwldwmkrU6rdujIoVSrYy90l69WKV69e0Z+OOO8hlweUVqUNYUxFUzdPg8Nz0GWRmErUsh+EWHTZYbFLx1yGZ/8/e2/SY1uW5Xn9dnOa21j3Ou8iIyMiKxsyK5EKMSghhMSkABUSErOaIDEuiRlfgM8AI2Z8BYaIQqKEgFGWyMqozMiI8N5fa81tTrdbBmufY+aeGe5unpKP7pKeP3vmZveeu8/Za6/1X/+1/j54jsc9k5s4djLHeZwczktGTZpn6P1u+w6PI1GAQAq6gOkibZ0pJGs7A/NRgG0FOUvhIZsKcpI5qylibYX3ns16JfQi7+nLQOVaZ8axIk4dCrCmExmXgveA4EzT5PBukoi26J/FGBesZiZC56LXouZpSQqUygttLGeY/Ih/ZHUkBkklJi/", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-45", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "eP3vmZveeu8/Za6/1X/+1/j54jsc9k5s4djLHeZwczktGTZpn6P1u+w6PI1GAQAq6gOkibZ0pJGs7A/NRgG0FOUvhIZsKcpI5qylibYX3ns16JfQi7+nLQOVaZ8axIk4dCrCmExmXgveA4EzT5PBukoi26J/FGBesZiZC56LXouZpSQqUygttLGeY/Ih/ZHUkBkklJi/8xRBnmkkoU8TCMlIwllbjueIqmGQuvz8xuInJO3yKZGNRVS0940pJKtu2sFlTP7mibWo2V1es9juqpgajsSlTrVZU6zXNdktz7NDaEFJk13WYpsW2LcPxwDSMdKNE/9lqcbpIU0YqDleA4vTDuiPK78wUsJjy/TosIzZnAj4l+i6ClEmmwJUbBiX+lWePpW01hog0LwNaUzfSyVS3MsjIVBVaGeYIhfzwowgtKKlQOLkKkwBlSmdeofygizafcEnJEvnnXOYAP8LW6zUxKOq2xtqGum0wtUVbw+3QMw4DvvZkZZlUTdaGrC0kB9FT05CDxWehbQVTMfQ9Q9fjJyeYd8lUlNULI0Mw2YDRDjdOQkNU9y29pgy4l8Nuhm0UTYEFtVKkGIjBL2Mw/TjJwKRxAOdK08L06Cl9uujjyZS/sjdmVK3MVlZKBmnEEPBKY9xEUDJFzxtxjq40QHgvHWg5Z4wxRRvtCednG1Jp8be2KkN2xEFbW839eA/iizLTIqXFpzjvRJUkeEKKjMXp9mWuiyv7fh7ruLRRf4d955QxsfvNIN8zGGUKtlu6qCqhdYi+VeH1ln0UoielQFP3xBg42wjtJ8bI5Ea8d9yu1/RDR2sUh+OR6GX/v/zyC6JPHPcHbt5ec9wdMIioY06SamqVHjhdWcXKFIpJGXKcyCijUFZhaoWyoHR4dFCXkc4W6ZETV2GNxSiNH0Z6s2fVVsQUUWGiygFtwapElQP9/ppurzjs3jKNEzmBG1uG/g5bN1R1i5tGckq8fPcSXUV++tFzGmt4enHB8cl", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-46", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "tJ8bI5Ea8d9yu1/RDR2sUh+OR6GX/v/zyC6JPHPcHbt5ec9wdMIioY06SamqVHjhdWcXKFIpJGXKcyCijUFZhaoWyoHR4dFCXkc4W6ZETV2GNxSiNH0Z6s2fVVsQUUWGiygFtwapElQP9/ppurzjs3jKNEzmBG1uG/g5bN1R1i5tGckq8fPcSXUV++tFzGmt4enHB8clTnjx/D7JIBMUpcLjekaeIjtBay83bt/wf//v/xnqzZb3ZUBmD1Ya6bbHG0vZNmdMxyyLe3+f0Q6Lcsi4R8Dnhc8LlyJQiLkZCSoSciyuVg1uX3n6VM0m6TEg6gS6OuahMyzVJ5880OZRLbM8vqJuGs6sLTGXlEFIycCSXCoQutYd5tIbW4sALaXMRg1SopbDjfJDrsy1K16CqIio6b6bHRbpVvZL50c+vsLqiadZLRPWXL/8Nr169ZL3eUtUtZ5cvyEoTSneVJmJZ4Zxlun1HThlbeN2Hmzt8P9Aagbm6bs9uf4cbJzYrgcsqa4nOcdP3rNcboRC6Huc8tWrJWhG8wk2KYdjzVD1lvaoxOmNMwk09Y98x9UfpmvQTyU8QPYf9gdE5XOxJ+ZHNEZUBnel74dZPo1ueH2kEMdTFJaYU8E5UVWY+7axAPA/+MUoKgpvNls1mzfZsK4oklcVqK1lOKl2kfsT5e+qZqFFLgTSVQABT5NpDZCoyTpOfyijUqcCZvhSM5wAryYGRZd7Ld3W5fqvTzaUNLasSJehUurAga6ExLG9QMCdVChRLtAGF5ytwheAwGZ0k5VFkjNFst9siqTNS1y3OK8ZYM7iB25vXHA87jvtb/NSjkseQqIwsEhRFCaWXzSa0DYXGLteiDGgrjleJtNGjg7o5ShMYIxFDUaDNMkrOB7+44/7YSXtvGRSSUy6TqiQin4eyaw3T0ElE4eRGxxh5+eXnRNfx/EwYFrd3d3RdTwiJFAM5RW5ubqhsRXfs8W5ivW5wU8JNXuYbDArdNihrUdQ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-47", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "S1y3OK8ZYM7iB25vXHA87jvtb/NSjkseQqIwsEhRFCaWXzSa0DYXGLteiDGgrjleJtNGjg7o5ShMYIxFDUaDNMkrOB7+44/7YSXtvGRSSUy6TqiQin4eyaw3T0ElE4eRGxxh5+eXnRNfx/EwYFrd3d3RdTwiJFAM5RW5ubqhsRXfs8W5ivW5wU8JNXuYbDArdNihrUdQolciFcZGXHILlv+nBdx5j0rKalqh2PphSCTflc0oUqlMmo1GqbKT5ucoJPc8/zsVFz1QlnTG2wehaGmCaVpSP7f1ovfurfvjV1z/L/J4gcJhKSniaSZpa0GAwcr9iIkUWAcbHHkf7fcc4OqZxIuhIDJmm0LSGvqPvOmJI1I2narZlaIrF6CyHT4qkAFPfE2NCx0x0nuBEfUSXgexuHJiGHu+8zFSuKq6uLhd1bEnBNW3TCC1t1UpBK5eGJGNoaktlFXWlycmwWbfCLEq+RL01Z9s1TV3hY0YZQy2tQY9ak1yUnr2XrDAsXX5KZq2AKI6oObMoQ5lKvYCsHtwH+VoUHYTRkBVURgrzlZlVZkr2FYuDLdhwmrOuTGHsKJRRBQIT2HAOCoXKJnCcDMsvVMOcHzheGW/6D3K6JCl0JAIZUxhjsoGSTiWqVAugrQsM8WCJxcFqjdJQ5ZpkoojAFcqOqTR1qmmamhgj27Mtk3M8f+/IV69v+dXHL/nk159we7uj6w+E6KmqgDWZtlZL99YsDW+tUEHqSh7uphbNMFMqmAsxX8G9UOL3t3nWKMgJPE1D+cyaw5tXpBDRWmAMW5mCP8ncUaNN6R1XC1WnrmpsVXG4fbvgm4MbCTEw7F6yXrV8/utfsWpXrFbnHPuew75jmDomN/LVly9RSbTRyJEXzy9wrqHvbElOZJ2q2rBqhWKnlTj+lFXJwvX9tCTydz40f58FH0ohxguDJKeFdqWMkbGFSgvGqs3iWHM5vGSIyhxNFry3vPZchGybLXW94uz8gqquqdqGDMt", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-48", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "cUaNN6R1XC1WnrmpsVXG4fbvgm4MbCTEw7F6yXrV8/utfsWpXrFbnHPuew75jmDomN/LVly9RSbTRyJEXzy9wrqHvbElOZJ2q2rBqhWKnlTj+lFXJwvX9tCTydz40f58FH0ohxguDJKeFdqWMkbGFSgvGqs3iWHM5vGSIyhxNFry3vPZchGybLXW94uz8gqquqdqGDMt8DWHTzJhCLuspPPBc+JNaSexMkVtKWTEV2feMEWVatKjPxkmcbhZB0sce0B9//IU0Dt0eUVl45+tVS2Ut++trjrc3HI2lblZoW1M3Le16i61LQ0CKJF9axZ0DH6X46GUcp9Xgh57u7o7D7S0xRJ5eXvLk8oI//8d/xn5/4NWrVwussFm3SEd2JZlFup9tK2NRDUbXNLVCf/iU4M5w4wXz3ExdUvPNmSImhaozj2UXBi9ByDCOjNMkPODiU3LOmKSLmgRS+KbM+J15vUvBSrLpnOR+d8PIoe/YHQ8YlTBKgjJTOtkyLAdrSAJ/pfK7KF263sRPzNCC967oKvpFAkqe1/sAI5QC36xU8g93uhS8JZb0L2fBzJZyYyHtlgVIafZm91m7ojhYZDjLPMSFBFkljNZLi19KkcyWpglYW5NVRciazarhervieFzjg8PagDYZW7HIqyull8KL1pq6THCva7uogoqksmGZ6vIDMunVSlJza9YFYx4LsT/QqYR34N1EDAkfJhSaQY9oJdG4DOWu2Z5Jqq+1IYbI7c1NIYeXkzQnlPO4duRNu2G1WnN2JnSa7njEhYEQHTYXwT4jFCplW+rG0DTidFUGO5PgC9WHuQjKfIDc46yPLRYtlgq3OEvGY8pmzlmyolwig5wSKupSEWbJGJQSbaslSi3P16zmrLWhqlqsqakb4VZao0t0kdBJMPasFVElsrrX05NeE7U4lwLpkpU8CzHLARRT2YBGZsgGPxf4Hl7T97euG5nGiXfvrqHQ35padAdleE8WnqvSTOOAsYaq0oU", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-49", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "G0DTidFUGO5PgC9WHuQjKfIDc46yPLRYtlgq3OEvGY8pmzlmyolwig5wSKupSEWbJGJQSbaslSi3P16zmrLWhqlqsqakb4VZao0t0kdBJMPasFVElsrrX05NeE7U4lwLpkpU8CzHLARRT2YBGZsgGPxf4Hl7T97euG5nGiXfvrqHQ35padAdleE8WnqvSTOOAsYaq0oU+mei7kRQ9u9tb3CRyT7J+Ze0UHPd7Momx70kpcthLpDsNI0Ypri4u5EBKSVqvlULX8kFEgkYXjnMt9C0lAYsFmQW8k/kL0YfSlQrrlRwSpvnu2bHftKEXnvuqNoQQ8N4vLBTxFRlvRDx2mhzzDSyNhovCc0qiT3jf/SWDlJwP6KJNUlu573VdLQexnB9zpMvidBNC+dLFeXpfsNwg/Po5wpXW6FAK+OJ8Ze5uec1/KLygyqjCvKSDGnSWNHG+YCTkB5lGllKRPVei3DuzHOYuGEqaOfPapNNE3i9nkWZJObHdRtabc84vLnhyccb19QW73S3OTSjtJZIska4xeim8SFR9X4201bzh5H2NsaA091prj3totpsWaw0X21VpGQ0yHKPv0USmEfa+x8dAP44z7Xaxzfacpl1zfv5EZnpqzTgO3Fy/Leq0blm/OEwMdY3VFavVmnFKOOfo+46sPJmALTIvtpJCQmNWKNVC3nKvJif3YZYuijEtjjdzn2ItD+IPMQHIBBJQCmXE4Up0q8kpYQq/VpeHd77nZJlDIEUWMTmIpVtKl0q7QuAYW+5tZYxE18WZ6xhFXkZLrUEVGtgsi11Vlu12Q11X1FVFRryvbVZkpZnC/QaOPhJLYTSEKCNDH7kkx+NA33W8evVaIsUs7alGa8Z+gDK9K2UYh56mnVviZS2P+70wEt6+ZRpHhq4vM0sK57iuMJUtNC8n3VaTx2pNdzzSti1PnzwpcI/AeWhIaqZApbLWeqnm15Ucftu2wfuJHIRl4BVLar5Zr1htzqhWawmgHmH", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-50", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Vlu12Q11X1FVFRryvbVZkpZnC/QaOPhJLYTSEKCNDH7kkx+NA33W8evVaIsUs7alGa8Z+gDK9K2UYh56mnVviZS2P+70wEt6+ZRpHhq4vM0sK57iuMJUtNC8n3VaTx2pNdzzSti1PnzwpcI/AeWhIaqZApbLWeqnm15Ucftu2wfuJHIRl4BVLar5Zr1htzqhWawmgHmHdsaNtasK6EYqnd9JSrIU6lpPGK0MykneppROu1JAKtjszYIKfMxZQU6ECFlplXcn9X62acnBTHPYMoqniC1TpoNSgY6GticMNRV0mlaaNVCSfZic816ZywYXhu5+T74x0VS4Ot1ympIJ5qUJqVTrT8hxZFP7jnLqr+WlVS+dGIVmisirp7lz5lnbanDPZJoytaFcN55sNw08+ous6vHekPJGJKBWWTzmjk3NUouYUEw8FXdRqTl3uMeDHBrtXl+dUVnO2aSVKS5nz7aZwGoU61ncyVrIbBoIPjJNf5sbGCORId7hjGmXzT2UUYQheTv4StU/TiLWW0U3CM3zzmrqytG3F+dmKzaalrURryhq9JB1zNfghVlvuwAIDpRIFq5lfC8td/gGQLtvVmhAlcorl9LfGEHNFSvcYWM73TAalZqdbcP+ZvlVmAcw1gplHmigDtoNgkdkL97ZRkJQiGUNQUtDLan5eS9GkMCCM1jS2Yd2sJPNCCU1LKYxNZQE1VJFcVYuW1nFwhEcOd9luL6iqBv+BhySNQ3M9uqob4ZanhDKGdntGu1oxQytSeZeDp21XkmEmlmHjSmkSMDmZuCaCAIm9P9D3A1obmRW8XjNnrAv2rWetNrVAbkabAtPp0nkW8G7ki88+YRwGuuOx4PKW9z/6KZdPnnP59BlV/bg2+hgCQSuCDyXSFWgn6UTUskfnwM35UAqghTmlTCl8ZbwL5DI0R0yhioCBNFklfJDD2xX8W5VuuofZ3D2WK74pU57f0houLekCfcUFkpINlokL+2aGHdP9xvu", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-51", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "9P9D3A1obmRW8XjNnrAv2rWetNrVAbkabAtPp0nkW8G7ki88+YRwGuuOx4PKW9z/6KZdPnnP59BlV/bg2+hgCQSuCDyXSFWgn6UTUskfnwM35UAqghTmlTCl8ZbwL5DI0R0yhioCBNFklfJDD2xX8W5VuuofZ3D2WK74pU57f0houLekCfcUFkpINlokL+2aGHdP9xvud9h3ClOUhzGWHZqH2yD9zcbTlzFDcF9EWJnkp1RSnuji5VGSKdcHZ9Ox07+GJjCh9tquG8+05KSnGcRTHFKXnO8apbKjS1plzOYXKv0mEMJByIAT3wOnKdeof4HS3mxXWaNYrqbGqDJt1u7x3jIlxlM64bhiYnKPvB7rjQNcNHI+D4J5uLL35DW6a8G6S09X7Ev0rfPDy0HiHNYbD8cDZds2TJ+cYLU0njZUoQR6gctvmNfxampNLYjIXGpECX3G0xVcvqNFjralqrBGKXEgJFSMpayyZmHQZsCL82nn03cNxglqbkh3dD2CZP4NQcmRwdmZOrwUmMSCttVp0vyg48owIz45dKITyDBulqUxdnC5kPWPaMwOnpOEq4Y20Jw9OJN0fY+1qjTGWyyu3OF0pzklGF2MkIgGIrmtMJdX2VAo/wqW1UjBUmlzS2RjCsq8lFfbLszf0vSgaQOm+2ixrPUd1UkEuDtbcZ4W6jE+UzDDhp5GvvnrF0Hfsd7ui8FBh6jVZVVTthiY+bk1E1brMdwixNElANsjM5ZJZa6XvqaBzzUbpAitkfDlkUiyMBq3RSQqxCmG+hMJ+DCne+6Y8l4/LvS8Zb9amZIOUoltRlikyViKzFMvvUNY0LwHfzH5I6btrIt/qdAU4hhRLVKjMQo/QUSp9db1GK4MpffZaqXJRchJklYVBUKKKrDTKVEs0nFIg+oQpHURwv9kEM85LkWy1KrIseS34Z3KlIHdflFnkRWIk5YiLIzF4pkn0qVSWvvElp36k1XUthIlMwSNl/BxaYbTIgq+", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-52", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "LvS8Zb9amZIOUoltRlikyViKzFMvvUNY0LwHfzH5I6btrIt/qdAU4hhRLVKjMQo/QUSp9db1GK4MpffZaqXJRchJklYVBUKKKrDTKVEs0nFIg+oQpHURwv9kEM85LkWy1KrIseS34Z3KlIHdflFnkRWIk5YiLIzF4pkn0qVSWvvElp36k1XUthIlMwSNl/BxaYbTIgq+algw8yxQAPuImkSfq+xHnHYeuL5y/gWkCpc4IXrquhBgig0ysMaw3khJvt1vapmKzrlmvGsQviRPLxeGoglXPfjPNDJR8z1RIueQteX5AHtb4f4DHRXrqxVEYwdbLH7mGosxc0rP77p37a50PiqUQWyIKeYhnhkoux6X0zFlF0fB7kLlkGe4Sy2etyvzVnGSYztu3N9zZI9ZcS1SkVFH8kJrDDFGRNSRVZJYSu/2e4B/HSa3qGltZmlaaLlSmRKuKVCCRlOV4SEaykKwo4w9TGdQSha6VojBjUl7aoJc1ylkw35CWhoOlcFzWhpnRU6Cf2enOTk3YH/OsP03KGm1XvHj/p4TgeDYOCzx38fQ57faMiJJJdo8wXXyDd55pnBj7QSbwVZacJByaoY7KCmastFmytGWvF+w0xdlPSJZsrDQnKa3xQSil/uCWbFt+UDIrbYSTLYXkkvWUgp7UVkQxPAVX2DmFb6ZLC7nmPhtjVt7+B/J0FyJ7+VsE7AoRWINCil9ohSq8ypQzaJnRKlVi+foeO1XLqZsLQP7w9SkRzPyzwOKMTZHdSSVy0ZGvOduHfyslGzwpidZNkEaAHCOziPUPMaP1fSD/wPRcJESG9MzFJPnRjG8DwYtUjHNSIBinEe8nUpTOmnmgyFz4aRqZcbpZt9S1UHbq2tA2FlsZ9L1/Kv+R5ov7bIPlYFgud8YP8sP/zwJJPIh3H2WzQ1/WQN+X+w3lPiWZS5u/xhoph/Hyr9lRp+XalvO3RCrq4W+q+0KZUhqTWVRtAVFHzsJSSCkR3IT3CaO", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-53", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "PMaP1fSD/wPRcJESG9MzFJPnRjG8DwYtUjHNSIBinEe8nUpTOmnmgyFz4aRqZcbpZt9S1UHbq2tA2FlsZ9L1/Kv+R5ov7bIPlYFgud8YP8sP/zwJJPIh3H2WzQ1/WQN+X+w3lPiWZS5u/xhoph/Hyr9lRp+XalvO3RCrq4W+q+0KZUhqTWVRtAVFHzsJSSCkR3IT3CaOC0M2UkgLJ8oLFGWHIWctr5czQyyH5GBOKnMZaPW93SLoopdwzN3I5JJISip1ACJlsK/m6qSW9LXtvfsbzkgbHUvyL1FVcjtBlTZW6P5SUkuYPpUor9NdW88EfQBmadkUVZebBHHnXTYuxthzgj1qSYg+CoxDRpszdeOCvZjqbyroQUtSD5zQ/WL8SMmRARZmjUaoZqRxQ3ody0JeHRSE1ECxktTB3BG5hgRBSnCfQxRKhewpOg1r2qTSlzM1i36Qo/n2mfgg96GQnO9nJTvbD7PGTTU52spOd7GQ/2E5O92QnO9nJfkQ7Od2TnexkJ/sR7eR0T3ayk53sR7ST0z3ZyU52sh/RTk73ZCc72cl+RDs53ZOd7GQn+xHt5HRPdrKTnexHtJPTPdnJTnayH9FOTvdkJzvZyX5EOzndk53sZCf7Ee3kdE92spOd7Ee0k9M92clOdrIf0U5O92QnO9nJfkQ7Od2TnexkJ/sR7eR0T3ayk53sR7ST0z3ZyU52sh/RTk73ZCc72cl+RPtWjbT/5V/9vxlEB0hE4OKiGf9NHS1VpL2VVqBEPdMkg0qzThh4Exc9ohgS0ecH2pCmyBdlYnB0xzu6rmO3u+P58xdcXV0tgnza2KJBJkJ3BsrXohtVrqioc6Z7dU5Vru8bl//f/hf/2fcWBfuf/uf/MX/4/of883/2zzHGAmpRDu36I8Mw8Ju//nfc3t7w8d/+NSkGyInVumW1bvn0i8+52+8YQ2QYJ3778ccMw8jx2InwnzZMzuG9X9RntbnXZEopFeVXkTOv62pRzpU", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-54", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "0xzu6rmO3u+P58xdcXV0tgnza2KJBJkJ3BsrXohtVrqioc6Z7dU5Vru8bl//f/hf/2fcWBfuf/uf/MX/4/of883/2zzHGAmpRDu36I8Mw8Ju//nfc3t7w8d/+NSkGyInVumW1bvn0i8+52+8YQ2QYJ3778ccMw8jx2InwnzZMzuG9X9RntbnXZEopFeVXkTOv62pRzpU1MdR1zWazWkQntRY11YuLSz766Cf8y3/53/Hi+Xt88N5Hcp0XgOUAACAASURBVN/uBdWYhaqUqh4llPbP/qv/OBtj2KxrpsFz3E24CYIHZUeUjrR1ja01m0uNqRS2Vgw7Q3+niUGJ6GqeRF0V80DdOBDzRNSOrAPbi6dYa4A9Tat574MLktO4zvD28z27Nx1V02Krhu3Tj0gJbt+9YeyOHO+uFy231aalbiq2W40xECJMU+TubuS8annarvn3P/oZP33ygt97/oJ10/Kf/w///fdel7/868+y1ppV3eIjTD7TTT2DG0VOnowxNRL7aGKCkGAKERdE1rv8FBlVxDblmbA5UeXE+aqhrS39eMQHz9FlUtagGwY/cRyO3B2uOfY7bvvXjO7Ivr8m+IlpOKJUwqjM+1dPeO/qCZf1llZXjONIAmxTczge+PSzz9hsN1xcnPPe02dcbM+pVY1Rmv/mv/4X33tNXJjyrGX40GatNvU79PlSFN24WVstuLAIyypEfDWEQPS+KJWLEGoOEXcYUFph1y3Vdkt9dUnO6l41W90rmxlAazDq67q1CoU26u+oZasHXz/8+XW7/p1r8p3ClItM9gPHtXzYBxLay5suuu9FbJL5KmdxSnktXb6dFlHKDLkIT6pvXoMID+asl2uaVW/l5WelWyWqnkVlcN5ci4x8Eb38hmzfty3B37FpHHFuIqWE1rMY5PwWSQQzdcaaTNMYQhTZ7ao2VHVFs1qxCgEVIxjDar3BVDWr9QaFONfD4cgwjozTSIzpXjByFuVDlYNQ3avnLp9", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-55", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "M9gPHtXzYBxLay5suuu9FbJL5KmdxSnktXb6dFlHKDLkIT6pvXoMID+asl2uaVW/l5WelWyWqnkVlcN5ci4x8Eb38hmzfty3B37FpHHFuIqWE1rMY5PwWSQQzdcaaTNMYQhTZ7ao2VHVFs1qxCgEVIxjDar3BVDWr9QaFONfD4cgwjozTSIzpXjByFuVDlYNQ3avnLp9JHkBSprIWa035fUUKETc5umPHeDaWTz8LET5cn/y1h+n7mJsS1oA3Ce+SHKoxE2NGqYQGlNYorUWYshzS2ogiOIj0N1kERV1MRSdSyxpEERhVypCjIWHEoefM1HuSN4Qpk0Ikp7xsvOC9iCeqjDZgKkVKLN/LJHyIIr2dRLzS6IzRCU3Eu5GhPzIc1uAeJ0wZUwk2YiJECCnjY8LHJMrGWgQ81SzGquUvo5QEEGWfpOIe7oUmRSy+CAjL/1HyWtZkElpELqeJYdzTj3u6ccfkOnwYyNGjcqSxhspYVnXFWbtiXTXUtsJqS1XXZBRV2zA6hzKGDISUmLynHwfGMPI1jdHvacv+/cb3Fnso8DgL1xaF5JRSEe5Ms8MpwpRF2booKsuayt7QTSUiwFoVReni1+SNAVXWOKOLwrT4m6/7sDnweahKO1/37MO+j+TktzpdebF7HXelNSqJgulyMim1/ExeHG45xcq6aC1Ph6hhi1S3QdYkirAn6oFzRT/8cCJxboxGG118+gNHo5jF3kk5oTD3PhZYJHOT3ByRSM7zc/po5/L27Rs27QrvHeL8rNx0pUjJk+JE3QTWm8zT5zXew+QSm23LenPOCxSby4khBoZxpA+Zuqp4/vQFMUS893z++Rdcv7vm5cuvGMaBFMP9PVAKrQx6kayfD8SwfJm1InnP2facJ1dPGIYB7zyHbuD2+pZPfvsJGssf/P4f3m/4sq1/qE7p7bsRazWhhxAC0xTw3hNCQBmFtpqz80qcnhpFBru22EZTrw0qryBXWKWIwXFzd5BnxlT", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-56", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Sm23LenPOCxSby4khBoZxpA+Zuqp4/vQFMUS893z++Rdcv7vm5cuvGMaBFMP9PVAKrQx6kayfD8SwfJm1InnP2facJ1dPGIYB7zyHbuD2+pZPfvsJGssf/P4f3m/4sq1/qE7p7bsRazWhhxAC0xTw3hNCQBmFtpqz80qcnhpFBru22EZTrw0qryBXWKWIwXFzd5BnxlTEEAiTQ1c1OjfkYUVAc3eTsVUku6OoxkbN1DuICZ0SOgb6w4GsQOtI1SjW5xXTlHAugg4kIvvDBDljrUKhaWtDayK1ntjfviLv7lC3e9a2ftSaDBFUzLjoiQl8UhwHx3EYWa1qqsqwtlae+RjJKGyS/RQSqHI4aGTnV6pkdlp2gEGjVS4BgMUqzapRpJzx2eMON3z15m94t3/LXXdDzgOZgA6eSmkuVmsutmd8+OwFm3rFulmhMCg0q3UGrVmdbVDWYr/6gmgUfXS83d1yt9/TXe+ILjxqTbS+37sPfQjc+5pZ3Xd2tikmUojypyglxxiWyHcJNub9bAxJA5W8l1k16JQw0WNMxqSAMhWqSL0vwV853MTpamQ/5AfXfn+d95+B8j1KEKS+cw99q9OdU4CU0jcWZn6zb55Y93Le9xdSTgEyxCSy7DFImhATMWRSyotsORhyCCJ5nEKBNGKJIsvrPkhF0hJYFwddTqMHAfdyLQAqaxG8LgeHeuRRHYsjEdnr+4gFYJx6+n6PmwaC88RgiL4iuExKDUq3VHWgyQaVIlpbzrfntO2K99/7gGkc6fuB9AFcnl1xdXXFMA7cXL9jHEd2+z0xRkKIIkdvZsl3MErLoaU0Ck2OGaMNbd1ilSXWgaw0TVVzc33Lk8s73OSwxooU97xQjz2FljvvSUnjnCLGVGTNI8qmJUKI0aECpHEiRQNZEUODMpZK1xhV01hNjJ5+cHKIqkwyiqoyKFNk3ZMjZ4Uun3M45AJNGaJaoVpQTYO2Fav1Odpa6jrjXUd/VEyTY5ocTas", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-57", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "2+z0xRkKIIkdvZsl3MErLoaU0Ck2OGaMNbd1ilSXWgaw0TVVzc33Lk8s73OSwxooU97xQjz2FljvvSUnjnCLGVGTNI8qmJUKI0aECpHEiRQNZEUODMpZK1xhV01hNjJ5+cHKIqkwyiqoyKFNk3ZMjZ4Uun3M45AJNGaJaoVpQTYO2Fav1Odpa6jrjXUd/VEyTY5ocTasxRuGcRAdVXUFUJKdQGUKMdGEk5UgdFI3+zhjlaxZSJqfI4D0JRcyayTuRTU8J80C/XJ5JieS1BmsUSZVMr/gpY8TJGpXLfVdLtJ6JJBI+ekIMDNORvr9hmu4I8UhmBBVROaEyGGM435xxsTnnanOJURqDYQoBnxLjNBFzIu1v2R32HIcePSnMqBmVwaIZbo9E/zin+7Vn5mEmDQU6yITg7xOvOcot+z+jQMu9VgYoPmrJzpRaKlVzhp6SQDU+JDSBkZGgAklZdIlurVFYo1mt6gKBfs15ldfj7yTGiy/8hs/7Nvt2p5slpJ+x1NnuHa3CGHGWX3O96T6Bz+XDZzLEQE6RnAR78S4s6bPWBqU0iooYPMFPRO/F+QZxvvNrziHqjGbAjOVmtM7LaTNH3vdAgkKr5aK+9pm+rwUfCD7gQ8DYiDIWrTRaQd/t2O3e0ndHpmHCjxbnFNOoWW9WKDY0rUJZR02mbhxPrjrOz8752U9/zmF/4PbmlhdP3kcphYuOcRr4q1/+FdfX14y//jXDMOKcw1YaC0s6ZGqL1qDRGCXOqNIVm2aNaQ0KMHWD0oaXX73kbHPO0I00dUPT6LI+FMf26GUBJlJSjFMgJQRWsBFjZjgk40NPzAnciDYV4zFTNw1V09LUK2rbsF3XpOgZ+oEQJkIc5M6pmpwVKSty7EkJKiM46PEuglZoW2HqFfa8QVcWU1WcX72gaRsun2yYhlvurhXj0DGNPXWjMEYxjBIsrNo10cFwF1CTZxod0XmOAY5xj82PWxgfI947jvsdKA2mYnIeHyM", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-58", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "26GUBJlJSjFMgJQRWsBFjZjgk40NPzAnciDYV4zFTNw1V09LUK2rbsF3XpOgZ+oEQJkIc5M6pmpwVKSty7EkJKiM46PEuglZoW2HqFfa8QVcWU1WcX72gaRsun2yYhlvurhXj0DGNPXWjMEYxjBIsrNo10cFwF1CTZxod0XmOAY5xj82PWxgfI947jvsdKA2mYnIeHyMxViSTBGohk1OAcnAaDdbAjMEZo9BKUVlxuhrIChIKnzIhRaIKhBwY3IHJDdzt3rLbv2Ic3hJjh9IjC5gUFJWteHrxlCdnlzy/eI5zjmma6F1HN43c7O4Yx5Gb/Y7JTRy6I+SEJmMiqJQJx6nAPo+z+4x4BgdkL8aQSCkyDiOZjCm1DEAwxwIbZKXRVuoYJR6Vn9FyKM9+K0ep50SViDkzukByiTREjlEzRo1NAUPmYlOzbmuaxqKUJqUoENVs5XC4j9Huo/Q52M5ZfOZ3baBvdbohhMU5zX9mhzantnMUvDj7BdxWy4kRo0SWMY7kFMjJ46ZJore7A30/lBNe06xXJBKjF+cyjQNdd6SpK7QpjtnY8oDq5YOrBxHuQ3wh33tc9IzjlJuutX50YNeuGppVg60tSitSioToy8MyMA4Dzh0JYaKqHahEJBLSnv0hs+9GnA8LNn6xPudsdUajKw4+Mx4Hri6uWK/XDHGirmt+/os/4P0PP+Sjn/w+h8OBm+tb7nZ3HA9H+v5IDBI96aSwWj6nMobu0PPq5WsqW2G0JlqFMnD99po352949fIVm/WGzWYDC26YyQqevffsUety+aSGrFGxIcSM9wlbZ0wlaReAzgpiRqkKRY02W1DnoC9Q1RZTNzSXl0DmMlq8n5jGHqUNylSgZRPaUjgMLhJjYho9MRsCFS5GQkxM44BzsFEVul5xfvmCqakIvsPqCgtszi11ozHmHDKEAfqj4xj2pRYsGy2SiCmjHulf7nbXdN2Bj3/zt+Ioqprz8yu223OCV2gi02jQKhP", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-59", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "GrFGxIcSM9wlbZ0wlaReAzgpiRqkKRY02W1DnoC9Q1RZTNzSXl0DmMlq8n5jGHqUNylSgZRPaUjgMLhJjYho9MRsCFS5GQkxM44BzsFEVul5xfvmCqakIvsPqCgtszi11ozHmHDKEAfqj4xj2pRYsGy2SiCmjHulf7nbXdN2Bj3/zt+Ioqprz8yu223OCV2gi02jQKhP8yDQ5DseOYz/S9SOXF+es2hUX51uUMTgfUWSMgikGhuDZjwODc3RDjwuew3CHdwPd4R1dt+O4e0uwgmdHNLE4hDF63uxuOHRHbq6vcdMkeP/UM3nPMPTkLIei1pqsDePgGLoei0aj8N1ICo93uvAgRc9zPQZSysR4Dx/qB5lF1hLMzQ46SwxWakKZmGdUVhG8BHPeS0Y9Ok8InqnvGSIcHVwPgf0YOU8ja5352YfP4WxDvtoSQsSHiLEKbTTGGgnYSj74zSxf8OECj6Z5J/1u+1anG0P4O4D3HE4/TA2+FgXzoAKpBCXJKZJjJPjpgdMdmYaem+u33N3tAI3WmtV2K85MZ2KKxBgZh55jZamqWirxthYnq8Xx6oLNaG3IOT68Ar52YboUjVLBkM3fBfS/y5q2EcaANaCFuRC8x3vHNIxM44B3PTE6jPVklalyJESFOwb2e8c0Raw1VMZwtj1nU6+olYWQmLoRe2nYNGtUMlS54cOmKtGB5u5ux8uXr/nss0959fIVzntiYZVQop+spcA2DCPJJ+qqwlpLe3mGBu7u7ri9ueXm+obgQslM5E8s6dpjne7ZZQ3JkP2KEBLTlKhaRVUJLpkT+L5kQTmjdIPmDKXOQG9RdoOuG6rtU7QxbJPBTw7TddimpV6tZRNoTVMZtFKkGAk+0XcOHxVj1By6jm4YGMZACp6oDMrUrNYXaAWr9S2EAH5ku6lZrQ3bbQsZDm9G8qhIgZmAI5FLFpiAB3DA97HD8Y7bm2t+89tfkZXG1DW//9Nf0DQ13mk0iWkSB+PdwPF", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-60", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "aRVUJLpkT+L5kQTmjdIPmDKXOQG9RdoOuG6rtU7QxbJPBTw7TddimpV6tZRNoTVMZtFKkGAk+0XcOHxVj1By6jm4YGMZACp6oDMrUrNYXaAWr9S2EAH5ku6lZrQ3bbQsZDm9G8qhIgZmAI5FLFpiAB3DA97HD8Y7bm2t+89tfkZXG1DW//9Nf0DQ13mk0iWkSB+PdwPF45PWbN9zt9tzt9vzeRx9ydXHBulUoawlulIzFKDo3sRsH3h12HMaeXXfAOce+uyG4nvF4Q/AjfuoE02wrvLLyfCjNlAJ33Z59zLyZAm6cxPF6J/vOBypr+fCDD6mMAWMJIREOPWiDURo/TKQfEOnONrNr5ow1JYECZmxUKVN+DoFQVF7gxBSLs01Sp/ExkDLErPCTJ/iAc44YIr1zpc4wsp8S74bEq93ATTfyfh64sPBsvaLRmhSkWDcME1VjqVoLxmCUOPQlx1f3zncOQmff+LsYGLN9J6arCqg8R5O/KyPPufh3tcRLBW5JxOJsDQFjFU294smTc+q6pm4q3rxZEXwgZ6mWSkCUBBsMHu8ctze3PMhJ5INrDVqjjcUYi7a2YJQao00pwFmJ/LTGaI0xWjY90GTQJWr6vqasBqvJgHMTfddz8+6a3d0tX3z6W7rDLdb0kAMpjuKAXOT27i23u4m7nce5xNl6w2a1Yf3zlqAmut2B/c0dt2+veXb5hLQ9Z7Ve0xpFnVbyYCZFVa1YtWd88P5PGMeJ169f0h0PvHv7JUPfcfPuWghI2qCrCmUrxuBJznE39SijqZsVH0wf0m5athdbrp5cEQuG7kMkPdK5APzpH/1TlNJYVgWXS6QcyDkyuYkYI34MxJDonSdiCLmmWa1p1muunr5HuzqjjzUpQmzeRzWwPc8YW2Ntg7FyuLaNxWhFTMKQ0E6Ks5sMq/2O8XhEDSP97obXn/wNd5XBvfuCGB3DcGQcjkxDx26faRvF0+0anRX7dyPd4HCTJ0cJDgiRHNJ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-61", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "athdbrp5cEQuG7kMkPdK5APzpH/1TlNJYVgWXS6QcyDkyuYkYI34MxJDonSdiCLmmWa1p1muunr5HuzqjjzUpQmzeRzWwPc8YW2Ntg7FyuLaNxWhFTMKQ0E6Ks5sMq/2O8XhEDSP97obXn/wNd5XBvfuCGB3DcGQcjkxDx26faRvF0+0anRX7dyPd4HCTJ0cJDgiRHNJSaHmM/ev/81/RdQc+/fhvQWtM3bC7uebzy9/y5PIJbdOirSamRD90dH3P7c01+90d+/2OT35zxXa95tmTSyprSMFLpG8rvMlMBg5+ZEweF5xgsGkkuoHx7h1+HJiOB1bnZzSbDVWzlujRRxpb8WJzhvcjh/0NqR/x48jk5BnoD0fONlv++J/8h1TGMnY9n+XP8W/vaOuWumpQ6wt+CBYlKHTBYhMLIyE4TwgR5yM5K5yLpJwJIRJCwMeA84GQItMgjrTreybv2A09o/P0o2cYBtzkGfoeHwJjgTHd5JiSoo+adVuxbio+/MlT3j/bsLKK5AOvX75ldxj4/Mt3nD875+LZOb//4VMuz9b3WfSMHT/8TI94NL6DMva7l008fP57I8U5hc9lYRMRRcIUwLqqNOv1is35GefvzuiHHjdJxKWrSpxuSlJsMFoW3XvBaXJaCnxK6A5ShTQVxlhsSaVNiX6tFUjCaGE/GG3KminhxT7SweivVTwTzjmOhz13Nzfc3tzSH+9oG49WEfKw4JtDf2B3e+SwjziXMV6wMWKCmIn+vmgYy0PWmBW6MkJlypmcFE0NcZVZrbYA1HVN1x0xOnM87HFjIKdUCgSarJVQloLHTR60YpUVLni00VRVRdu2hOCIUZMJheHxOHv+7CdopamNRI0pZbx3hOgZpxEfPG4QaCUeelxUxGhQ2qDJVHVD3a7oBkXICkyN0RpbGbS2aF3fZzSVLUwWwWJtJe7QItmAUZb1akMcjtxev8YTeRtGcs74GJmmkWmSg2BqoPFgsqbfO0YfJOL", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-62", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "87HFjIKdUCgSarJVQloLHTR60YpUVLni00VRVRdu2hOCIUZMJheHxOHv+7CdopamNRI0pZbx3hOgZpxEfPG4QaCUeelxUxGhQ2qDJVHVD3a7oBkXICkyN0RpbGbS2aF3fZzSVLUwWwWJtJe7QItmAUZb1akMcjtxev8YTeRtGcs74GJmmkWmSg2BqoPFgsqbfO0YfJOLKaUldyZBU5pGQLp9//gnTOLDf38qhb2tUzAxdx9T3tG1LAkKKHPueYejZ7+44HnZ0hz1Tt2fVNhxuz7BWk2MUWlhVkxtDag0jEacSWSUyGa080Y2MwxHXD4zHo6ybMqxMS20z0XkaNGfGMCnF6B0uOLRzMAVyCEyHAyuluVytaeoaZyru2ndUKBplqI0tDJrH9VfNmfFcNEtRYMwYItM0EWJkmqJEragCE3gmJwfL6Lw40l4YOfvDnsE5brsD/eg5DhP9MOCco+96vA9MXgI4H6R4lnTNB1dr1udrVuYJ29aQc8Z5z27f8e5mz+dfvuUiBI458/Ryy2bVYI1Z6GO/67N9nTr199u3Ol1thE6UYlzSeXlxmMPNrznAYikGYgxSZNOa1arCqoqNrUgpMrmBlBpQmaat2GzXtG1hJ2gNJQW6r2pK+ryE9llo4rI5hHYmvisXEnVgmgY5PWOhlM8YrtaFKaHZbs+w9nEV6Xa9olm1mMqSYsJoTXCOfn/g1eevePfuNX3/CqMDl5eK7XbDs6dPaGxg2yjaqwtSqtBZs2oatusV27MNV0+u0NayubjAec+73Q2XK4vNNd0wFKerub294+XL10Ib04anT5/y3osrnj97jxgCx+Oe7rDn3ds3vHv3hnfXbxlixAdPzE5wyVHRD0eO3ZHLq0uqxqJMRkeFi14W85H285/+GUZrmtqWQ08tRPZY3v+zl59xfXvHx29+wzAFBpchvIY4YEyFfgF19QSjKvrCnfVeCq85D1hrMdrQO1sOPgVKk8paaG1pzq9ozy5YtTX", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-63", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "94+XL10Ib04anT5/y3osrnj97jxgCx+Oe7rDn3ds3vHv3hnfXbxlixAdPzE5wyVHRD0eO3ZHLq0uqxqJMRkeFi14W85H285/+GUZrmtqWQ08tRPZY3v+zl59xfXvHx29+wzAFBpchvIY4YEyFfgF19QSjKvrCnfVeCq85D1hrMdrQO1sOPgVKk8paaG1pzq9ozy5YtTXT4Y7f/sW/5nD3jk8++5SUIkKikedp7C1NrTFHqG1FVa8waLJzYIpzbyqEVvz46P/zT3+LNZqLzUrw4QT7d294+9UXfNm2KGOJaKF4lWJxDJ4cAjlG3nUd5MwrK4wGbSXAULai2basLtY05yuqVU3V1kIfm3qGsedwfYMfPW5wkDqiy/yjZz/hanPOq5svWGfFH15+wFh3nPWJQzVwsD3740A3jNwcv2JMWjBcZVlvzmiblsoocgr4aWQYD8RHBi3OOUKIHPuB4+i5O47s+p5uGNntD0zTxKF3TM6z6zqGceJ4OAhvfRgZ+x7vHFPfEUJgGHpCDKWZCGKSJpKIFDJTzoQoz4m1FReXV7z3/iVV3VAZw8FFXnUTQ38HaAyKd9e3/M1vPuX4cUVvKv7L/+Sf8O/94gP+8U9fsG1rZseaH9AWHjIwvish+k5MFwSHu8dQBbGI6QF/t0AQukAQ3juCd9hKuqUaY1BaqrDSNaUKDUZe1xhdaCELw3eJRuf306Wyq0A6qLLQSFJWhOJwQ8pEhTAZopJuniyFnKzy8vuUVDHFRFLxcQ9N8DjvGKeR4Dzee9w4MQ0DQz8y9CPHw5Gqilxc1ChVYW2gMpHKZCpTQW5RaJqmJsTSXUambhvOLy84dB2TdxhrUUaaA1LOUkRUgl2nhPAW41yoMhijOb+4oqpqUs5CkFeZfW0Yx4FpOsgBRsIHTzd0TNNISAHnJrz3HI57vPe8eO+9R63LerVGa0VtLcYoicwKg2SaJoFjvGeaHM5LGpliJriJOPX03YHVcU9zcY4xRrj", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-64", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "c1ChVYW2gMpHKZCpTQW5RaJqmJsTSXUambhvOLy84dB2TdxhrUUaaA1LOUkRUgl2nhPAW41yoMhijOb+4oqpqUs5CkFeZfW0Yx4FpOsgBRsIHTzd0TNNISAHnJrz3HI57vPe8eO+9R63LerVGa0VtLcYoicwKg2SaJoFjvGeaHM5LGpliJriJOPX03YHVcU9zcY4xRrjJgCrFrOXzoxZu5hwESINBRqlcoDBNsz3HGM3Fs/cw1uBchxtHhqEnOil+tu2Ws82Ki/NzmqqSSHQcoD9CKQRbW2GNReKuRzreHAvrwMhnjYkcAyk4pjGRlSZiCiujkPORzZyV0CGJGZ+DZLNJoVQEE9G1YRUTm7phs9nQrlcolTmmiagM2WdykIM6hkyYIlbVNNUKo2vIlmn0uDEQpoR3UZpaXAKX0VljlaGyFVVlJcrTeilWJWCcHD48jjKWktQNvI8c+5GXNzuu93t2x467uz3jNHHoJybnuT0cGMeJ4/EgTUnDyNR3ROfw40CMETeNkpWkyMxezkaTlCaVwnnMWoqxWmBBifwVKif6aURrxT5LJ1+lNDfHgV0/cB16bqPiN6+vqTYr3rs6I6bMWdtIY8s3b/eDzOjb7Fud7mG/K6/G8jBba9HGMAwjoZzOSinqtsVYQ1VVHA87jvsd682auq6pLtfoypCtREDb7ZqqLsWhUtvy0UsXU5o7j+zidGdbnK5KzN28WklEnrShVnohyqkS1UjhQIEueHRWhCh4szHm0bzUr16/JKP45LNPUSETuol3r9/w7tVrjncdw1E4oNZmnj3fcLZVbM8GpiEyHhOVbbH6gma7RWnN25t3DN5Rn29p1muePHnB2k34EGnWa1LO3NzuyBnW6xVPntZsNhfc3e05HjqO3cj+MLC7vaWqK372859y8WTD8/ff5w/iRIgTn3/xGXd313zx2a/pugN313cMU8+XL79gtW55+vwptzfXHA8HPv3sY/q+4w//8I8ftS6blbR", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-65", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "mnj3fcLZVbM8GpiEyHhOVbbH6gma7RWnN25t3DN5Rn29p1muePHnB2k34EGnWa1LO3NzuyBnW6xVPntZsNhfc3e05HjqO3cj+MLC7vaWqK372859y8WTD8/ff5w/iRIgTn3/xGXd313zx2a/pugN313cMU8+XL79gtW55+vwptzfXHA8HPv3sY/q+4w//8I8ftS6blbREa60Khc6gjTjFt9dvub694a9++Utu7/b0hwHQrIzlGAPjGHj5xecc93v+4M+22GbL2I9kSkquhDJlrC18SjmoZ9quMZBVAgQayIBeb2g3W/74n/6nmDCyGm958/IL/u1f/gWvXn/Jm7df8ad/+if8wS9+xk8/+ghrDLe3d3z+1Zd8cv2GWKrqT59ccH62xaeBxOMO6GfPztBK0VQGN3qm5GgqTWNbJi9c3aiEpN9WglnXtWUaB6ZxJDeFh2TEKQ/Bk2MmTYk6W67qDX/+e7/g9376IduzLTFG/u2/+/947Q1fxJfEVGNIJKdwKYPaoOsr1HZkcJ6/+KtPcEPP4faGYy8Rbh4jxMRZteXZ5VN+770PqNuG4zCCNkxZo5RAebfH1/T98Kg1yQliyIzO8/mba/6vX/6Gz1++5vX1Nbu7A9Pk6IszP/YDMQaBFt0EfsRMI6q01s+1mbmpStroLZoKZUA1kk0YWy0F9+2qpqotSknDxes3L3mjFa/dlqQrNqsVx8ORd93IV4cjLw8d/+u/WfP/vLnjGBO/eH7Jf/TzD2kqg7EzbUCV5y6T4ncfzN/hdG/lQ5V2U6UVxlZoY8XpeuHbohS2qTGVpWpqhmNXeJaBurKctRqrKlANWgvR3WihneSshGvZ9bhJqo0KhTEzQ+G+Ha+upICirPBi5+phUpqsNUkb4TvmhA5CjlG2FtxJm5nciNJCJ3/QR/m97fr6DeTEL3/5l+ikUS6zu7sl5oSpKppmhTJXrLewWl/Srhra1YazCwVZ0dYfUtkL6s2amDNv3r3FxcD", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-66", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "bohS2qTGVpWpqhmNXeJaBurKctRqrKlANWgvR3WihneSshGvZ9bhJqo0KhTEzQ+G+Ha+upICirPBi5+phUpqsNUkb4TvmhA5CjlG2FtxJm5nciNJCJ3/QR/m97fr6DeTEL3/5l+ikUS6zu7sl5oSpKppmhTJXrLewWl/Srhra1YazCwVZ0dYfUtkL6s2amDNv3r3FxcD17Q2mP2IPrRDoU2J7dlYaF5DPkBO5pKB1Zdhu16QoEXvcruVA1KrQZIKwR5Th/PKSqqlIcaDvj2xWW548fUoi42NgmEbu9jvubm949folx+74qDWBr8NLQuuZ22sTu67n7nDk7nbPYb8nTE4oWmimvsePA2PfYY0hugltauFoo0kpY43BWoOW72BN6Y83Gq00VWWEB6v1AjmlEhWvz85o9ZYX1SXaKL766gu6/sDhIIVZ7z13+z1aKW7u7jh0PQqD1YCxpROytJA+EtSNcSIhTsZNnmFyWNQSRVpjUPUaayu2mzPqSlpy++5Ir4/SwpwSUUl1ROlKYopa0dYbrKrwg6e/O+KPgpvfvtlxuBkIkyJ6RfKKrCH5zM1Nh+LA4DLJZ3zfM3ZH9te3jKNjnBxV0OgMtjVYLMEnUJFx8PigULoF04A2aFtjqscdRDPcqYB12/Dek0uIgXVlOKxWTNPEruvph5EvpoEpSD1Czftaq0JYu+eWa22wVUXTtrTrNaaSg0w1LdpYzlcboMANtmY/TWxK4fXJdkVVWeKUGUbH3aFjGEd208ToHTl5uu6I3bX8+t0NUWf+5P0nnFNzbusS2N0zML6PfavTfffqC8jSbpgRXERX4nTH4nTD5OQtrUHXFdV6RXSR6CIqe6yBy62h0hvUhVCtmrpCaQtoclKEkHn39prj8Yh3Dq00bb1aII25OWO73dDUFayaxWkDstmspBXRT+TosU7obtX6DKoa6locLlm6A7UiGrV0t31f++STX/Hq5edcv35NpRtqveLZ+RUXmy3teiN", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-67", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "zML6PfavTfffqC8jSbpgRXERX4nTH4nTD5OQtrUHXFdV6RXSR6CIqe6yBy62h0hvUhVCtmrpCaQtoclKEkHn39prj8Yh3Dq00bb1aII25OWO73dDUFayaxWkDstmspBXRT+TosU7obtX6DKoa6locLlm6A7UiGrV0t31f++STX/Hq5edcv35NpRtqveLZ+RUXmy3teiN8Y7thvbGcX1yxXjdszras1y0vXrRsNx9Q1WeYtmaYJu7GnmEY+fjTT/Ep4XPEhUDMmRfPX7BerXly9RRjpD3WTRNDd2S93nJ5fkbwgRQTl+drKawamJxj3x2wlcHWhqfPXlDXlufPznHjwO72jtVqjbYWFwOHvuPlm9e8evUlv/r1rzjs949ak4c2QwpZR4ZpZHATr25uefXmhtevrhmOe6wWeGGYAt4FvI901kLOuLEHbfFTmZmg5bkyphH2i1JUpsZYvRRM26ZwuI2VwmXK9GMkZzi7uORi0/JHHz5lvd3y5uUrpmmgO+5wk+fm+obrm1tijNztD/TDiFYV1jZUtcZYTVIe9ViSLjCOO1KKOOfwLjJNnpVuaHTDenVFVa1Yn13RtGtePH9BW9WctTWH2xsONzd03RHnHGOJsdeqkkEVtmXdGqy27N7u8fuesesYx5FPXr3hOIyMR0WMunR8yuicTz5+w7vriXq9RoVIuN7T7W65/uorUkjkkNialsZUtPUKmwz9bkRVgdvjwDgpVHUOtiIbTbM+J+vqkasiAKFWmueXF/wHf7Ki+/AFY9dxLJju6+sbUsuKZwAAIABJREFU3t7eEY87djmQhsg8L8E2FaZkBzP909iadnPO+dNLrt57RlVL0TWbispW/PzJc1JKXO/3fHG3469fv6VeKVZJcXnxnPPtht/sjnRDx998IvswaU3wE1YFxuMdUSf+70/XfDH0/PlPXvAhW85XtdDX7hmXX+uw+132rU7XxiCk4DzjTBkVFYpEozNNJVXmrACrwVp0Ywh", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-68", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "U3t7eEY87djmQhsg8L8E2FaZkBzP909iadnPO+dNLrt57RlVL0TWbispW/PzJc1JKXO/3fHG3469fv6VeKVZJcXnxnPPtht/sjnRDx998IvswaU3wE1YFxuMdUSf+70/XfDH0/PlPXvAhW85XtdDX7hmXX+uw+132rU7XxiCk4DzjTBkVFYpEozNNJVXmrACrwVp0YwhKE7UlR9BaOkuM0VRVTV3XtE1LwhBTKZplFrxSukCk+DLb4XhkGHr8dMmqbWniFqu0tBUrhJdaV9BUJDeSoycPDq0MlaklUmwaVBLy8lIUKb/7GNMpkYKnO+6p7Apfa95btVw8f8aLFx+VirvGVorNhaWqLO2qJkVLjpZMg0MTnScr+Ed/9MeEuXKbEzEnjv3A6AQXzyly8/aNNAFMpWuo63j69BkXF5fYgk+1JdrLBpSyYM6Wbh+jLAbLql7TmJp1u6WqG1bbM1CGQ9fT9QN9PxFiWqabPcZinFvFSyt2pjS2CB6XEjz74Of/P2nvtSRZdp1pflse4SoiZWkQBNFsmx5az91czwPMA8y8wjzpGM36ihyySYAoFEqkCuHqqK3mYm33LLDJBIK9YW4Gy8iKdD9+zt5rrV8xDieGwxuYJ2w+0XcrvGtwTYdvOjQZVRJWw5Iy07RgNLTFghbHNecd3ln8BTNovDBTjBXHt5QpaRYWyATnNPEdgbvHI8p6NrsbPvv8c9rVBkwj83GV0K6jVY6XpqGyEdFWWkYRdDztusQ41ZFWRikx0kkxMOWMVi0la3bPPf2q5/b2Fmc0jVKcUVfud4wLvutQzrNe3ZK1I+gWqzPFZo5LZFgWxqMATY+PE8M4MxynynvNgodoxeHukWWYWa16HNAvgT5nsnUUncFkWqVFYINiniP/8JvvydZxWDIfHs4MSyGUhDGJYVqY56eZAJ2HkRCEzmWd47PbDXndktMtyzQTY+ThcOKnu3t+fDjw5mHP2b5HFxGnOHMxSqqzpaLwvmG", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-69", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "ud4wLvutQzrNe3ZK1I+gWqzPFZo5LZFgWxqMATY+PE8M4MxynynvNgodoxeHukWWYWa16HNAvgT5nsnUUncFkWqVFYINiniP/8JvvydZxWDIfHs4MSyGUhDGJYVqY56eZAJ2HkRCEzmWd47PbDXndktMtyzQTY+ThcOKnu3t+fDjw5mHP2b5HFxGnOHMxSqqzpaLwvmG3veHFi1tev35Z5dHCgrFa8+KmOhRqzT6DO45YHzEu03Qt3apns82clkiKgVQ33ZyF8hjHgaLhcb+n9ZbfH04Ya/iLmw211AV+Jvj4E1vKn7fpYqR9B8gRBXij0VbjtIgEilEUY8jeEjVEC3G5tIPSSl023aZpiElan4vLldL6aoZxMbSQDwKH/Z6Hh3vImc1qxU5ruSjzct10dd9ikieHiRID5TRhtKF0vYB8KaGLQmdEO6iQg+RJtwyokikxMJyPOF9IqsV1LTcvX/KLz79ms9rgfEsuhdM8yu/XIsstGYbzwBIW5mXEOcsvf/2f0KjaTgtg9HA4cDqf+cPvv+V0PHL/4aP3wjLPjOMIOaCJrFc93jsa26E0BFUw1uK7TkC+EDD1f63rUQ10fYexFu09Reu66crDGlN+8jWR76lcudpFayiKZZ44n/ZCzyrw/PNfMo0Db+NEUUdCnNhtt9xun5FSEdI9Mo93WhFiZp5nvDOkLPxtpTXeWZrG0bWNiEx8pQlqLRzfVEhLZCmBOAVOk2I4nZiOR5RrWG9vUCxY70EbchSKmHEF52G3M6L1rx4EIQaU/lM6o/9x5bTUTVeAPqMhpVC7uQ6weO/pu56b2xtMKbAsAkYvCyHIv712W3y3YvX8FVF7TrSUNFPyxGk4EOeJ4bAwjzP7R1F6no8TF0THWhHHHO73DOZE2nT01rLzDpsL3nkoFztW6R9nFMsc+Yff/kAynnOxzMvEtBRcjhhdGOdKQ3zCGsePm+52t+H1boX3Fmt0pRrCeZr4/t0df/f7t5T+Az/", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-70", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "hpVC7uQ6weO/pu56b2xtMKbAsAkYvCyHIv712W3y3YvX8FVF7TrSUNFPyxGk4EOeJ4bAwjzP7R1F6no8TF0THWhHHHO73DOZE2nT01rLzDpsL3nkoFztW6R9nFMsc+Yff/kAynnOxzMvEtBRcjhhdGOdKQ3zCGsePm+52t+H1boX3Fmt0pRrCeZr4/t0df/f7t5T+Az/FgioajcFZGfkoU0daStO1Lc9vbvnsxXO+ef1SxB1ZTI2MgmebjfyZ0rwPGd+fsWbAuOW66a43mW6YxXYgJYoxIuoqmTQP5BJ53B+w3vHd8cS6lcNa+AAXVd0f6wj+vfXJTXe5ewAQZYgx4By2dWhvyUWqTOs8zjlW6w3KWopvmENhDlSpY2GzWrNZr9ltbmkbT991DOPCEmdpiZW68mqHYUQBbUvlkDasVh0hrHBWBh2+FFxKLPMsY1prcMXTWAM4OQT0IlSfeRFamVLCI0lFWhOjMRuFck9rj9IEzarn669+xe72FS8//wW/+PIrXr18xRACp7sPjHOUtn0cWEIUazxkJhiXhRQjjw/v0Qp++uEHrDZQCm2dSdnGY6zhL3/5S7n+pRBj5HQ6CQNgmqoCT7PMchP/9OYtIUbOS2K12fHq8y/pu4ab7fqq2Z/HQIqRw8NJZNTW0vc967UiR4UuQsky5j/ubS+zrerXJGJ0VBowaWTbblg5i/3mL8gxkuLMN19+w9dffs1hf888jUxREUth2zVQMh+WkXksDKbgdEHnQGwtVrUYbzEZXFTkGJnnhbDMUiUOIzllXOMqX1uBmohdJgfI2eBcg7WefnWD0eYqBrJGX53tLoZC1mr+LU76p1bjW7z13KyekWIkTAtUtVtMlozh7s2PTKcj205sFZ3VnMPAqCLBCBunWfe0qw3r7Q7lOm79miWMTNOR2DekuLDd3DJPMzlpzscjKSwsYWKeR7TKZJ3JWZEinOKZYISy6EuhVYqipVNS2qKMpdndQLciNT1LMYx", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-71", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "dJgfI2eBcg7WefnWD0eYqBrJGX53tLoZC1mr+LU76p1bjW7z13KyekWIkTAtUtVtMlozh7s2PTKcj205sFZ3VnMPAqCLBCBunWfe0qw3r7Q7lOm79miWMTNOR2DekuLDd3DJPMzlpzscjKSwsYWKeR7TKZJ3JWZEinOKZYISy6EuhVYqipVNS2qKMpdndQLciNT1LMYxBkbX4WexuNvStgxyZp+FJ18Q5h9aKzWZF2zZX6X65tOcUsZtsWm6fv+A+KVanWTw3MsIP1gouoKZSGN/gfCMy/3SpVhM6C1ddlYjKWWbVaJy2tNbQWc2SZk5hZFaRaBJKFZzRuMZXLonhqApLWchx5HE+8d/ev2WfA+ve8bxpedW2rJyh1foj/eoT69OUsWGoQJOMDlRKGAPKKkoO0rxWp6vWGZS38tIFlLS1GmgaT+Nb+kq0bpuGJRTkPJUllDAlDkOAz428QWvFh7YRsrxWMu7QIJ60SoAxQ8FqhTIahSEqBVmq0qIgTbXaDBnlLFhL03dPV6QVGVlsNzc8f/6KL774mttnz+lWa+4/3DEOI/vTyBwDx3FgmheOw4BRMscqIZJi5P2bH1AlQ8pYbdDAerthd3PD9vaGftWz3mzwztE1zZV6FUJgmqbra/8odpCH45FxnDgMMzHDy1ef4a1ls+6F4FMyabGUlFmmUIGZhFWO6BMlg1YGY0ylSD1t/XyWdZlzlVJlv3lBlYnWdhRnQd3UBwy++OoXfPOLX3L/vuN0OvDT23tSKLTeMhkNOdZZtibMlmgghpl8EQvUQ7vMgXAeCPUa5XmBUtBWNPMmK5xOeK9oGsMSLM46rPHstlucqziBVjIislJVSWVlqkjjaZuudw2d77jd3krlrUd0AZ0Vw5QJMXMYzqhcGA4HUuMpTSMObQqUURhl8E1D2zZ0vsU2HX61ZpjAsBCMIiVLNA3eL6zXD5Azh8aRy8I0JRQJsWyULyeEgI6KRYE1YnpfTB3TOSdjwu0", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-72", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "eCPUa5XmBUtBWNPMmK5xOeK9oGsMSLM46rPHstlucqziBVjIislJVSWVlqkjjaZuudw2d77jd3krlrUd0AZ0Vw5QJMXMYzqhcGA4HUuMpTSMObQqUURhl8E1D2zZ0vsU2HX61ZpjAsBCMIiVLNA3eL6zXD5Azh8aRy8I0JRQJsWyULyeEgI6KRYE1YnpfTB3TOSdjwu0Wmp7oW8iKFDOqbrpd17FadazXa6lQn3Sj1G7Fe6y114PsQmQC5PobS9euaLsJ162uFrBaVUluCZSSQSm0tddgg5KS0PJyRGWNEJpEVEUV8mo0VmmcVqSSWNJCIBIr/1+k5vb6hk5V7UcOTHHm+9MR7Qz/+Ljhq9W6HhwenEWln5vi/Nvrk09XW1MiElk8bsk4o/DWMMVELlkkplpULVYpvHHVbi1DiRQUjW9Yr1a8eP4KZy2UzDSn6+wFhBOcY6RtW7k4upBLZJpHtFa0XVdTEjQpBzQK560guwpSicxxFjpJFP5eSRlrNHkxhDCR5kicg/iyOseubfBPbKb/6998SSqa+4ef0N6zefEFRRlOU+D927dM4yhAg1KoppEkByWJDlZbwjRR5oklBHIInE8nckqcjydevHyBUophGtHWcNzvictCDIGu6/j666+lNRtH2rbFOccSRLb5uD9xPJ24e9hTgOnLL3i2W/Nsu8JoKSP2pTAZxzJnoWM5j9WWOEVM0TS24dXuC9JqedI1ATk0iyofHeZKYYmBKcxQ55maEYqmUULyt8Zx2L/jN78dKmg0cxwmUI7nLzagNNvGUEokjSdmVdBL4KAMSzuSxowzlsY7QfljQukGmh7tAQpZiWZflYSyhn7Xot0a390IWJcLxrU0bcer5zsa71n1bcUhzJV7rq4ez3/++uzVl3Su4/X2M8bDkYd9oNUKrzTP1pIYMXQZYyyracbHTBcKHs/r9QvIAaUKN3aHyz32nGnI7G41d2Pg7XDkw3HPMI1E5cgFbjY9nQNvJk6", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-73", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "mVdBL4KAMSzuSxowzlsY7QfljQukGmh7tAQpZiWZflYSyhn7Xot0a390IWJcLxrU0bcer5zsa71n1bcUhzJV7rq4ez3/++uzVl3Su4/X2M8bDkYd9oNUKrzTP1pIYMXQZYyyracbHTBcKHs/r9QvIAaUKN3aHyz32nGnI7G41d2Pg7XDkw3HPMI1E5cgFbjY9nQNvJk6nI497VSt6R9P2aGNIZJzW3PieZ5sNf/H6c3TbYNqWaBzFGGzXU7RhLI7zuPDmwwOUFujovMXozO1uR16tnnRNfv/DG7yz3G43KAxa2ethIF2SEvXZUt0FY6r+DtJ1lKoUVEXAtQJkFYnLRFgMyyLc95xTLcoU52Ek5sxpnpiCHMZOQafB5ogKCw93D9zdHzhPE23T0rUNjbF4lTjsZ4Ywo1MgLzM/3X1gP515M514uV7z+W7LL7qeF97zsmhaFP/nq2f/7jX4tCKNWrXwUVQgLK0LebuCYLmwxExWkaJDdcrPokpT+ipucM7jjJFTSP9xu3axfzTmoxKtFPkd0t4ZjJXWNyuZn6haZaMymUIsMsNRWn1EN7Vk+hRVaixHuZ58JYvs9inrm29eMM6ZH9+LvWJRQkWZ5oVhnBjHEd8LGRtbZa7ayElcCdqoi1pL5L6lRgoppXHeE2IiLgvv3r1jOJ8Zz2c2mw2bzYZhGHh4eGC73dL3PTGKscc0z8zLIh4K6eLWEskp1DQFOYjCEi4lKDHI/1dFophUKXjbkNV/rNJV5ZL6IWVuSlKFpyzeuqXEKsmWzczqTAoT51MWvXwM4pZmRc7tnaPznhBFcZRTTaOI4u0Ql4h41qtri6q0dBRoue/EAEmKB6XEMco3AUVGMYmq0PgqI3eVa26kw7Lu6s/8UVnz56/t5havpZrWGHSGxjp652iMR6PpEaCrSxlHpC2KNhYkvkoa3D5kTA6ofMbnhF9Z/PGAPx0xpxNqmrD9Gqyju9mRUkvbFtbnjqYzWOux1tGtN2j", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-74", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "mRc7tnaPznhBFcZRTTaOI4u0Ql4h41qtri6q0dBRoue/EAEmKB6XEMco3AUVGMYmq0PgqI3eVa26kw7Lu6s/8UVnz56/t5havpZrWGHSGxjp652iMR6PpEaCrSxlHpC2KNhYkvkoa3D5kTA6ofMbnhF9Z/PGAPx0xpxNqmrD9Gqyju9mRUkvbFtbnjqYzWOux1tGtN2jniBSs0jxzLc82O559/jm67TBtS1CGrDTGWQoKFwpKn/H7ajKTRRqhCjj7MVLpz13705muadisVh9tEMulO5KkkXESel2svtvyxVbLgXzBDsrHWWoWw5s5BMbq8yHWjgWjFCcrqr+YpZK1SmGqL3GrNd6YCtTlqloUappRCqdB5wgxoGKAGJiWmahhsYohJ84U5mXhznnO2tL/iRvl05uuFaqYURUooxCLAELBeKIWnqOJBTUGOC+kuK86dfnA3jpx3ponUYpVj8xLBaGVtHRiAymo5CXnLMbEsgSM0TSNo+86vDXyPhC9fCqFjCKqQkkLm9Ua7zyrtkflQraGrBWL0ZSQYIkoxDMThWw8T1j/9//1f3D/OPPf/u4DNM+x6+ekoFhi4TyOnM8DG+MwzqGVpWRRi5Vcr19KVwpczpmwLLRty1dffcXX33zDr379V/z2d//Cm7dv+d2333L3/j1v37zh2e0tq9WKb7/9lr/927/lxYsX3N7e8vXX39B1HQ8Pj6SY2K3XbNc9nTecjw98Nwr/dhoHhmFGKcPts1fMS+D+cc96vWG32zENJ8Iyo7JH81QaEBhdQclKDs+lMI0Dj48PnM8CmnkPRluschhVcEZQ+nEqnM8zS0gks8JZTTYNTWf44rOvuX944P2HDxRaMD3GrzGuoShLRhPzBeQSR7tU8vVQt4jsu1w2MV3oOk3f9ey2tcvRDQrF3X7C6sRhfycYxHaHc75m0D0dXvyrv/oblnHm/scPzHOCDC82N3xx84zVnHEpo51UXiVnmBZYRlRMYrSTqt9", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-75", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "PnM8CmnkPRluschhVcEZQ+nEqnM8zS0gks8JZTTYNTWf44rOvuX944P2HDxRaMD3GrzGuoShLRhPzBeQSR7tU8vVQt4jsu1w2MV3oOk3f9ey2tcvRDQrF3X7C6sRhfycYxHaHc75m0D0dXvyrv/oblnHm/scPzHOCDC82N3xx84zVnHEpo51UXiVnmBZYRlRMYrSTqt9IPogBTNEErXjvFHNcaOJEnxNZKbr/9J/pnj/jL/7rf0E7Q4gj0zRyOh+x1Zdk+/w5rm1ZqreGr4qzvulAG4o2LHP1/phFmLCeZxQLaXpkmc7M0wm729J0HSWHq3HNn7u+++Ed61VP1wqHdrOrTX+B0zgzzQvffv+Wu8OR43hmClM9MkutbKWLTuWjefmcFh7GyJwmDtORSzhMCFKwbVYdjbXsfIfXmuedpzdgVOCr3S23uxv+eZfgHPm7vKCyJqeFHBdyWVDHA2YcYbcVu/jziiUGxrSwHwb+cDjw99bgrOavup61sfw/n7gGn3YZ0yL5K0pc7BP5mlUUlSICUYomQiqEZeF0OCAIVqHvWlRb/wIIuf9ndKSPOWdiquy9pyRPqYDGzz18lVLia2nM5XCsJ2UVJjoB+0QkIbMpVbiacitbq+KiMEo23UB5cqXrjKXvNC9fvmZRGxbtiXEhRDEcn6eB9Xoj7vy15dbIzFYVhbpIBetKWTK9tJYqzFrLRep6yYSKMV5f8zxzPp/p+56maYCCs47tZktOCe8dzhiG85FQb4Tzac8wDBwPZ5SWGWFBjIe0KZQSSGkmxAmT85V4/pR1kaFfcqK0ku4lJXmIc8yS2WYMKZbKe5WZrFaSKVeKAFYSOlnQRtP1PeslMM8T1orufx4HSoo4lcGa+n1WF7ws95MqVcxTXea8FTzAaIVW8mcXbmUuRmTTITOHhfF0gAKNb0TlZMyf5F7+m9dEYg5Fpo6AVcaJsY+dB3RMlONJklRqDlwJ6WomkpN0QZccwSXX59Bo6R5", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-76", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "1kaFfcqK0ku4lJXmIc8yS2WYMKZbKe5WZrFaSKVeKAFYSOlnQRtP1PeslMM8T1orufx4HSoo4lcGa+n1WF7ws95MqVcxTXea8FTzAaIVW8mcXbmUuRmTTITOHhfF0gAKNb0TlZMyf5F7+m9dEYg5Fpo6AVcaJsY+dB3RMlONJklRqDlwJ6WomkpN0QZccwSXX59Bo6R5yQFOwWqPOZ0rXQ0oCMLaddFWlyDXQhsY5rNaCdZSCqv/uFKIoN4FlDqSYWMaBnOSeW057yjIQpjPzcGT2FqsghuXJjnRzmGmiBUplGMRrAbbMkWWOxJpzt/KOTevZtf460zWVCaLq1pVz5fAW8dOYQ7wCuEsU32offe24L9WrVLpaFTpn6X1Da2QMqHOUEWVOcLnGQcyAVAyoizVCKTI/joGkIWJQxTBZi86fPog+zV6wddMtEFUhlozJCZ0yY5ENNxWwCpYM+8OZ7373HSgJZ/z81SvU7Q2l5DrQTlAuyQTyZV0eln7Vo43CW4EzdTW+UYrqF5s++js0VkYcSTxkNaC9BWdIxrIUhXNWFG9GquispX3XSsvJrw3TvBCfaO4yHGaU7vnLv/xLHkfFT3s4nUfO04nD/p7heOLFs5dYpQVpRZPMBZ7N4uNbRMacED8CZ8MfpRbnIjxla4VSdPHDvZjHXDbglBJ9v+LZ82c8u3lGyeJZnPLCT9//nr5r6LqGx/0D4zDw5s1bxFMgs9lsefnZFxUciizxzDge8dQN6cnrI0P8IgVWQEmFNMl73XTCQDmcTrVqET8OowUwStmgbIOyDp0DWjs2u2e0TcPNtuX+/p7T+cDjh4OMbF4+p2k8ftVDnbteZJkX9kFjfEXDm2vgoDA/XFWSyqgipcR5TpxOZ97+9D2fvZ4qt1xc2H7OG/9z1zwn5iUTgKQNxTbo1Qq33ZAfjqTjieF335KnGWrNmPO1jCCV2hFVB7FUQe2inCRIqIxF0WhN+PEn8jBx/OWXNLstq+c", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-77", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "yDp0DWjs2u2e0TcPNtuX+/p7T+cDjh4OMbF4+p2k8ftVDnbteZJkX9kFjfEXDm2vgoDA/XFWSyqgipcR5TpxOZ97+9D2fvZ4qt1xc2H7OG/9z1zwn5iUTgKQNxTbo1Qq33ZAfjqTjieF335KnGWrNmPO1jCCV2hFVB7FUQe2inCRIqIxF0WhN+PEn8jBx/OWXNLstq+c3GAxOyRhDFQjHI4HCMg/kmIiLuHuFuXoPp0hYJlIKLMMkowRrOI8j6XTPfDpy2D+icyTOK0JY/odU3z95TZYzIRqUhpgix/OIqvPacZhYlgAl463mi9sNxhqGkBiXwDQH8fZQ+pqAvSwzMQbGebjaQEq8V2JOEVB0oSMpRdXY4QGnCk4nVo1n23b0RtGqjF1GrAYVZ+F6hxE9DrhxQM0jKjYS7WGUXJ8UUTHTaofXGRWWa8rNv7c+uemqvkMVGW7bCnqZpkahVBMWox3OWrq2YWqkyipQwROhNYlOXtWzNIsBDZVMXCtZrT8qjOR0rpHs1egmiTe5bLZaLniuMKGcYAZjHDlDpLDkVGXGmqQKiy6omNFBkHqjJTY9P7GCOTxEiknENjMHsV1MKRHTjFIJbWpln/M1mVYjBj8/l/JeKF8pfhw3XHSE/9o441/Pvi8MhmEYGMdReLtRKuMUJOZmns+ossIaQai7ruP5i2cUwDotQGUOxCRZdTHM5Bxp+1ZMeZ64rNG1Isu1chEEuvENojwEb+QA2fRcrfZKhoJkoRkNqAgspGWkmEKxDmUbus0znhvHZjtxPh5IMTINR9JiMXnBO0vbCGBkrK1zWcNm3eOdY73uoVAjYQTdjpX8ntNSQVdonKXt1ihl66Gc/mwjk3+9ZEbfcPv8JaFdM7kOfMsxBPQ4UMYzyzSTQ0BZjXzyatxUqPcM9bkRVz1UtUotGa0ToYiUnuOBkjOP//TP2PWK4WZLjolUWRwAilhZLJOMtuZYv/tUnfwiMct9EOYgc3JtZfMaR9I", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-78", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "BO0vbCGBkrK1zWcNm3eOdY73uoVAjYQTdjpX8ntNSQVdonKXt1ihl66Gc/mwjk3+9ZEbfcPv8JaFdM7kOfMsxBPQ4UMYzyzSTQ0BZjXzyatxUqPcM9bkRVz1UtUotGa0ToYiUnuOBkjOP//TP2PWK4WZLjolUWRwAilhZLJOMtuZYv/tUnfwiMct9EOYgc3JtZfMaR9IwME8zx8NBItqvoZBPuE+sxllD4yp7ASlCKOC0RlvLtmtxWjGPIztnedW2LNYRfGbdNiKMaRrhEy8z4zTzsH9kCpEpBPHiLgqTZISXTo+MxvBhPrLEhJoX+tXEjc/EeeB0PhCDiKq0lj3Va7GwbHTDbrPBe89qd4vb7ui6nmwMs8q1yoWbYumTZUPB/gkTrU9vuvVG1VmAMGMMynupROYZkwuucTTOs+palraldY5cNBlBZW1tz7RRfFQpA3WmKQAHdbb7kSPqrPkomFBys10rQX1xm6rpnUgqr9WOuVaAU7wQlwV4W8gQIypEok2iXkKqiqesx7tEMZGySowoUrbEFIhpQpmMtXITlZxIUZyytKqGHJXGllMQIr/R16r14i9xAaEuwOLPndxUBeDE0m7EGMP5dOLc9xAELEvLmRBmplk8dttGVHG+9fTrTmwEU0YZiGmRGJN5ZgkjuQSpjn3zpGsCsrHnnEUYgRyi3nnaphOgK4O3DX3T0thGcreGE7ES0LWu34UKlKKI85lsIJeetuvoupbt9gZdAu9/+o7hdOTD3Qe0KhBa+r7D6Q2tU3TO0rRChbrZCR/0drclxsw4BjGUmTOZRM6BHAZyLnjjyI1ntb5BG804zdfIl5L/eCz05yyhmnk22+eEcWZc3UCa2M8LDGc4HynzLEkWtmGhCChTEksuOG2wWmEp6NoyU4StU5QApRKxplCHRzif+ZAXVOPx2zUqF7kvkhz0pFkSMFKoeEK6etpmlcgqUVgoJdXYHEXEkqw892mZmKeJeZk/Og4+sSvyTkx", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-79", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "UmTOZRM6BHAZyLnjjyI1ntb5BG804zdfIl5L/eCz05yyhmnk22+eEcWZc3UCa2M8LDGc4HynzLEkWtmGhCChTEksuOG2wWmEp6NoyU4StU5QApRKxplCHRzif+ZAXVOPx2zUqF7kvkhz0pFkSMFKoeEK6etpmlcgqUVgoJdXYHEXEkqw892mZmKeJeZk/Og4+sSvyTkx92qbFWSvfX8moXGhMDSVQPa3RjPsj2XmmrpMDAMWz7Yq+8azWa2H6zBPH88D3SnMcBvbHk0Q2lYSJRXycl3sGMpMWnwZnDKtV5LnPLOORQ9GEZSCnuVIGFV5DZxytMjzb7Ygx8vnzF/TbLc/7DUHDIQemkhlz4fNsucGhU+KaJfTvPSuf+uGihMy8Xm3FZLrrxHvBWsZpJuWE814qKdegc+L8+mVlCcB63dG0Td0oEqkUdD3V0sXSsKZ8dm2LNVrmdErMuZWWALoU0x+V7MrUOLokv0vlUtkNFqMtuRRsripBrclKeLwlCrtCV6MblS8b3Z+/vvrmOSG37FPDPCbG88CyLJQUefH8GdxmulWHVpZQxHwlZsm1QsES53rTSjU4DEe0UuLYVp3RpFP4yOi4vEWtP7a7zosEdllmzscjD+/ekcIMaZF0jkZD3kiHoTQoLYIWQGlDTJmHx71cxwIog3UNU0rk+DRwEcBqTUZilqS1N6zXW26ePcf++APjFBiGQMny3UmwZIPT8tCJx20h5EIhodSCwpLTQpgKJUW8VVit2Nw8Y7Ve0ThRvY3nI6fDzHDas9tu2azXvPzsC4zz+KYRTqgRruk8TyzLxLLMcr/ZhlUnCsJhmIg5y3X2nq5rsc5fD7unttLWOkARY8K1HZvNM+LhnnTaM9dZbzCGSOZYAgOFRzLJKKJVmEVm7LdK4wv0UW6MohLZZLLJdRABJmdKWogPC8Vowt7JKCuVOtLi6s6VSiYrRbSOpA3JWbIWkyAxmgLvOpS2ZNuAc9D3uBRZx0XuPWv", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-80", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "RTqgRruk8TyzLxLLMcr/ZhlUnCsJhmIg5y3X2nq5rsc5fD7unttLWOkARY8K1HZvNM+LhnnTaM9dZbzCGSOZYAgOFRzLJKKJVmEVm7LdK4wv0UW6MohLZZLLJdRABJmdKWogPC8Vowt7JKCuVOtLi6s6VSiYrRbSOpA3JWbIWkyAxmgLvOpS2ZNuAc9D3uBRZx0XuPWv/iGf7567desW271g5Rc4Ly3kS/+acSUGq+WWR6KAy3qPGATsccL7FNR231tC7jAkTqEKrEk6PnMwBp494vcd2ooQ9moEUA64csDrTePDG0FjLN1vN65WG6R3jdKKEkdYqfvXlKzbrNV998QVpmYnLzLTMLDkRc5Asu/0jRYHJiV4ZemNwOpPUXJOKP72pfDqYktqabja0XU+/2Yp81GhWc910nROQqyjSvOJmt60ULXBW45wVbl01FM71tJZkiI/tvXVW7Bpj/dKbVnwvq1l4KZmLaDiRpBqsNxQpXyOFtBZ2Q/XmomhFRrx0s9JCc1OXFFb+JJH5X6/bZz1T9Iwng5oF2EohUFJmvV5h67wwF/E0zUXqe0lwLaQcCSlcK4V5mfHOX2e6PwcOL/PHnydhuEpOb3yDdeI1MM8TD/cfiPOELoG29dibdaW+VMrcpUmtrX9MmXkZ6/hHMq+01nJQPNEjFS6bbibV92utpe1aVqsN2kj3My8JpS7AiVhGVxU5MYaa2prJZBQLpYjqKcfCkiNkS7GGVdsKw2I5MerCeX9fJc9zpTNlbl++hkt+ntE1BTqTrvPwUMURBu+lMBiGST6MElFO23ZVKKJqksQTN11jSEWM2JumZ727lZSGeWbShoQiWEl/uCdwVoV7CsVYijGSmhxFxt0XhY9VNK81qWQi1UMYrlHk6TySVQ3PLUXEGCihUiZ59qKCbAyx70kGonOy6WpFMWKfabo12rrKeW4w6zWOTF8ybdPgvKvqsqeBrn3b0HtLYwpLWkjzsVbhWZgTKRNm8R1", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-81", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "MElFO23ZVKKJqksQTN11jSEWM2JumZ727lZSGeWbShoQiWEl/uCdwVoV7CsVYijGSmhxFxt0XhY9VNK81qWQi1UMYrlHk6TySVQ3PLUXEGCihUiZ59qKCbAyx70kGonOy6WpFMWKfabo12rrKeW4w6zWOTF8ybdPgvKvqsqeBrn3b0HtLYwpLWkjzsVbhWZgTKRNm8R1R4YgOAy7saV2mNYaVcfS6kJPMwb0tFDWxNgPGnLHmQO803sCxjOS0sNEDzmRWDXhr6Kzj2cqzbR3384G4zOhsaa3my5e37DYbvnl9y/F45nhSGKfJk2Q2xhgIw1lwg5wrHc9jdCIr8Qb5U8fQnzAxz5iV5/UXX9BvNmxub4lBpKSqcnZzLqQYGQ8nmZ1tN1w0D4WarBAlkfN4HrFGU0pimicKHxH6eZqIy8J4OmKtpel6iTe3HqXl7xkl6D9RHoAYK8/0GmZXmQoKlLFilF5RS1BEDdkI4ayU9Ccvzr+1xvT3LHlFUn9RyTKCzi9TxhYotoBEagosAAAgAElEQVQJwit0DU4ZnLFQIuSI9gbTGHRsIIu8FwXr9RrnPYXCZrPh9evXzOPIZ5+d+PVf/orVasUvvvmGvutom7byURXbzRoDPLYtC9KmeXvJLa0vVXnLUT6x1hanFdr2dZSRK6oPSyzwRN9YAGcVJWvxo1OiBDL1mFzmzDAG7h4OkuqshasdwkzXeLrW03cNrffAwhIj9++/Q7sVq13GtWtct2E8vOE8HHn34S1hmVmwhJQZplkq0ZI4DHegHvjxbmS72fK//c1/5tnNlv6LVyhgvW6wNmNM5v279xxPJ968u2MYRu4/3NN0La8+f8nzZ1u+/PobNNU0Rj89OZoCKUTOxxPT+czh8cDGGlarNXzxJaHvedt45nng/vwIbcP29hbjG4z3DO/eEo5HHn56zxADvetxKBxFmDGVQJWU3NcK0EUcS+WLvpQpFSuodXHUmuIc7LbovsPf7EjGkI1", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-82", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "/5tnNlv6LVyhgvW6wNmNM5v279xxPJ968u2MYRu4/3NN0La8+f8nzZ1u+/PobNNU0Rj89OZoCKUTOxxPT+czh8cDGGlarNXzxJaHvedt45nng/vwIbcP29hbjG4z3DO/eEo5HHn56zxADvetxKBxFmDGVQJWU3NcK0EUcS+WLvpQpFSuodXHUmuIc7LbovsPf7EjGkI1mOg8MIZC7DdY1NOsdpm1odztuWkfXeZyz/2FGx1qfWLHQzL8hjyfY3+OqZed8knidWA+XF66wWwVubZ3vm4Uu3+OWgmFGq4xLGqtmVP8A/gTbA2tvaI1mGgdKDvRmwaiC1ZCTJQZH1xi8a0kKnCv8Sm8JNPzv/8uv6ZuGm82Kv/vNb/m7/R1LWBjDwjCNFIXcw1VBS46VnF4FNKX8SSD605SxJBE5TdvKXK3vmaeJoARYUyDo589oUMZaqnV45agimUbzwvF0klmtKixhAT6KFWIIkl6wBPkAl3lRNbVQuvrUqKuITarfK/hUy/r6M3UhUyu58a6ewOrixVlQKvHU+cI0fyCkhRhfk5MVNkYCUjVbofrq6wymyOZjNSWLm702UlEqZ8lRYme00VgnrVpKEeccfb9id3ND0zRoCl3bVVpYZvpyuY4hrIYUA9ZaYjUN+TiayBWkqterbsUKXUUKqh56H2fGKaf/0MOkjVS6OSVSDoQpczydOZ7ONWpl5nA80XiH840IQ8Jy/a5d/fwxRpZ54fD4gGsW2nYrc0DdM+eJEgbGwx3TOBLcmlQUc6wpIhSWkGWc487EpDkNM42fOZ6GSilaGMeBaRq5f3zk8XHPj2/eMgwT5+OZXc5oYzDOYZ0nB9nQ5TB/+jEtz0EkpkKMhbbrSKaBrkWFFWnVEXRhCQO27ek3N9jGY5uGMI7kXBjVe8njq0WlKfyRJenleIWPqSoiTb3w3+GCA2Zgpv73NbzUdn31IzES8RMytmgcNVtQS+VtnKfrOpwTZWjOP/+X/7zV6oB", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-83", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "GSilaGMeBaRq5f3zk8XHPj2/eMgwT5+OZXc5oYzDOYZ0nB9nQ5TB/+jEtz0EkpkKMhbbrSKaBrkWFFWnVEXRhCQO27ek3N9jGY5uGMI7kXBjVe8njq0WlKfyRJenleIWPqSoiTb3w3+GCA2Zgpv73NbzUdn31IzES8RMytmgcNVtQS+VtnKfrOpwTZWjOP/+X/7zV6oBXCZMyJp4w4R6rM1ZliAM5JGJ0Vf7sMFqA2agKSWWsSmgyRs2okomhQJpp1ID3A60eWDlLYxRBj5QcaU2U4ybDEjJTAnvREThNsYZntgPbc3v7Uqphb+m9FQFVVcFdchtTSqBFrnwpWkBfHr3/uUo3zCNhHonLRAoNJS5XaV5Evsl5nFjmmeP+UFM4F2KcSVGkpFpr3r59x36/5+7uA33b8uzZDte0+KYXxDQuHPZ75mlEpSgtn/NgZA55dcOod5hE1uRrsq8o5WoV9/MPXmlaVLMMBThra8xLZhpHkTg/Yf32n/6FrF6w6K+YwhYXdrgU8LlQ5oGgEpiAMgnrE1iHrnQ1ZSzeWBpr6XYdqWs57l+w3W7wvSfkmfvHDzjX8eLZC3IpLHFh1Tka57hpd2y7G26a11gt3MuH/TsOp0d025DDwBxGQnSk2bKNCwsTplhUqUBX0ZU/mihqEuQbUKUF5eq85enesU3bcDqPfP/uPd//9JZ//Off8ebdO+7uHnj39kemaeDND9/RdR1ffv01vpWZ6WmYOZ9H7vd7FHA8nDifBv77P/6W1XrFX/16z+df/YJ171j1ntbfshzfY1zhHDM5FdJcUxhSRruWtm1oti+xmxse8orjQ+Af/vn/ZTgdePjwnvMwcB4GQhKcwbcdbdvx1//lf+XFi+f89V//GlThux9+oneKxir6rnuyJ4VWCu8cL292xJQJS+YwHvlwuMfkhdxo7lYNgynockvbr3m+eUFWiqxhs3uG9w0ffviBlBZ8mem15tY5TBEPh1KEex5rPHku+po", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-84", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "5FdJcUxhSRruWtm1oti+xmxse8orjQ+Af/vn/ZTgdePjwnvMwcB4GQhKcwbcdbdvx1//lf+XFi+f89V//GlThux9+oneKxir6rnuyJ4VWCu8cL292xJQJS+YwHvlwuMfkhdxo7lYNgynockvbr3m+eUFWiqxhs3uG9w0ffviBlBZ8mem15tY5TBEPh1KEex5rPHku+pobmClEZAyRSiYgwpGpZEpYQBs22vLqZUssipgUd6eZ8/FEe16E/z0u+K5hM4+c+obh3LFZr2nb9mqa/5T1opnpTMAOB1ZxxjcjpkxoAg+5IUTHQ9hgTEvXbqSNbwqNtihVC7pS/ZqXhf/+7TtsCdw2hi9WiS83o3izlELyF+AwklJhHjNL8ExjD3qL8pp+vWXtNzzrv0DbHtduCcvM+bQnh4k5LAK+OY9XGldAhSiZgs5iSsIWjUF4+U7Cwj+5Pj3TnSfiPBOWWcCVkrFagbVCDK4D+utuf0Hdi8gFS23NDoc9w2A5HU+s1z3WKlZrsK4VWkuMXHgE1/BIpFq7zHOpyjMBOyun9VLV1j+/tlHlEphZObpFVf5j3bhzAVVVQE/kX07TAjpQmoI1mnXfYMqWVevI+URGZNBKK/p1I9JS10kChDLE8UznHA5FmAO3twNt24hsuqYdqwvbwVpJUG4aSeywFuuh64THrFFi6KLBOof1TjiQylCUkX7jwoooVAQ+kUuAK1ItVz5pkamKXfbTN92CdDT3+wPv7u754c077h8eOBwk92pZgpjRxEj/8EDXCyJttcTvSOZeZhgnhnEipsQSFk7nE+NwYplOWNdgGke3XoOG+RxIIaF0giwzN80FyBOqYgiBkBbe3z8wnU8cTwPzEpgjWN/hneXm2S193/P85UvW6xUhJkKYmecRs+1wxj+5IwLhl1MuTb18p4VMyIGQRR59jgtTilJ5KrHENKbK3ZuOUjLJWGalOeRCVAVDxhSFyZcKtsjGinDnJUlXNt2kCqn+LCF", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-85", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "SFk7nE+NwYplOWNdgGke3XoOG+RxIIaF0giwzN80FyBOqYgiBkBbe3z8wnU8cTwPzEpgjWN/hneXm2S193/P85UvW6xUhJkKYmecRs+1wxj+5IwLhl1MuTb18p4VMyIGQRR59jgtTilJ5KrHENKbK3ZuOUjLJWGalOeRCVAVDxhSFyZcKtsjGinDnJUlXNt2kCqn+LCFz2wUFWl0d+0JO5Jp/Zr3Ddw1t0woNr2/wbYNv3RVA08bWZOanmwBZAjrPpOVMSQvEiVJmcgmEYFgWw7SAsTDHn7GUFGKkdaVRwjTD8RxxRPoSiC6hUh1ZFqFsksW7IcVMCJUeF2NNmEk17ktjrEIZIAVSmFkmCcC80Dl/nogmwGSpxlvS9duKZVypfZ+8Bp9Yp4c7nIbh8EjXeHRO9I3HrBzDaSAsgWwMWZuaiyUD+6K0tCohEGPgu/cfiFGSHG5udoL0v0x43zGPI8s44q3Fdp3o/2v7mXMm5Y+UnVRqHpFKVwRMKT5uuFfalRKnpsI16yqjri3i1XVrnv8kkflfr2WKaJdoV7BZNWxubmnd5zjrmaY98zLx/v09xig++/w5SltKcWgtr+Pnn0ss9/7EOM5sVq+ZpoHD8Z6uX2HcRa0TiSiCMkTTYowlagM+49dQ0kLOgcRAZKbfrMFkMANkiyq9xLsgCQIkTVqWKpMdgARquRpdGzNhtK8P/5MuCSAjjuPpyP/329/x7Xc/8I+/+VZ8SXNiiYWYYToPHI4nHh8PbDYbPvv8NbvdjpvdtoYLZo7DyDQvrLcbtFbc3+/pug9stitev/6C9eYG/dWXspH/8BY1zAxLIuVISjOWBq2KhICqyHK8Y55G/vD9W0opGNPQbG/YtB0vX71ku93wl3/xFX3bYLTidDrx2999e+3y/K++ZtU+E2nxE9kLw3msm20V5miDNgVlk0hc55G3+zvCkvBGBBhhXujsiq7paaxlcg3RdZzUxLAs+FS4VxqdZIIl6deycSY", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-86", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "stitev/6C9eYG/dWXspH/8BY1zAxLIuVISjOWBq2KhICqyHK8Y55G/vD9W0opGNPQbG/YtB0vX71ku93wl3/xFX3bYLTidDrx2999e+3y/K++ZtU+E2nxE9kLw3msm20V5miDNgVlk0hc55G3+zvCkvBGBBhhXujsiq7paaxlcg3RdZzUxLAs+FS4VxqdZIIl6deycSYKs9J1zlvncEZXzqKqacIKYwVgdK0Y15yWszB/nOP5yxuUueWmclPXq5UwZroObz2N8xXEdRclx5OWjScoE+fpXu7hOKGIlJLZPzr2g+FhSRiX6LyQOmOo4wEuYaQCp4/TzP3DmYYz3XDilGbOBvmeShKefEni15FgmoTBRJggjpAHxK0sUvJESpFhOjAMA/f39xyP++r5kT4WeiUjonJ5NUYoitaLHesySgf1yWvwqR+qeqOVJCAQSeSC5ML5eGQcRw4Pe5Z5Zv+4J6cIOVdDaScVbIFxqLQqClrD3d29IMvasn98YDgd6bsea4z4r2rFOAwirV0GLiqjkmr1q0Q+eslacz/b9EMQKtrpeCbGRKw0qVwJwTLrkk3XGf1k9VXfb8D0aA1No9jsLN44nPY42+MXzcP9HcYUVq3Q2LRytaWHPAdUnDnmPSpPWHXGmZnWRawaKfEgcs8CcU7krAi2UC7CgRgoy0JKEykujOMD8zyQwkSOYjBEDecDS0pi/iyx8wGtF4w+X0blVyqUUhGFJkZhGjx1KTLeGV48u+Hh4UDTeMISiLHIA6rARbmBc05M88zd3X3dNDL9qpccNG1ROlTBiczRTuczx+OBm5sb+tgCMlrabm5xPhHMDU2I+DngfId1Laa/Bd+RTYPylvWzLzBG0bQNbScjiHXf0TSO8zAxTYtcz+HEYf8gFXPlh6saePpUytjPr84F4bXW0OIJ3pFjqGMwAYVTjsKqUYIFzMPMMAwSc640OE/RilB574JOyO92Wuh0zkiAnKoBstZatDVoa/DeVDq", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-87", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "BiczRTuczx+OBm5sb+tgCMlrabm5xPhHMDU2I+DngfId1Laa/Bd+RTYPylvWzLzBG0bQNbScjiHXf0TSO8zAxTYtcz+HEYf8gFXPlh6saePpUytjPr84F4bXW0OIJ3pFjqGMwAYVTjsKqUYIFzMPMMAwSc640OE/RilB574JOyO92Wuh0zkiAnKoBstZatDVoa/DeVDqhRG75pqVtW3a7G+mqKjNJG03fdpXB0VS+scNooWResuhSTPXB+vPX3f1ADCPDfk9JgZwCJc/kHHhziJzmhn0cUcYRjp0wnSq19AJYiNw7E0Lk8O4ex4TxJ9JxZnycxHw8J1KapVOIqioPYWUTt03GrSdUGpE+T1N1rSgCWonYx1tDYzWqiCSbqiRtvBfbT++uXtjD+UxMkWUa/ycVaYimucRwfcWUKWXh/u4Dh8c9P/7wI2GeiUu4nozWWYxx0k6Wwul4ZBjF7DiGQOM94yhx5e/f3XE6nXn56+esVmtKLszzxIf79wzjiePxQVRGxkhEdMokgqQKb0Rt1DcdjfcY65jHkWmc+O6775mmSTbdUoMoAZS6Rst//volXfs0IcDu5iWZDSPQdopnzy2mOFR2dO2aZXb8qCNaFda9JME2zknLkyJpGChmD/FHyjJhcqBRCdsFvDqQpsC8zISQiMualC1TPEBZOIU7VI7oEljCiRhHDseJcYwsYyAtCaMMWjsUDRRHjAaDq+KRGa3OtM2dGHvrvkqOI7ks5JxI8woVn+4ypkmsGscvv/6S83nknzcrzueRcQTvO2x17IoxMJ3lwD4dDozngWEY+fqbr/G+FfqNEee0VGmGXd+wvu+4fbZltfZQNFYbXjx/zZQd7Bpi0SzFyAN0OVGUJjuNdoVXv9jSNI7NpmfVObrWEqcTcZl5//5BYpCGA3EZCcMjfd+x3ayqctDUMc3TWR0f34uAmN47vFeU1NZYIsVSMkuYCU1LNkWc0zSczyceHx9JMYPSmLYBrViMCC+", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-88", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "dDozngWEY+fqbr/G+FfqNEee0VGmGXd+wvu+4fbZltfZQNFYbXjx/zZQd7Bpi0SzFyAN0OVGUJjuNdoVXv9jSNI7NpmfVObrWEqcTcZl5//5BYpCGA3EZCcMjfd+x3ayqctDUMc3TWR0f34uAmN47vFeU1NZYIsVSMkuYCU1LNkWc0zSczyceHx9JMYPSmLYBrViMCC+MFezDWouvxvfONxhnpDJ1ktLSVIpX3zV451j1HdYKVe5iYWnsz0cGH4UPF2k1QKohsqW6CC7T02XA3/1w4HA88U+/fVvFUYW4nIhxIud7SjGEIvzm7y7dax1TXsRUl4BZSoG4YFRibyLfGelwYop1xj9LB5XkvtDK8Itbxd98puk3J3bRibuYsogJlsHogLNKGDWNY+UdKidSWOo4QrHuZN/xjRciQQg83t9zOB4Jy3Q17Pr31qf9dL3HWQdFHNgb72tqaOAP337Lm59+4rvf/Z5lWdAobnY7vvzqK/q+o+u7qnapM5aYGKeJeZpZloW3b96x6nvmSVrcb77+hr5b1Sywmbv37xmGI6fjPTe3z9hsNhyHM/MSMC7jvcWpDq/BG1i1jnW/4vjwyOmw54c//IHT+Vw1cLLpmnryp0p3a51i6ton3TSng5DaTzExDHfc7/+eEhpK9NV0I3DYv8N7uHtb8M7QWEVOMzlOPD5+zzA+EoYfKEug052ICgq4aMmDRcWCjnB484FxUuSssDax6UecWfBuQuUzJo102qGdYdU6nLFoq8jZk6OYtaQMcUmUErBxweqAr7aXc1VZ5QLTLPEwy2CElfHEVWJEU7hZeV7dbvj6i8/44ae3TNMsD3X1o7VRjIhiWFgYWULk8XGPbxpOp7OwThArz5QT0zhzOg08POy5v3/AWk1KBaUM2xvwuuXl2pDwRGXr9FRVZrSq/FxojcLoiE0nlmNiOUjySIoSn5NCIEwTyzxxPA64axqBkTFVyk8e6/pGDnSluLJoShWjeN9SMrx", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-89", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "wy2CElfHEVWJEU7hZeV7dbvj6i8/44ae3TNMsD3X1o7VRjIhiWFgYWULk8XGPbxpOp7OwThArz5QT0zhzOg08POy5v3/AWk1KBaUM2xvwuuXl2pDwRGXr9FRVZrSq/FxojcLoiE0nlmNiOUjySIoSn5NCIEwTyzxxPA64axqBkTFVyk8e6/pGDnSluLJoShWjeN9SMrx48ZJhmDjsz1hnCWFhmgcUmVwizmu++cWX5Jzp2rb6QXw0R2p8UzfduoE62Ty999cK1VYjdmetdIbOXmXoF8e/+iZFwZku3sEfVZExJkKMzBeP25wl5++Jm26/e022a3YvAilIysMSNoQwU7KmFFiSmFnlKz6UMNWONVcxVYrIvuQMWhWyzoRSiDkTizAOYvFS6VYfEKc1ixL7AHU1oSqV46wq5W7BEmhU5KZ3fPZ8x/f7gUVZVl1P34rXbgyRt2/ecB4GjueT5KrljDcG4z4djPDJp8s7V03HPyKx4zCzTBNvf/yR3//uW/7lN78lhEDjPPOrV9zsdlijadvmysG9GLVM40guhePxJCYnWsxnGt8Q5rn+/cQyLzw+3DPWTXfd95jthnmeGMeJttNYU7BauHfOQNc4NquWkgLD+cS7t2/ZHw6gL5SajPOephX1kdKK3bolhvlJN81wgpDgMEXy45709gNxcqRgMSqjdaZpJvpW8XC34G2hsYkcTqR45HT6gWl+JE1vUCnT6GdQLFk5oZ2NClUMKmqGu8ThWJhnhfcF97JAO+P6AcoZVUa83qFsS9vIA2mdJibHMsvJnTMsMZJTwSwRZxKbXpOKIsxyZYqC4STy2LQ0T24ZAXISKs+qsdxse16/es7j/sAH9VDTOaRtzUZSf7XWH12slgXrLOM4cnOzBYo4ygXFEoIkYhxOHA6HmvqQquKtp2kLfb8mq0JSWjZbpci5ovh1sywGKAkVR5bpImcNpJQxWtyq4jKJi9s4s0sZY71Uy7lUn48ngkbe/TEAp9T", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-90", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "TwSwRZxKbXpOKIsxyZYqC4STy2LQ0T24ZAXISKs+qsdxse16/es7j/sAH9VDTOaRtzUZSf7XWH12slgXrLOM4cnOzBYo4ygXFEoIkYhxOHA6HmvqQquKtp2kLfb8mq0JSWjZbpci5ovh1sywGKAkVR5bpImcNpJQxWtyq4jKJi9s4s0sZY71Uy7lUn48ngkbe/TEAp9T1mXDWQwM3t7d4P7IsWWh/KbAsFXcg4Zzh889fY62lv1ao1e/XOZrm46artUaZi9+JuYLS5lK51mtz+RQffatLBZ+F7/7zt3xxiluCeDaP81QVolVa/MSTqF3fknTLejcTl4UwTejQo0OAC10xBlIp0lXnLI5fZEpJ1+t3ocLpKoMqWqLuc8kkErkkIoZCJqpKpVNC6RRMqKaJ53wVlgCoEtAl4IisG8uz7ZrdZs05KbpGuOTeSRd/9+GOx8Oeu/0D/WolXcVqVV3yPnFffOqHL168out7GbJXe7h5njns97z56Sd+/OF7Doc9KURGpTFK8Xa7gZJpvK/heoH94579Yc/peKonGDJ3iZHbmxu26w3/P3tvtitJlp3pfXuwwaczxZARmVmVVcUqsouk2EBTILqhAVA39AIaIF3oUm+iJ9KF9AQCBAGipK4m2RwrK+eM4Qx+3N2GPepibXP3E5kVGScI5NVZQIQPx93M3Gzbv9dew//fXN8IYU1W3KxvuL66xLsdftzhx57oBqwSQpKL0xmztmLRaGoLrYGZzbQ247od2/U13XZHv+tlZWekq8Salvm8wdYtxloymdEN9xo0n376Bc4vudpUkrhqPMQKosUUdrWTZWJtPC+/uqLSntaMKDq06lguHXUTWM48ZBg2t6RsybEhR03CkFVDzhXNbMYsWwIGLLgMcezoA+RgSLFlGGpCMKQQsQYWyxo3Ck1e1w3cDoF+aAnB0KaOSifC0OBGx+Xla5pGYpybWxhGw/PnF7Sz+6kBAHx1/TucG7m8vmSMik8+Oac", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-91", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "6Bc4vudpUkrhqPMQKosUUdrWTZWJtPC+/uqLSntaMKDq06lguHXUTWM48ZBg2t6RsybEhR03CkFVDzhXNbMYsWwIGLLgMcezoA+RgSLFlGGpCMKQQsQYWyxo3Ck1e1w3cDoF+aAnB0KaOSifC0OBGx+Xla5pGYpybWxhGw/PnF7Sz+6kBAHx1/TucG7m8vmSMik8+Oacfd2QFV5eyfFcYSbymhFJCVh+8jJHtdkfXd6TkBUxmC2JWoDaMznN7u2W77VguhY2OrBiGLUolFvMGrUeM8VCANyapWBmLSOFQZHzGQXhmh2GQVEgWYPcxcrsdpIpEzxhc5ubmlqtZhc6ORxfnGFPf+7ygplzTtFzXGFXR1IrKNjx/VhNC5Pnzj/dJKWu0dIUZg1IKWzq/BEgV2hzqqvfeaiGImvYl+xUvESalZlnhTNUsxx12b1ZmqoklLlPKCjXGVNR1JmpJbut832kIxu4lQ9cz9K+J3ol6t/eoAuSpVBSpLKTipENllMpZMiNZ6mwPleelQYpc6uODdJ1OwpKIrJcGWpNFcTp2ZA9p8ymMV6J4jmHYXuLGwK7zrNeRzTZxujpDz86FCS94fvObf48PkTFExhKqrCvDrJV28x8q6Hgr6M7nc+q2xWhbbhJwztF3nRCtbLfEEKSdNwaGvmezXrNcLulPTvBOutfGcZSwwjgScyamTAoiV7OczUizGcPQ03c9uQD7MAyk4MgplIB7ECYqo5g3NbO2orFaWo1NUYDXmRR9aSP0RO9RRmZDY6WttqoqmrYVJdjsiPdcHt1utjgH65sdqsqYmUflCpUNRiWMTtQGyD2726+xamRmRqwZqPSAfmpQK03diPfnQyAHyLHac4YmJcXYaClYdwl0UPSuNHtkIFpyVAy9JkZN0glrFJVVBA/kSAiOMY+MYyYEiyaTtWK3M3Rd4uWrHbNZYrWybDYJ5+CD5y26mt3rnABc714xupHL7QuMWTJfPub0bMb59kSSQSU", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-92", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "H65sdqsqYmUflCpUNRiWMTtQGyD2726+xamRmRqwZqPSAfmpQK03diPfnQyAHyLHac4YmJcXYaClYdwl0UPSuNHtkIFpyVAy9JkZN0glrFJVVBA/kSAiOMY+MYyYEiyaTtWK3M3Rd4uWrHbNZYrWybDYJ5+CD5y26mt3rnABc714xupHL7QuMWTJfPub0bMb59kSSQSUDTEaaD3LC2oqUMjpmfBjJPtL3wl61altM4S5IMeFGX7qVIo0RcUNRbvUoHAaolAgJihcjYb+YIjlGouvxo2PY7RiGgX4YULpI5jiPDyIpnrMCY/E+Mgw949gzOkvKp/c+J1OJYsnhAiUpB8WJgflcbkHxMEXBQ5eig0lHzBhTksClAkcfkSPtTe2bYqBwTWdJGqc0EdPIv6mn4bgJRgBX7eWWtC5HX95XpbHHIvI6KaXSIHU/2A1uV6R1dnJ/h1JBFCM5SIJbAscZnaYS1EOVhEwb0sq/n2LKbKGUSN2jS9ljaX4KhdoKhNJRkyF7cnIkdyPEQe0NMRtif4MbI7tdYLvNbLZgqzOWVSv8Md7z+vVrCQfO53sqg+NCkR/Kzb8ddBdL6tKNllGsN1tevHjJ5599zub2lhg8F+dnEqNzwk702Wefse06rm+uWSwkEeGHgVSkYXKIeB8kQ2itENVoxe36FpRhsTor3KWib1bVDTlH+m5LTuIFnK5mLOcNZydziVtpTWMyKjp0lmVu0caU+KDWzFohNlku5sxWK6q2Zdje7Js43tlKolOhqYua7l4Fg4DKCe87+l3P7/7+ayyJRaVYtppFW+N3lmZu0CfnRAzXrx3WNKyWZ/gYcTEyRk2ImuthZNtt+c1f/yM5RU5XM+oq0daRZdMwq2tmdoFWht12ja0U8+UJPoALUM2E0vFcn6P0jFoLAN1eb7l1A1+8/pzZbM5yWJKTQSuDt4HQ3D9h9P98/n/gg+dmc81yfsbjs484fXLO6dmHaJN4fXnLixeX+JBK+ZR", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-93", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "YtppFW+N3lmZu0CfnRAzXrx3WNKyWZ/gYcTEyRk2ImuthZNtt+c1f/yM5RU5XM+oq0daRZdMwq2tmdoFWht12ja0U8+UJPoALUM2E0vFcn6P0jFoLAN1eb7l1A1+8/pzZbM5yWJKTQSuDt4HQ3D9h9P98/n/gg+dmc81yfsbjs484fXLO6dmHaJN4fXnLixeX+JBK+ZRI8mRRYJSQQAx0XY/3gXZxTspGeBsShdxaiHIWq5XwSxipCnFjwJPosxM1W23IWegku97RD56vX1wxDBIfHt2IcyMRS8qaUErrUoyFvNrRpxm9WaKen9E2dalyud9SOpL34DbFdSfvUCpK2FOYai1gNt20B9KUqe4z7+/mg+zUVIM9uZzFm0aew11wVmXnaWqIOfK+JxkcCglUjFH+rgXMrbUYDLWqDyAoVGT3srHr6Lc71peXpRogo5IXwvCp+ikrzNFMJSyE05QxRe0LR2CpN9b54PXm5At4F9BNE52sZqYThoiKA3ipSCB0xM0ZMVvC7Q3Xm8Dff+v4zRcb/sNXO376qz/h5OKJjBUf2HW7Et5a0gdPv7mlr4Wsa3V2Tv3Piek2tYCiLu2pxx7uRCq+Wi7RStF3PW4c2W539N2O7aamKSq+J6tFWS5phkEkOdATVaHEbgbnqIaRahakUyglITWvJQPbtg0xCylxXdnyTzweXZQOph55a60MFq0xRsp0mqqisoeEzjRb3zf7WlUtKTdYIyGXSe9LCFxK9LiEulIyJAykmujBZ1jnhNkm3C4SMmxvE02dwRjpCsqaMRp8Utxsdmy2PZvNRkqtwsjpqeH0pGK5SqzmGZtFaWDbSYlVP/Q4pwheU80a6qoWtVRtqYxIcE96UT4FbPS44KUxWmt6N2DG+ykkA6yHK0IMdHGLCprGtcyUoakNi1WFC3Oub7akJK3jck8Z6ezJGW2LEkCMshJK4v1pbYUOM8UCIFrOuynx1gzD6Ikh4cZAVYusujUVKI3zkdF", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-94", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "sdmy2PZvNRkqtwsjpqeH0pGK5SqzmGZtFaWDbSYlVP/Q4pwheU80a6qoWtVRtqYxIcE96UT4FbPS44KUxWmt6N2DG+ykkA6yHK0IMdHGLCprGtcyUoakNi1WFC3Oub7akJK3jck8Z6ezJGW2LEkCMshJK4v1pbYUOM8UCIFrOuynx1gzD6Ikh4cZAVYusujUVKI3zkdF5dp2Mu250eB/wIRNSIGVFoJANRV9KIweGKjIOmjyRBnH//ojveJIIcbvKUzv23ZpopQ5jU6ljbzVPGLr3Zies3fP83vGwjjaaD6/fPP68b5Mv30eaiY4BX00bLg9ZS2hnv717nhS5Z6wImeYSAilddVOII8XD/qejz0weLSUdWRQjlBJCdzgC3Qw5SaNEFt21TGH8KzW7UvsfS5NXwHUdMRn8ODIOkc3gWXcj642sjBbeg6mZJkGtpBY4esc49AQ3I/lauILrt4eh3gq6J8vCIhRH+o1nu9vy6sU3vH71EpVhuVjyi1/8Aq01l69fc3N9ze3tmqHbckvkyfmKk/mCv/jzPyPFxBdffM3LV6/5zV//R5RQ8tCNA1FrlrdbRipyvaDvR0bvqeuG1ekJH338Mc8++IAvv/yaYRhELaCpaZoWMsSYCVGRhoBt5ixOz9BNjRkG5ouaWdtwcSKEMioG/NAL23zf7SXf39UeXfyMYahx/RxtixDPtFzTkgVFW2w95+LRT6lVzVlzyrjZMmx3vPj8G267Ld/GDqc0lbWcna34ZT7FthbbtgQ1wyvNP/zT33Hz+orxdiRFz3j9ko8fX/Bf/Ouf8PFzzZOLzDdfbLi5Sqw3sOsjX339NSlWBDdntrxgMRNl4lxCMynDEBw+R0xdoawprHAiBfPlt19hLu8bqYNXwzeAQlnNOlyyvr5iXn1NY5ZcfPhrHj95TnCam+stX3z5lbRlG7uvK1Zak23ADdJI4L3UKFf1nOB6ghuLs2fReobSLZlE72LhULjl229fc3F+wcn", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-95", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "f/Ouf8PFzzZOLzDdfbLi5Sqw3sOsjX339NSlWBDdntrxgMRNl4lxCMynDEBw+R0xdoawprHAiBfPlt19hLu8bqYNXwzeAQlnNOlyyvr5iXn1NY5ZcfPhrHj95TnCam+stX3z5lbRlG7uvK1Zak23ADdJI4L3UKFf1nOB6ghuLs2fReobSLZlE72LhULjl229fc3F+wcnJCRePLqjrlm6MbPvI67Vj9J7Rg9I1qmokDBUDw+6aGEZcfysF+34gDmeY/Bg3/vWp9g4AACAASURBVAFGF5HTf4595+vljQKcMcsiWCSl8h54xXudUHU/1MomppBBphRy3wXyvZc7vaHf+O7Rx4oZq453sAdk4NA4U4LHSud7Vy+cn3+IsVvOT3tG7xj8SPAjwTtClvps71xJvk8td4VNMEuibB82IheCJTBGzoOCfXhCWvCk0WiCagM8nYNzCaIkX300fPnimhgVtYnc9IprZ7gdM9sh0u12LOZr6uU5WU1MYonU7xjWN1y/fMGysSyaitPVktPTt4ei3gq6s9oQc2Ycdmy6gav1lm++/prrqytSilTWCOOQFpb++XzO06dPOVnOWS0XPH18ztnZKaenJ8SU6buecRyojdAvTjyqPgR2u46sKkzd4sadJDe8wfkgMuPey4XxA7frNW6ocMNQykdyGW6K9WbDMI7UTcVsMaOdVbRNRVNbjFVYJXR4SsNyMb/3oBkGhw+KZja1bmZJjWdRuRUMjtg6szo7Z2ZbHs3PuSHhhi3VDOZG8eFS5IXaumK51Dx+1GPqGtPAukv4PkMQgutaVWQFIRhwirCL2KDEk8RRqYTWDaBkMKVcmiFKOZCWmtWcAzkmYhQ1gaaWWk5jpI5Xa413I97d06WD/SpEK7tf9nZuy6gcjb2mTtC2lvm8oWnbPVsdStpjMZakFNpII0XwnkymqipydFIOVGoiQ8hUUcAzkQgoAp6o5oy5pg+WrTPUaAYPQ9CEDCFGxrEvHUgRN4j", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-96", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "QgutaVWQFIRhwirCL2KDEk8RRqYTWDaBkMKVcmiFKOZCWmtWcAzkmYhQ1gaaWWk5jpI5Xa413I97d06WD/SpEK7tf9nZuy6gcjb2mTtC2lvm8oWnbPVsdStpjMZakFNpII0XwnkymqipydFIOVGoiQ8hUUcAzkQgoAp6o5oy5pg+WrTPUaAYPQ9CEDCFGxrEvHUgRN4jUix+6kgvYlhWFYygio24chWT+vutoOGoHzXeATLzdN7q5VD4s9xG40/Lx7yzhJ+dVZbX3UA8gehxcOLw+7L0Il+e8/9TkSKs73/7OHu/ufP/0fpNRSoqYNCEZQtKEqIi5EmGCqsaYhM8DOUYSXojYi4ov6sAuPSX6JtBVU4YNUBjxdFWphoh6v0pKCkLOhKzxSbHpoQ+By8stKWsuTi0uWMasiaqQALmRoeuwsyUqS9hTxUQMfq9jl6N4zak8vs3eCrqrRUU/el5fXfLZF1/zH/7670vB+0hTt/v+bNEdM5yfn/Hhsw94dH7CxdkJ52dLZrOa09NTYhKiihQ8i6ZiDB5fJMiTcry+vqbuR3rnSNHRDyNGw2ZXcXu7Yd42dLtbhr7jK7fdx6KkMybtEwHr7Ug/eubLOVVTMas1TWWYtxKKUDpjLOhKszx/hPmB8o437er6FmMblqczfEiMLpBzA7nB1g3GNmSnMbZlNZ+zmi14dvaY36me9fZrTmeKM2P58NcrZivLaiaEOFqtUaYFk/jH3+3Y3Q7Y1FGTsHZFIjCMnrw13H7liKeG5kxT+UQVEpoWcmQchP2oqSqMKnpS2qC0ZRgC0UeC61A5spqflT5/Q1VJVvxmvcO/B4l5zOKnpcyejet2uMGFSDZLZmrLYvURWi84uTmh63q2m3IdCyOczlJ0n2OkHwesMbSzlhRHMhnnvOhojR5jW6q6JSmNN5nUNKhFxWga8fT7GustIcDoEgGDC4F+e43rN7hhy9j3pWuyNNBEt+d8NkYqA7a7HeMonU33dXb", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-97", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "0ZRgC0UeC61A5spqflT5/Q1VJVvxmvcO/B4l5zOKnpcyejet2uMGFSDZLZmrLYvURWi84uTmh63q2m3IdCyOczlJ0n2OkHwesMbSzlhRHMhnnvOhojR5jW6q6JSmNN5nUNKhFxWga8fT7GustIcDoEgGDC4F+e43rN7hhy9j3pWuyNNBEt+d8NkYqA7a7HeMonU33dXb3y3+O9O848Ie8aRN3si6T5dTx/71wXzBTo/fHpd4EXnVY/UskQd5P+zngMBkc//37d3YIaxTO1PeYhmB08q/3msEpRgdKNyhrsY3Akde3Et8dR0gRFRyT1NfE9WDLERjEiZpoATKg9y3bRdQzltBUAqpMJDJmTRc1L24S693Ip7+7wRhL217QB0OXNEFZTFUz9B0bnZkvT1AIJ2/KQaqqvJPKiZiKivKIH95eEfVW0LXGUNlSlK0UQ9/T7XZ03YBZKpK1BO+k5jKFfRIgly6n282artdstreEmHj96pLtbkNdV5IMCFnKX3RGJwfBiExLaQ8MXtqBr6+vUUUFNoZAUEJMPUkBVVYXuR9NM18SEpxfnJFTZlZp6cVXsbSWemanZ1SzOYvl4t7MUc3MUNeas0eJYRjJmx2kGSQPDKSosTpiKzhbatpqh64dq4trnvmRuhZSkUdPA3WbaK0UfYfgJQxACzvPeOmIPpCIxCoQiQSj2bnAy6stX78MNM3IzUbRDQZUjbGZ2ewUo2uaaklOmt12xPsbfNCs19c470jBY7VmOVvi/IjrB0lSWli0LXD/6oWUpqVoWSRrqGohEnHpipQHThZzVosV/+bpn3Jzs+Hzz77m8vqG62tRsJiacJKGHAM+e+LOM/Rdaah5QbeTOPdsvmB1ci78EqYRVVnvcG7DkDO3rwUkplzD+voS7wbGfkP0Y2HC80yabkCpZFEobTg9PeNnn3zCkyePOTk5QWtdyq/e3WKQz+cc9yC352EtaDrFcQ/VCQd/U+wuEEoYVZbROUPWB8G", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-98", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "HThZzVosV/+bpn3Jzs+Hzz77m8vqG62tRsJiacJKGHAM+e+LOM/Rdaah5QbeTOPdsvmB1ci78EqYRVVnvcG7DkDO3rwUkplzD+voS7wbGfkP0Y2HC80yabkCpZFEobTg9PeNnn3zCkyePOTk5QWtdyq/e3WKQz+cc9yC352EtaDrFcQ/VCQd/U+wuEEoYVZbROUPWB8GpiQD/GAwnb1upKfig9hPAndBBZk/sfwDiQ9XDFDvNEldAvMZjXZN3s+1mwzA4qsqQotSRByUr3pPTU+qqpmoqnBvZ3lyjk6ORzmMhDJqOZ4qzKABNLooxKYMJO1QMDIMT6gDTIGrimiEGvrzN9DHxzWUkOIUPmlgt0LVByg2T5AlSQhvNMPSo6FidrlHaYFTGh8B6fcO264lJdAv9MJL7ThoH3mJvRRyjFZU1zFopvB6Hgb7r6HY9s7oh1jXB+1LoHsmF6FnkeTzDOCAigJEQIuubHdvthroR0E0klFElmRJQcSQ6KZqf+FaHQbG+viaOPbWVhoo4tSZWRfywabCl28xULcrY0iuuaY0mR48fd6W5ouPk8RNmqxXz+QJr7we67czSzAxnF4nt1jH4rczGOeD91D2maBvN+UWLUQM67Vid32DrkdOVpm0sizZgNFgVCT7S9SORmpgTeesZr0ZhQ1KRaIIwRWnFzgdeXm85f+mxtibGJS7UEpqxitlshdE1lZmTUqTbOtabnr4PvHjxLSkFHj1aYrXwqKYQ2PTSg05SnKweUVX310ib8jlapX3na12aIsbxmjHfspo/YrWY8+c//2NevbjBKAgp8PrqWjyZLIQmGvBBFCyC73FDLx1AL17y8uVLrq9vaGczHj9+TNO0LJanJZxicf3AMI5stptCOXqNdyP9blfAaupqKmEmhUy8SpGYal4tp2dnfPLJT3ny5Amr1Qo4Dhe8m6Wp5Kmwtkko9EjiXlFkm6aE1uShHt67s8fJbS0e555HhCOXdmLoQ0B4ShR", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-99", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "//2NevbjBKAgp8PrqWjyZLIQmGvBBFCyC73FDLx1AL17y8uVLrq9vaGczHj9+TNO0LJanJZxicf3AMI5stptCOXqNdyP9blfAaupqKmEmhUy8SpGYal4tp2dnfPLJT3ny5Amr1Qo4Dhe8m6Wp5Kmwtkko9EjiXlFkm6aE1uShHt67s8fJbS0e555HhCOXdmLoQ0B4ShRLdY3+DkiqA/ofknbHf9/XA8s9l0igpyYUEVq9j+12O5yP1JUley3sZ0WiabFYMF8ssXXN0Pe43Y4qRxZkrM0YI4CfUaUNWK5ZUpakZ8SUiTljxgGlIkNwpABZLwQTrGXwI1/2ntfrRJMD58bQGE04mWMbQ1aKEDODk1WPMZpx6IhDkgohW6GR6o7LmzXDIGWwwQdp8Bp6aZF9i70VcSQzKp1oZycrfvqTj3jVvOb6eo014N3AV19+jkLR9R3WWra3LTdtzWxW7aVY6loSclVTcfHonPl8RZ5EKvV0YRsyqtAiaprZz2ibiuWsYdE2tHVFVUBXahxFP80YUReVHnIlihFqIsCBSiP0bmEuradupF2dUM2ED0Ddk5ruyZM/wFpViI4bTubC0WuNxY8jKTrIOxYLxS9+usKaiGEgOE10S3LsydHhdg6CwlQ1qqqYVy29bxlH2IyR9S7SLj6gWVkunj6m2+34p7/5O8aQuFxnvnyRCClxdjHHVkvq2YxkIPYR5yO322t23Zau27LZDoxjYNfdSn1rflSqLixuHIixZ96csFrNOT07ey/QNVUQ79aqfUF9nlo3i3d1tfuGmBwvt7+hWZ3yn/9nf8qTJ2c8Oj/nnz79nKubmxKn1EjxgVTb5hJHTIVxbrvr6PuB3W5XpIGqAlB6TwXqgy/SLyOT0i17ANH7WKhSSug3rWWxXHJ6dsrPfv5TfvHzn/HrP/ojLs7P7121MNmBu2B6fYjbyi86PId9bcNhGX+0HbgLiMfNEcBdUvOc7/xTSu2f3zm+N6R29nwL6ug", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-100", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "hWZ3yn/9nf8qTJ2c8Oj/nnz79nKubmxKn1EjxgVTb5hJHTIVxbrvr6PuB3W5XpIGqAlB6TwXqgy/SLyOT0i17ANH7WKhSSug3rWWxXHJ6dsrPfv5TfvHzn/HrP/ojLs7P7121MNmBu2B6fYjbyi86PId9bcNhGX+0HbgLiMfNEcBdUvOc7/xTSu2f3zm+N6R29nwL6ugYp+1OlQbIUjojOof3zYlU8xU6Rs6NotOBPCrqtIPUc7JJzF1LGzKjGyG/ltBS7KQ13VhSDkWfUyabqBqqtub07NGeR3h8vWPsNlzfDAwBzj85ZTab8Xi1YNd3XK9vuL26JN/2PHs047xSqHGLAcadIo41lgWtNSyaGud6Qoh0XYcxlnEciN5RqYxtaxZ1xdOTBRerOWczw8kP3D4/ALryaLRhPpvx6OKcMHrhwC08k323RTr3gjQw5ERKIz7Y0q5oUFqUAWZtjZ1XXJy3hfNUgSosYFHKxza7HlNVnD06p21qlvOWprLU1ggrmJabZFqiyWvJXmoztTru87gYJaQ9KlpibPC+oZrPMXVzWObdw2azRxKDJWB0oqlGmlpTVQqvHClkchqY14rTVUtdJYxKhw6yrcONirBNUlpmRFJeWUP0mjFGBh/pxki1OKNqZzx6/FgmLquIGbo+s95mqgqquaFRtmifQUbjvGOz23F9fcPNzQ27XS98tslRWc1qJwxS1uhSjJ9p6lp6y2dzqh8oefk+0zbKtbBakiWxtJZmYKqZDRuMVtz0X/DBmebjD/6QYfC4MfPq8orNdofPEYp6g5T/VFgrG/JBkZE6b5eTSM/fyecfA0aJoZaGjJxL8kUXWRUtZO6iWtGKcvDZOU+fPuZXv/wFP/n4Y54/f/besjQgMdrikx6aI94A3b1Nia07bx0A89jeBEbZ1l07Btnj7RyD6Z3vl+cTiN+ND08ee+HqUOK9p8JVex8zzRwVHVXohcNWZWocJnc0I9SxwmSDCYF53uKTx8U", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-101", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "/VFgrG/JBkZE6b5eTSM/fyecfA0aJoZaGjJxL8kUXWRUtZO6iWtGKcvDZOU+fPuZXv/wFP/n4Y54/f/besjQgMdrikx6aI94A3b1Nia07bx0A89jeBEbZ1l07Btnj7RyD6Z3vl+cTiN+ND08ee+HqUOK9p8JVex8zzRwVHVXohcNWZWocJnc0I9SxwmSDCYF53uKTx8UBTS2JvlzOWhKNFrTGqsy8KTSmGbLWhJwYBk8XFBempm7mLFenRG3IvWOMa1yf0WhabbCxQ7lEGBUpKIxSVEbTVhZfVu/OOYyJJSwVaWzRAzSW03nDyaxmVimaH0gTvRV0tbYYMrVNPH38mKpq2Pz0E3a7jmGQHmwfhAhbK5GcaduGqjbYytLUlfSKNzVGG6q6xmhLZUSzXquMkm5o/OjxLvDi9TW2rnny7Dl1XdG2tdCBMuV9QSm7T7UqEkydKJNIH0BKBXSFN8IYKYRvalPKpPK9ARfg+rZGqYzRlpwMKc0IMaHGROxbchypqpbBwevXLXWtaSuFH0fcOLDbgBtB+Q0pJLaXjsE5rrrEzTZyeRv4/OuR7Zj48MmHLJZLWm3w2rKcLUnJMzjHZmjRfcu3//ElKb0k5VNSVAyj8Apstht2ux273Y6JFq8tYoIqZ9q65unTx7RNw6xtaZtKKlFs8x0P6F2sruXmDnHSzVNkX5N9jQoalRSRke2w5f/7/P/m2e03+Dxw8ehn/NsP/5yqsnz6u6/467/9R7a7jhgm/lmRibF1Sx293OhBHmXwZ8kgM8mmlKFRmL2sPWIdK+kYbQxGG1bLFbNZy4cfPuf09JRf/4tfcXZ2xkcfPRfWOiXlgCmn74DUO1sp29rX2qo3wPb77B1387bJ4Bg83/Raj1nEoNSsFpCOb2Tec+n2mkA35QTakpUWXot7erqLZ59wc/mav/nN33B7c8PVq1d8eJp5vIDb9S09ieAGFImzOqPrhGqE00Qrx2xRU1WSN1FkYhiALWwuSxO", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-102", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "GG1bLFbNZy4cfPuf09JRf/4tfcXZ2xkcfPRfWOiXlgCmn74DUO1sp29rX2qo3wPb77B1387bJ4Bg83/Raj1nEoNSsFpCOb2Tec+n2mkA35QTakpUWXot7erqLZ59wc/mav/nN33B7c8PVq1d8eJp5vIDb9S09ieAGFImzOqPrhGqE00Qrx2xRU1WSN1FkYhiALWwuSxOFYqxuGU56Pq88Y7SoMGJTy2mj0dHiFzXfLOdsBs/u8U/YLFp2X/0denTM5omkWxZtKyV8BtxuQ3COoR+E8FxFzhc1vzx7RmVrmqphPqtLX4MA/tvs7dSOU9G/UjR1zdnJiVDDLeaMw7gnKwEBIVsJwNrKyPOqKpnxat+0oJXBagEuqRoU79hZh68Du77HWBErrOuKpq5LvO9Qsze5DdJRqY9yD0WSekoETP/KfxotxCdTQfghEv/O5oLexx/JlpwVMUo7cXIGkjRthJh58QrqKtPWGjeAG2AcpAJmbhUpZta3gd2guNxkbnaBq42jGzwhJYzOVEahkzDUV1WDd9Kuu+sT2MBuK7zBVot0iyvdVn23w/sRchRWKiPdc1VdM5/PWSxmLBYLZm3DfDY7sE5picPe15Q6kKXkPCVsFORC4AOQlXAd91tuqite3X7DyZMPWK5qPnz2mBgyX33zgpwTu11ftNvsnaV0TpGoFClJx9QEiBMD1ZSRV0WUURQ7NMqUx6KoXFc1jy7OWS4WfPjhc87OTvnww+cslwtOSgy3BEZL/PS7HucP2SEEWwD3yNN89218/3feDDUc23HY4Ycmi7ukN/k7fzv8PR1AVxVVlu/5zg9ZUBUuG3Yedh56r+ijZihKGDYHnBNsiCoLMZYx0qWmNbUyGG2JRYRVJmGp7TVKYTSSJE2JkMGnzK7rsday3u7Y9VJ91ftInzObBE1SDFFajH2S8l4Rfs2Y0nIdU8J5j02aqlbUxrCayThqqoaqNJJJA8Y/A3RBZmVjDDNtaeuG1WIuXWQ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-103", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "kN/k7fzv8PR1AVxVVlu/5zg9ZUBUuG3Yedh56r+ijZihKGDYHnBNsiCoLMZYx0qWmNbUyGG2JRYRVJmGp7TVKYTSSJE2JkMGnzK7rsday3u7Y9VJ91ftInzObBE1SDFFajH2S8l4Rfs2Y0nIdU8J5j02aqlbUxrCayThqqoaqNJJJA8Y/A3RBZmVjDDNtaeuG1WIuXWQh7D0LUNLppSXgrpS4prpwmuoJFQttn8oCXFplpKg709iZyGpEiQGqFFBJobF7BYJUQhq5ZBFsJR1m2hwgdEojKK32XkZKEneX3MQkTMnvqcV5u/lCejSO0kXoHLx+tWF9s6aprFykOJCzh9xJokYpvOvxbuDR+cBy6fnVzwSQ/u6fOnZ9zWa3Ykg1fWwY3BURB7mHWBN7g/KRk8UJVy5ydX3Fq5sNWQViEQ9dVgMK8H6QbLkKNI1ldTHn0eNTlss5jx49o21mLBYnhUO1OpQ1KUVSU+jm/qArfA8KovA+aa2IVSTpkWjlWsYRctKE0UgSwv2/qNSybE75l3/6B/yrP/lPSDHy+Vdf81d//Q+44snmbMlUIjeUshClpEispPogF09XQgkCuvtMfKlqqeqGtm05OVlxcXbGxdkpn3zyEy4uzvjww2fi7bfNnfhnzgfQEpKe+wPmFNeeQmHTPH+cNIO7Q/HNmO7x54C7ybGJ8KaA7PTaGHN0DHf/HXdhfl/8d3r/sN985z2R4kp7RZf72GffXrHrHOnsY6w9ZWZWbHSgS5HWCt3iut+ITtmrG8nLWE1di1zQqZvTVJbt1bdEP6IT5CQ1s7WRxHmiIiTDp9vETRf5h7/+LcZa/q+//ZQUkpDre0cMgdpoLuYNCwKzyjBPDX2uGKNUJ3WdY9PtWO96aheoK8Npu6IxZk9hMJWmhZRZ70Rs9W32djXgPfPQ4QKIVLcqPKMHqrhJonqqu5T9mn3JiXz/kH2deFOn40ul9Vc+nckxEH3GKWGol2VP3g9aXQAiK1F", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-104", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "T5CQ1s7WRxHmiIiTDp9vETRf5h7/+LcZa/q+//ZQUkpDre0cMgdpoLuYNCwKzyjBPDX2uGKNUJ3WdY9PtWO96aheoK8Npu6IxZk9hMJWmhZRZ70Rs9W32djXgPfPQ4QKIVLcqPKMHqrhJonqqu5T9mn3JiXz/kH2deFOn40ul9Vc+nckxEH3GKWGol2VP3g9aXQAiK1FFkA0Vld9DEc2+XLq8PJyP/B5LRaDrNyKP4hEmqyzdXmiNqYVlaBwGUQaONTkKl3DwiuAVs4XBhkygRatEIBFzQ2RFSA4fBnxUhJiK5FBkNpuRS2Y954wPkaiK5GCS8xiUzPTWyoqiaRcslg2LRcv5xQmzectqsaSyQk2nCu3fFHedluNTU8V9zfWFiiRPNIITd2yRVlIKo6tSRylXykfHtr/h9fpbzufPmbdLPv7oA5SCr75+wa7r2XV92YMh60Mhf0pxD4ZJC2OZKuMHKITcak99uFwuWSxmXJyfyb+zM5598JTVyZLFfC4Kt9rss/7fZ+8T25U4+1FM9yhUuq9XmFZuP7B5/T2JrzdDBcfHeuzFvmn3+y3FkVEH7/19Ii0APkRCSPgELmXGlBhjgOzpAJUym+2Id45dnzBaUVuwMWNDwqlAXQNqiapmmEL3GNRI0pqoNTELp663img94xhQIZF7J46d8+ToySlyvetIOaLn4hx20dABY/K4EAhRSJZCFNoDayacA3LpkMuJHMX79+/Au/wDoFtiPWk69ROoSawUXbpnShjiUHN4tJE9uQZMgofTjScOlQBjP454F6ReNUNOkTBmuq3IcoQY9+TLpiRsjJkfSeF8H+NRiUcdBRP09OQ9B82r159KdYRu0abG2DntwoCdc3p6SmUt169TmUkXDL2nH3bEpIloRgwjgWznoMA0GpMaTDondbcM/orBKfrRMzoJWzz94CmbTcc//NNnUmscAkmlMsfIb4xhwFYVF2enrFYLnnzwiPOLE87OlkWexeBHmUh", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-105", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "OkTBmuq3IcoQY9+TLpiRsjJkfSeF8H+NRiUcdBRP09OQ9B82r159KdYRu0abG2DntwoCdc3p6SmUt169TmUkXDL2nH3bEpIloRgwjgWznoMA0GpMaTDondbcM/orBKfrRMzoJWzz94CmbTcc//NNnUmscAkmlMsfIb4xhwFYVF2enrFYLnnzwiPOLE87OlkWexeBHmUhj5OChFKUAbewhlvseK4DNlTDRVZUp1w9M7TG1xAiVUtSmLYKaFVlHIp7X62+ILnG+fMKsafjzf/lrfv6Tj3l1ecXLV5d8+vnXR56cHJ+3VsoSvS/ton7vwYlkuGI+n9PUNY+fnLOYz3j+wVPOTpY8/+AJZycnnJ6cCGev1kUQdPIwD6j4vsAymSpehdTmTo7LYeV3vPl9GGzyrqcxe3QHi0y9/g74vhkieNObffOzdwF3Oq7p/tXf+dz0O0DtQySHxO39LIaI84HdMLLte276DcF1BD8Stx3ZBeE5TsKuYIyhaWpUL92L1Wakqit+9skntG1LbaSKwo0jKSvGpBi9l6ac5Uu07lHqlpQTXUIEKkNERSnz/Ormlqu+op0/IyvL6yHRpcyN3zEEqQRy3uOcF1GALGWuqKkpS0JbMQVSSHReFKbfZj9YvZCPk8PTkNx3gOT9hUCp4oWoOwNIvnb0YgoxqGn+FEtlSWVsJUtyq/eObV2nUjNnirctJNiVNYfYFdNAVhzdN/ukcN77Xsez9Ht4LkmWDzEFFA2YRF1LHWFdOayJXFy0OKe5en1DziPeDzi3w/uO9e1IjJHPvlConPnsq0jXWdbbLd2wYzvcsusGQgAXIi4GQk6EHPBxABWYtSUrrqBpGurKcH6yYta2PH78iLZtWS0XNG2NVjUhKFTMpChRdF3agpU2IiK6TzKVc/geYKNiC2TGlPcDR1dFE1ErFIYqG1Q2kCOprJBG37POV3z+4h8Yx4Fnp7+gaTX/4pc/43S1pC+5A+/jfjyG6m5FwTSetFE", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-106", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "sq0jXWdbbLd2wYzvcsusGQgAXIi4GQk6EHPBxABWYtSUrrqBpGurKcH6yYta2PH78iLZtWS0XNG2NVjUhKFTMpChRdF3agpU2IiK6TzKVc/geYKNiC2TGlPcDR1dFE1ErFIYqG1Q2kCOprJBG37POV3z+4h8Yx4Fnp7+gaTX/4pc/43S1pC+5A+/jfjyG6m5FwTSetFE0VUVdWU5WK9q24fHjM+azlscXF8xnLacnS2ZtS1XZ77hreUp6la3+c0E3xVC2G9kPS533v/1oz/sHxbSS1Ae3kuLVqkngUh2O9ThufPR4+N50ho5/5/eN++m974L9fqGYi+xVgqy0hBjvmUh7fDKnMYnLmUXHijDOcBa8M/RjwMdEKO20MWdIBq+kgUUI2YGkeXpxxtnpKSfzlpwSoxvJWZFQjD7gnDTVpOi5XYf9ymji6hVnUmhmQ8zc7kSxeqcDUVmcAZ/zXhFCJvdAjhqL0Ig21oCWUraIhOeSUj+oAfD2mO4dwH3zzbR/nWEfNE1Zmir0RK5xvIzPRyGFN9ZSGchKib6a1rT1sdbVYQaW2Fg6LNeQcqRplpaYrToC2sMR78fRNGm8h0en0yict0BSIyp72qaWomndo7Xh/GzJOFiur16Sco/3O/rhlmHYkBnpukgYFDFkfvvbnl0H17dKusNCj7WJqoLRRwYfcCngkmf0HejAYm73x35+vmCxmPHRR89ZLpc8ffqs3LCqcBVE/CixcFNqonVlmTTElNIkZQRwpwTY/U8LJrbEnBjTAAhRdD3Le4UQrQwVFSrrIk0OOSkGJ1SL//DVX/Hq+ltWf3zGqr3gz/7kD3l0fsbNes1217Hd9hLvzyLjAlBVdZGfqagqQ11blvMZ81nLxfkp81nLo4uzwnS3Kh2MB1BJadrWd+OWd5ds73NGIIawz/5TtpZ0OPgEE1BS7qKyG6vMvktuf/coRc4RdQTYd2Ozd6aLw0+Y3jl6fsfurErz8cN3Jp2EgBqF2Cl", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-107", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "GJ1SL//DVX/Hq+ltWf3zGqr3gz/7kD3l0fsbNes1217Hd9hLvzyLjAlBVdZGfqagqQ11blvMZ81nLxfkp81nLo4uzwnS3Kh2MB1BJadrWd+OWd5ds73NGIIawz/5TtpZ0OPgEE1BS7qKyG6vMvktuf/coRc4RdQTYd2Ozd6aLw0+Y3jl6fsfurErz8cN3Jp2EgBqF2Cl4f5TXeTf78NGKdZW5XFgMLTEuGQbL6Bp8NxJ8IMYg5y5FiBpSwFipZEkaVF3x4ZPHPH/2lGePLiBnvHOFtEnjfGAYR77++guG3Qa8E9FWbYQIPSVCnsaSMI9d3/ZoBTGM2LphttJkDUlJI0RKSbjDg6YGGq1pq4qkLVEZERjIMqH+UOTyBxJp07I871/m/aP8Te8vaAkxHAHedKnyFBNiorTLhwRO+UtTNVRWtIuUEkLgCWy1lmXg9Gl9RGAsYHGII7559Me3i5q2meW4E/cbMACPz57JtowmZpH+DjHg/UjfjXgf+HwcGfqRL798wfa25/WrW7wf8GGkqjzWJL6qJD57cxvxUTF4udHq2jAr3utmvWUcRrz/P1EKVquG01WNeXZKZUT99fR0SdPWNMu5FG67HqmL1eRshCdXV3KldInLTRMZpbPorvv/XjZbSrNGHjUkkXy3jEAoPK5RqimSRgcBuxAP8+p2tya6wG+//lvOlx/wsyd/yk8++oD/8l//K/puYLfrianQPk6ga6Ve29gpRCI6fnVlmbUSy22aCq2lRAz1xoRSxqvRqnAuyHhIWTrX4Eig8T3Oz+SB53z0unioxyNzussOq8N411kpE0D07s6gzuTvjvuje+84H3O4xIfl354Ufdrkd6oXDochR6xk7GQ5nvchAfrFR0/ZnLQw3LDedVzeriQnozSvX71kc7vmN//+N2y3W3b9uD87kpmX8kGlDb/97DMur6/45uwUa0SR29YiFjlVVfVdJ40MKUhzzHEFVJSwaaMNrTUSqsqZ7eCog4L", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-108", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "oxyNzussOq8N411kpE0D07s6gzuTvjvuje+84H3O4xIfl354Ufdrkd6oXDochR6xk7GQ5nvchAfrFR0/ZnLQw3LDedVzeriQnozSvX71kc7vmN//+N2y3W3b9uD87kpmX8kGlDb/97DMur6/45uwUa0SR29YiFjlVVfVdJ40MKUhzzHEFVJSwaaMNrTUSqsqZ7eCog4LKkXWhulSZptIsW8tyVrGoLW0lVAQpZxHLLd1we4azt9gPlowJyObDbKymC7EfKuUGVkej7PjvRzGk6QKX352mrAiqtOMeYl1KgltFndQcFZq/6ZmUC1IqeeXP36UKSUzecNlDPnT/3McW7QlKIxr3fmDXb/FeQgjb9S19P3B1dUXXDbx6ec1uO3D5ektMnpQCSkXUUT1xKOcwaSvKrlVL2zS0VcUwdPTDlm13w2LW8vHz58zqiuWsprGG2mqWqzm2sgStiDkz9J6UNDlOLGMWrWqJ16p47DrJDTRdmGJTV9R9rWoyOkKMihw1JINWAYX8XqUyPntUMqikpSc+IHkBoxhiT/SBF1dfEWPmVx/9GbNmxWo+Y+hH+q4npFiWenLNjCnlZHvmb7Un8DElsQsyucc73bJvgJ6SJf2kNnsYE6mssNQdD/Odz4kVee5j71nAfL/2OgLz4qXmw12w94jLb5smheP0sH7LpDBVpUxcDXcrJfKePGZfDne8jXJvTmdiAtw7Y+Y96rmfXpywqBXj+pyzRctq1lI3M2zV8Go14+b6ms//6R+J3tGPjn0ZIpSQmDgML169Yn17y+16TV03LGdLJatoFgAAIABJREFU2rYRQdwY8MEz9L14wEW2J5EPUdEMKisqram0lrKzmNiNgagMrQ8lNiaNEo3Vog7cWFprqEtCjYxUqEyyQu9QzfFOxAPpCO6mC3AwGR6HWVOAOmbh2pxiJ3k/wKYs6DToDjcH8hsFbKeKhCn4e8eOpt8yRMmHbWaE83IaNAnxjlRi/1mVKcTY9wP", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-109", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "LJatoFgAAIABJREFU2rYRQdwY8MEz9L14wEW2J5EPUdEMKisqram0lrKzmNiNgagMrQ8lNiaNEo3Vog7cWFprqEtCjYxUqEyyQu9QzfFOxAPpCO6mC3AwGR6HWVOAOmbh2pxiJ3k/wKYs6DToDjcH8hsFbKeKhCn4e8eOpt8yRMmHbWaE83IaNAnxjlRi/1mVKcTY9wPdb1/cEGOg67aMrqfrNyLr4ga2mx7nHNtui/eO3a7Hu4Q1Bq0gKU3GAwldaPm0jpKU0InZzHCyaHl6dspyNuPlqwEXMienc05OTvj5xx8IebsR+kRyEqaz4HH761KkpG2FUhVa2b0nceBDzYeJcn9G78b07mvRg1KG1aJF5YQmYhqNrqRFGxTRRGKKOG9RvqJ2C1ISEUmnEl7D17yg2zmeXfw9T06f8ovnvyrVH0LvGYu89lR1MYFkTKIQsA8Z5FSqb4p67NHNkKfxsH995BWWO/xQ8ys3+vt4u//u3/1XsPeg1T7jrTjc+ftrcPzwxpAUv+T3Owg/dFzH5V/fV9/7e2uQ85Evq9QBeN/Y7n2srhv0ieaTX/wS7wPD6IRe1Fo2m+fcrG/48rNP+eKrb1jvesbR0w/dXnBz7FtJVl++3jtRRhsaK9uwlSEnqSp4fXnJOI7icGmJkecCKZXSqCzcNNZkxkL7iBFeaR8zlZFeg9PH5zQ6c7poaSvLoq2xWu2vpVHQGIXNGj9Jbb3F7q21XVYWxaZwwd3X8sGSHEsHWjupAzz2NuTx0BbJ0ePxIJ/CE0fHcVwVoaYJ5vjHHh1Z+dvkWE832H1VIwDGMeC94/ZWQLfvdwyDLGN2u0HUa4eBEHzJrpei/jz9Hlkq3OkSKksYoxVVZWibmnnbUBlFitBUhlltWcxbIfbRptQ0R0J0xJj2HrMxJf6tJrAoKwCFrOOPVG2nZ3fP+/uZOEpqP8FoonBv6JIUQiF+vbR766RQ0aKChiSxs6wzwzDQVTu6YYNfntI0tdD", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-110", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "kWE832H1VIwDGMeC94/ZWQLfvdwyDLGN2u0HUa4eBEHzJrpei/jz9Hlkq3OkSKksYoxVVZWibmnnbUBlFitBUhlltWcxbIfbRptQ0R0J0xJj2HrMxJf6tJrAoKwCFrOOPVG2nZ3fP+/uZOEpqP8FoonBv6JIUQiF+vbR766RQ0aKChiSxs6wzwzDQVTu6YYNfntI0tdDxVZHgDCEI6KacpeWzxGZDTKgoxEo5xtKffxTGmgAkT07E3cXx3ftkGqSldnUC43uC7gdPn36nkqBAF3fH8xvAe38se2f7vpbid/1sVpo0ebrvAbjAvn54sVgSQ2TWBinXM4bKKIyC5WLBbNbuJ4NYZJSUFvWZFA80BDEG6YbVFm0OlR05Z8ZxvNthdxTrUQqMUnsZoOIVyXcpISYl9+6s1sytYjlraArPi5RcHybOkiogTY7jW0y9b1/5gz3Ygz3Yg93f7h+UebAHe7AHe7D3tgfQfbAHe7AH+xHtAXQf7MEe7MF+RHsA3Qd7sAd7sB/RHkD3wR7swR7sR7QH0H2wB3uwB/sR7QF0H+zBHuzBfkR7AN0He7AHe7Af0R5A98Ee7MEe7Ee0B9B9sAd7sAf7Ee0BdB/swR7swX5EewDdB3uwB3uwH9EeQPfBHuzBHuxHtAfQfbAHe7AH+xHtAXQf7MEe7MF+RHsA3Qd7sAd7sB/RHkD3wR7swR7sR7QH0H2wB3uwB/sR7a0aaWPHpMNZ5LsVRYyz/CuaQm8K6SkRgEscJJxF9nzS+qV8L7NZbxj6gbquSTHy1Rdf8urlS/7yL/+Si0eP+OnPPuHZh8959Pgxp+dnNE2DMXeE2u5sT4QJj4/p8FnRzEr0nWMYPK9evWIcHX/xF79+Z/Grv/pf/udslGZR1+iqwtQ1um5QtuJ3L1+y7ToqH6mMZXV6RhpHxvUNbrPB7XYkC4HMl9sdN73j7768ZHk2449+/YzPwpq/Hy7JQeSh7axFKU0YMyopasxe8+tm69j", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-111", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Si0eP+OnPPuHZh8959Pgxp+dnNE2DMXeE2u5sT4QJj4/p8FnRzEr0nWMYPK9evWIcHX/xF79+Z/Grv/pf/udslGZR1+iqwtQ1um5QtuJ3L1+y7ToqH6mMZXV6RhpHxvUNbrPB7XYkC4HMl9sdN73j7768ZHk2449+/YzPwpq/Hy7JQeSh7axFKU0YMyopasxe8+tm69j2XmZNBUkFlFE0jUVrkbDXWlR4b9eRccy0c9EsSwG6PvLq1Vg0wBLNDGyVGftMDJnf/u3uXoJg/83/8D/mGAeG8RW2/YBm+YcobQCNSqLCamtLM2t4+uFznBvY3Fzz5e9+y1ef/Y7zxx8wXyxZzWHWGj56NqPvO7788mu++eYbvvj8C57/9Jc8evohf/yf/hsunjzhZ89OGfsNn/7jX6G0xdo5Z6cLlvOW17eXbLodf/f3nzL2ARMbvBvod2turq5Y39zwp3/yxzx//pxf/frPWJ2ccHKyIseA391ijKaqLMYalNH0LhET/E//7X//zuflf//f/teDZOCxKvCRTRpexhhR9uVYCfgwbuU+DKI4yyS6KSra1pi9omUmUlUVZ2enZAIZh9E1RldotSBnwzBsCdEz+p4QPG7sMVqJ2CKAUlgratJZiajn4CMpyW+wk9py0Xn7t//1f3efsZIn3bjjc3Ks1XZfAc7v+/wdDcKjz+VJWjgp5FSKcG5wjhgCw67HOU/X9QyjF03EcSDFSNsaZrOGT/7w59SzGjurOcaX432pt4jPvRV0J0E9pXLRMBcxQKVAay3vf+fHHglPKrUX2ZtUTicZ9/2JKOJ/Q98TQyB4T06JqqpQShFCIMVJVTiRUkJrI8Ce39znQV5aHtTRPg5ChaIiq7BV/Ybc5Q+btrUMOFuhjEVpg0KURRtTEesGYxJWG6wxBH33BiKD0orWWmY2UluDztD3AbJioWZswsDgA5VOIpMdyoSlNSklQsg4lxjHhJlELZuD9qRcO1HIJWdCzMSU8J4", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-112", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "UTicZ9/2JKOJ/Q98TQyB4T06JqqpQShFCIMVJVTiRUkJrI8Ce39znQV5aHtTRPg5ChaIiq7BV/Ybc5Q+btrUMOFuhjEVpg0KURRtTEesGYxJWG6wxBH33BiKD0orWWmY2UluDztD3AbJioWZswsDgA5VOIpMdyoSlNSklQsg4lxjHhJlELZuD9qRcO1HIJWdCzMSU8J4CyGY/cYq2nohFpqgw5r2UtRmHkZQ9MSTwHjV2UFSIVRRl3ZoabTPkSE6BGBwxeHmMIlEPBlUk/+S3BmK5/kVHvMihK2KIxBDJKaFVFodAtH5JRzf2XgSyiBzK+4mU5XPWGmxl9wKiWusyvhWqPNd6Oqp3t/1YzPxe0D3cY4pcpOB1kQxVRXVVPpOZ9D33uJEzOStSVkVFWyTZtT7cV7FIg6ccqa1Mut8rJ1+EZCcR2JRkf8fgJeNlEj41Im1+z3Pyffb7APL7PvPm6/fReMwpk0IihUz0gRQSYRwJITB2Pd4HXO8ILhJ9JIeMShliLiAtkuuT8vl97a2gG6NozJZDLY9y0Yw5yKfvLwjlhlUyyO+oYuYDSO5nhiwnIIbAq5cvGYYB1w+EEDg5OcFWlu12u1f19N7v5bD1ETJMg1Ce57sAXLzfmDIxQcjgoiKhWZ1dHP2ud7PZyYWo3U6eiVIyCWTF48UJcbYkVBoFVDEyxoBTZq/Kq1BY4OlixqKyvJjvcAm+/WKDXjX89PQDfrP7hm83HbOZxhrLwlTURqNnFucD221gfeO52QxYnbFWcfqkwhqFtRBCZugjPiS8T3gfiSmxGxLGwKPzFmPAakPMiUTEO02McHICVXX/oXT56jVKJ6zxpOGWuP0cpQxKGVIUgfj5YsnKn5DDE5LrcP2aobuh362ZLxdUVqGWK5Q2pJRw3rPZbBmGkSg62mhjMMailWLY9rihJ/mA0YbKglKRlAIhBnyIxJBIMaFQpJhwQ0/wIyk4gveEGJivWlanc2pbkzzgK4zWGKM", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-113", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "xGxLGwKPzFmPAakPMiUTEO02McHICVXX/oXT56jVKJ6zxpOGWuP0cpQxKGVIUgfj5YsnKn5DDE5LrcP2aobuh362ZLxdUVqGWK5Q2pJRw3rPZbBmGkSg62mhjMMailWLY9rihJ/mA0YbKglKRlAIhBnyIxJBIMaFQpJhwQ0/wIyk4gveEGJivWlanc2pbkzzgK4zWGKMxtkIbQ6MU6Z6nJe9FhcsEcASw04TgvSfnjDHiWRqtscZQGUsugOmdnL+6LR6tFschZlBZkbPBKlnZVJXCWk3OgRASo4OsA0pFLk6UTC6DKDTv72ilBMCjKCWXd1FJYa2MWaPlnEdAa5mgfHDkHL/zu9/FJkn7t0nC55y/A7b/HFNKCd7EyLgdcTtPt+kIztFvO2IIuGEgxUyMiRQVMSqspkzokawjYfCiarxgD7z3Oc4fkGA/BsqDCvX0qFQWD6CcHFWk0FMsAz8EUkqydNKKqm72MuSTGvkkQX672bDbbtEoQoy0bUvKmb7vcd6J3HgIGGPIVfW9y5Ljg917timXk1K8vAyoSMpeji8nYPnOJ0wmmmm2l42maUFoNEwq5+TiX4m3hFIkJVLN0xE3WvPByYytj1w5jw0VswBLDCtds+sig0qYVkEFsUkFJGWJCQlba6pKPFSlIKZEjJkQwPvpn0w6WosnZEymruHszOC9wjmNK8Csjha397Gu69AqY60nKggKULIKiFkmbx8cKUV2mzVu7Ol3G9zQEfwAKaBUpqoslbWkGAk+MI5OZNUzyMnVZRmu6HYCusEFKmOpyBAiAUd2gewjwXmCj1QmknPajxsBjcAwjvjg8MFjlCGntPeKpxs/FwX1+56XmBL7hV0BGaYQglJ3nIWUM5q7ntskFJ/IKAW2rsk5Ecoktg87qIyxCqNlrOWsCCHLhOsSujLoyoAxKGMxRhOT2ofjcjxIwE/7DyGhtAQQtLHUdYvWFqXr/UrBe1mFvo99H0gde68TKL/rtt4WYjh", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-114", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "2gewjwXmCj1QmknPajxsBjcAwjvjg8MFjlCGntPeKpxs/FwX1+56XmBL7hV0BGaYQglJ3nIWUM5q7ntskFJ/IKAW2rsk5Ecoktg87qIyxCqNlrOWsCCHLhOsSujLoyoAxKGMxRhOT2ofjcjxIwE/7DyGhtAQQtLHUdYvWFqXr/UrBe1mFvo99H0gde68TKL/rtt4WYjh6g5QS3nm6bUe/7hh2A9EF/DjKijokSKBlcSFfSxmVM9ElglH4YcRUhrIg+V5P/W32g6C7f5YP/5SClCYAnk5QRmmFSuCcx7me3WaLd462banqipMLi9bm4AEfxYYvLy9ZX1+zWq0wWrNYLNj1Hev1mmEQ79d7j1KKqq4B7ni7bx73BLoxHfTrp+iLwhNTTz90Jab29J1OlhyyLjePFjDVuoQrMskW72HwslwrS0ZlDNkoshbvXyu521qt+fnjFZfdwM3LniYGVkPiETWDnXG5XtOFiDmF1CYWbU1MiawikDAq084UdasxFpRKBJ8IXuGdwjlwHpyXm2s+B2MEoJtasVoahgG6XeZmPdD3CTB8X7z8h2yzvkWpjK0g5gGXN2Q0GUUq17oyFUO/Y/38GW4c2Kyv6He3+LGDHDAa2qamritC6HHOyaTrAhPgKm3Q2gBwu77FDVtc72htRQ1k7/AukwZHHj2+H/E+UrctMcUyAWuMsYzOs9v19GPHzLdoNLp4owK2kRRNgb77x1xCKPHaKVRhzCFam+V6QYnw5CS/73tBN6K0pm5bUoqELpBzPIROVMZWGmMkTpkyjKOscsYxUZmKSlmUqdG2wlQGkxUpl1BLlHOSlZJQBEkmDOS0N23FcragbhbUzbJck4G+G8g53Pu8HNvbwPddv/99HvHvCzvEmHCj4/Z6zfrlDXEM5HKdFGDQTEihsyorFXFwZLKLDLuugG4Bw3vaD4OuArL+TmxHYkMSnyyhNnnfgBv/f+berMmuJLvO/Hw8wx1iApBZlTWwKM5mMsl", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-115", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "KSlmUqdG2wlQGkxUpl1BLlHOSlZJQBEkmDOS0N23FcragbhbUzbJck4G+G8g53Pu8HNvbwPddv/99HvHvCzvEmHCj4/Z6zfrlDXEM5HKdFGDQTEihsyorFXFwZLKLDLuugG4Bw3vaD4OuArL+TmxHYkMSnyyhNnnfgBv/f+berMmuJLvO/Hw8wx1iApBZlTWwKM5mMslIytT91/up+6X7gWo1TVKRYhWVlQNyABDTHc7gcz/4uTeQZDEzI8Wi6GZAIABExL3nHN++91prrz2z2z3yd7/8W+5vb9leXbK9uODf/+Wf0zQtUi0buxSMsbRtQklJLoXH+weMtVxfX2O0wRizYG015T9hcqdA+g+xsnN2UqhZ1TCRcyKnRBGZQuSrr1/z+HDPw+MjIQT+3Z/90fe+YELV7FZYQ1my1xM4n30gp4RcbmJRklwyqUQyhbwA3KJQM14psEJz2ff84Y80owsc54EXxrJpLceUuHeO6BNzgdHOIAumhfWFQjUG20mkFvj4VLrGADHV7EkIaLtaukpVSKVw9xCwVnGxqaWRsmXBcgWzK/jwA3CyUgvTmAp5yea0sUhlSEJThEAhUbpZglkkxXq9SiloY2mbFqU1UskKCaVMjHkhcJ4OOiElAsHkHONx4M27dxyHI87NdH1L01geHx7ZHQbu7+5IGcyLFSnnWokVQcqCaQ4chpH7xzuEAlYKLQS2lDMmKhTIJft/Lv6fYqqPea4VoRTy/FHIgkSgtVnIsYKQsuLtCxNaUqkZsDIopVDKIkXBGBBEBBEpDSCJQZCXSl8I0Fosm7vQmIzVkTB9RpozOXlyzjRWoBD4qGvWW0rFr4UkpsiJmim5cgRwOrTFGfP+IQf0Pxdk8FurXfhGED7HhiU+VYLUYIxCa0EOhSILaqmeRA0cC08gEKq+ZykEEYWy6swZPNEMz3s/3xp0yyl/Xl71KdA+KRhEBduXNyWXNxZC4Hg88puPP+b1p5/x4oNXvHz1ij/", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-116", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "I0Fosm7vQmIzVkTB9RpozOXlyzjRWoBD4qGvWW0rFr4UkpsiJmim5cgRwOrTFGfP+IQf0Pxdk8FurXfhGED7HhiU+VYLUYIxCa0EOhSILaqmeRA0cC08gEKq+ZykEEYWy6swZPNEMz3s/3xp0yyl/Xl71KdA+KRhEBduXNyWXNxZC4Hg88puPP+b1p5/x4oNXvHz1ij/7t/8WY+zTdy+gtcZai1Q1gO4PB7q25cWLFyitMGeCo+JiaSHThJDnC3rGcN/bLFBJinmeiCHg3EQpgSIcX3/5G968/Yrb2zucc8+6YEuqitCavGBcC5Vcs4YYkSmfCNJzKVbIlIW5Ot2oE7SijKFdr/ji/pG7/Z7L9oK26/jcjxRdeHM34nJhdh7dSmwj6IREW03RFTucjoWcK+GWUg1+J17EWoHSipgiKcHjLtF1hVWvKAKkykhV76/3v53w+e51Im6WZ0GB0QrdNATRkJGoAkqbhYhIlFRJsMqI1+dAa42UomJq+RRwT8RSDbxyeQCdCxzHiduHR8ZpJAbH1fUl682aw+HIfrdnv9tThOTqKpFyIS/EUy4S5wLjOLPf7zFW0+kNjdLoE1F7JrQkYtHePGedSLIiCqVIimSBpaiYqqxKgIKklLxk4FUtgaxZ66m8l0ojhEFI0AookVziGWqp16u+RKWWYCEqwdboTCMjfn5LLjNSKorQGL2FLElKklJGlAopIEDkJaHhlImfsup6OCslzofR/6r124LddxFwUkq00cshtsQvSYVmAFJZkrNUr7UQVcWiJBKFMGq5Jsvm+gFb5duJtHP2VEsYgeKkUVqQJmQRZ1zjJA9bbVZ8oD7kj/70j7m6vuL65prtxQVd36K1guVrAaRWWNHw4sULcsocdgdizIzjTKHQdz0lZaZhJPiA1gYpFI21aL0CxCnmLVf9dLErfmm05LA/8utf/y37wy2P+zc8PDwwDiNCaJ67kTQ1U1ELq63qKUNJqWZusaovCoKYBZk", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-117", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "b5duJtHP2VEsYgeKkUVqQJmQRZ1zjJA9bbVZ8oD7kj/70j7m6vuL65prtxQVd36K1guVrAaRWWNHw4sULcsocdgdizIzjTKHQdz0lZaZhJPiA1gYpFI21aL0CxCnmLVf9dLErfmm05LA/8utf/y37wy2P+zc8PDwwDiNCaJ67kTQ1U1ELq63qKUNJqWZusaovCoKYBZkEp/+LXIrtUlUPon4mhcQozYvNCqvhbvDcPx74g+2Gn1yu+Nv2jikmnIeQMjkUnIv4EMljzUKGIdSAVyQ1NzN1A0nJMHpSTrgpLKRi3dDjmJCy1F8KtFHMzp+zmucsqU8Ytzo/D5vLKzaXL4hySxEWkTyNUWjbkFLEGIvSBiF0LXt1Q9N2WCOY86FuDFnJOAApF5WBUiAFcwgM88zdMHBMnklDXLfE3DApCFZBoxHIimkmScwQsyRmxeQyRXru7u9BwNX6R2j0mawVMmOMxTQdIf1ApnxR3QAVb+UpO5NS0hi9/KxvqlwENRikkglLxr/bHQDxBPWd9hCn+/5U7ebE+SAjDbj5SAhHcp5Js8QFwbvjgBANjV7RWIU1CttUmZxzilxA6HoQpFRVJjnMSCJGFTbrhhT/FUr9a27zW5eUkqaxNJ2l6RticMRSkHqpaDRIodBSY43AaoFSGiEkc4wUpbGtRVvzVN4/c32nZEwsOE9VJUgogiKW868UEBW3q++zvlNtDb1ccfPyBVJJLi8vWa1WSxYjT7ICKskkkEj6fsVqtUYIScoF5zxqkfLkXPDeE0JESo9tGnLO9Ov+fJUr9HLmEoGTDjDhw8zj7o67+ze8u33NPDliSKzXl2htnnXBJO9l+UJQFly51BqsZnBLZlYWsgJZsyUpBKeT6YlUEWc8rbMGJTt2Q2XWL1TD2gou1gbpBS6kKoWKGR8zPmRIVLY1ZlKqAVgIULogS/3+IRRCzMxzJqeKvVfMt6B1wZgTXCMWKdLzyRFjVcUE0fU5KAWjLU3bYfS", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-118", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "mEoGTDjDhw8zj7o67+ze8u33NPDliSKzXl2htnnXBJO9l+UJQFly51BqsZnBLZlYWsgJZsyUpBKeT6YlUEWc8rbMGJTt2Q2XWL1TD2gou1gbpBS6kKoWKGR8zPmRIVLY1ZlKqAVgIULogS/3+IRRCzMxzJqeKvVfMt6B1wZgTXCMWKdLzyRFjVcUE0fU5KAWjLU3bYfSWIltU9lgF2mhiMChjaganNEJU+Z1SeiEFl+sjn6RTlcRU5wQjlUyiEJd7MovMLApOFJKRFKvRnUWgkFojVKjVhhAUoUipEEJimmamea7Zoq53+UQSK6XRSi/ysmdfFmCBed7TiJZSzoG9ksxVsVCfo3Im8fKyV8IimfPe1SdQquW5EWeCTspKVotF8icldf9YQ3SOFKGgKEUTA3gH+0NAKwl9QimJtZzlckrpCoOZk/IGKJlSEoKCkoLGarL6X5fp/sP1/u0pS+B9qtPrQyMEFCnRWlWIQUnKUuVVuR5oJbBG0BhJYyRySc4ihayqgkY8lfvv6T2+3/oOTPeprDiVebIIRHmK8Pl9okrW0kNKhW0kP/7pT7h++ZK+7zHGoLVZ0v0n2KLuJ8H19TUUUNri5pnb23u6vmWzXTMNIyUndrsBHyK2bbi8uuA//Me/wJimZlcnxcCi03F+ZnITr7/4jN3+AUSmbVsuL19gX3Ro3fLTn/yCrl99z0tVV05x2Yyq4johQPSI4CnBU0IgL6Sd9lUOF4EsRa0Lc319ZVE/1HOsQA6YUrDCcN1aRMkMe88sMqutBgWzj0wuc9wnpjnjA2zWhaYB02hSKkxTQoiEMgHvF2WCr9IopSrDXSjElLm7daxWgs3mSQuaQtXzPnd99IsfEULisHfMs2McJo7HkSJ3NJsN2jZstxv6RnF1pZgaQ06RYYoMTiBNT0yKGDKSQoxp0ecuD4kEoSom6sKEjGDWhrXe8tPmDxCdwmwsTdeRm4bV1tLFiLnqKbHQlJbjTnGcRmICkyp", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-119", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "WhaYB02hSKkxTQoiEMgHvF2WCr9IopSrDXSjElLm7daxWgs3mSQuaQtXzPnd99IsfEULisHfMs2McJo7HkSJ3NJsN2jZstxv6RnF1pZgaQ06RYYoMTiBNT0yKGDKSQoxp0ecuD4kEoSom6sKEjGDWhrXe8tPmDxCdwmwsTdeRm4bV1tLFiLnqKbHQlJbjTnGcRmICkyp+V1JkHmfmYSIET9SWjMIaS9c32LZHmYZQfIWRnrHOpfcp282lQg3iCXrIS4C0CzmcT+950QunnIg5EmLE+4SQGmPbSoLFTLNe0bYN24sLrLHoxtZgLiRJFAIQ3UtyiFil6iuaRtLk6MpbSg7E4nEZisv4nBfooAZfoxuM0fStxdqGpmnJMZBTxKzaHwhF/W7XGXsX7312wklyVbg0SrJuG9R6S/QR5xxCZNo2Y4yg62vw1VKSvSAHQXSZLAUl12ryh/bzfjum+48uaK5C7KWUWeQLT/83L+F0EWAb01AQWGtR6r1Svrz3cflWxlps0wCVYQxhQirBKvWEGJBeEHwghEhMCdtYUkoolc9i/jPwXzLOOeZpZpomYohY25DzilIKXbulsT1XVzd03Slb/n7rROQgJSKlWselDCmdpSWnbPt0oBS14MCVPXtfc8dJUnHKcHLONEqyMRoXApKMKmK5TFW3mWLFUyRULFaf5Eb1t/qt8/JaK3RQ2epvEo7O1QcshpoFSAnW1BP/uWt90eNdZJ4jIUhKAe89chwRZkIIg1Ed1iqa1lJKol31dKs1/dohhF6y8ZkUCm72BB+qbkApmsailAQKsxspIhNzoIiC7Rpkq9CdQTUGYTXaSiiGtYQSC8wCUzLtcUOWmqIt2TnkgrWfdASlVGw0F1Ezw0WBgXov4/ueS0pZG3rEQqItyYlaOApOFeOiQCh5ISTLE1YoRd34KEXKVR9uTNXwZpWxbUvT9+hmhTCGrAwgiAViLricmZPEJ4NF10PXGGR2NL0j+gk/J6LLTO5", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-120", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "ilAQKsxspIhNzoIiC7Rpkq9CdQTUGYTXaSiiGtYQSC8wCUzLtcUOWmqIt2TnkgrWfdASlVGw0F1Ezw0WBgXov4/ueS0pZG3rEQqItyYlaOApOFeOiQCh5ISTLE1YoRd34KEXKVR9uTNXwZpWxbUvT9+hmhTCGrAwgiAViLricmZPEJ4NF10PXGGR2NL0j+gk/J6LLTO50MCuMrjCD14k2gVEaJTNF50U/DEqZH1Jd/27X+QWdiJ5TjCqQS5WzRg8loWRBKwFK4qn7VmmBNqAM54aTs7RuSTxLypSUl+TpvcP1e67vlem+zxKKZfOemgNqoKvBorJ/VTaGZGGSJDnXIJxL7VAT72ngTlpF03dY78iiMPuZ3cOOmC5Zb1aUMdcNLCuWO7oZ7zxu9iAUVur6oFIVCykG7u7uGYYjx8OAEIIPX31ECBnvI9vNJV3Xs173KK2edcFM3yGUBm0QKaNCJDtHdg6RMrKApmaxqXbBklQiRwNJn/dZiVW+hJanWpkpBkY/sdWaG23IYYeMGTlX0ikETwyJ7AKtFahOYjqJUOCjJ+SMdxmlBVJLdH2ZNFaSE+wePSkXlJGkCOOUqR1Ggq6VtK1gs9I/aCP9+GcvGI4z0+iIMSCA4fDIYb/jwhVW6wvshw2rbs3lVYtfG1RjKbpDtpfMg2O/O4J7QIlIzgfGaaSUTL/quLi6oOsspUS+/PJzhBLEMFNERmowQmClomk0TadpW4PSks5cEWNh3AVWU0S+eoUfPW7wHN++IYwjsmtBmyq7IxOcpwhJkoZOJYzKSNMinxl0tVKUBUbQWmOMobEN2piqC81x6SAs585L78NC2smFmBas2hYQrIpEaI1pugpXScn66gV2tWbMhjnDYUr4kBmmSIzpLBkMMWPEiJHwwc0GY1pefWh4eNzx7tGz2x3Z7Q9QMlLA5rJHa43Qmm2n+fmLnst1x/U2kpdKd3u5wZjvDCH/4uv9WnrRMEDO5BiIbsYPR3IYkWJ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-121", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "0tVKUBUbQWmOMobEN2piqC81x6SAs585L78NC2smFmBas2hYQrIpEaI1pugpXScn66gV2tWbMhjnDYUr4kBmmSIzpLBkMMWPEiJHwwc0GY1pefWh4eNzx7tGz2x3Z7Q9QMlLA5rJHa43Qmm2n+fmLnst1x/U2kpdKd3u5wZjvDCH/4uv9WnrRMEDO5BiIbsYPR3IYkWJGyoiQiYxDCipeawpSL+3BueBdJM4FUk0ow+TQ1pBCRiqJ0P+MQff9/ugzC1gWWcXp7/kmlpIXAa8sS29zOX1N/fMpwxcL7ifKkvAqWd/AgiueGiukkHjvqwoh1u/V9F0F93N67zXWfzsRDNY05C5zff0CKCglz4Lx7eaCrusxtmI5z1lC1y6EHCPi1IEQl4y3pirvZbiKkjIiZKIPTJNnpTVGqkWiUruKziL5pQzNqsquGiVosyC72lU2HiN+6YEXUqB0wbtMphBjISVIJ9UCtUWzCFmxuJxJgXoASmqmlDIpSrwTNLa+Zm1+WBuw1KCNpO0t3kfatn4MMaFVwJiA0Q6tDFJGYnQcDo8Mx/3SJFHL5SF6hIjkOJByXO6TpWlbjNZArsFASbJzpFRIk6OhsFppVsLTmUKUmiIVut/UTqvhSBKSRAFlMVbRX91Q1hu6jaFZdShbsdvDuKO4gTLuWLsNTdex3b54T3nzvHXSkuaSq0pAPOH+p5VTXghMsUiRyjlrE7I+cwqNXNQ+xhisbbBdizSG6AW+wJQyPidcrtd+nhPee2KIhOIwqpC3miQK03HP/nDg8Tgyx0xWDSlWGKXBkKVB6Y4xZb6+O+JmT/Yztmkw1lLK5kxy/mtYAs4yrtPnNeBkSook70hLN2IpESELLL+kPrV6V7liCLXFPIVEnAvZA3lRtaREjjURQoB4Js7wHW3A32zxex+PLYue7QTkn7DfUuTCz9eTglwz29OXFqrhiyhL4bZ8ndICbRVSVUlNCJVpV0pxOOzZH3YcjzMg+ZM/+1O", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-122", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "BkKVB6Y4xZb6+O+JmT/Yztmkw1lLK5kxy/mtYAs4yrtPnNeBkSook70hLN2IpESELLL+kPrV6V7liCLXFPIVEnAvZA3lRtaREjjURQoB4Js7wHW3A32zxex+PLYue7QTkn7DfUuTCz9eTglwz29OXFqrhiyhL4bZ8ndICbRVSVUlNCJVpV0pxOOzZH3YcjzMg+ZM/+1OapjLgVbf7RFAU6gGx2W5Z5TUffPABMUamaSbF2qm12Wxo25aUR/Jz2xhNNZ3J3iPmGTVNVGFsPMufBJVgS62FKaFdYDpM3N0fENstq1ZhZfUYSKeTSQBkZMnEUm/mSktKlqRdYpoCd7cTKIFqJciCMvB4H3C+YFtBToUYakCdx0TKkZQXCZEAP5+IvUq+kQvRF5yErq/3RplFkvTcJQvaCrYXK0QRhCkyHAeYAn0XWPeOphkwWgBrpmnHl198wv3tjof7HVK0SDRzdlACyY+0XcOLVy9p2oambUimalr7foU2lv3BUaaAuzuw8R3Xa83VVWTTCr7MLWPRmO0rSgR/JLV/AAAgAElEQVT/ruDKxOQDOleN5urHV+hGsWoKK6Mwq44wTNzdfcUYA8cYuNhcslpt+Df/pmG9vnjWJXm/Sqw684TLjhjTmewSYiHL4pKE8H4/f83VhFAgFUpV3Xrbtmw2a7bbLR5FQOBDZiyZQ6r4fSyeOSSOQyRMI9HPiDRgVSJdVS+GL7/6mtvHgdfvHmi6Lc3qgjAeyCkQ7AraDtldcRj3fP2br3ixKuy3ghc3N2y3G65uXlCbaf71rHOEOquZCuRIiZ4w7mvAjTOQQGVQBaHBLFh4LpoUEzlk3Ozxc4CpQITOrJAyU7wnB0/2YWn0el62/yxM94TbZvJZa3rS0LJ8XshPjPOigTsx/nLJahVALsgUKbl21+TkEc7x8uYGkWHYHVFKMgwD81y7k25vb4mp8Id//EdVDmYMWqvlNbCk0JXQs9a+p++sdX7OkFPV3ZWSub2", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-123", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "ZCuRIiZ4w7mvAjTOQQGVQBaHBLFh4LpoUEzlk3Ozxc4CpQITOrJAyU7wnB0/2YWn0el62/yxM94TbZvJZa3rS0LJ8XshPjPOigTsx/nLJahVALsgUKbl21+TkEc7x8uYGkWHYHVFKMgwD81y7k25vb4mp8Id//EdVDmYMWqvlNbCk0JXQs9a+p++sdX7OkFPV3ZWSub27xbmJjz768HtfsFTKAsZHSJGy+FOIk5a4gl2VkdcWbCJby1QyD+PIbnQoqfj51Ya+MfStBgFRFUyUZGHwKTOlgMqFVAQ+1FJRKwG69ti7MTGPMI1pyeAXXWepGW8InB3X0sm4pNT7VpaKQS0pbQwZN9XsqsTasfbctd60eJsIU8bPftF016y/ayXrlaQxAaMDSiRUicgc0MVjiltea5UTSinotlvatkE3DbEUwjjRXBhsY9leXGBsQ3h4QGeNvjDcXEo+uEw0XYORDba5IMgN05uZ8eC5/7tbSizIJJFCIYVk2rn6vF4YRGc5XhWiLAxtNQjSKSKtR8iZnI4V+37GOku7lm2UU6n9pSUANdjmRYJ5CrWnHZepmbFAknJGCUnftrRdy+XlJdZYkJqQClMquFQIpWCNrAe61KgC3iWsshAF8xgge+7v9+QUeNyPDJMnISu/K2qTjMiFaX+PnwylBIgB2xtcnnj7cKSIgg8zFzeXz09a/gXW+S4tlWOYPck7vHPkGEjLfoop42dIURJChU2PgyOliI8zOUZyCjRZYZC0umqZU/C1ynL+3F/wjZ//Hfjc9w667/c3l0x1QpJP/+d0wmRRFjnLwspLqvwGkKVwpplyhpDIvqb7OTuEd7y6vkZkuH17i5RqCboz3nvu7u5wLpBiREiB0eYsQ6ulWVkqMok1DeWsutAYbc8boHY7Re7ubjkcdsCff8/bWWVKImdECsuvVE+VRRWQc65qCq0RxkJOlKZhLoXdPPGwmwmh0CnFzbpj2+r6tbIa0BSpGYJjiIFVrvC", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-124", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "0wmRRFjnLwspLqvwGkKVwpplyhpDIvqb7OTuEd7y6vkZkuH17i5RqCboz3nvu7u5wLpBiREiB0eYsQ6ulWVkqMok1DeWsutAYbc8boHY7Re7ubjkcdsCff8/bWWVKImdECsuvVE+VRRWQc65qCq0RxkJOlKZhLoXdPPGwmwmh0CnFzbpj2+r6tbIa0BSpGYJjiIFVrvCDD4mYElrXE0tIGIbMNGVizqRSyHPdukpByrUF+ESspeUGaVWvT1n+Qi2+GSnVoFuSILon86LnrPW6xZuEHyPjYT77PEigbyWrXmJNxKiAItVfOaIJWOHxKZOyJEuNkJp+u60EbNPi5olxmtCXW3Rj2VxsadqW4UuLKZ5eGa4vJa8uM7nVZNlh+wuc2DD++pH97ZH7X91hjOXi8hJpFVIrpv2A9w4jekSRDLkQZGFoM7pkDBmlPEJJUj4Sf3B8OUEGizpBLPCBFOS86GvPv70feJddlTJFFkzb0K1WXF5e1K66BCEn5pRxOdegq2Vt8zZALhwbgW0MumhymvAucfewJwbP7jgyuExGVbJQ5sUSFI77h9ocYCJaKkxvcccDh8cdEPF+5OX+1Xuv9l/POvFmlFr1BeeJzhFmR04VNphdwvlMioqcJTkpQkjs9o7gPdM8IUVEyczWGIrS0GTIiewTyRuSD6jmeZJT+N7eC+/3N9fIdWKWa5uuqDrDnKtlXwrkGInTTA6Ro1/kMgsUEFJi3u0Y7u/x44HoHXq7QnUdL3/y+xip+eL1l8zzxH73iJSC1rb0fY+SnnEaGYaBeGq5FfEbwvKnJRayb+kuEhVzjd7hnON4OHDY7593xWKkxEiZBkSIC15aM/jaUSSq7aO2CK0RUSFL5mbVIl5d8dgOTHPg03dv+fxO8vXukot1y0evtkQgqIKKteNM6qqftb2i1ZpWGiaXONxFnK9aXdNVv4NT6+w0F4QqrC4kyUMKLPZ+BdOdz4PFQ7S2i1orGQ6Zw2P6zlP6n35", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-125", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "bwvKnJRayb+kuEhVzjd7hnON4OHDY7593xWKkxEiZBkSIC15aM/jaUSSq7aO2CK0RUSFL5mbVIl5d8dgOTHPg03dv+fxO8vXukot1y0evtkQgqIKKteNM6qqftb2i1ZpWGiaXONxFnK9aXdNVv4NT6+w0F4QqrC4kyUMKLPZ+BdOdz4PFQ7S2i1orGQ6Zw2P6zlP6n35UKhu8WbU0P264Xr8ihUKJ8PLDnm6taNpI27R8cH3Nyq5oRMc8T7h5xgVJSIK3+yMxF2zXYoyibSwnH+SLly/pL7ZcXW8xWhN/dEWYNG5wYOHrQVLsNaV9yWFsGXzi3d+95Xg/EB8DamUoG4U2Hf2mx00BfOHw7si4G0jZ0fSK9dWaRs60smCLQJNAHVmOr++9bNOdYYWzPndRKjy5bMFJiCbewyMRS6u9UnSrNU3bcnN9RdM0SKlwITOEzJ1L7GPiMAdCzDBXvDH5wOwC0+Q4zhPZzzzuH/FuhumhErOpkLA0TU/b9rRtR2MVIs4M95/g3cDd8AlKydrQMs/EccQPD9xZzbpfc7G94M//4n/7Yc/M72AVeGLQ8unPilQkxzEzT45xGAlBECJIaYEqOY0xMYzVJGhyBaMlRguSkhQpqvLhpIJwjvlwQBpFc/H9DbPgWZlufUtpad2c5pmcE6u+r36k2AWsnoh+JvqZNDqyD5TRU0LEO0eK1Z3ocHvL45uvmI97gp/pXt7QXV3x4vf+iL5taazFuxnv3NkwpzGWnKrz2DgMhBBQyqB0gsWm4v0llm65emik5Y+lmobE6jIW4zMNO3KuAHrw1Sgj5yXLPSnAJEJVRycWuz6Azmiu1x2mZAYtebPb42Iho3E5c325qg5IslQWldomKhAoLVF5MTRJMI81wz15niotUBpSqnpbpQWmWaoLBMQq7FdGVOxcV94vRzBW0HaK/UNiGvJi3vMDVklIIWgbw6rpeHm5RWFRaPotaJPw4QFrDOu+w6oWVdrampwjk4P", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-126", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "NO3KuAHrw1Sgj5yXLPSnAJEJVRycWuz6Azmiu1x2mZAYtebPb42Iho3E5c325qg5IslQWldomKhAoLVF5MTRJMI81wz15niotUBpSqnpbpQWmWaoLBMQq7FdGVOxcV94vRzBW0HaK/UNiGvJi3vMDVklIIWgbw6rpeHm5RWFRaPotaJPw4QFrDOu+w6oWVdrampwjk4PZZ8qbW+YQEEajtaJbpGJSKjbXV7TrNU1j0RLW6w6vI2RLlrBzApFXCHnBPEvcEDm+OzLej+Q5UUymJIEQGmUsWhoUivHoKDKRdWBz1bK5uUHZKs1SSdRNxvzsTj1tm+paJiIntzIhMoKTFwPvydCent4K3dVMWEpB01i6tmXVt1XrvkjCppQZU2aIhRATISTyXN3ZhsERo8c7hx8HwjwwDEe8m4nDHnJGSYPQCt2o2iygNa3ISJ0hjMTxgbDfI7TEdB05QgmF7AVOS3aPO35wx8jvYL3v7XuGdIVYzPQVPsA0Zw4HT0ySlCVKVXz9FHS9T/hQ/UcEtVrLudp65pN/cCnkEIizI8XINySgPCkn/qn17eqFBa+pOrX6nV5/9hu+/vpLfvPx/8DNM//xL/+Cq+0FP7q8Znzc8eaTT5h2O+b9DpUKshTi7CtTXqrdnXeRYf/A/uEWnwIxR/xXrzHrDVN/SZaS6EYaBS8vNtimRSnLdr0mpcRf//Vf88WXX/Djn37EBx98wE9/9rOKRy4KhvMSy+cl1wtWUn1PQmBMy+//4g8I0T/vztYLQo4esfi1JlG1ArkUkIujmKqEYn3PUIREKM311RU3V7Due4bZ8fHdI188znzy8MgH256PrjdcKklvFdEYSoxLh15kdz8x+2q+3TQSZSVeFLwvqFyVC8EJ0qK77XtBdymYxmqCI219PuahLM75tcNQqerP0LQC52pm/NwlS/VP6C+3fPjy5/zB7/07+qbqoYf5jnHa89/+6/9DTgmF4PLigh/95CVJSqKUfPbpF9z", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-127", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "EKM311RU3V7Due4bZ8fHdI188znzy8MgH256PrjdcKklvFdEYSoxLh15kdz8x+2q+3TQSZSVeFLwvqFyVC8EJ0qK77XtBdymYxmqCI219PuahLM75tcNQqerP0LQC52pm/NwlS/VP6C+3fPjy5/zB7/07+qbqoYf5jnHa89/+6/9DTgmF4PLigh/95CVJSqKUfPbpF9zd71CPOzSFZtWxXq346Ecfnp9BnzIxRPZvvkKSUDGQo2Y3XxG1JdJz0/yYi80LPvv1J9x/+chwNxImX3XMOeOnkXkAqRMpTIjs8McDKc64Q8TdNsj5wMUHHfmjDsUyLeV4j3mmHnO9vai4bUpL15mmsRpjFDFU05mTdroazQi0kIs1auI4TYScaJqWrmvpGkMRkikGhlzYicIsFEkoep1JRfAQZ+bJcf+ww88TbtiT40SJDpE9tlSCM6eANSDkDjE90KkbdP8CSiIlzzzsOe4eeHy8RWpJt+7YbK65uHrJ5eaSvl8huwu8+GGKjt/lypwIdUAKmn5FLgplR5ACH31VHSiY/FxVUKnq36fgqrEWGZ8lJUrmEFFFEPTSoKMEJXjmYcBOK2KISK0WI6/vXt+d6S5SlxA8zk189fUXfPrpx3zyyd/j3cwvfvIheR5pvWe4f+Tu9RdM+z3Tfo8REoUgnRyLtKleryExeY/PmTklfEo8Pu4Qk2P9+edIYxjdhFaKxqhF1VGWwJp53D0ilOD29pa2bflxjJUckYsZRfnGmzhPEciLdOrkWNX1K2xun3VDS8kV1E6JkmLdICe98qk0l/LJ/QfO42oEAqM1WkouVx1GS1bHI9lF9oPjOCt2o6VrLa1dTDkWyUdOdVpESrXZwbQC20mSy8TFBSqfPBTrT108jGvL56mPIy+KBUE1uxaC84EqpKhjmZ4ZXABYsm6jFH3XcX11yXp1TdeuedyD0hIpDSlWCZNtC82qxwtJQuCAMXpccLjgkUETY1iM9BdszgdCjCRmJBm", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-128", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "J/QfO42oEAqM1WkouVx1GS1bHI9lF9oPjOCt2o6VrLa1dTDkWyUdOdVpESrXZwbQC20mSy8TFBSqfPBTrT108jGvL56mPIy+KBUE1uxaC84EqpKhjmZ4ZXABYsm6jFH3XcX11yXp1TdeuedyD0hIpDSlWCZNtC82qxwtJQuCAMXpccLjgkUETY1iM9BdszgdCjCRmJBmLIkbwqaWYDtFegGzJGebdxHB/IDpPSrGW8jkRvSPMEm8ghxlSlUGlMFN8QBI5PijMRrHK62pkLgXTPBOe6R2rtQUKRdVn02iNtQa7+EfURKDesWpTWgm17CM5g4vgY2GYAyiPD5GCwMXM5BOji8xO4AM0JaNKQopSpVCLGiaXVDkVDSrXJpA6cSOiSqwEgfCE2RJmW6WO6TQWSKCURWtFYzrabsVqvaFfb+m7NUhF+udKdP/R9ym//a+/89s8MZcntqnKUTVKV9e7OtGkgr4x1Yo3xeonHXNYSN1q3iSoQwEC1dA8CYmQukpVQyTFWOWjSn4j0/62PfTtmW4q5JIIceL160/55d/8F/6/v/4rfvXrv8HPE1pKehm57DZci54wzoz3e0JMxJhp+g7dGMzlBtVa2utLBAvOGT19cMwPD8yHA//9l3/DfnjDf7t9oGksV5ueVzc3/N5PfoIvR0oRHB7vOewPPDzeM/uZ//ev/hPH3Y4Pbl6w3mxZrTdLgH0a1XIKtiGG+m8lU0rtNJKmRfC8jYSbwDnENJKCx88zYrmZom2qZ6NWCF01lUqAcZ4SEiKBFhkhoWtaGtvwl23LYZ54fX/P28PMf/74NX/y4Q0fXW74cLuhRSGLJMeascpWsLqWXL6QrC8E797COGbGoQbdtqtwg23ALCOcgi/MY2E81qBNFnS94PpaklJiOES8r11YUv8w36gcQu3iMpESAyWFOsnAKqztsNaBsMx+5DdfvOYyZOQHP+XdMHA7TPznv/8VX7x+zf7rN5QQ6ZsWqzWff/z3aKk", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-129", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "4fX/P28PMf/74NX/y4Q0fXW74cLuhRSGLJMeascpWsLqWXL6QrC8E797COGbGoQbdtqtwg23ALCOcgi/MY2E81qBNFnS94PpaklJiOES8r11YUv8w36gcQu3iMpESAyWFOsnAKqztsNaBsMx+5DdfvOYyZOQHP+XdMHA7TPznv/8VX7x+zf7rN5QQ6ZsWqzWff/z3aKkxUqO1RknJ1cpgtGYsa+ZiOeYLri6u+ehPP2I+PLB//RX3n37Kw+tb5tlTikDZhhwi4X7ED4r5QWPSjEietHtHDB6pBXNQfJUmwqpB/aRnfdnTdIa7/WeUaX7WNVF6yQLziQ+RxCwrrKMs2tRD+OSe5kJ1TXs8ena7iXf3A8Pska8fWfWWcQxYU70p3jweeX37iAuZmAvXN1fYpmG77mm7BtFeMo2WYS/pTKFRBfe4Yz4OfFlqs1Gcp2oWHwPzvGP/+CVNs0JJjW57rmzLz37/T+i6nuubG0zb0PQtlFohhem4TLX4Haz3JV/vrZpXPGk9cq6woUQ+9QMsnZ5nCwNAaMn6YkMImaYbCXEmJMfkpjqZJKV6sMe0eBJXF7WAwKRC0oWVtqSs6ESLiALmmTBNhGlCKInQilNrwrfl/9+Z6cYYOBx33N6+5dPPPubtu6/Z7R4hZ6zW3D08EuZEtoXkInPMdSZXKfSyumc1xqCMRRiLFBplde3OSpY4TvhpJmlDkpJxGvFuIk0HRPC0ApRqEEqz2z2yPx6Z54mcM68/+4zNasX93R1Kmxp0c64Wfjk9zcla7CDTQlqUfNJO1kaK56xpPJKdIxyPiJxRpWCEWJyv1CJkl09Zb+GUYlbFBou3LQt2pzUra7lZdcRUmOY6eWGXIptQfXjj8hCwNHk0bZXFpVxQukID3tWfNc/1sJQCoqr+ofnUrRzrx/oSC0qdrBMXcq2p5jg/BKYTS6aWY2KeJu7u7whJMLrA4XjgOBwZY2LMMKdM9o7muOPRBXbBE0yBlUZftjW", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-130", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "pPJKdIxyPiJxRpWCEWJyv1CJkl09Zb+GUYlbFBou3LQt2pzUra7lZdcRUmOY6eWGXIptQfXjj8hCwNHk0bZXFpVxQukID3tWfNc/1sJQCoqr+ofnUrRzrx/oSC0qdrBMXcq2p5jg/BKYTS6aWY2KeJu7u7whJMLrA4XjgOBwZY2LMMKdM9o7muOPRBXbBE0yBlUZftjWAC4EviRxGNAqLZtU2SGNIWSOSIBZFkAbbdzRdS9dopvvAtDtiDKzWFm1Vze5TJUKS84QiIEmULiiRUa0BK1CrBtlZ1HVPc3OJWq8orSUZSVSy+mc855qc2uTFU2/U2eOCihWmkihF4GNgcoH9MLPfew5HzxRqFdNYg5CGmECqJYCngkqF4iuuOB00wTnwiZgFwUXSPFfssVRydRon5nFcfIyX2YPLbLwUq2mSaBRKNazWFwgKfd/T2AYpJDlF5ulYu04RyOy/oSr+F1tLC3WtxGu1kEud+FASlFKx11xB9GpqlUudbKLAKLFYn1Z/7phi7fTMmRTL0k8hlmAOnozM4LyoFdbJQyaEOvLJOXTXospihfk/JRmjMM8Tr19/wt/96pf81X/6v3m4v+O439HYDlEkn725Y9V5xhddPf2MJap6+l71K7quI7RrVGsJpkVrTWMNizsl8zKmx2y3tALm4yPTMPL2/h3vmpavP/4frLZXdKs1n+12HJzDLZaT+9tb/Djw+z//OVJpbl6+ICyzsfLi9pXOGW85mSqezXvyuWXu+6/7d1/jppm7r96wbjtuLi5YaYtp+uq1qXWdFLH4/ZacySlSUqCOpLHVEyHFpf9e0DcNv2iuuG5bPmgtX+TCF96hjvXlTVEQikSqjO0kq0tFSonDLlez6lUNdlMu3D3W8tI1ENeaEuUZlkipanmbBrQuCBURqQaHfl2zg+M+EX+AibnSFokkes/t7TvG+Zd06yuadsPkHHPwfD17hlI4CIGeR7784jdkpSlao172XG1e0f+4JQb", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-131", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "cySlSUqCOpLHVEyHFpf9e0DcNv2iuuG5bPmgtX+TCF96hjvXlTVEQikSqjO0kq0tFSonDLlez6lUNdlMu3D3W8tI1ENeaEuUZlkipanmbBrQuCBURqQaHfl2zg+M+EX+AibnSFokkes/t7TvG+Zd06yuadsPkHHPwfD17hlI4CIGeR7784jdkpSlao172XG1e0f+4JQbHfDgSRsfu9ogOEuslbSMwWpGkIYuGY+yQtufyR1dcXq/YWMnbw5Hdl2+5uu65vOzwc8K5yNs3j7jB4aY9aU44mbDXW3Rr6V5egFU0H72guVix/tkHdNcbulcXlOKJJRBaRcnPEzCfKLJ8Fv0USqwMeKS2xY8JXMzc7WbG2bPbTzjn8a5irkZrbq6uWPUNUtnqtCYFjW7YmpYwDPhp4G4cSUiK6ilFVGVCToQUmEqilMjuzVfMxyN+GpbAW2rWigFaBB1tc0Xfr+muXiJFJqcZcmIe9zg3Ms1HrJYYLbm+uMSY50umfvA6d8AWRKnTU0rxVT0UIYVUR+r4qepxKUipWG0vKVmgiViV6ZvajDItcxfrhJJQk7Fy6h48/byCVImkCkcVSNFglEbmjCgFPwy4wxGz6jC0Z0nst63v6EgLTNPA69ef8/Wbr9jtHokhoJVBUB/CJDRJG8p6RcyFaarzzHJMOFH1vGbRhCskRinazlZX/JDPXgkleUoKBO+I3i/dbIuzmZaoxpBLnQ5bYkZSmXgpKmZTRCHkVEHxBcf9hgkzS2fGe5dFvDcz6/uu/e6B6AJxcgRp8VnSNC3pYrsw+PJsBH1yY8onnCkXRM7IXN83ZRnuWV8NEokWivVKIUwd855iRpaEyLE2UcTMNC4EXiq0XcVicxFLGVUlbCHAPFWryVPQVQqKFKQMIQiOhyXLSdSHSDx1Mz936aaDmJmdx5cjQ5Lo2SPNkUP0uBy4y74Sf0oSyEQ/Io1Bogl4oozkRoAxGLNCb1raTY/yoD20TYs2hogiI4h", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-132", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "sBH1yY8onnCkXRM7IXN83ZRnuWV8NEokWivVKIUwd855iRpaEyLE2UcTMNC4EXiq0XcVicxFLGVUlbCHAPFWryVPQVQqKFKQMIQiOhyXLSdSHSDx1Mz936aaDmJmdx5cjQ5Lo2SPNkUP0uBy4y74Sf0oSyEQ/Io1Bogl4oozkRoAxGLNCb1raTY/yoD20TYs2hogiI4hG0G0ML350gTWCcRzIpepZX768om0bksvsDyOHeULYDHaNNAJlBauXl7SrrjaoWIV5ua2Z7lVLagtzGMjJUaIjuZHsngcvxOCXSiWfNaMnwiwLQ0ZwDBCSBG1oe4VtDCVWhUzXKKyR3NxsaBvDum9RcnnelUHYBrFp6eaJw5wJCXxU5FRQISxqoUSYJ7yfCONInMY6ly4GytJ8IdCUIChRUaInB8fkqxmQxqMV9EZhhaZRPd7NhNkx25kUfkfNESfhPyza2mp0dWqmSmkixcA4HPA+MB0nok9Lhj+QgyPmAlLSba7IqTDsxqrAiDDGxBQT8zTiZoeb3KIukU+9BAs84XNGpMLk6t+FNqBKQQHRedw40flQW/6/B5n2nUF3GI98/PHf8/r15zw83GOVxSgLRZFzzTqSacgXm6pzo4AHJDXolkKbCjqDQdJozapv8C4wx3iWb+VQDZL9NJK8ryOPWQT/xqC7lgzEFCEmhJDVYFpJtFW1SyaF2ixwcgN6D1A/ozsnpu3UxvzMoPt4f0cJiTR6gulxSdJ2K+L1JXqckDnXZg1KfU8pVDa11DHO8gQzpFjJNb3kQ1kiSnV42m5a+rXlImvcHFA5IVI1yAmhcDwEchLnkenalNpYQEboKgWLYSHf5rA0b1TJGKXCDG6G/MDZwFzJjBCFHCU5Pa+MBtBNTxSeMeyJLhCOE1EcSbLlnmpRSW8RehmSKDK4IxaLRuOZicqTFnW+NQ1aSlqpkb4gXabLoAsMoyLGahKvrgyvfnaNG47sbt+SqVzC7/3hL7h5cQkB3t3", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-133", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "lImvcHFA5IVI1yAmhcDwEchLnkenalNpYQEboKgWLYSHf5rA0b1TJGKXCDG6G/MDZwFzJjBCFHCU5Pa+MBtBNTxSeMeyJLhCOE1EcSbLlnmpRSW8RehmSKDK4IxaLRuOZicqTFnW+NQ1aSlqpkb4gXabLoAsMoyLGahKvrgyvfnaNG47sbt+SqVzC7/3hL7h5cQkB3t3e8/FXnyFngbnU6JXBrAzbD1/QbFfYizWy0bCxJJlxVG8I53YUP9cGnnGgTNOzrokLc1Vd5EhOhZIKSlfj9iwtsSj2rnY5KiPprWTbKawUGCFYd4bWKjbbtnqNcEoZJHa9prvOWHfDEAL2MTD7KofKIZKnEcmRMDni5JkOB8LxSJyOuHEgpoAQsRryC00OghI0OcwkJZinPaJEepOxreFitSZjiK3lrZuZZsekJrx6vg/oP9V09Y+WODUyJWY34+aZnGZKDvh5j3cTd2+/ZjiO3L67x8+eefKU6UjxEy4mipDYzQ3RJx7ePrBZbbi5ekm2HUlZpnFgmv0y7w2kqIblFReummqfCkiYVI0ZvguYXLPuOM+440BygRLzeTjBt61vDbq7/R139295/fpz7u/vyDkTaz1KcBEhIy9JGJFQyaFjwkaPFRKjDcp7ig9gDEKBERkZAv7xQIiZ7DNxDgRXTwlZBL1pKVIjpcFoWUeml4IuBWU0jbFEPyFL4WTkEksklEgo6axfPdmwnfwjhKCKxs/QQ14MeZ4XdNe2raWGsKjGYlKq9oW7fZX/nOUAkIPDB88hxUUUDypnZEzVbpCCEItOeFFFlJSZp8BcCk3fEJRAmYLtMt2VrG2jiyphgZHJGYwWZCPPJj61Aijv9fgvb6A8/TmlJ7xRiurnUNtPn3VJABgmT0kJ0bYIFEJYvLSM0jBR8GSEKdXrV+Rzc4BLmSgUsXhyiZX1zeLJItMUpBHolaZtejpjaVMLWJS8qaPbX2iOOXLvdjwMdxz2j/z3zwSb3ZpuvWF", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-134", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "OeFFFlJSZp8BcCk3fEJRAmYLtMt2VrG2jiyphgZHJGYwWZCPPJj61Aijv9fgvb6A8/TmlJ7xRiurnUNtPn3VJABgmT0kJ0bYIFEJYvLSM0jBR8GSEKdXrV+Rzc4BLmSgUsXhyiZX1zeLJItMUpBHolaZtejpjaVMLWJS8qaPbX2iOOXLvdjwMdxz2j/z3zwSb3ZpuvWFOgc0fv6InI7REdBLRKPS6QxpN0YJAYnb3xBxwYSIFT/QeGSIiJtQ0VznBM9Y8TRUfZJlz1vQ0fY9pWqyuTQ5XCRBgdR133ppFE1oKCE0Ugn2oWuuUThNXJHMsjLGw9zBGyZgMsdTOtSwLQSSyzAhZ6FqDzh2teIHrV4w+EfxMzgestqy6DavW0neajfZYCWajMcqwXXVYo+gaUx23UmGaZ+7uH2oTy/8kvPCN+YaLWXsudQ/HGAjOM45jDY7TiBIBUQK7u68YDns+/fWvOByOvHlziw+R4COdyDSycHH9Emks7776Au884/2evN7S5khsNiTbs3s8MM+eGKtaQytTh5+qZdy9qFhwFjXoCgrBe2QpWCVJk2PaHfDDSNo4lF4Mzr9lfWvQPRwe2e8euL275Xg8VqBZVAcbHyJSpKUEyYjkkSlhYqCVhk4awjRX4sC1oAUyJwiFGBw5C0oSVR4TIqSMyIVG1zetpAYqJhtLIZSaQWq9jBFZJjXksgDhORFLqi2xp0w3F0LwCFhcxiLRLSX/Ut4/lwdolSWoUiceSI3KCek9YpooWpGVqtMjgBJ8ZYZzQpWMEnUznQ2tqXDDWVKdl15xH3HA3FWJijQFbat9YUxV/mUMWAthIcdEtVxFK3GWp1VT+Rp4zzG3PD3sKdUNLsUiIoczNPPcNbt6sChtEFIjdEOQhllKfFHLdId8VotULfkyfTbFSiiJpdwV7w071QWpVZ2lt+lp2h6lVkjR0NgbbNOgNpL0mBjSxDEOHNyBL++hdQ2X4hVoTfvjDcIoTGcpRpA", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-135", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "dl15xH3HA3FWJijQFbat9YUxV/mUMWAthIcdEtVxFK3GWp1VT+Rp4zzG3PD3sKdUNLsUiIoczNPPcNbt6sChtEFIjdEOQhllKfFHLdId8VotULfkyfTbFSiiJpdwV7w071QWpVZ2lt+lp2h6lVkjR0NgbbNOgNpL0mBjSxDEOHNyBL++hdQ2X4hVoTfvjDcIoTGcpRpAtZCmroXj0pBiZpwMhONw8EEMkeY+KlbCyISHj865LCB5O4nyt0LbFdD2271kbi5HyrC60RtVpD7LgTsEtSmIWzClQi6MF9ymZMcLgC4dQmBLEWFuKqwV89UcpolYxjdHYrkWJir237R4hBCnMtE3DZrWit5LOKlqVsCqwai3WGC63q2U0lUCkTBaJECLDNNEs04n/OdZZaZSX4bMp4VztHB2OR4ZxYByOWB0RJXD/7i37x3u++OQ37PYH3ry9I6ZMiInLRrO2ms3FJVJp9rsHwuwJw0ArCr5tiLnupWkYmJyv5KCQ1dtELZZdApDVCB4BPgqMFBViFQKRDSl40iiJsyc5j+475Hd0R3xr0P2//s//g7dv3/L6s8+Z5mkZWhgoJdIKU29mCMhhYP78c2TMrEZPIyWNUMhhrEz0VwpnLXdffk7bNFyue4yyWNWwenjH6rhjt3uguJkX6yuEVEQyo595HI+E2SGPY334kcy+9lCX5Nndv+XLzz/h4vqCD378IVJYpFT4OTBPI5/++ldIYNM1OJ9wLuHn2l+9PzwQYuAv//d///0fjgwojd12mAwyFbyOJBUwPiAR7KUg5Myb/SOH3Z77d/dsBWzUYmxSarA9ZbcsOF1Kmcl7ctNWV7A4EEk0FwprNNxXdcFwKPhZYAwEX4Owd3lpjqjdTNqkJcCWJbtdxrssgfiE2ypdVQwpQsgVkvghmO7weEQZRbtak6QlyhanYVTVehJRB/4hCnkxuVfnyCppVIOQFqOq7aWSCq0ljbU0TUPbtnSrFaZpMe0aqSzK9iA", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-136", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "H3Z77d/dsBWzUYmxSarA9ZbcsOF1Kmcl7ctNWV7A4EEk0FwprNNxXdcFwKPhZYAwEX4Owd3lpjqjdTNqkJcCWJbtdxrssgfiE2ypdVQwpQsgVkvghmO7weEQZRbtak6QlyhanYVTVehJRB/4hCnkxuVfnyCppVIOQFqOq7aWSCq0ljbU0TUPbtnSrFaZpMe0aqSzK9iAEDxzwq8L6J1d4mylXhuO855BnojdY3WBvLkBAYsbHgD/6s4ctMVBSxI8HUo6QPKqALOL88XJ7iSnfntNbQ7QAACAASURBVMH8w7VZbVFa061W2Lal61eYps7XMlJWnwNOumiJT4WjzxxD4RgLQ4r4zEIMF1IQ5JhJLuDmyDx5nHfEFKoJS07gHVpkViZUSdrLK6JfkYLni6++JgTPqlN0psGqSxqjWTUKq2Q1z+/+f9beq1muLLvz+213TJprUQDKkt3VpKiQCcWMFKGJ0Js+tV70ONKMRuSIEzTDZjuWgbsmb2Yet60e1s4EmtGsJqp7V9wAUAncm3nOPmsv8zetGMVqaXcdjyMnO6CYMz4lNpsLPv9U06/a6nn4x1nTNOGXheMwCKGk7tm+71mvV6gXz3m4+5797p7/+ve/4PHdWx5f3QNwc/kM7RpM22JLwlJw22co41jyE91qw//w9desGse2bTnoDYPquN/tSdHjTINWUD1BMRpSNSdNNUGavWDuh+NAaTw9mew9aQmMuydc3+O2a8zvyf5/MOh+881veHx4ZDwe8THKBLw6qm5aI86pBXSIpMcnVMq4OWC1xihNGSXoqpIoxjLHBH1PvL7Eug7T9Oh5wgaPiZ4cAm2pMA1t8ErLBDglSgiYUgXCq45DUollHHl6uGO/2zEe9rhmjdGOELxAl96+RZdM3qxYfGKa49mV4H73pvpO/etXERYB2jXoXFBEMIrsFDkkSBD9gs+ZZfEsi2eaF1aNRVVV/1OPGTjTKLUWeNFZuIdCKRFIVWdYoF+CQpC", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-137", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "oqpIoxjLHBH1PvL7Eug7T9Oh5wgaPiZ4cAm2pMA1t8ErLBDglSgiYUgXCq45DUollHHl6uGO/2zEe9rhmjdGOELxAl96+RZdM3qxYfGKa49mV4H73pvpO/etXERYB2jXoXFBEMIrsFDkkSBD9gs+ZZfEsi2eaF1aNRVVV/1OPGTjTKLUWeNFZuIdCKRFIVWdYoF+CQpCeLVkEqlNEDsTK4VVaSWCtTMJ/jrw9IZg+bD2VUodoMkH4qGsCkCpwP2axN49Jk5AMFivuHqVKTKpqaa2VwVRSi7VyDZyVQZPTBmssXdPSdh1939N0K2zT4roWZYTCmsgsMZBsxm4aGt/RqDXDw4HgM0uZKUVhrXREU4yEOOP9QvJiAa+iqMVlP1NKkmteDIoTMQca19Lqj5Pwa1yDsZau72nbllXf4hqLqfhtpTjDjFIlLswJpgRjhDEVfC74IN52cSmi7zp66ZvPnhQX0TpJXm5g8GKoaYsI1Wg5aLMCRUKRaKwCrWmtwxkRyDkHm2q6eoJWSu9XYa0IeseU6Loepavo+48Bdf+zdZq/LMvMPEnvthRpXZzE310j2Oz9rgFlOB4n9k9H5mnBWse672hXa9rNhrTMlBTwWZNLYYmJxlouthtWzopkqulArbDGnPSq6p78wPVcUfvopbr8ZCKi9Z20IgWpQGIGPy/4eZZh2h+S6f6n//s/ysDrOEnqngpZzKq4ub3idrPhRjXowXP45ltUiLQ+0ZjKvimCT7XzRC6ZR/1PjNsV/vOXNP2Ktt+S4kKTAxsSvkT8u1cobSirDktmq8GFiB0nlJIbsHdGtBSyZ3f3jv/8H/4fxmHm6XHHs+efsV5f4EzDcb/nb/7qL1ExcLtqCUtgmRZi0xCN5m73Fh8+DtwdtKFYh3IdBmiahLneop9dkO6OxCkQlpmYMm0sHEJgPx3pzQVb3eCNJqOklM6K1md0HVwZxAxv0omsoUtW0AV7zbJXPL5JKAPrXoZoKYjmcUa", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-138", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "mq8GFiB0nlJIbsHdGtBSyZ3f3jv/8H/4fxmHm6XHHs+efsV5f4EzDcb/nb/7qL1ExcLtqCUtgmRZi0xCN5m73Fh8+DtwdtKFYh3IdBmiahLneop9dkO6OxCkQlpmYMm0sHEJgPx3pzQVb3eCNJqOklM6K1md0HVwZxAxv0omsoUtW0AV7zbJXPL5JKAPrXoZoKYjmcUa0bCngMLJNSrXrsQhoHFErK6WI/bfUoORQYcS11WCa/EP75V9eCULJDE/i5lAULOFIzhPtF1foixbWLcoqrAarLZ1uaFyDs41oDZCJKgh1s13R2pZtd0nX9qy6Na51GGtIWqRnppQIOTLGmWQSbA2d26CuLFxElmlkOo5My8K484IcCYG0RNISBBBfCkad+v8BFGil6VTLSq1q5ptZ9xvajzQxbZoG4yzaSl/e2URrDY09HauKnAReeQyJYyzcxcyYCxOwpESMifm4EHzk+DCSfCBOE61W9FaxchbbOkzTCmkgClhVp4XD4YmHu1ccnu4Yj0/YMqNL5LrP5KyYF8lc/RLomobUtKjZ42IhVTJtThnnHOu1kdZdgZcvP6frNzw83uH/AHLESUQrVejWd999x2G/5/bmlq7vubq+PtsbnZDO6/UlIWSc24DqGIdE0zrWVz2r7TM++fRTvnv1it3jI//4D//EPHuGpyc+uVhxfPGE6RrWfcv1y8+5uf6M128fpLoDtDK4ppWqUGtCyUCu9vaZORVIihC0eM9ZjcczF+geHqBxXH/xOe3vcQD7waA7HkdSTO9prFWDVWsjnPSMNBV9wHpP4yNtSKJzqypMi4xDYFytKnKSOEtW4GP1PIvxDKUiRDAWnS2mvkETRUWs6FyzOZkq6qKI3nPY7Xjz/fe4tmM6Lmy3l2zWF0zDwH63g7CQD4qwSGO+9D2lcczTSEgfNxwJSO8pxEhjNMEaYkgwLMzDSJwDhtoqmOdqvyJBwiuFPenO1Uwn64o7VFl+b96noEZptCo", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-139", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "9wHpP4yNtSKJzqypMi4xDYFytKnKSOEtW4GP1PIvxDKUiRDAWnS2mvkETRUWs6FyzOZkq6qKI3nPY7Xjz/fe4tmM6Lmy3l2zWF0zDwH63g7CQD4qwSGO+9D2lcczTSEgfNxwJSO8pxEhjNMEaYkgwLMzDSJwDhtoqmOdqvyJBwiuFPenO1Uwn64o7VFl+b96noEZptCo0SeGSZCMFAYDnKFbq8IEUo/qwd/sBOYNyppuWoioG8VzZv7dty9SS7sewIwwoQ1Q1kMSFMFex58MirQ0HurOstisaY+lMh9GS6aoirQdjjHD9uy2ta1l3F+KQ0HZoq1FGUUqsLgyRmALeL0QfCHEhZE8qkaJFoDqX6tIQFiGzxEiO0jOkYi1RAgWSXrjGaovNBpMUec5kn1C9+NF9zMo5EZfEYZlpGsuxb7jcbtmsVlgnGVvOmVgEwZAp5JMgU0j4aSb4SBgXok+UmMShV1uaRtO1ms4ZnNWkIjoOPkdi8ExHmcfc37/FT3viMmBdRqt8bqvkJAezdRZjHdpaYhYxlyVVOmxKWOtJ5eRIorjUmrbrWPXrj3bT/q0to94TRlISf7xxmriptWCIgmVPVbualDgeDkzDxDQFpikwzIElK+xxwLQH2m7FuD8yDRPH/RG/LJhpwphCfHwkrleSpKBFXc22tMZJHFAVVarUe1z1SUagFNLpOS4ZmxNTDMSSCSnjx4FlGIghkGLC/AB07AeD7jKIKGuLxiAPrXMt1jWUJTDngSkstDFyURIblblR1fqibfAqkVWmNQ1aaVYXt6irS/RXXzDOgeE4cfCLCHuEQIkRFT2WRKMabA0EeRzJuTBrzVwgJ6mtG63Ii+f+1Vt290/83V//DZ9+9gWXl1d89dVPKKVw9/o1fjwSDrv6kCa6i0tc36Ma93snjf987bOwVw6HyLxe49ZXzHdHlu933O0eCNHz5cUFpRS+f3oiek+nLVFpHgqkGOhKxmqZitJKw6GQWFIiOXG", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-140", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "XXzDOgeE4cfCLCHuEQIkRFT2WRKMabA0EeRzJuTBrzVwgJ6mtG63Ii+f+1Vt290/83V//DZ9+9gWXl1d89dVPKKVw9/o1fjwSDrv6kCa6i0tc36Ma93snjf987bOwVw6HyLxe49ZXzHdHlu933O0eCNHz5cUFpRS+f3oiek+nLVFpHgqkGOhKxmqZitJKw6GQWFIiOXGd0EWcAhSF26hYlOLqUjOOmWGfq3Gixq402gqfHkDrfG4fxFBIqZydNcRXC/xSKkWX+qWYZxFC1+bHWZyaZk3RitJapqNn/3QkHxfK5CEW9MqiXvasrjd89vIrurajcy3L7FnmBVXkfTerNa5pubx6RtN0rPqtCAhphccTS2QJ4o47j5HFzxzGHYufGadB7FhCYFnGWv4JZC/NCV0kc1NizHEmL+gih5HJYre+NivMWDC7xPQ04mdPbi5g/XHthXE6sj+O/O0//hMohXGGP/v6J3z1xWfcXN+I+4lKJCAoTdHgTKGECb8fOb47MA8LeRFUR9tt0U2P2zxjs1JcbBQbk2hU4vXdA+M08/SwZ79/4le//jnj4YH9wytuNi0Xq4am69Ba8e5uR4pZkAubFTfXN6JPUPWrvfc8HQ+EEAg+yHuvML627bn65FNuWsez57cfjf75XSvljA+Bw3FkfziSScQcePt4xzRM7O53pGFPHvYEZZl94Nvv3vDm9SOv7/coDa8ORy6+/Y6rppeDI2eO1dXlT5eZq1Ezpz3t9TOWT7/g8nPNtrviwm0YzIF9PAiZRAksTEdV4aeJULUoUmsIBnY5MsVAmCdUDKjoae46cslM+wHXrlm1/b/4eX9wF7WuQVMwKZHQRAzbi0vW6w0X1tFpTWvA+Jm4g4QGbdCrDrtaU0ogl4gzUtLGvkGvetz2gqUcScNCUkr8K0FKO2vBGEIucoLEdO5VJiUK941RgKWvSl5LgVTFUI4P9ygfOGwusNax6TrmFNmPhpASQ0rgI0UHWNJHN6V", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-141", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "5MsVAmCdUDKjoae46cslM+wHXrlm1/b/4eX9wF7WuQVMwKZHQRAzbi0vW6w0X1tFpTWvA+Jm4g4QGbdCrDrtaU0ogl4gzUtLGvkGvetz2gqUcScNCUkr8K0FKO2vBGEIucoLEdO5VJiUK941RgKWvSl5LgVTFUI4P9ygfOGwusNax6TrmFNmPhpASQ0rgI0UHWNJHN6VCzMw58zBO6JC5xkIuNLmw0pCtZd33FApbH/DWMhvB7U7zwCaLNKatsn2h7+Qa50ijNGvjGHwi+IS5dhilydaSMMRZ2gGqDsdKydVlFmI9man9pJP3W6WhS+5Q53YnAsQJTqZNzXQp/FgjAK2tuB0bDU5TOi085GwJ40JZFpQNNGiabHBBCzJgypipnF1IGutoSsMq9zhaGuOIKhNI5ByJRTJbH6RVtPiFZRQxctESEJGck+xmU5mBKVaLqFRQScxTVc3cpIormKJw2dBmS6MMnbMol1iSoVWutm7+9cs5h1aKZRqlD5gj37cNKUT2x4l+tWK96aWU1YZhCdwdRg77mcNhISwFVQxt51DKoG1HURpfretTiIxlwZTA427HNM083N0xHA8s4564jKgSKUksaEIQtTOlRMqx63usdSIitIhQ/jiOeO8Zpqmqb5XacpFDHmMrBE4Gbx/JjP4X9o7GOcf1zQ1GG3Z37/B+4bs3bxmGkafHPcwDLCM+gw+Rd6+/5bh/QpcFXaQSbH1inSsGD+iduCd/FjLrnOjGhGl64nHm4e09+9X33D88sN8fGNJI0ieTAo2J4tiRciaUSEFcomcFYYy0GnwDvcr0KjEPR7RtmY8DzWpkdf0jg+6m7zEU6c0qSzYtn372Bc9uPxG8ac5QFtJw4PhO0VpFMQ3mYkOzvUTFWewuggCG87rHXqxZ3VwzpkJ4PIhCl6pUSa1QtTe1xMwSI9Oy1H6OIilFqZqt1hg2TYNPmewDs/f4ENjnQtgfuFxtWK3X3G63jEYzjwcOKbPLC8o", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-142", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "FYYy0GnwDvcr0KjEPR7RtmY8DzWpkdf0jg+6m7zEU6c0qSzYtn372Bc9uPxG8ac5QFtJw4PhO0VpFMQ3mYkOzvUTFWewuggCG87rHXqxZ3VwzpkJ4PIhCl6pUSa1QtTe1xMwSI9Oy1H6OIilFqZqt1hg2TYNPmewDs/f4ENjnQtgfuFxtWK3X3G63jEYzjwcOKbPLC8onyIE0x7MOw792LT5x9J5Xux2ln7mZMttObHeM1WAbbq8uKChSMUzBcwgjT09PHPc7rgpQFC0aZR3D+gJDZhUiHZbOau52B5bZY/+spekbYtvgDxa/F00LrTiLKcukVVUOfRX3UJD1SSj7pAFAzXplCKc+eFrE8aLCdn500BUx+2KhdBayk6hmFPPrA3GeUbOjjZrWG5qk0D6ix4iZUnUAMXTZ0YaWy+sNJjuUbigszDkQkidkzzSP+MUzHAf8MjMdnojBi4ZzSeI7l6NAEGvvJIb4ftAo05G6q+pBVYQR2BRDHxu2esVFu6LpNKMa2JgWx8eV0n3bcTQDfjgyjkeOw4FlmHj9/TtuP/+U9cUFP/niU5qmJWrD437gm+/fMQ2RZUysLm9p+xXr7QZjLDE7lhAZx4HjYaD4gxhohgU/HQl+4t2rb1mWkWV8IscFQxGd66BYZsGfGtNgjGF7cSFT+cVzHI4ch4FpmgkhEOtGsMae2wBJabJzZAXGWla9wf64wui3ljUW1Wq++PxLjttL/ur/+j959f23/Pv/+B84DiOHYcImj0mBpeodjE8HSsqseifJSopcZsOzOMsQVhvMdoNThs/mjEkBjgtGN8TuwLe//obDEb759juOhwPeBJKGME+YYnDJkauGTFQVzmgSlESaRjoDh15z2xle9obx6ZG4JIaHHbbpufni9l/+vD90Mf78p38CKaGWgaId2XSsL9YYZwS6gmRNKE3jOkzK1W9J/m6ZFlgyZYlCP+0K03Hi7TffcjyM7KeRMXiWnIgVK7oP0ttakgw", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-143", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Ths/mjEkBjgtGN8TuwLe//obDEb759juOhwPeBJKGME+YYnDJkauGTFQVzmgSlESaRjoDh15z2xle9obx6ZG4JIaHHbbpufni9l/+vD90Mf78p38CKaGWgaId2XSsL9YYZwS6gmRNKE3jOkzK1W9J/m6ZFlgyZYlCP+0K03Hi7TffcjyM7KeRMXiWnIgVK7oP0ttakgwaihNxdM7ao0oCPkDOgq+cZ6JY4aJTRMfAtH8izjMHpQl+EczwElAxEVxgAhoL5mNxupsOnRt+smp40a/46uKSprEYp7n3AV9gp+SEXX9ySR4HHu4HYmNIXcOTsSzKsFcGZQ2xdRgys5Z+nS5wGCeOi+du8LhUWJyC3nC5djL4yEks2pKImic4YwNzLpXCqKq+sEarDLrgOmFJrDaSCVsH1ilsozg5Gmijf5R7xLpfM5VE8pMMpzphK5ZeoXOP9Zbuusdd9dw/PnDbXfBl9wy9vkKvDNo2KG0JSRTglmEWu5RLRazTvngYWaaRx7dvmWdpSUlPd5I9EiNKFQylIjHK2YBTVZkNGWKqs+iQoGPkxZykl6gTbLoVn118ijWaw+i43lzQtB/XXtg97RjGAdsaetXiXGHTO3pbWHZ3hGHHPx4P8rkxlDo9/2S7or/tsf1GhKIaRc6Jw35E+4D1A9M4MBz35Hkgh5llfCT6iXnekWIAonh9LQtPy8xRFa6vr8VVudqmP+52xBhZ/IL3AR98pdrm8z5QSoTxQ/AUAzoajscjDw873oxPlBT4d//b//5xm+UsBnWaz0iV0zUtyTWo/QNm95Znyz0XRMIKrNZY1aKKRZUW9bwX0fyuo9WajdZsJs/FMHPcjSzHhTQNMkNaZowqqMbgfeTpcccr9R13w8Lj7pF5XogqkhUk6yToZke1ziVpwTyjqp/jPBM02Kjos2PRDSWPlKgYH++xzgA/+Rc//g/uoi8/+5QSPen4SNYN2fTgWorVpCJ21ilJudK4FkMVA9aWbBp", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-144", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "m95Znyz0XRMIKrNZY1aKKRZUW9bwX0fyuo9WajdZsJs/FMHPcjSzHhTQNMkNaZowqqMbgfeTpcccr9R13w8Lj7pF5XogqkhUk6yToZke1ziVpwTyjqp/jPBM02Kjos2PRDSWPlKgYH++xzgA/+Rc//g/uoi8/+5QSPen4SNYN2fTgWorVpCJ21ilJudK4FkMVA9aWbBpK1LAUyhxF1DvDvHhev72Tm+wDPkViKSQtfbs5FVKBRYuLrXWu6uCm9yIUKf1W0PU+VP1PUCmhUmQ+HlnMLBqkOVNSEPhalHIBpWitxZaPCzDtqqMBer3med/zYntBsZpsNLvjSEqZPdICud30TCoQHgvRaFJjObqO6WRcaQyNM1gU3kjxalCM1jCh2E0emwvBKlSr2a4sPsAcC0GLvudUMjEWgSDxHoqm4H02VwdmrhG9iqbTYtvjFNZpjNU4kzGm4Dr30bb0IFldioE8DGDBtJbc1Gw8d+jg6G7W2K7lab9nm1uuVitWzZrercF1FG14fFpYSmScRpRL2OxIWSyS4nFm2R/Zv75nHEe0kSl7VMu5b6K0IDXq4Fl6uKUG3dOkUZ2cPqruRTppIwgtXJfCatPz7PoZPs4oXbhYrWmajwu6+8OBcRpFI0Q7+kaztpbOFB6PO5aUeLrfkzH4bFhvVjx/8YzrdcPz2yswjqINE4I8GeKAjgsujQzLkfG4J04DeZmYx3tSmAheSEwi3B8JIYpQdwq0fY82hmbVknLiab/He888C4mpnC6RUjT2fYabs0hBKqexyTIOA0+7PfdvvsPPH0eNrj8C+GCvKlGAaJ0lWYMaDthhx23cSwXcdVgn0p6dzjgNvXWCgOlXNEpxAbjHA05FXj0sPB73zEUqbGUVyhpU2+ND4Olpz5tseTXOTN4Tk/RuUZC1xWBw5aS7a6ha/2S8XNfFkzQ0CSYNoTWouECAeb/D/Z7D+Qdf7dBkZViso297VptL7PoC3a1lkFMKah5R4yT", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-145", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "PH0eNrj8C+GCvKlGAaJ0lWYMaDthhx23cSwXcdVgn0p6dzjgNvXWCgOlXNEpxAbjHA05FXj0sPB73zEUqbGUVyhpU2+ND4Olpz5tseTXOTN4Tk/RuUZC1xWBw5aS7a6ha/2S8XNfFkzQ0CSYNoTWouECAeb/D/Z7D+Qdf7dBkZViso297VptL7PoC3a1lkFMKah5R4yTg/mFgedyxbFfEy544N+Qc6FRP17f8yX/33xL7npddxzx7xmnG50hIkYc3bxiPA7/51Vt8ziTnSLrgdZGTOEYapeQNLyIcHn1kTDKAKhURMMVIzoXh7h1FKaYY0AXaItbvrsDli0tWN1dctR1Of1yfbnt7TcqZJXi8sXwXM/v9yLB4rBIWy5vHkYhC2YZxHHl42ItYcg5oFdHa0K9kI7VKC/9du7MexF4rlkZxt9tjraFxhZvLlqZ/IQ6mMRN0JKrMbhnxOUo/VWkRIqpUSl0fIFMxsE3XYLTCOgVkcvbSE9ZCSDghS9SPCLpfffklT/PEw/cZT8KbzNEsTMZjN06y6JUla8OgPQMLhzTT2Q39egVNT1aGsB+Z4syDf0DNhm5YSMeF9DTx/c9/xf2rt3z3zXfM04hqCm7TcvGnzygWsinvBzu54iW1MLSUUehSBBebhG6doxcVOi9qazpZjHK01rAyLdvNhuflGdttx/NnV7QfaUL4m29/wzTNvH17T+eEUlsQE1JnExnP02Gs0p2W24tP+ekXt9xcXXFztUEpS1GKYYn4qNk+70mpJeUVv/rNxOPrR4bHt4zHPSkOlBLRWoLksnhyEiRRCGL8+bjfM84z/TABIvVYKgNMVSGeU9A92ViFIJ6H8zwRkmdeRn7Jz7l/cyd98z/CIA0qYCAumDTxJ9vCp88c/2O6EmGsqn9d0OgcUaVgq1xAaCwqBcx8JA8P5Me36ONIM3lUFlGhg5ZKejyM7Gm4KwOvbyL324WxohJIlcWpjPxaDEpZlHaCKFEVN18yKiY", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-146", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "ksnhyEiRRCGL8+bjfM84z/TABIvVYKgNMVSGeU9A92ViFIJ6H8zwRkmdeRn7Jz7l/cyd98z/CIA0qYCAumDTxJ9vCp88c/2O6EmGsqn9d0OgcUaVgq1xAaCwqBcx8JA8P5Me36ONIM3lUFlGhg5ZKejyM7Gm4KwOvbyL324WxohJIlcWpjPxaDEpZlHaCKFEVN18yKiYWDarRbDQMrSGTKMZz//1vGIfdD37WHxYxj9UZIZ9OP4NrGlzXifkiSgRbtIb1hikXhmEgOQMGglYUA2iNtpZ2s6ZdrVBtx+Q8zlgJujnhh5GSFbp5kky2ccSSiCUwlcJShWSUEtB2RpFLPLvhFmlbEmrrISwTqcAUA6aKhAhpQNM6x2rV0/crrPm47CVViHnMhbkkSvHcDRNPx5lP1gI2PwwzYyyMSUvP8TjXfylkB6U0RWtMzMxFWhHaNhW2JIM6X4SdZrOia8Va+2rVk4oMGIP2BB1Rc2BJCuN0fXBcLavTWdbXOvFXa9oWbUQkp+REigrpZOqa9QmX+MfImF9sL0jG0LsOpwq9Efhb1pDbRLEF3crPISqyhaAyRSusFUlMlCbrQlSRKU+oqGFWxMNMvB8Z3uw4fP/I9HrHMk8Um2ivV2w/vYKipVdQY0A1BqhzxVM5q9C5Bt2YSKEy0nyiZIWpw0WjFVYJ1nzV99gGLrcXdO3HWdPsD3vmRQwiNVocCEzG6CTwOF3IaSHFQi4WVSJ9Y+kaQ9sIlA4kAJoCtrcVhml50yhUmkj+yDLvKcUDRWjySeB0Ai8zoAQmOPt6yNRrFIOI1ZwRh6epK5xNAGKMggXPmRIFgnfY78kh07hONDv+gHXiERQQpabk6VWkNYlVZyBmEW8qVaY1VZGqCp8ckshllmWizCN5GsFLRatrD38yGU8Wtl/JHEth6lf4piE6S1KKFCJkMMog8o4KpR1aR5KCouSzyzA2SVVeNIs3LD5ilcLkxHjYkcoPiwD9YMT5m7/7hWD", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-147", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "rFIOI1ZwRh6epK5xNAGKMggXPmRIFgnfY78kh07hONDv+gHXiERQQpabk6VWkNYlVZyBmEW8qVaY1VZGqCp8ckshllmWizCN5GsFLRatrD38yGU8Wtl/JHEth6lf4piE6S1KKFCJkMMog8o4KpR1aR5KCouSzyzA2SVVeNIs3LD5ilcLkxHjYkcoPiwD9YMT5m7/7hWDXnGbVj1ztJ9aXC6vNxPXtM5qm4fi0IwxH5uGJZZk5UPCHvcjHPR1gWlgtgcZ7nv6/v8a0LfQ9IYp+w+QDISa0sYDiT//kS3xO7JaZh8Oeu4cnZh8ICVpncNbQGXGrmn2iEIkkitYUo0nWErWWlkUphCI9vKyMOPVqzapZcdlcyDT4IyFj/8e//0tQ4pu1bRuebTrudgP7YWH75z+j7XtCGtiPI3/3zVtMSax1prXQ6EJRIkv4uH9i8olfvpG+07pzbFeWy7Ul+EBOGf3lJX3niBe10tEi1EMWabuYZ9adYaOaGlANVjsRzsmnaZFoFyh9eoiFUKFwWNXXTFcTYiGmImiRH5G9rLqOmOH55obNes3zm2f80+O3vNq/5VV4x5g9/dWWru95efU5V3qD1iuSVczMmKQpyqBWhdIUBj2TiOzf7Zl/88Tw83vmd0fsPvNJ2hJKw93DG0qYmN/usduW5ro7Dw/lFC4VsZGIVeMjLp4iDXFShW90uhHacXK02WKC6AyQI9tVjzYdz69v6JqPC7rTOBK1oXn5gmla2O/2tECj4fmzK7q2p+8NxkeGyfO0P/IPP/81L188Z54D/WaDsa7qniRWrT4zp1rjcWnAmUDTQCmWUjLzMJFLQWEwWgs5YxkpSvR1SaC9RyuFMaZ2WyTSFj4UispnaVStDV23xhiDM5bGiKh50zSYj0xafuc6tcbiRJj3/PpXv+b43W/w3/6KHDMxasGlp0JIIuivVpcU6/C9w5XAJjyxHga2UxAaddbYqChKk3qLMoWtghIUy5zxpkCj6b/", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-148", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "qniRWrT4zp1rjcWnAmUDTQCmWUjLzMJFLQWEwWgs5YxkpSvR1SaC9RyuFMaZ2WyTSFj4UispnaVStDV23xhiDM5bGiKh50zSYj0xafuc6tcbiRJj3/PpXv+b43W/w3/6KHDMxasGlp0JIIuivVpcU6/C9w5XAJjyxHga2UxAaddbYqChKk3qLMoWtghIUy5zxpkCj6b/8AtqWv//rf2AeZwhZ1Ay1aPFqbSlVqrVULLfJ4DVko1mZzLYB1RiM0zy8+Raz+wNowPf7oapxGOaQSKkwF8sUCrbt6brEFITZ47UmWgt9f0oXUJ1AXIpxFKNZAJUzJXgWH5hnz7wEYsy4egPdqkEVTacSdjYiEANkJdbmprE0bYcqiqg9xisRj619ulIxUOWkq1t7eMbZs/ZATJlliaLh+5FBd/QC9DdG4zRMXrPEiE+RJWf8iYUAlVKaz6Imxqha7oq6FgiduWiFLhpToTymBsd+vaLbWNq1iNegC7ZkYsnoaHC5RRS9DE0rfHmrnQD+q0B7IVc3i5MrqpAXFJV8oWvQDULzdFVB/2NX2zS0IdFaS28aNrblyq0JzZZH/8gSFxE1KtT+nBPJzjpptkYGp7ZTmCQsJR8C8biwHAaW/YQK0KqGti1EZTkcG0rWxDmgW4NK+X0AOWnYVhJA9kEyopiqfYiq6HONOVki+Ugk4JWQLYQYoM7v0f0rjQdPa7XqBFdqN2TdiPBPWCgxMAfBZodYnU4KLEvg4XEn0qA5s76YadqWZZ6hFFKrhbiiCtM4UNKCJdNqgWum9B42qKu+xfuipZx/KaW8dzf4MOCW35ZEBcmUT4Muqy3OOlEQ1Eaqoz8GD/j8FqVCezzMPOxGHu/GKj0gGsE5Z3wSVwcWS7GOMBp6Ii9YKD7RZ0UAAjIfUUrhOgsGlBIab+8Lm8ZSVh3Xz25R/Ypful+zIAPFpArRSNBVKgqWXwmL8DQvUFnhi2IJVhJHU8jWEPz0h2W6P787UChECq0", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-149", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "ipZx/KaW8dzf4MOCW35ZEBcmUT4Muqy3OOlEQ1Eaqoz8GD/j8FqVCezzMPOxGHu/GKj0gGsE5Z3wSVwcWS7GOMBp6Ii9YKD7RZ0UAAjIfUUrhOgsGlBIab+8Lm8ZSVh3Xz25R/Ypful+zIAPFpArRSNBVKgqWXwmL8DQvUFnhi2IJVhJHU8jWEPz0h2W6P787UChECq0zbFrLqruja1u+2h/ZXmzpWoNBY6+uMEpxrQ1JG5KxlKoeZhc5WZuLNbEkhmXi6eGRu4cjOVaW2TxiteZlYzDOcrldMYWJpjFkJUG7325Z9x0319doFM04U45HbIqEysjRzmGsI00jsQZsYy3rzZZYmWTvHvbcD5G5Ul8/Zl3d9rXvpbHasGgDXUOjFA/ZM/kBYxTbzvH5JxeYUliR0Uay464XS/FjyPiQuNxucFZzuWnot2vWl1uM7dG2obty0jaod6kCwqR3W+QrVyrrKWPRtUdstdCBFYhuRha2l9KKxjVSOoZYHxw5CChVYvjjLgkAt5cXGG3ojcWmQtrPfNZc89knlzx++yCU3KeZkhTL1UxpV/Q3PZduwzN3IYOLonjctPgwkw6Z8WFk980b1NuAOQQumy2rrqPXmhQDKXtGvbDfTxgNat2c7ZlClGqhSGp0Zl9pZ1GNQVuLw4qwvherqcc3O4ZkiGrm2faGZZlR2mJVZcx95Prv/+Iv2Cf4r4Mm2BV9d0F480+Eh1d8d3dPniZ01TbQ1vF4OPJwOKJ/8UsUipubZ6zWa662F1ithbpcIkuZmR7fEocnNimyNooBg1eZSQvF21orJflJUQ9hRr8tLwAAIABJREFUgxqtKguwSvkXGaCdWggnxILRsi+appUMtwZbV40qtTb8UUC6H6xcFEtU/P3rmV9/M/I3vxjxKKJtUaaAKUT/gYEkikLmmVX8T2vN105x2bTMBOYq7+is4ubTLcaCCRNPx4QtgZtPLglffsGf/tt/g9pe8vP/8gvC5Hl6e0/", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-150", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "ZSQvF21orJflJUQ9hRr8tLwAAIABJREFUgxqtKguwSvkXGaCdWggnxILRsi+appUMtwZbV40qtTb8UUC6H6xcFEtU/P3rmV9/M/I3vxjxKKJtUaaAKUT/gYEkikLmmVX8T2vN105x2bTMBOYq7+is4ubTLcaCCRNPx4QtgZtPLglffsGf/tt/g9pe8vP/8gvC5Hl6e0/UhdBWb1mqFZc6qd8pXBb4ataGpxF6m7m0PRvXUPyAyn+AtGNShkIhpISKmUn595i9In3VySexzvAzrbNsekcyhqw1vuJJrVIYpYkoUhGPLJ9PliXSoy0xkoDdfoe2ltI2LH7BOidOwKWwxMhxWbg/HNFKMS2BwXuiDMkpQMyFkDPKiImh0wpjXeW7ForS4urqA5i2urD+65ey3Vm0JSU4xkxUmtI0ZF2t2HWR0zJlFh85zgttZ+laS9ODs7X3mUHZHqOgbS2r7ZaLq2ts02Nsg+uNBGurK7nhBKJVKAwaTYi+Bpqq16u1bJTf6rVJ1m9qVmuNk3+jQxX50Gc7I61+XNDtnPSzcwzsh4ndm3uurrasNyvKktAe0hwIRjMMB/bG8jjv6FF0xmCzkBgelnse/Z55OrIMA3430YxgivTi+7ajN4YULa5psCWJs0bNZk+A9lTtmpSWXq+xTg4kY6TlsGSWZYZYUJMnzZHDww5XHM5Ka8AvC6ve0jgrOhYfeU1ur69Q44K5uydaMUR1XYv75DkaRRwGwt1r8omCinyJCH9mGI6klIWerbU4G+eATzNpGEk+VneSxFIivnoDfogPOCEyVEX+CLxb2i4Kee3D7Pac1RqL0YbGuSpIZGuw1ufv/+MsTH/HOiXhRZGL5pBgl+AuFUIRta+a4JNE0bXSkgslRxoUu2B4VIZ7q3koikGJPjGtolk1FJV5mgv3KfMqZfwwEZ+euD4caI1DW9BOEQhCqc4n/bcT0qWcg26qbuJaaZaUGXxiiYVYisBWf4/Z7Q83ZFxXJ+E", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-151", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "yVEX+CLxb2i4Kee3D7Pac1RqL0YbGuSpIZGuw1ufv/+MsTH/HOiXhRZGL5pBgl+AuFUIRta+a4JNE0bXSkgslRxoUu2B4VIZ7q3koikGJPjGtolk1FJV5mgv3KfMqZfwwEZ+euD4caI1DW9BOEQhCqc4n/bcT0qWcg26qbuJaaZaUGXxiiYVYisBWf4/Z7Q83ZFxXJ+EBnwNlqeIWqoA1mKbh6XAg+IV5eGS7WqG7XmyegUMILIuvugGaNohd8TJ5xpDEnK8IrCeHSE6R/eERtEb1HclY2r4l1OHR0zCyH+Bht5fbrTQhJXwWgZWiFEuM5FJo2laCfcm4OsEuiK1PjImUApvLrXDgP2a5LWhBBIzzwtMwiK1258hGk1UhaFgoDKPnaT/y5t0Tt1crbq9WrDYdvbVcbK7BWLZFkbK8p/X1LbfPX9D2Pda5swKZq/z2VA0uc07n8m7yEzFVTWLECRjOOuqS6RukxVIRCihFIhFRAlNTiljetxV+zKO0bhuOxhDmie+/f80vfvFLfvb113zx2eeEo0eHQjQLIUbuVm/wcQKTmS+PTHpCV4uaf3z6jqf5yO7pHfPdiH+9x80dTVmz6nouLja01hKCF5nEFGmIqFiI40JIsYp5CzvOdV3N1FpMUZioWIaZ5Tgz3O9Zhpn4dCQunuFpoLUtahM5PD0xDgdubtasVr1w6T+ylP7is5fYd/e4h//CohpCc8HFZ5+z/upLlusbwuGJd0+vCGEieS/UQGPFVDVnYZZNM9M0o5TC+4WcIjEsuDDRhMgcJkIKHDOEIgpYp949dWiocsHUqF5KhUfVTVIqVvbUZtJaY4yh7zqssVWMSFUK+XtXhFJJOn8UmbG6MpqA4S5pXifF9wWxeK/jiaqJRC01ZaiYEyYp3iyFBkVjDPdZMwCq0+SVobtcMYfAb95EvpsT/7AU/Ls74hJpfvJTLqcRTMK0hYVJCEjJATJcLkXkUIH6c2VGhDYMMWGmyBATcy7", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-152", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Yp949dWiocsHUqF5KhUfVTVIqVvbUZtJaY4yh7zqssVWMSFUK+XtXhFJJOn8UmbG6MpqA4S5pXifF9wWxeK/jiaqJRC01ZaiYEyYp3iyFBkVjDPdZMwCq0+SVobtcMYfAb95EvpsT/7AU/Ls74hJpfvJTLqcRTMK0hYVJCEjJATJcLkXkUIH6c2VGhDYMMWGmyBATcy70PvL7iIs/zEhrLTmL2+2qablc6apwr8lhYTjsSYtAONZtjzMGP088jg88DhPHZSaESJNr2ds0QCFFTwqBtCzkIJTNHMV5IMVERkoI3baYvoLGjWGpYhIKOfnbrsO0LdvNmtkvLN6zVH0F27ZorcT9VltWqw2n+2f7DbrpWK23GPtxg4CH/RFrDet1z+7g+fbVE85ZnLMsU6RrLd4HFh+ZvAz4tpdrTGPxCXZjJpvEl59s6VdrdNNQKlts1ff1EFCkOqVVIMQP5IGp+CehZCJTVrGMruIcUTblKWPRSsR3BPFRFcQq4D2XU2ZjyCq/13T/EamukoEyzmli8TwMj/zqu1/zeHjgfnxkyjPzJ6DXlriS/v6h77i3CpzoC+eceFweGJaJeJzg6LFDAZ8Iy8I8jQxOE6wlpYR1ttpHRZTRaCP3wSmFaR3aVCZkLPinkXn0LLsJf5zxhxkVJIN6sb6hu27Y/mwrdqlJ88nNDcZourZltVphnZUy8yNW17VsVz0vP7nmu7tHvvv13xKODxzf3dAai0oRp6XVkbPoDZcsAklagUEQCDmJN1tKEzFElnFijgv4hZA8MUuWmyp5SALjye02nrPYk3B/iqesTRTVjDEYKzhYZx3GGKyVNpip1/DDwFu3kMwJ/ihBV97f0/GJu90dE4FowXbS/jHOoU9481rBzYsX+cfWsbIG2zWMOfKdD9ynzJgSnbF4o7n1mcOS+Nsh8XpO/DoV4jCRY6b9279hfXHB2/u3HMbDWVQrpSBoHky9jroO/GQwX0xB4wTNpKpXpLYkEr+", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-153", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "2nrPYk3B/iqesTRTVjDEYKzhYZx3GGKyVNpip1/DDwFu3kMwJ/ihBV97f0/GJu90dE4FowXbS/jHOoU9481rBzYsX+cfWsbIG2zWMOfKdD9ynzJgSnbF4o7n1mcOS+Nsh8XpO/DoV4jCRY6b9279hfXHB2/u3HMbDWVQrpSBoHky9jroO/GQwX0xB4wTNpKpXpLYkEr+vZ/mDEWddg26jDdebls+erYEIKlPCwpQSpTiMVqz7HqUyfpl59+4dv3n9likKpKvJur5ZDbVctlpLORoCpESOQUgMUdwfvPc0BVZtV4OuxseE954IooPZ99i2YbPdUg4HlhDwMUIurKD2KjPaGFardS01NZvrW7r1hqb7+PbC4/5A2zSYpuPp4Pnu9R5jRXDb+8Sqa0gxknJmDBGtNduLNVQ3gKcpEXXmp/0Fm8tL+s1GNlKNdqpOSYWyWtPVs6UPgATJcjpxc7Xyzqd/J305EZeWv5ui9DTPGU79Gblwvi+lbqb8Y7x65FvKzNUaYg48jjuG5Yh9beTe6oJRHc43xAvLogqHVYNyBd9GBDecZOg2L6RhogwBOxZUiESvmOcJ4xS+9iuNc7icccaD1ShrRTHLWdp1jzGWNHtCWhh3I+Pjkd23d4TDTDjMbFdbVm3Pyxe33Fxf8Wf/zZ+TYmb3sOfm9gZrNF3X/uhMt20a1uuel59c8Xj/hv1v/p7p8R12c8Xty89p2xarFcZaIZaUTLXjEmF7JKHJFUmQ4kT0gXkaxMU2iLZJzolSD2bnXHWizuLKcUaxyAA11+wQKnrBqnPPtmlEZlOCrj1nvueg+8Hv/6D1z9ExdT/ujk/cPd0zEyXo9mJ13rgafNFo0wCamJ7IJbNer+idw617puHItA/cx8SYEyuj8FZx6TOPc+Jvj4m7kPk+I9CyaWT8u7+lbSWbDzFU5tnJONaiVCHnDweGSqQLAFMqK1dpmWMpIxT037N+MOhe99L/iq5ws215cbkGLRbfgYa", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-154", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "PtmlEZlOCrj1nvueg+8Hv/6D1z9ExdT/ujk/cPd0zEyXo9mJ13rgafNFo0wCamJ7IJbNer+idw617puHItA/cx8SYEyuj8FZx6TOPc+Jvj4m7kPk+I9CyaWT8u7+lbSWbDzFU5tnJONaiVCHnDweGSqQLAFMqK1dpmWMpIxT037N+MOhe99L/iq5ws215cbkGLRbfgYaCResWazTb3op1TvR0Vibofl6YfSBh0Erjmo4SC35J9F1H5xz9WjLS4XgkeM8cRXjDVxcDExL5hBUs5WT2IkI3IVR8HyyLCAibGqCjl8zZWiu43RhEQi5nlLPkEvGL+WgiwDQnQvRo9cQyHXF4Ua3KCp1XaByuNcIcm2RIEl0+l3Ixw+Iz1rX0qw1N28k3zuW8qd/vxMpdrZnAuYem3jdeu74j54acRWDjxMh6XwaKO8QpIJ8yoIJgMbU2aCM07lL95X4U4N042rbn5bNnvHp1yVpZoYrnjHYKbQ2NatBBM353JO0zerHYSEUuRAqJOE7k0VP2nnKMZB9xWJqm4fLimqvba7pe2l4oGONIQGE2De62I87ibTZ8+0CcPdPDXnrJTzMGzaebWy5fbLlcXfDlF19wdXXNn3zxOev1mttnt5QCy+xp2pau79lebs+Dxo9VvMw50bYNP/36a0zTkLXmH3/5a7757peM9+8w1rHRBaMU1phz1cEpizz1XGMk5cTx6YngA8M4S0D9AKWS6v6SYakEjVyENHTaQaoe7EZLf7breiHotA5j7PlLhrFGZD7rjECek/fPitZ/PORCqXOdsHhSiPzkiy9x2rJ73BNilGoVUeXTiFqTSpHWGn761afcrtf85OqS3fff826ZIEi1+2qJPJTCw6snxhB5mwsjgBGT2FJgnkdyTrx4+UI+l1JM08x+f0Ah10+rUq3ChDikTgNJKzADHxbe3N8TBs2Xz69ZdX8AZGzdWDlhjbQa1l1T8aKFJTkyBmsarNGsWktKkYV8Fh1OIRJ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-155", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "SiPzkiy9x2rJ73BNilGoVUeXTiFqTSpHWGn761afcrtf85OqS3fff826ZIEi1+2qJPJTCw6snxhB5mwsjgBGT2FJgnkdyTrx4+UI+l1JM08x+f0Ah10+rUq3ChDikTgNJKzADHxbe3N8TBs2Xz69ZdX8AZGzdWDlhjbQa1l1T8aKFJTkyBmsarNGsWktKkYV8Fh1OIRJ8QJMx2uDsKfBEsksyDrJSEk7TBFoTgYAiFLEFDzEKrCrnun1OAUWRUyJSVcZSlGKg2r1Iza2qi4WqrsMRHwLNPKGNuBt87AA2BAlMk56JwWP1CYJELQsLzlWd1JBrSfK+bvdLqPg/g7Gy2RUKdBYkQR1UKKV+64H6MJCeUP/Sa9L1W0s2c5ZmLFTPqXL6o3y/UihFrHxyybVk1OcStKT31/mjltIY23Cx2bJdrVi5hmleBHTuLFrL6E8lTTwEdF7wbmLZtixXLcIXFIpl9oEyRcqcyDGDEVnEru9Zb7as+p5cMl3bEXXClAbrGlrXkacIS8Y/DvjjyPRuTwkJExRdv+L2+pqXz1/y8vlLvv76p9ze3vLy+XO6rqXve9lX0neR2cJpgxQ++qrknDHGcHN9zew9Xx0PfPvtNyz7R8b9EaUM+upCzB2birNWp/vPmdpesmjK+nnGhyDea/B+b9WfBdJm0qUerOX9AVu/oUAFtRZ4pmtwzlZ7HtmTuqJZTs/YqX+q1AcDtHO74ccF3d+6jkpVyEymhAAx8my7xV8e2TaOOReWEMTRgpPIFjQl0WrDJ9s1z7YbXlxfUp4eeTIVmaEU+yhkhv3TjM+ZY4Gg6qVQcgiFiuBpG6l6l36uZrYVbkk5O7FYp87XRoJvplSPxsMYUb7w4pPr3+sw8oOvvry9lDI0F5zVzMsiP+z0oBMJYSGUwrhPMq2PCT8t9LbFarHCKPXeGa1EuzRDXhaOKTFW07uTaM0URTw5lcyyLKQYzwOgzggk6VTmKcTuxCGKaG3jWLcNzho", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-156", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "U+yhkhv3TjM+ZY4Gg6qVQcgiFiuBpG6l6l36uZrYVbkk5O7FYp87XRoJvplSPxsMYUb7w4pPr3+sw8oOvvry9lDI0F5zVzMsiP+z0oBMJYSGUwrhPMq2PCT8t9LbFarHCKPXeGa1EuzRDXhaOKTFW07uTaM0URTw5lcyyLKQYzwOgzggk6VTmKcTuxCGKaG3jWLcNzhoRw9YC4KZSP4MPDPNA2mXceKRfdT8oNvy7lkYUvYZpIpVMv1mJoLtSbFYt676hay2TEosTCWaZVdvQO4d2mqZ1TNOe47Hlk/UXH/TJ5GaejDVNZeA5iafi+FD/0/8cI1l7XfKcSm5zEl8+/d18xoPV3KeU8wNV6v8vufyooJsKWOd4+fwFf/aTr/l3/+Z/4e//4ef85ptvZcpeCp3XGKz4UQVDPM6MDwdKk9BWxNPjHElTxN8v5F0iTon1dsNms6HrepxrmKcF7xf2ux2H5chuuRdZvn8CfxgI40yfNY0y/MVnX7NZbfj8s8+5urrm888+Y7u94GK7ZbXqaRoxYFS6Po3qPYsLpd63ehQfHWRyNSjNRUp6i8YoUU7z00xKmXcpYI2hPUGznJT5TdOINKTWHEdpJyzLQsrlfEieBLZPARZq2yhn0aeu/09XCJituOi+XWGtq3Awe87g9DnQIsnAbwVXXTNxdc5yP/bZ+RdXyagYuBzekZ6+5c/2v+Z2fIdrR4a0cJhn8c1DoZKBogm9+AB+ev8t3c5gvjFcPB74YhiYUyIZxV2EOUhlXZSi6JZSIqaEc3ut5ExOiXmeBSFS/ejOJBsKzqkqhenO1z6lzLIMKNvinGJOScZumyvaq6sf/Lg/GHSdPVm/SK8x5SyumeVcsEhZk7PomdagG0NAV6iYO9mBazmVVfUrLTkTQzg39NNJ3UgplEEA63Xi7ozBGk1j7Hk4JBAqySSapqFrW/q2oW8a8dmq/aimEeymXzzOvd9gcILSfFyASTnVrFMyC+usUAOrqlf", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-157", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "OJBsKzqkqhenO1z6lzLIMKNvinGJOScZumyvaq6sf/Lg/GHSdPVm/SK8x5SyumeVcsEhZk7PomdagG0NAV6iYO9mBazmVVfUrLTkTQzg39NNJ3UgplEEA63Xi7ozBGk1j7Hk4JBAqySSapqFrW/q2oW8a8dmq/aimEeymXzzOvd9gcILSfFyASTnVrFMyC+usUAOrqlfKAjQxWp8zXq2KUEutpus7upUwp+SBeb/RT1jZ8lsVvnr/wJ+OaMr7rPg0VK296ZOTb+FkCFM+KAWrDsE5qFa4d+3lnl5TPyLoQh1uti1Xl5d8/uln3N098PCwY0hi9xRnoVrqxlF0Ji2ROHmWo8E6OXzznMhzJk2RvAixoeQs/c6cRIg+BmLwLPPMUj21soJiFdpDVxyX/Zp10/L588+5uNjy5ZdfcXV1xcuXL+n7nq7rKt5U1/iqapBS5wQPOAe3969/5KqA/+gD4zASQhTETtU88N4Tja5tIEMuGWqLzNYhb0qp6vG+V9p7/+1/G/L14Z9PVF6FBEpnpX3gmkYgYcaeM9z3QUbWh8Oz0+9PhIs/NNP9cKn6nikZFz1dXLiIC5TIl71jLJkn7+sDIdUxRZycjYGtn7BFoUOmGRdWMdEiztk5ic60r8+OrS2SE/wMOB+Ky7LUoBvPeOVzW0/99pc8o2IRr1xL27Vs+4aLruXy2XMu/pCgK7P+uiEpxFw/dDqVrnKixpiYplFovYtn8oESIttG09n+7EGk0izJlYAFP+hWygBGOcfW2XPActbSuIa+bWmb5tzct1UBqVBw1tG2LUYbjKlkgRM2UWu6rqcUUaaPMeL9+2nuaTL7Metpv8cZw8WqR2uLcYZpXhj9Qno80g6eTz/ZopTi5vpC/J98xGhFVPDpl1/y7JNPuLx6RtusWJblt/q4p0EaQDSmHlg1i61oA4Um6XIO/IVzonvWB5ZrW874XAnysuk+JFOceiMxJhm0lY/Ho4Jk4cpo2lbzxRdfcn1", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-158", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "2LUYbjKlkgRM2UWu6rqcUUaaPMeL9+2nuaTL7Metpv8cZw8WqR2uLcYZpXhj9Qno80g6eTz/ZopTi5vpC/J98xGhFVPDpl1/y7JNPuLx6RtusWJblt/q4p0EaQDSmHlg1i61oA4Um6XIO/IVzonvWB5ZrW874XAnysuk+JFOceiMxJhm0lY/Ho4Jk4cpo2lbzxRdfcn19w8XFFS+ev+A//tX/y5u7d+y+vUNbw3q7wvQWOztBqoRUGYmGsmTinJjfecqY0QuMh4H7cofrGkL0dM4Rg+fw8MBhODAeH1ivN9zcPOOzr17y/PYZX37+OVfbC16+kNbBeis01t91zz/UHS6cgoC8dtrnp8D7UddEm3r/Pa9fveYv/9N/5ttvv2ea5ppNyVxCRUGoWGuJKZ6Dbtf1aGvOs4iTp1tB/Y4g+37OJ5lYpXPngjXCdFz1vQzMXF9RC6ds1nxwDU7v/RRsTxmuPme5p/VjrsnvXpKtt6mwioVPiuW63fAnX/+MeZrY7XYcDgeOxyMqW1Q2AgmkiAllrId0kGuzUopOKTSSIFnr6lmazpKn56BbRF/i9evX5z+DxAa572LrngtoI4PxkjMhROZ54uXzF/zsZz/jf/2f/y1/9pOf8unLF3Rd94Of9ocFb3L64MK/Bwe/P/kERmGMBrr/v71ze27jSs7479xmBgMQvFPX2Ka3rKxdqVSqso/531Nb+5TKc55SlU1srWXZokyJBDHAnEse+pyZISXZS8nlJzSLBQIkB5hz6dP9dffXWBcwVkpSG+9ZZDKaGEVBxgxeh2KdDaenxphsmTpRrHUlytdaS2WlK6jJlomkApEBbYO1bvhMxWIuC8JaQ0ryOQtr/mi93H95aGNAa/oobmdUkb6Pwm+byznl3kAcBUUwsomMMbi6oW4kD9dYI591OFTVaMTkEykqsttX7FLZGcWSL4eWzhSVaTjMSk+rEUYgZaUSxxO84H7iVpXk3vsPjFjdZcwdbTvnwYM", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-159", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "dDaenxphsmTpRrHUlytdaS2WlK6jJlomkApEBbYO1bvhMxWIuC8JaQ0ryOQtr/mi93H95aGNAa/oobmdUkb6Pwm+byznl3kAcBUUwsomMMbi6oW4kD9dYI591OFTVaMTkEykqsttX7FLZGcWSL4eWzhSVaTjMSk+rEUYgZaUSxxO84H7iVpXk3vsPjFjdZcwdbTvnwYMH9N7z4scfMdrww+tXeB/YXm+wXgh5TOWINpCqKPSf20TsIv2qJ3UR1UWsslR2LTyxua2PNYbD5T7trGH/eJ+9vSWnxyc8fPCA46Njzk5Ombcty/29oaPsu8Gf0XId4JXh5/L67bLY+0pKwnEcgmxUMWQNWmf3NnNd+MJ5kUBr6VBtjMH1Lvc0CwPUN8xwYgItlOejlasy7KXNWBRTih6mFuwtKxYGz+pdq3Ycu9+0/BcFyuCOzmi6juVaOo0YBW/fXHL1t++46p/z/OItJgpOm3JFXVBReBl8ImGIlWUdpVAhKolPKC2EWMRCjJVx6snBFUIcdEq5PynHlq7VSmkxHnNsqUCAzjraecuDR4/47MsvOTw8pKo+IZAWoh+wmyGfL2OlOkc1C/5XFkCMgRDEakqF5irXaPhQWlXqobqlGixYmXCrkFSdqhqYhUpppzJyvZgHxVqTF7TgMBLBtcPJXwYPwFiNtRFrw7AwS+njfaSetaSUWHkvCqQvrWESbSvM/Dq3MTEIJZzSClM5qqahblrquqWqGymASL7Et0bFmz/80MRhvJni8AzvQ9k0jBsjpnirpLNIyX4IPgzqJoSQmeTisFnvHzISq66Mu7Uaax3Pnj3ji/NzKlfzf99+y7//+c9cXl5yc3mFrR1sEnprcZuKVCeiSYReYIfNRUfoeuJNT+w9KvPDOqupK4M1jq+ffUXT1Jw+fsD+csnx8TGLxYJZm/k/QMi3J5trqkTKvY55rGUtlHWthnH5GBGynaxcUaAEV63rmt7rDDFsiDHkrrqKje5Zdxu", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-160", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "isFnvHzISq66Mu7Uaax3Pnj3ji/NzKlfzf99+y7//+c9cXl5yc3mFrR1sEnprcZuKVCeiSYReYIfNRUfoeuJNT+w9KvPDOqupK4M1jq+ffUXT1Jw+fsD+csnx8TGLxYJZm/k/QMi3J5trqkTKvY55rGUtlHWthnH5GBGynaxcUaAEV63rmt7rDDFsiDHkrrqKje5ZdxusNby9vsZZizWZrrNk8NyBE6aP0/VsjbSxH1uZZwy3BG8hZ8wwYrUwrKnb36MVXOQ3U7zakFzN/Jt/pfnDN+z/y7+hQsCkyPPvvuW///M/+K/Lv/CXi//FpX5IqUNBMNmYSIZZM2MxW3K17Vj3Pf12K+XfRryHqITMJmmdYx8l7308ZEwZ68xqqLUYEQV6KH8fgkB8dVOzPNjjy2fP+Oc//QmXg5K/JL8CL0wnNE1eK9ZTmQjBbFMSJi/nct1/PiHIeaW+/GtWEEWRj66MFHMIvJAzJ/LJkhAOBZSSsmTIf0PG+qSjaHGZCs9fsSCkLJRBSYdMzJzuqXS3G+lEu43FdZHyYlJks90SQuCHXEJztdrgKsNiIWz9dWVlwmIk+DCA8rJ0dbFjB6xgyKwYcKT8M2Mwg6xop5kNU29knL9yjdH1HDqdTp9/pJIpCl/r4UOKJZAST54+oW4aXr9+zaufXvHX//krffD0a89GdYRtoqo92kgGTPQB5aWKShsFKdL3HdvNmk23ZtaIEnl89oR2PuPw9EQghHaGrQwoiTzLUGYM7z2gScllLS7zVOne8jrgnQPs75GUD1GthRK1nbfUdZM7AQsUJu29Jaf6tjEgB+G2wGkwpE2GyRyVCsNbnt7kcHFZ4RZo7i5OW9ZQSVe8e71y73fz2VNK7/EcPlLy++tZi7KOZGvpWQeY1Rr2Dlm7hp+9BM6L35cUBC2WqDWOen7A/OET4s1b4fnubqTIIXohIY8pl7uX9V34JxQxQlVVPHzwaLKWSxWerI0XL/7Gtt+gjYxNVVW", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-161", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "7AQsUJu29Jaf6tjEgB+G2wGkwpE2GyRyVCsNbnt7kcHFZ4RZo7i5OW9ZQSVe8e71y73fz2VNK7/EcPlLy++tZi7KOZGvpWQeY1Rr2Dlm7hp+9BM6L35cUBC2WqDWOen7A/OET4s1b4fnubqTIIXohIY8pl7uX9V34JxQxQlVVPHzwaLKWSxWerI0XL/7Gtt+gjYxNVVW4yuHqSrpU5y9R7R/ZDXhUupGU1OQxSRIx5A0+shoppbFGvo2WnD8y90FQ3Aqec2dzC1rJgNsKlVrm0TUaWzmU0qw7qdCp62pYqFLaG6mqZui3lVJivV5L+pb3eTFpaUvic0J+uJ/SXXdbaXVNorKKtrZoC0RY3XT4kLi+usbHxJuV5+iwZblscc7S1BUKwe96HwZrTPgPClBfNsX0uRoODMibwBQvgswYNXX9JBApczdhj8pdUn1ubf2+oMvHFkcUPLpQ/U034xeff86jhw8hRV58/wPd9ZrLyzf89OqCrrsh6Q5XbdDGEJIcWE7lteAMENhuO7q1tKlZLlusazn/8nMWiwXz/SWlRlS4u8IkwKiEHOiO0h2x2njrtfF2Mr8w4xzct5AmZVjHGENd1yyXe7z++S3O1WhliCZmUhopVghBytNlfiLbjcAMLivMck0f47B3jLWDBwqlDDgr1aIUjFjLqDEDoQRiyzyV2MG0GOLW7N49wCfjcl/50P+YWQtNgrlAf9oY9GpN2j/mxrW8li17K7umRzDr+cxxsjzi4IsvUW9eYa7foH76gdBFVBTWr5gYPAU1iZamJDwq1tQ8fXKeLVszHFbSEmrDyx9/xIc1Tgsc6nIQraor0FJar1NAxU9QupLjJxkEonBLPmcaFqtScvFblStZCQQfCCmQfCQpRTDivmgYErULlmbyokhKE1IiFLpGBdvgiT6h+y0JlfMUFX3wguk6C1pjcgWazz2iRIn4fLpFjMquf1AZnrBEfU/2KKNwWjMvHfmC4ma9ZbvZDpkDEvx", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-162", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "eLVszHFbSEmrDyx9/xIc1Tgsc6nIQraor0FJar1NAxU9QupLjJxkEonBLPmcaFqtScvFblStZCQQfCCmQfCQpRTDivmgYErULlmbyokhKE1IiFLpGBdvgiT6h+y0JlfMUFX3wguk6C1pjcgWazz2iRIn4fLpFjMquf1AZnrBEfU/2KKNwWjMvHfmC4ma9ZbvZDpkDEvxQgvsqzaypaSpH5QxSfCZWdsjurSpBDRn1cZMPyrdsOJV5ULUQnivhjFSI/r6Lu8n8FGsRYk7yJmO5pdLplvINH2fpvs+FH/Bmo6nqivPzLzg+PqJyjovXr/nuu79xfbPm5uZG2mYnsM5Q1xVPHz2icpbGWa5X17y9esvp2UP29w85P/+co4MDlvtLXIahSGL5FG9heg+JdwM+osQMaYJfF8Ulh1whOhmtvfsrmDGdazFv+fyzp6xWHVdXa6EQDFJpGELAaOFc8MESgxRDFPxQZwWhZOqw2g6ufiExKj3NCj9ySTlzVpq4Cl2ozl6UjJJVpbSVIelfT+8ze1Ufuu9PwbrvippgrNMMjbquOTo5Y3lwxKxdsu5WuQecBMgen55xdHTMs6/+yHK5z8HBIbO2Yt3tE7Xm6uotL16+YL1e8+bNm1t2nhgJLnvkjqqaQbJY0zCbtdRNTVPX/Hx5wWazISVPjD19L62yQuhZr29Y3ax4dfEj33//nHa+kKDl/MkH7/VXlG66s4VHmGEccBmgKfZTsEfpTZYISpRuzJFxy6gahuCBkkUV0ZDL61QmLemDVKlJ5FYsA4nGJwmQ2DGnsFTpFKA7hEgi5gNED4tVrHN9V0f9qkhATDGvDb2HjU9st5GbdY81UqmStBHLNInCrCo3pKuJwZLwIfd8ywpBLMw0uG0wVQJlo4zBM5Uy/JCmeZSTuZh8ZnEhR0yu4FYqu7TFyp3CEPeVu0p38uZ5bhUnp8fsLRdYa3n16oJm1vLzz5eC83YbfJRKxXbe8sd//IpZU7NoGy4", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-163", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "pFKA7hEgi5gNED4tVrHN9V0f9qkhATDGvDb2HjU9st5GbdY81UqmStBHLNInCrCo3pKuJwZLwIfd8ywpBLMw0uG0wVQJlo4zBM5Uy/JCmeZSTuZh8ZnEhR0yu4FYqu7TFyp3CEPeVu0p38uZ5bhUnp8fsLRdYa3n16oJm1vLzz5eC83YbfJRKxXbe8sd//IpZU7NoGy4uLvjh5UsWe/u07YKz01OOjg5o2plgbQV6SQyHkEoTWEy9CxWAeGejIXFbsZb5m75+f3ghHwYpCfZ8eszy+UvaZoZC473JWQmifGOUxP6gDSb6gSkt5SrBAgXYSY52yeU2uZKtENaIsnU5LaxwJowKVykwea+q4VENkFYa9sn7Ld7pnH+qyCEn6ZvDOqJY8o753lKqN5uWru+IvgcU2lgOD494/PgJ33zzT+OYWEW9bYSQZnXN9eqGFOAyvpkYBbIsrLUY4zC6oa5nKGUxpqKu57IW25rr67cochWaGS3+Er/yvme1WnH55pKYIs59QiDNZaZ8ibYadJ5Ahc6KDLyPKFUCNwWQl4BSDFFJADt5AAAD1ElEQVS4TMuk+hwIUyN+pPKGFKJgndvwJLa+F0WhFd5L4UXvJanZubyYsMQgZZvFFS/7rHQ1DWEMGiUCqC0+E+f44O/tTp8ctMJmf7Oh67a8XXWQAm0Vqa3GKPBxg1GGx4ctp/st86ZlPm9p562coHUFA/1bAvTQ6SFRcn3F05BHsXTE+kpDvE0WnijozU1H8IFtJlnXRmeWKDsoD0mXGssZi3IvubxS/qmG9LP7yBTCGA7REvYr1pQxNLOGh48fcnRyzNPP/oGu29B1Hd1mQ4iBWSMFC8dHhzhjcM7QdRtubm5wrsZax97+nkSI1ZiXPHyOkoAxKWeeHtTlwCmbZopNlvVTYIFB2fJhl/iXx0QOe5UibVPz8OyEz54+wG82PH/xkuvVGm1mhBBwmcTHezsEj0OBP+LtQ6FkIwxWbrEd8v0ZLbn", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-164", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "CGA7REvYr1pQxNLOGh48fcnRyzNPP/oGu29B1Hd1mQ4iBWSMFC8dHhzhjcM7QdRtubm5wrsZax97+nkSI1ZiXPHyOkoAxKWeeHtTlwCmbZopNlvVTYIFB2fJhl/iXx0QOe5UibVPz8OyEz54+wG82PH/xkuvVGm1mhBBwmcTHezsEj0OBP+LtQ6FkIwxWbrEd8v0ZLbnr1hjx7qSMFKszR27m063y2pEto8TTzM9ieW0iHwsn/L0y4MvZaFqtOrZbj7EVi70Dzs4es9n2bDZRcFw7Z3/5gL3FCc7NcxrXhuAdOinmM4VKNbVbYG2PqLsJRKcVjx4+Yn//gKdPz2maGbN6j8HKbgyu0pyeHXNwuKCdGza5yafWorDPz8/5+utvePr0Kcvlkr29vSG/+kPyi7/VQ711cdoGcIAC/BeIIKWQLYc8ebFYl8N9ymlNdvfkzicLO1+TEXspuG+MMfd9kiwFax1jpB9CiNmCHK0USR1L2YUXiUkUevA5wyLEdzbtr0ltDT5BH3IakPc4HbEmYbV05Q25g0FVGWpnM4GIw7lqEji8HZxMOSm+3ICkkZXiiZSx9IxJDWEi0Jkweb3u8H3Per2WbAljqFwlSrdMduZETUksIm1z+2vyWw8Y8r2G5JbcTdYfrpWvrZMUiFR1zXw+Z7vt6fuebrshBFG6zjkW7UxYrrSmbT17e3toLVi9rSR/+d2pm+CO01+mscru3fzWD8x/4h0ei08RoUusaWcN83kjhQpaEQsTWg4Oi8IPBdAnRV2wuEFRFGtWFKwqoYH8PjnPfcCBy/4VY0nnaxilxNItew41WY15NN9j7U6x3d9c/Sol86ZyXYAPYoAoyYap6watLeLViSEoB3GNpIJKdxuSQEfWVDjr0dpmGFQzGjuCazdNw2Kx4OTkhKZuIJkhvc9o2Qt1XeGc4fj4BO+3+LDJWQ2Wk5MTDg8Padt2CFr+mqWrfktcZic72clOdvLL8mmtPHeyk53sZCf", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-165", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "By/4VY0nnaxilxNItew41WY15NN9j7U6x3d9c/Sol86ZyXYAPYoAoyYap6watLeLViSEoB3GNpIJKdxuSQEfWVDjr0dpmGFQzGjuCazdNw2Kx4OTkhKZuIJkhvc9o2Qt1XeGc4fj4BO+3+LDJWQ2Wk5MTDg8Padt2CFr+mqWrfktcZic72clOdvLL8mmtPHeyk53sZCf3kp3S3clOdrKT31F2SncnO9nJTn5H2SndnexkJzv5HWWndHeyk53s5HeUndLdyU52spPfUf4fUbyyl6yAM5IAAAAASUVORK5CYII=)\n\nPython implementation:\nimport matplotlib.pyplot as plt\n\nk = 0\nfor images, labels in train_loader:\n # since batch_size = 1, there is only 1 image in `images`\n image = images[0]\n # place the colour channel at the end, instead of at the beginning\n img = np.transpose(image, [1,2,0])\n # normalize pixel intensity values to [0, 1]\n img = img / 2 + 0.5\n plt.subplot(3, 5, k+1)\n plt.axis('off')\n plt.imshow(img)\n\n k += 1\n if k > 14:\n break", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "183-166", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\n\n$\\color{blue}{\\text{ : }}$\n- number of training examples= 8000\n- number of validation examples= 2000\n- number of test examples= 2000", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "184-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\n\n$\\color{blue}{\\text{: }}$\n- We change the weights by training the neural\n network with the training data set. To avoid overfitting the network and fine-tune models, we must feed the validation set into the network and check if the error falls within a certain range (This set is not being using directly to adjust the weights but used to give the optimal number of hidden units or determine a stopping point for the back-propagation algorithm)\n\n- Validation set is different from test set. Validation set actually can be regarded as a part of training set, because it is used to build the model, neural networks or others. Validation set actually can be regarded as a part of training set, because it is used to build the model, neural networks or others. It is usually used for parameter selection and to avoild overfitting. ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "185-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "et actually can be regarded as a part of training set, because it is used to build the model, neural networks or others. It is usually used for parameter selection and to avoild overfitting. \n\n- If a non-linear model (such as NN) is only trained on a training set without using validation set to evaluate the model, it is very likely to achieve 100% accuracy and overfit, resulting in poor performance on the test set.", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "185-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n## Part 2. Training the Classifier\n\nWe define two neural networks, a `LargeNet` and `SmallNet`.\nWe'll be training the networks in this section.\n\nYou won't understand fully what these networks are doing until\nthe next few classes, and that's okay. For this assignment, please\nfocus on learning how to train networks, and how hyperparameters affect\ntraining.\n\nPython implementation:\nclass LargeNet(nn.Module):\n def __init__(self):\n super(LargeNet, self).__init__()\n self.name = \"large\"\n self.conv1 = nn.Conv2d(3, 5, 5)\n self.pool = nn.MaxPool2d(2, 2)\n self.conv2 = nn.Conv2d(5, 10, 5)\n self.fc1 = nn.Linear(10 * 5 * 5, 32)\n self.fc2 = nn.Linear(32, 1)\n\n def forward(self, x):\n x = self.pool(F.relu(self.conv1(x)))\n x = self.pool(F.relu(self.conv2(x)))\n x = x.view(-1, 10 * 5 * 5)\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n x = x.squeeze(1) # Flatten to [batch_size]\n return x\n\nPython implementation:\nclass SmallNet(nn.Module):\n def __init__(self):\n super(SmallNet, self).__init__()", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "186-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 2. Training the Classifier" + }, + { + "text": "= F.relu(self.fc1(x))\n x = self.fc2(x)\n x = x.squeeze(1) # Flatten to [batch_size]\n return x\n\nPython implementation:\nclass SmallNet(nn.Module):\n def __init__(self):\n super(SmallNet, self).__init__()\n self.name = \"small\"\n self.conv = nn.Conv2d(3, 5, 3)\n self.pool = nn.MaxPool2d(2, 2)\n self.fc = nn.Linear(5 * 7 * 7, 1)\n\n def forward(self, x):\n x = self.pool(F.relu(self.conv(x)))\n x = self.pool(x)\n x = x.view(-1, 5 * 7 * 7)\n x = self.fc(x)\n x = x.squeeze(1) # Flatten to [batch_size]\n return x\n\nPython implementation:\nsmall_net = SmallNet()\nlarge_net = LargeNet()", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "186-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 2. Training the Classifier" + }, + { + "text": "Explanation:\n### Part (a)\n\nThe methods `small_net.parameters()` and `large_net.parameters()`\nproduces an iterator of all the trainable parameters of the network.\nThese parameters are torch tensors containing many scalar values.\n\nWe haven't learned how how the parameters in these high-dimensional\ntensors will be used, but we should be able to count the number\nof parameters. Measuring the number of parameters in a network is\none way of measuring the \"size\" of a network.\n\n$\\color{blue}{\\text{ }}$\n- By adding the number of parameters, we can calculate the total number of parameters:\n\n- total number of parameters in `small_net`= 5\\*3\\*3\\*3+5+245+1=$386$ \n\n- total number of parameters in `large_net`= 5\\*3\\*5\\*5+5+10\\*5\\*5\\*5+10+32\\*250+32+32\\*1+1=$9705$\n\n\nPython implementation:\nprint(\"The size of tensors in small_net:\")\nfor param in small_net.parameters():\n print(param.shape)\nprint(\"The size of tensors in large_net:\")", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "187-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "2\\*250+32+32\\*1+1=$9705$\n\n\nPython implementation:\nprint(\"The size of tensors in small_net:\")\nfor param in small_net.parameters():\n print(param.shape)\nprint(\"The size of tensors in large_net:\")\nfor param in large_net.parameters():\n print(param.shape)\n\nExplanation:\n$\\color{blue}{\\text{There is a another way for calculating number of parameters that is easier:}}$\n- By using `.numel()` number of parameters in each tensor can be calculated.\nTotal number of parameters in small_net= $386$\nTotal number of parameters in large_net= $9705$\n\n\nPython implementation:\nN_p_small=0\nfor param in small_net.parameters():\n N_p_small+=param.numel()\nprint('Total number of parameters in small_net= ', N_p_small)\nN_p_large=0\nfor param in large_net.parameters():\n N_p_large+=param.numel()\nprint('Total number of parameters in large_net= ', N_p_large)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "187-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### The function train_net\n\nThe function `train_net` below takes an untrained neural network (like `small_net` and `large_net`) and\nseveral other parameters. You should be able to understand how this function works.\nThe figure below shows the high level training loop for a machine learning model:\n\n![alt text](https://github.com/UTNeural/Lab2/blob/master/Diagram.png?raw=true)\n\nPython implementation:\ndef train_net(net, batch_size=64, learning_rate=0.01, num_epochs=30):\n ########################################################################\n # Train a classifier on cars vs trucks\n target_classes = [\"car\", \"truck\"]\n ########################################################################\n # Fixed PyTorch random seed for reproducible result\n torch.manual_seed(1000)\n ########################################################################\n # Obtain the PyTorch data loader objects to load batches of the datasets", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "188-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "The function train_net" + }, + { + "text": "ed for reproducible result\n torch.manual_seed(1000)\n ########################################################################\n # Obtain the PyTorch data loader objects to load batches of the datasets\n train_loader, val_loader, test_loader, classes = get_data_loader(\n target_classes, batch_size)\n ########################################################################\n # Define the Loss function and optimizer\n # The loss function will be Binary Cross Entropy (BCE). In this case we\n # will use the BCEWithLogitsLoss which takes unnormalized output from\n # the neural network and scalar label.\n # Optimizer will be SGD with Momentum.\n criterion = nn.BCEWithLogitsLoss()\n optimizer = optim.SGD(net.parameters(), lr=learning_rate, momentum=0.9)\n ########################################################################\n # Set up some numpy arrays to store the training/test loss/erruracy\n train_err = np.zeros(num_epochs)\n train_loss = np.zeros(num_epochs)\n val_err = np.zeros(num_epochs)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "188-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "The function train_net" + }, + { + "text": "#############################\n # Set up some numpy arrays to store the training/test loss/erruracy\n train_err = np.zeros(num_epochs)\n train_loss = np.zeros(num_epochs)\n val_err = np.zeros(num_epochs)\n val_loss = np.zeros(num_epochs)\n ########################################################################\n # Train the network\n # Loop over the data iterator and sample a new batch of training data\n # Get the output from the network, and optimize our loss function.\n start_time = time.time()\n for epoch in range(num_epochs): # loop over the dataset multiple times\n total_train_loss = 0.0\n total_train_err = 0.0\n total_epoch = 0\n for i, data in enumerate(train_loader, 0):\n # Get the inputs\n inputs, labels = data\n labels = normalize_label(labels) # Convert labels to 0/1\n # Zero the parameter gradients\n optimizer.zero_grad()\n # Forward pass, backward pass, and optimize\n outputs = net(inputs)\n loss = criterion(outputs, labels.float())\n loss.backward()\n optimizer.step()", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "188-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "The function train_net" + }, + { + "text": "# Zero the parameter gradients\n optimizer.zero_grad()\n # Forward pass, backward pass, and optimize\n outputs = net(inputs)\n loss = criterion(outputs, labels.float())\n loss.backward()\n optimizer.step()\n # Calculate the statistics\n corr = (outputs > 0.0).squeeze().long() != labels\n total_train_err += int(corr.sum())\n total_train_loss += loss.item()\n total_epoch += len(labels)\n train_err[epoch] = float(total_train_err) / total_epoch\n train_loss[epoch] = float(total_train_loss) / (i+1)\n val_err[epoch], val_loss[epoch] = evaluate(net, val_loader, criterion)\n print((\"Epoch {}: Train err: {}, Train loss: {} |\"+\n \"Validation err: {}, Validation loss: {}\").format(\n epoch + 1,\n train_err[epoch],\n train_loss[epoch],\n val_err[epoch],\n val_loss[epoch]))\n # Save the current model (checkpoint) to a file\n model_path = get_model_name(net.name, batch_size, learning_rate, epoch)\n torch.save(net.state_dict(), model_path)\n print('Finished Training')\n end_time = time.time()", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "188-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "The function train_net" + }, + { + "text": "ent model (checkpoint) to a file\n model_path = get_model_name(net.name, batch_size, learning_rate, epoch)\n torch.save(net.state_dict(), model_path)\n print('Finished Training')\n end_time = time.time()\n elapsed_time = end_time - start_time\n print(\"Total time elapsed: {:.2f} seconds\".format(elapsed_time))\n # Write the train/test loss/err into CSV file for plotting later\n epochs = np.arange(1, num_epochs + 1)\n np.savetxt(\"{}_train_err.csv\".format(model_path), train_err)\n np.savetxt(\"{}_train_loss.csv\".format(model_path), train_loss)\n np.savetxt(\"{}_val_err.csv\".format(model_path), val_err)\n np.savetxt(\"{}_val_loss.csv\".format(model_path), val_loss)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "188-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "The function train_net" + }, + { + "text": "Explanation:\n### Part (b)\n\nThe parameters to the function `train_net` are hyperparameters of our neural network.\nWe made these hyperparameters easy to modify so that we can tune them later on.\n\n$\\color{blue}{\\text{}}$\n- The default values of the parameters are as following: \nbatch_size=64\nlearning_rate=0.01 \nnum_epochs=30 ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "189-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\n\n$\\color{blue}{\\text{}}$\n\nBy looking at the **`train_net`** code we can see that with following code for each ephoc optimized parameters of the model would be saved to disk:\n`torch.save(net.state_dict(), model_path)`\n\nIt is obvious this line of code save parameters of the model (weights and biases) in an file and the name of the files depends on model's `name, batch_size, learning_rate, and epoch`. For example, for epoch=4 the name of the file would be `model_small_bs64_lr0.01_epoch4`. So, 5 files would be save and the name of them would be as following:\n- model_small_bs64_lr0.01_epoch0 \n- model_small_bs64_lr0.01_epoch1 \n- model_small_bs64_lr0.01_epoch2 \n- model_small_bs64_lr0.01_epoch3 \n- model_small_bs64_lr0.01_epoch4\n- model_small_bs64_lr0.01_epoch4_train_err ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "190-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "r0.01_epoch2 \n- model_small_bs64_lr0.01_epoch3 \n- model_small_bs64_lr0.01_epoch4\n- model_small_bs64_lr0.01_epoch4_train_err \n- model_small_bs64_lr0.01_epoch4_train_loss \n- model_small_bs64_lr0.01_epoch4_val_err \n- model_small_bs64_lr0.01_epoch4_val_loss \n\nMoreover, We can add following code after `torch.save(net.state_dict(), model_path)` to see what would be save in these files: \n\n`torch.save(net.state_dict(), model_path)`\n`for param_tensor in net.state_dict():`\n` ---print(param_tensor, \"\\t\", net.state_dict()[param_tensor].size())`\n\n\nBy doing so, we can see that following parameters would be save in every epoch in every file:\n`conv.weight torch.Size([5, 3, 3, 3])`\n`conv.bias torch.Size([5])`\n`fc.weight torch.Size([1, 245])`\n`fc.bias torch.Size([1]` \n\nPython implementation:\n# Initialize model", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "190-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "ery epoch in every file:\n`conv.weight torch.Size([5, 3, 3, 3])`\n`conv.bias torch.Size([5])`\n`fc.weight torch.Size([1, 245])`\n`fc.bias torch.Size([1]` \n\nPython implementation:\n# Initialize model\n#model = SmallNet()\n# Initialize optimizer\n#optimizer = optim.SGD(model.parameters(), epoch=29)\n#small_net=train_net(model, num_epochs=5)\n# Print model's state_dict", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "190-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (d)\n\nTrain both `small_net` and `large_net` using the function `train_net` and its default parameters.\nThe function will write many files to disk, including a model checkpoint (saved values of model weights)\nat the end of each epoch.\n\nIf you are using Google Colab, you will need to mount Google Drive\nso that the files generated by `train_net` gets saved. We will be using\nthese files in part (d).\n(See the Google Colab tutorial for more information about this.)\n\n$\\color{blue}{\\text{answer : }}$\n\nThe total time elapsed for `large_net`= $111 s$\nThe total time elapsed for `small_net`= $98 s$\n\nThe `large_net` network took longer to train because it has more layers and more paramerets in training process.\n\nPython implementation:\n# Since the function writes files to disk, you will need to mount\n# your Google Drive. If you are working on the lab locally, you\n# can comment out this code.", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "191-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (d)" + }, + { + "text": "process.\n\nPython implementation:\n# Since the function writes files to disk, you will need to mount\n# your Google Drive. If you are working on the lab locally, you\n# can comment out this code.\n\nfrom google.colab import drive\ndrive.mount('/content/gdrive')", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "191-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n### Part (e)\n\nUse the function `plot_training_curve` to display the trajectory of the\ntraining/validation error and the training/validation loss.\nYou will need to use the function `get_model_name` to generate the\nargument to the `plot_training_curve` function.\n\nDoing this for both the small network and the large network.\n\n$\\color{blue}{\\text{ }}$\n\nplots of error and loss for `small_net` and `large_net` have been added:\n\nPython implementation:\n#model_path = get_model_name(\"small\", batch_size=??, learning_rate=??, epoch=29)\nmodel_path = get_model_name(\"small\", batch_size=64, learning_rate=0.01, epoch=29)\n\nplot_training_curve(model_path)\n\nPython implementation:\nmodel_path = get_model_name(\"large\", batch_size=64, learning_rate=0.01, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "192-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (e)" + }, + { + "text": "Explanation:\n### Part (f)\n\n$\\color{blue}{\\text{}}$\n\n- By comparing these curves for `small_net` and `large_net`, we can say that when epoch is bigger that 15 error for `large_net` is smaller than error for `small_net`. So, `large_net` has a better accuracy\n- For both curves, we can see that when epoch is less than 5, underfitting has happened\n- Overfitting has occurred for 'large net' when epoch is more than 20 as the loss has ascended for validation dataset \n\nSo, the curves for `small_net` and `large_net` are different regarding accuracy and presence of overfitting ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "193-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (f)" + }, + { + "text": "Explanation:\n## Part 3. Optimization Parameters\nFor this section, we will work with `large_net` only.", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "194-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 3. Optimization Parameters" + }, + { + "text": "Explanation:\n### Part (a)\n\nTrain `large_net` with all default parameters, except set `learning_rate=0.001`.\nDoes the model take longer/shorter to train?\nPlot the training curve. Describe the effect of *lowering* the learning rate.\n\n$\\color{blue}{\\text{Answer: }}$\n\n- It takes longer to train the model because by using smaller `learning_rate`, our weights would be updated with smaller step size. \n- By lowering the learning rate underfitting has happened. **The corves indicate that the model is capable of more learning and possible enhancements, and that the training process was stopped too soon. The model couldn't learn the training dataset properly.**\n\nPython implementation:\n# Note: When we re-construct the model, we start the training\n# with *random weights*. If we omit this code, the values of\n# the weights will still be the previously trained values.\nlarge_net = LargeNet()\nL_net=train_net(large_net, learning_rate=0.001)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "195-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "raining\n# with *random weights*. If we omit this code, the values of\n# the weights will still be the previously trained values.\nlarge_net = LargeNet()\nL_net=train_net(large_net, learning_rate=0.001)\n\nPython implementation:\nmodel_path = get_model_name(\"large\", batch_size=64, learning_rate=0.001, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "195-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\n\n$\\color{blue}{\\text{ }}$\n\n- It takes shorter to train the model because by increasing `learning_rate`, our weights would be updated with bigger step size. \n- By increasing the learning rate, at first overfitting happened for smaller epoch number and then underfitting happened for bigger epoch number. **The curves indicate that the model was unable to learn the training dataset at all.**\n\nPython implementation:\nlarge_net = LargeNet()\nL_net=train_net(large_net, learning_rate=0.1)\n\nPython implementation:\nmodel_path = get_model_name(\"large\", batch_size=64, learning_rate=0.1, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "196-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\n\n$\\color{blue}{\\text{}}$\n\n- It takes shorter to train the model because by increasing the batch size, as it is equivalent to taking a few big steps, instead of taking many little steps.\n- By increasing the batch size underfitting has happened. **The corves indicate that the model is capable of more learning and possible enhancements, and that the training process was stopped too soon. The model couldn't learn the training dataset properly.**\n\nPython implementation:\nlarge_net = LargeNet()\nL_net=train_net(large_net, batch_size=512)\n\nPython implementation:\nmodel_path = get_model_name(\"large\", batch_size=512, learning_rate=0.01, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "197-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (d)\n\n$\\color{blue}{\\text{}}$\n\n- It takes longer to train the model as decreasing the batch size is equivalent to taking many little steps.\n- By decreasing the batch size, underfitting happened and the erorr was hight. **The curves indicate that the model was unable to learn the training dataset at all.**\n\nPython implementation:\nlarge_net = LargeNet()\nL_net=train_net(large_net, batch_size=16)\n\nPython implementation:\nmodel_path = get_model_name(\"large\", batch_size=16, learning_rate=0.01, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "198-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n## Part 4. Hyperparameter Search", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "199-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 4. Hyperparameter Search" + }, + { + "text": "Explanation:\n### Part (a)\n$\\color{blue}{\\text{ }}$\n\n- I'm going to choose `large_net` becuase the error for this network is lower than the other one. \n- `batch_size=256` has been chosen because when batch size was lower than 64, the model couldn't learn at all. Also, batch_size=512 was so high for model to learn properly. \n\n- `learning_rate=0.005` has been chosen because when learning_rate was bigger than 0.01, the model couldn't learn at all and learning_rate=0.001 was too small and Underfitting has happened. ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "200-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\nTrain the model with the hyperparameters we found chose in part(a), and include the training curve.\n\nPython implementation:\nlarge_net = LargeNet()\nLa_net=train_net(large_net, batch_size=256, learning_rate=0.005)\n\nPython implementation:\nmodel_path = get_model_name(\"large\", batch_size=256, learning_rate=0.005, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "201-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\n\n$\\color{blue}{\\text{}}$\n\n- In this part, I'm going to choose `large_net` again becuase the error for this network is lower than the other one. \n- `batch_size=128` has been chosen because when `batch_size=256`, Underfitting has happened and by decreasing batch_size we may get a better results.\n\n- `learning_rate=0.0075` has been chosen because when `learning_rate=0.005`,Underfitting has happened and by increasing learning_rate we may get a better results as bigger steps would be used ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "202-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (d)\n\nTraining the model with the hyperparameters we chose in part(c), and include the training curve.\n\nPython implementation:\nlarge_net = LargeNet()\nLa_net=train_net(large_net, batch_size=128, learning_rate=0.0075)\n\nPython implementation:\nmodel_path = get_model_name(\"large\", batch_size=128, learning_rate=0.0075, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "203-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n## Part 4. Evaluating the Best Model", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "204-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 4. Evaluating the Best Model" + }, + { + "text": "Explanation:\n### Part (a)\n\nChoosing the **best** model that you have so far. This means choosing the best model checkpoint,\nincluding the choice of `small_net` vs `large_net`, the `batch_size`, `learning_rate`,\n**and the epoch number**.\n\nPython implementation:\nnet = LargeNet()\nmodel_path = get_model_name(net.name, batch_size=128, learning_rate=0.0075, epoch=27)\nstate = torch.load(model_path)\nnet.load_state_dict(state)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "205-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\n$\\color{blue}{\\text{ }}$\n\n- By comparing the results for all the networks and all htperparameters, we can see that `large_net`with\n`batch_size=128`,`learning_rate=0.0075`, and `epoch=27`\nhas the lowest erorr and lowest loss. So, bacause of low erorr and loss for train and validation set, this model has been chosen. ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "206-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\n\n$\\color{blue}{\\text{ }}$\n\n- test classification error=0.1415 \n\nPython implementation:\n# If you use the `evaluate` function provided in part 0, you will need to\n# set batch_size > 1\n\ntrain_loader, val_loader, test_loader, classes = get_data_loader(\n target_classes=[\"car\", \"truck\"],\n batch_size=64)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "207-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (d)\n\n$\\color{blue}{\\text{Answer: }}$\n\n Test clasification error was higher than validation error as we expected. We would expect test error to be higher because we use validation set for parameter tunning. So, our model is biased toward validation set. ", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "208-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n### Part (e)\n\n$\\color{blue}{\\text{ }}$\n\n- The main reason was that we wanted to have an unbiased stimation of preformance of the model \n\n- Using the test data as little as possible is important because the step of “choosing the best model” (based on validation performance) can cause a form of overfitting. Think about it this way: let’s say we tried a THOUSAND different models or model variations on our data, and we have validation set performance for all of them. The act of choosing the model with the best validation set performance inherently means that we, the human, have “tuned” the model details for the validation set. The performance value we see for the validation set on “the best model on the validation set” is inherently inflated.", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "209-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (e)" + }, + { + "text": "y means that we, the human, have “tuned” the model details for the validation set. The performance value we see for the validation set on “the best model on the validation set” is inherently inflated. To get a non-inflated and more reliable estimate of how well this “best model” will do on data it’s never seen before, we need to use more data it’s never seen before! This is the test set. The test set performance will typically be slightly lower than the validation set performance.\n\n- \nThe validation dataset is different from the test dataset that is also held back from the training of the model. So test set is used to give an unbiased estimate of the skill of the final tuned model.\n\n- A validation dataset is a sample of data held back from training your model that is used to give an estimate of model skill while tuning model’s hyperparameters and we use them during tunning.", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "209-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part (e)" + }, + { + "text": "Explanation:\n## Part 5. Fully-Connected Linear ANN vs CNN\n\n$\\color{blue}{\\text{ }}$\n\n- \n\" After exploring different hyperparameter settings we can see that model with\n`batch_size=512`, `learning_rate=0.005`, `epoch=29` is the best one. The error of this model on test set is 0.259. \n\n- Based on error for our test and validation set, our best CNN model was so better than the 2-layer linear ANN model (no convolutional layers) on classifying car and truck images- \n\n Error of the model for test set using ANN model = 0.259 \n\n Error of the model for test set using CNN model = 0.1415 \n\nPython implementation:\nclass simpleANN(nn.Module):\n def __init__(self):\n super(simpleANN, self).__init__()\n self.name = \"simple\"\n self.fc1 = nn.Linear(32*32*3, 100)\n self.fc2 = nn.Linear(100, 20)\n self.fc3 = nn.Linear(20, 1)\n\n def forward(self, x):\n x = x.view(-1, 32*32*3)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "210-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 5. Fully-Connected Linear ANN vs CNN" + }, + { + "text": "uper(simpleANN, self).__init__()\n self.name = \"simple\"\n self.fc1 = nn.Linear(32*32*3, 100)\n self.fc2 = nn.Linear(100, 20)\n self.fc3 = nn.Linear(20, 1)\n\n def forward(self, x):\n x = x.view(-1, 32*32*3)\n x = F.relu(self.fc1(x))\n x = F.relu(self.fc2(x))\n x = self.fc3(x)\n x = x.squeeze(1) # Flatten to [batch_size]\n return x", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "210-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Part 5. Fully-Connected Linear ANN vs CNN" + }, + { + "text": "Explanation:\n#### Tunning\n\nExplanation:\n- At first, we model simpleANN using the function train_net and its default parameters. \n\n- As we cal see overfitting has happened \n\nPython implementation:\nsimple_net = simpleANN()\nL_net=train_net(simple_net)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=64, learning_rate=0.01, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:\n- We are going to try different hyperparameters \n\n- In this part, we use `learning_rate=0.001` and other default parameters. \n\nPython implementation:\nsimple_net = simpleANN()\nL_net=train_net(simple_net, learning_rate=0.001)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=64, learning_rate=0.001, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "211-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Tunning" + }, + { + "text": "t=train_net(simple_net, learning_rate=0.001)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=64, learning_rate=0.001, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:\n- In this part, we use ` batch_size=512` and other default parameters\n\nPython implementation:\nsimple_net = simpleANN()\nL_net=train_net(simple_net, batch_size=512)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=512, learning_rate=0.01, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:\n- In this part, we use ` batch_size=256` and `learning_rate=0.005`\n\nPython implementation:\nsimple_net = simpleANN()\nL_net=train_net(simple_net, batch_size=256, learning_rate=0.005)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=256, learning_rate=0.005, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "211-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Tunning" + }, + { + "text": "e_net, batch_size=256, learning_rate=0.005)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=256, learning_rate=0.005, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:\n- In this part, we use `learning_rate=0.1`\n\nPython implementation:\nsimple_net = simpleANN()\nL_net=train_net(simple_net, learning_rate=0.1)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=64, learning_rate=0.1, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:\n- In this part, we use ` batch_size=512` and `learning_rate=0.001`\n\nPython implementation:\nsimple_net = simpleANN()\nL_net=train_net(simple_net, batch_size=512, learning_rate=0.001)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=512, learning_rate=0.001, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "211-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Tunning" + }, + { + "text": "e_net, batch_size=512, learning_rate=0.001)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=512, learning_rate=0.001, epoch=29)\n\nplot_training_curve(model_path)\n\nExplanation:\n- In this part, we use ` batch_size=512` and `learning_rate=0.005`\n\nPython implementation:\nsimple_net = simpleANN()\nL_net=train_net(simple_net, batch_size=512, learning_rate=0.005)\n\nPython implementation:\nmodel_path = get_model_name(\"simple\", batch_size=512, learning_rate=0.005, epoch=29)\n\nplot_training_curve(model_path)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "211-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Tunning" + }, + { + "text": "Explanation:\n#### Evaluating the Best Model\n\nPython implementation:\nnet = simpleANN()\nmodel_path = get_model_name(net.name, batch_size=512, learning_rate=0.005, epoch=29)\nstate = torch.load(model_path)\nnet.load_state_dict(state)\n\nPython implementation:\n# If you use the `evaluate` function provided in part 0, you will need to\n# set batch_size > 1\n\ntrain_loader, val_loader, test_loader, classes = get_data_loader(\n target_classes=[\"car\", \"truck\"],\n batch_size=64)", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "212-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Evaluating the Best Model" + }, + { + "text": "Explanation:\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Implement a neural network from scratch.\n- Perform forward and backward propagation.\n- Validate gradients numerically.\n- Build image classification models using PyTorch.\n- Train and optimize neural networks.\n- Perform hyperparameter tuning.\n- Compare fully connected neural networks with convolutional neural networks.", + "source": "03_neural_network_applications.ipynb", + "file_type": "ipynb", + "chunk_id": "213-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "ann-pytorch-classification", + "relative_path": "notebooks/03_neural_network_applications.ipynb", + "section": "Key Takeaways" + }, + { + "text": "# Computer Vision with PyTorch\n\nA collection of practical **Computer Vision** projects implemented using **PyTorch**. This repository covers the fundamentals of Convolutional Neural Networks (CNNs), transfer learning with pretrained models, and real-world image classification applications. Each notebook has been reorganized, documented, and expanded to demonstrate modern deep learning workflows for computer vision.\n\n![CNN Architecture](Images/CNN.png)\n\n---\n\n# 🚀 Repository Overview\n\nThis repository contains practical computer vision projects built with **PyTorch**, covering convolutional neural networks, transfer learning, and image classification through reproducible, well-documented notebooks.\n\n---\n\n# ✨ Project Highlights\n\n- Convolutional Neural Networks (CNNs) implemented with PyTorch\n- Transfer learning using pretrained AlexNet\n- End-to-end image classification workflows\n- GPU-ready training pipelines\n- Automatic dataset downloading for reproducibility", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "s) implemented with PyTorch\n- Transfer learning using pretrained AlexNet\n- End-to-end image classification workflows\n- GPU-ready training pipelines\n- Automatic dataset downloading for reproducibility\n- Well-documented notebooks suitable for learning and experimentation\n\n---\n\n# 📚 Topics Covered\n\n- Convolutional Neural Networks (CNNs)\n- Convolution and feature extraction\n- Pooling operations\n- Image classification\n- Transfer learning\n- Feature extraction using pretrained models\n- Fine-tuning pretrained CNNs\n- Data augmentation\n- Data normalization\n- Weight decay\n- Dropout\n- Preventing overfitting\n- Model evaluation\n- GPU acceleration with CUDA\n- Visualizing convolution kernels\n\n---\n\n# 🛠️ Technologies\n\n- Python\n- PyTorch\n- Torchvision\n- NumPy\n- Matplotlib\n- Scikit-learn\n- Pillow (PIL)\n- gdown\n\n---\n\n# 📂 Repository Structure\n\n```text\ncomputer-vision-pytorch/\n│\n├── README.md\n├── requirements.txt\n├── .gitignore\n│\n├── notebooks/\n│ ├── 01_cnn_fundamentals.ipynb\n│ ├── 02_transfer_learning.ipynb", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "---\n\n# 📂 Repository Structure\n\n```text\ncomputer-vision-pytorch/\n│\n├── README.md\n├── requirements.txt\n├── .gitignore\n│\n├── notebooks/\n│ ├── 01_cnn_fundamentals.ipynb\n│ ├── 02_transfer_learning.ipynb\n│ └── 03_hand_gesture_recognition.ipynb\n│\n├── images/\n│ └── CNN.png\n│\n├── datasets/\n└── models/\n```\n\n---\n\n# 📖 Notebooks\n\n## Notebook 1 — CNN Fundamentals\n\n### Overview\n\nThis notebook introduces the core concepts of Convolutional Neural Networks (CNNs) using PyTorch. It explores convolution operations, pooling layers, feature extraction, GPU acceleration, and image classification while building an intuitive understanding of how CNNs learn visual representations.\n\n### Topics Covered\n\n- Convolutional Neural Networks (CNNs)\n- Convolution operations\n- Feature extraction\n- Pooling layers\n- Image classification\n- GPU acceleration with CUDA\n- Building CNNs in PyTorch\n- Visualizing convolution kernels\n\n### Notebook\n\n`notebooks/01_cnn_fundamentals.ipynb`\n\n---", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "extraction\n- Pooling layers\n- Image classification\n- GPU acceleration with CUDA\n- Building CNNs in PyTorch\n- Visualizing convolution kernels\n\n### Notebook\n\n`notebooks/01_cnn_fundamentals.ipynb`\n\n---\n\n## Notebook 2 — Transfer Learning with PyTorch\n\n### Overview\n\nThis notebook demonstrates transfer learning for image classification using pretrained convolutional neural networks in PyTorch. It explores feature extraction with AlexNet, visualizes learned representations, applies transfer learning to image datasets, and investigates techniques for improving model generalization, including data normalization, data augmentation, weight decay, and dropout.\n\n![CNN Architecture](Images/TR.png)\n\n### Topics Covered\n\n- Transfer learning\n- AlexNet\n- Feature extraction\n- Pretrained CNNs\n- Image classification\n- Data normalization\n- Data augmentation\n- Weight decay\n- Dropout\n- Preventing overfitting\n\n### Notebook\n\n`notebooks/02_transfer_learning.ipynb`\n\n---\n\n## Notebook 3 — Hand Gesture Recognition", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "cation\n- Data normalization\n- Data augmentation\n- Weight decay\n- Dropout\n- Preventing overfitting\n\n### Notebook\n\n`notebooks/02_transfer_learning.ipynb`\n\n---\n\n## Notebook 3 — Hand Gesture Recognition\n\n![Hand Gesture Recognition](Images/asl-alphabet.jpg)\n\n### Overview\n\nThis notebook demonstrates an end-to-end deep learning workflow for recognizing American Sign Language (ASL) hand gestures using convolutional neural networks and transfer learning with AlexNet.\n\n### Topics Covered\n\n- Image preprocessing\n- Hand gesture recognition\n- Convolutional Neural Networks\n- Dataset preparation\n- Model training\n- Performance evaluation\n- Prediction and inference\n- PyTorch workflows\n\n### Notebook\n\n`notebooks/03_hand_gesture_recognition.ipynb`\n\n---\n\n# 📦 Datasets\n\nSome notebooks require image datasets that are **not included** in this repository because of their size.", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "h workflows\n\n### Notebook\n\n`notebooks/03_hand_gesture_recognition.ipynb`\n\n---\n\n# 📦 Datasets\n\nSome notebooks require image datasets that are **not included** in this repository because of their size.\n\nTo keep the repository lightweight and easy to clone, datasets are downloaded automatically the first time the corresponding notebook is executed.\n\n| Notebook | Dataset |\n|----------|---------|\n| **01 – CNN Fundamentals** | Built-in examples |\n| **02 – Transfer Learning** | Lightweight Flower Classification Dataset (downloaded automatically from Google Drive) |\n| **03 – Hand Gesture Recognition** | Hand Gesture dataset *(downloaded automatically)* |\n\n---\n\n# 🎯 Learning Objectives\n\nThroughout this repository, I explore how to:\n\n- Understand convolutional neural network architectures.\n- Extract visual features using convolutional filters.\n- Build CNN models using PyTorch.\n- Train and evaluate image classification models.\n- Apply transfer learning using pretrained networks.", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "ctures.\n- Extract visual features using convolutional filters.\n- Build CNN models using PyTorch.\n- Train and evaluate image classification models.\n- Apply transfer learning using pretrained networks.\n- Fine-tune pretrained models for new datasets.\n- Improve model generalization using practical regularization techniques.\n- Develop complete computer vision pipelines.\n\n---\n\n# 🚀 Getting Started\n\nClone the repository:\n\n```bash\ngit clone https://github.com/Miladsaeedi70/computer-vision-pytorch.git\n```\n\nNavigate to the project:\n\n```bash\ncd computer-vision-pytorch\n```\n\nInstall the required packages:\n\n```bash\npip install -r requirements.txt\n```\n\nLaunch Jupyter Notebook:\n\n```bash\njupyter notebook\n```\n\nOpen the notebooks in numerical order:\n\n1. CNN Fundamentals\n2. Transfer Learning\n3. Hand Gesture Recognition\n\nDatasets are downloaded automatically when required, so no manual dataset setup is necessary.\n\n---\n\n# ⭐ About", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "numerical order:\n\n1. CNN Fundamentals\n2. Transfer Learning\n3. Hand Gesture Recognition\n\nDatasets are downloaded automatically when required, so no manual dataset setup is necessary.\n\n---\n\n# ⭐ About\n\nThis repository showcases practical computer vision projects developed with **PyTorch**. It demonstrates modern deep learning workflows for convolutional neural networks, transfer learning, and image classification through reproducible, well-documented implementations.\n\nThe notebooks have been reorganized and modernized to improve readability, reproducibility, and compatibility with current versions of PyTorch and Python.\n\n---\n\n# 📄 License\n\nThis repository is intended for educational and portfolio purposes.", + "source": "README.md", + "file_type": "md", + "chunk_id": "214-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "Explanation:\n# Convolutional Neural Networks with PyTorch", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "215-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Convolutional Neural Networks with PyTorch" + }, + { + "text": "Explanation:\n## Overview\n\nThis notebook introduces the fundamentals of Convolutional Neural Networks (CNNs) using PyTorch. It explores the key building blocks of CNNs, including convolutional filters, feature extraction, pooling operations, and image classification. The notebook also compares fully connected neural networks (ANNs) with CNNs and demonstrates how convolutional layers improve performance on image data.\n\nConvolutional Neural Networks (CNNs) leverage spatial information within images by learning convolutional filters that detect edges, textures, shapes, and increasingly complex visual patterns. These hierarchical features enable CNNs to achieve strong performance on image classification tasks.\n\n## Learning Objectives\n\nAfter completing this notebook, you will be able to:\n\n- Understand how convolutional neural networks process images.\n- Apply convolution kernels for feature extraction.\n- Compare artificial neural networks (ANNs) with CNNs.", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "216-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Overview" + }, + { + "text": "tebook, you will be able to:\n\n- Understand how convolutional neural networks process images.\n- Apply convolution kernels for feature extraction.\n- Compare artificial neural networks (ANNs) with CNNs.\n- Build CNN architectures using PyTorch.\n- Train and evaluate CNN models for image classification.\n- Visualize learned convolutional filters.\n\nPython implementation:\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport scipy.signal as sg\n\nfrom PIL import Image\nimport requests\n\n#load image from the internet\nurl = 'https://i.ytimg.com/vi/BqKXHIRwGbs/maxresdefault.jpg'\nresp = requests.get(url, stream=True).raw\nimg = Image.open(resp)\n\n#ensure image is np.array\nimg = np.array(img)\n\n#plot original image\nplt.title(\"Image\")\nplt.imshow(img)\nplt.show()\n\n#increase data precision\nimg = img.astype(np.int16)\nprint('Image Max Value:', np.amax(img), 'Image Min Value:', np.amin(img))\n\n#convert from colour to grayscale\ndef rgb2gray(rgb):\n return np.dot(rgb[...,:3], [0.299, 0.587, 0.144])", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "216-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Overview" + }, + { + "text": "mg.astype(np.int16)\nprint('Image Max Value:', np.amax(img), 'Image Min Value:', np.amin(img))\n\n#convert from colour to grayscale\ndef rgb2gray(rgb):\n return np.dot(rgb[...,:3], [0.299, 0.587, 0.144])\n\nimg_gray = rgb2gray(img)\n\n#plot grayscale image\nplt.title(\"Grayscale Image\")\nplt.imshow(img_gray, cmap='gray')\nplt.show()\n\n#two kernels\nsobel_x = np.array([[-1, 0, 1],\n [-2, 0, 2],\n [-1, 0, 1]])\n\nsobel_y = np.array([[-1, -2, -1],\n [0, 0, 0],\n [1, 2, 1]])\n\n#perform 2d convolution\nimg_edge_x = sg.convolve(img_gray, sobel_x, mode='same')\nimg_edge_y = sg.convolve(img_gray, sobel_y, mode='same')\n\nprint('Image Max Value:', np.amax(img_edge_x), 'Image Min Value:', np.amin(img_edge_x))\nprint('Image Max Value:', np.amax(img_edge_y), 'Image Min Value:', np.amin(img_edge_y))\n\n#combine images\nimg_edge = (img_edge_x**2 + img_edge_y**2)**.5\n\n#normalize images\nimg_edge_x[img_edge_x > 255] = 255\nimg_edge_x[img_edge_x < 0] = 0\n\nimg_edge_y[img_edge_y > 255] = 255\nimg_edge_y[img_edge_y < 0] = 0", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "216-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Overview" + }, + { + "text": "img_edge = (img_edge_x**2 + img_edge_y**2)**.5\n\n#normalize images\nimg_edge_x[img_edge_x > 255] = 255\nimg_edge_x[img_edge_x < 0] = 0\n\nimg_edge_y[img_edge_y > 255] = 255\nimg_edge_y[img_edge_y < 0] = 0\n\nimg_edge[img_edge > 255] = 255\nimg_edge[img_edge < 0] = 0\n\nprint('Image Max Value:', np.amax(img_edge_x), 'Image Min Value:', np.amin(img_edge_x))\nprint('Image Max Value:', np.amax(img_edge_y), 'Image Min Value:', np.amin(img_edge_y))\n\n#return to image format\nimg_edge_x = img_edge_x.astype(np.uint8)\nimg_edge_y = img_edge_y.astype(np.uint8)\nimg_edge = img_edge.astype(np.uint8)\n\n#plot results of convolution in x, and y\nplt.title(\"Image Edge - Vertical\")\nplt.imshow(img_edge_x, cmap='gray')\nplt.show()\n\nplt.title(\"Image Edge - Horizontal\")\nplt.imshow(img_edge_y, cmap='gray')\nplt.show()\n\nplt.title(\"Image Edge - Combined\")\nplt.imshow(img_edge, cmap='gray')\nplt.show()", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "216-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Overview" + }, + { + "text": "Explanation:\n## From ANN to CNN\nIn the example below you'll see that to go from an ANN to a CNN we only need to make a few changes to our architecture. The rest of the code remains the same.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nimport matplotlib.pyplot as plt # for plotting\nimport torch.optim as optim #for gradient descent\n\ntorch.manual_seed(1) # set the random seed\n\n# obtain data\nfrom torchvision import datasets, transforms\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)\nmnist_train = mnist_data[:4096]\nmnist_val = mnist_data[4096:5120]", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "217-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "From ANN to CNN" + }, + { + "text": "Explanation:\n### ANN and CNN Architectures\nProvided is sample code showing the differences between a basic ANN and CNN architectures. Notice that the CNN architecture also contains fully connected layers.\n\nPython implementation:\n#variable that allows you to toggle between ANN and CNN architectures\n#True => CNN, False => ANN\nselect_CNN = True\n\nif not select_CNN:\n\n #Artificial Neural Network Architecture (aka MLP)\n class MNISTClassifier(nn.Module):\n def __init__(self):\n super(MNISTClassifier, self).__init__()\n self.fc1 = nn.Linear(28 * 28, 50)\n self.fc2 = nn.Linear(50, 20)\n self.fc3 = nn.Linear(20, 10)\n\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = F.relu(self.fc1(flattened))\n activation2 = F.relu(self.fc2(activation1))\n output = self.fc3(activation2)\n return output\n\n print('Artificial Neural Network Architecture (aka MLP) Selected')\nelse:\n\n #Convolutional Neural Network Architecture\n class MNISTClassifier(nn.Module):\n def __init__(self):", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "218-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "ANN and CNN Architectures" + }, + { + "text": "vation2)\n return output\n\n print('Artificial Neural Network Architecture (aka MLP) Selected')\nelse:\n\n #Convolutional Neural Network Architecture\n class MNISTClassifier(nn.Module):\n def __init__(self):\n super(MNISTClassifier, self).__init__()\n self.conv1 = nn.Conv2d(1, 5, 5) #in_channels, out_chanels, kernel_size\n self.pool = nn.MaxPool2d(2, 2) #kernel_size, stride\n self.conv2 = nn.Conv2d(5, 10, 5) #in_channels, out_chanels, kernel_size\n self.fc1 = nn.Linear(160, 32)\n self.fc2 = nn.Linear(32, 10)\n\n def forward(self, x):\n x = self.pool(F.relu(self.conv1(x)))\n x = self.pool(F.relu(self.conv2(x)))\n x = x.view(-1, 160)\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n return x\n\n print('Convolutional Neural Network Architecture Selected')\n\nPython implementation:\ndef get_accuracy(model, train=False):\n if train:\n data = mnist_train\n else:\n data = mnist_val\n\n correct = 0\n total = 0\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n\n output = model(imgs)", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "218-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "ANN and CNN Architectures" + }, + { + "text": "cy(model, train=False):\n if train:\n data = mnist_train\n else:\n data = mnist_val\n\n correct = 0\n total = 0\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n\n output = model(imgs)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total\n\nPython implementation:\ndef train(model, data, batch_size=64, num_epochs=1):\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n for epoch in range(num_epochs):\n for imgs, labels in iter(train_loader):\n\n out = model(imgs) # forward pass\n\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "218-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "ANN and CNN Architectures" + }, + { + "text": "chs):\n for imgs, labels in iter(train_loader):\n\n out = model(imgs) # forward pass\n\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model, train=True)) # compute training accuracy\n val_acc.append(get_accuracy(model, train=False)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "218-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "ANN and CNN Architectures" + }, + { + "text": "aining Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "218-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "ANN and CNN Architectures" + }, + { + "text": "Explanation:\n### Comparing ANNs and CNNs\n\nPython implementation:\nmodel = MNISTClassifier()\n\n#proper model\ntrain(model, mnist_train, num_epochs=5)\n\nExplanation:\nTraining deep learning models can be computationally intensive. Leveraging GPU acceleration significantly reduces training time and enables faster experimentation with different network architectures and hyperparameters.", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "219-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Comparing ANNs and CNNs" + }, + { + "text": "Explanation:\n## Enable GPU\nPyTorch allows you to run the computations on a GPU to speed up the processing. In order to enable GPUs you will need to:\n1. select GPUs in \"Notebook Settings\" found under the \"Edit\" menu option.\n2. setup model to work with the cuda\n3. make sure image and labels data are stored placed on the GPU\n\nAn example of this is provided below.\n\nPython implementation:\ndef get_accuracy(model, train=False):\n if train:\n data = mnist_train\n else:\n data = mnist_val\n\n correct = 0\n total = 0\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n output = model(imgs)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "220-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Enable GPU" + }, + { + "text": "odel(imgs)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total\n\nPython implementation:\ndef train(model, data, batch_size=64, num_epochs=1):\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n for epoch in range(num_epochs):\n for imgs, labels in iter(train_loader):\n\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "220-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Enable GPU" + }, + { + "text": "#############################################\n\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model, train=True)) # compute training accuracy\n val_acc.append(get_accuracy(model, train=False)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "220-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Enable GPU" + }, + { + "text": "aining Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\nPython implementation:\nuse_cuda = True\n\nmodel = MNISTClassifier()\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\n#proper model\ntrain(model, mnist_train, num_epochs=5)", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "220-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Enable GPU" + }, + { + "text": "Explanation:\n## Visualizing Learned Convolution Kernels\n\nExplanation:\nRecall what our convolution layer looks like:\n\n*self.conv1 = nn.Conv2d(1, 5, 5) #in_channels, out_chanels, kernel_size*\n\nThere are 5 out channels => 5 kernels, kernel size = 5 and in_channels = 1, hence we're using 5 x 5 kernels.\n\nPython implementation:\nimport matplotlib.pyplot as plt\n\n# Visualize conv1 kernels (i.e filter)\nkernels = model.conv1.weight.detach()\n\nprint(kernels.shape)\n\nExplanation:\nWe can also plot the kernels:\n\nPython implementation:\n#this line is required if using GPU\nkernels = kernels.cpu()\n\n#display first kernel\nprint(kernels[0][0])\n\n#display all five kernels of dimension 5 x 5\nfig, axarr = plt.subplots(kernels.size(0))\nfor idx in range(kernels.size(0)):\n axarr[idx].imshow(kernels[idx][0])", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "221-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Visualizing Learned Convolution Kernels" + }, + { + "text": "Explanation:\n### Visualize Feature Map\nWe can also apply our kernel to our images to see what kind of features it extracts. In the example we will use a new image, but we could also apply this to one of the images in our training or validation data sets.\n\nPython implementation:\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport scipy.signal as sg\n\nfrom PIL import Image\nimport requests\n\n#load image from the internet\nurl = 'https://i.ytimg.com/vi/BqKXHIRwGbs/maxresdefault.jpg'\nresp = requests.get(url, stream=True).raw\nimg = Image.open(resp)\n\n#ensure image is np.array\nimg = np.array(img)\n\n#plot original image\nplt.title(\"Image\")\nplt.imshow(img)\nplt.show()\n\n#increase data precision\nimg = img.astype(np.int16)\nprint('Image Max Value:', np.amax(img), 'Image Min Value:', np.amin(img))\n\n#convert from colour to grayscale\ndef rgb2gray(rgb):\n return np.dot(rgb[...,:3], [0.299, 0.587, 0.144])\n\nimg_gray = rgb2gray(img)\n\n#plot grayscale image\nplt.title(\"Grayscale Image\")", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "222-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Visualize Feature Map" + }, + { + "text": "', np.amin(img))\n\n#convert from colour to grayscale\ndef rgb2gray(rgb):\n return np.dot(rgb[...,:3], [0.299, 0.587, 0.144])\n\nimg_gray = rgb2gray(img)\n\n#plot grayscale image\nplt.title(\"Grayscale Image\")\nplt.imshow(img_gray, cmap='gray')\nplt.show()\n\n#select kernel\nk = kernels[2][0]\n\n#perform 2d convolution\nimg_k = sg.convolve(img_gray, k, mode='same')\n\nprint('Image Max Value:', np.amax(img_k), 'Image Min Value:', np.amin(img_k))\n\n#normalize images\nimg_k[img_k > 255] = 255\nimg_k[img_k < 0] = 0\n\nprint('Image Max Value:', np.amax(img_k), 'Image Min Value:', np.amin(img_k))\n\n#return to image format\nimg_k = img_k.astype(np.uint8)\n\n#plot results of convolution\nplt.title(\"Feature Map for Specified Kernel\")\nplt.imshow(img_k, cmap='gray')\nplt.show()\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nx = torch.randn(20, 3, 10, 10) # NCHW\nconv = nn.Conv2d(in_channels=3, out_channels=7, kernel_size=5, padding=0)\nprint(conv(x).shape)", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "222-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Visualize Feature Map" + }, + { + "text": "Explanation:\n## Summary\n\nIn this notebook, we explored the core concepts of convolutional neural networks using PyTorch. We examined convolution kernels, compared ANN and CNN architectures, trained image classification models, leveraged GPU acceleration, and visualized learned convolutional filters to better understand feature extraction.", + "source": "01_cnn_fundamentals.ipynb", + "file_type": "ipynb", + "chunk_id": "223-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/01_cnn_fundamentals.ipynb", + "section": "Summary" + }, + { + "text": "Explanation:\n# Transfer Learning with PyTorch", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "224-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Transfer Learning with PyTorch" + }, + { + "text": "Explanation:\n## Transfer Learning\n\nExplanation:\nTraining deep convolutional neural networks from scratch often requires millions of labeled images and substantial computational resources. Transfer learning addresses this challenge by reusing knowledge learned from large-scale datasets, allowing pretrained models to be adapted efficiently to new computer vision tasks.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "225-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Transfer Learning" + }, + { + "text": "Explanation:\n## Overview\n\nThis notebook introduces transfer learning using pretrained convolutional neural networks in PyTorch. It demonstrates how pretrained feature extractors can be reused for new image classification tasks, significantly reducing training time while improving model performance. The notebook also explores feature visualization, dataset preprocessing, data augmentation, and techniques for reducing overfitting.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "226-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Overview" + }, + { + "text": "Explanation:\n## Learning Objectives\n\nAfter completing this notebook, you will be able to:\n\n- Understand the principles of transfer learning.\n- Use pretrained CNNs for feature extraction.\n- Implement transfer learning with AlexNet in PyTorch.\n- Visualize learned feature representations.\n- Train image classifiers using pretrained models.\n- Apply data augmentation techniques.\n- Reduce overfitting through normalization, weight decay, and dropout.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "227-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Learning Objectives" + }, + { + "text": "Explanation:\n## AlexNet in PyTorch\n\nConvolutional networks are very commonly used, meaning that there are often alternatives to\ntraining convolutional networks from scratch. In particular, researchers often release both\nthe architecture and **the weights** of the networks they train.\n\nAs an example, let's look at the AlexNet model, whose trained weights are included in `torchvision`.\nAlexNet was trained to classify images into one of many categories.\nThe AlexNet can be imported and used as shown below.\n\nPython implementation:\n# alexnet\nimport torchvision.models\n\nalexNet = torchvision.models.alexnet(pretrained=True)\n\nExplanation:\nNotice that the AlexNet model is split into two parts. There is a component that computes\n\"features\" using convolutions.\n\nExplanation:\nThere is also a component that classifies the image based on the computed features.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "228-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet in PyTorch" + }, + { + "text": "Explanation:\n### AlexNet Features\n\nThe first network can be used independently of the second. Specifically, it can be used\nto compute a set of **features** that can be used later on. This idea of using neural\nnetwork activation *features* to represent images is an extremely important one, so it\nis important to understand the idea now.\n\nTo see how we can do this, let's first load an image.\n\nPython implementation:\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport requests\nfrom io import BytesIO\n\n# Converted your link to a direct download link\nurl = 'https://drive.google.com/uc?export=download&id=1ALAlZf9AQBadqLzK36FJy06wvW_RdnGU'\n\n# Fetch and open the image\nresponse = requests.get(url)\nimg = Image.open(BytesIO(response.content))\n\nplt.imshow(img)\nplt.show()\n\nExplanation:\nTo use this image we need to convert it into a PyTorch tensor of the appropriate shape.\n\nPython implementation:\nimport torch\nimport torchvision.transforms as transforms\n\n# 1.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "229-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet Features" + }, + { + "text": "plt.show()\n\nExplanation:\nTo use this image we need to convert it into a PyTorch tensor of the appropriate shape.\n\nPython implementation:\nimport torch\nimport torchvision.transforms as transforms\n\n# 1. Convert the PIL Image and automatically handle float conversion & CHW ordering\ntransform = transforms.ToTensor()\nx = transform(img) # Shape becomes [3, 4500, 3000]\n\n# 2. Add ONLY ONE batch dimension at the very beginning\nx = x.unsqueeze(0) # Shape becomes [1, 3, 4500, 3000]\n\nprint(\"Final shape for AlexNet:\", x.shape)\n\n# 3. Now pass it to AlexNet safely\nfeatures = alexNet.features(x)\nprint(\"Features shape:\", features.shape)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "229-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet Features" + }, + { + "text": "Explanation:\nThe set of numbers in `features` is another way of representing our image `x`. Recall that\nour initial image `x` was also represented as a tensor, also a set of numbers representing\npixel intensity. Geometrically speaking, we are using points in a high-dimensional space to\nrepresent the images. in our pixel representation, the axes in this high-dimensional space\nwere different pixels. In our `features` representation, the axes are not as easily\ninterpretable.\n\nBut we will want to work with the `features` representation, because this representation\nmakes classification easier. This representation organizes images in a more \"useful\" and\n\"semantic\" way than pixels.\n\nLet me be more specific:\nthis set of `features` was trained on image classification. It turns out that\n**these features can be useful for performing other image-related tasks as well!**\nThat is, if we want to perform an image classification task of our own (for example,", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "230-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet First Convolutions" + }, + { + "text": "assification. It turns out that\n**these features can be useful for performing other image-related tasks as well!**\nThat is, if we want to perform an image classification task of our own (for example,\nclassifying cancer biopsies, which is nothing like what AlexNet was trained to do),\nwe might compute these AlexNet features, and then train a small model on top of those\nfeatures. We replace the `classifier` portion of `AlexNet`, but keep its `features`\nportion intact.\n\nSomehow, through being trained on one type of image classification problem, AlexNet\nlearned something general about representing images for the purposes of other\nclassification tasks.\n\n### AlexNet First Convolutions\n\nHere is the first convolution of AlexNet, applied to our image.\n\nPython implementation:\nalexNetConv = alexNet.features[0]\ny = alexNetConv(x)\ny.shape\n\nExplanation:\nThe output is a $1 \\times 64 \\times 1124 \\times 749$ tensor.\n\nPython implementation:\ny = y.detach().numpy()\ny = (y - y.min()) / (y.max() - y.min())", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "230-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet First Convolutions" + }, + { + "text": "eatures[0]\ny = alexNetConv(x)\ny.shape\n\nExplanation:\nThe output is a $1 \\times 64 \\times 1124 \\times 749$ tensor.\n\nPython implementation:\ny = y.detach().numpy()\ny = (y - y.min()) / (y.max() - y.min())\ny.shape\n\nExplanation:\nWe can visualize each channel independently.\n\nPython implementation:\nimport matplotlib.pyplot as plt\nimport numpy as np\n\nplt.figure(figsize=(10,10))\nfor i in range(64):\n plt.subplot(8, 8, i+1)\n plt.imshow(y[0, i])\n\nExplanation:\nWhat happens to the output if you when we try to use other images? Specifically what happens to the size of the feature tensor when you use a cat image vs the dog image? What changes? What stays the same?", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "230-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet First Convolutions" + }, + { + "text": "Explanation:\n## Applying AlexNet on a Dataset\nIn order to use transfer learning with AlexNet on a new dataset we will have to keep in mind how AlexNet was trained. AlexNet was trained on images of 3 x 224 x 224 images from the ImageNet dataset. These images are of higher resolution than what we have seen until now and are in colour. Hence, it would take significant effort to apply AlexNet to MNIST data, instead we will use another datasets.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "231-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Applying AlexNet on a Dataset" + }, + { + "text": "Explanation:\n### Data Loading from a File\nTo load our data we will use PyTorch's ImageFolder class which makes things a lot easier by allowing you to load data from a directory. For example, the training images are all stored in a directory path that looks like this:\n\n/datadir\n\n /train\n /class1\n /class2\n .\n .\n /val\n /class1\n /class2\n .\n .\n /test\n /class1\n /class2\n .\n .\n\nYou will need to mount your Google Drive to make this work with Colab.\n\nPython implementation:\nimport os\nimport zipfile\nimport gdown\n\n# Google Drive dataset\nDATASET_URL = \"https://drive.google.com/file/d/1r3GaPtsYceOgUYHKpcsBvMxobKHoFfg0/view?usp=sharing\"\n\nDATASET_DIR = \"Flower_Data\"\nZIP_FILE = \"Flower_Data.zip\"\n\n# Download only if the dataset is not already available\nif not os.path.exists(DATASET_DIR):\n\n print(\"Downloading dataset...\")\n gdown.download(DATASET_URL, ZIP_FILE)\n\n print(\"Extracting dataset...\")\n with zipfile.ZipFile(ZIP_FILE, \"r\") as zip_ref:\n zip_ref.extractall(\".\")\n\n os.remove(ZIP_FILE)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "232-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Loading from a File" + }, + { + "text": "rint(\"Downloading dataset...\")\n gdown.download(DATASET_URL, ZIP_FILE)\n\n print(\"Extracting dataset...\")\n with zipfile.ZipFile(ZIP_FILE, \"r\") as zip_ref:\n zip_ref.extractall(\".\")\n\n os.remove(ZIP_FILE)\n\nprint(\"Dataset is ready!\")\n\nPython implementation:\ndata_dir = \"Flower_Data\"\n\ntrain_dir = os.path.join(data_dir, \"train\")\nval_dir = os.path.join(data_dir, \"val\")\n\nclasses = [\n \"daisy\",\n \"dandelion\",\n \"roses\",\n \"sunflowers\",\n \"tulips\",\n]\n\nPython implementation:\nimport os\nimport numpy as np\nimport torch\nimport torchvision\nfrom torchvision import datasets, models, transforms\nimport matplotlib.pyplot as plt\n\nPython implementation:\n# load and transform data using ImageFolder\n\n# resize all images to 224 x 224\ndata_transform = transforms.Compose([transforms.RandomResizedCrop(224),\n transforms.ToTensor()])\n\ntrain_data = datasets.ImageFolder(train_dir, transform=data_transform)\nval_data = datasets.ImageFolder(val_dir, transform=data_transform)\n\n# print out some data stats", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "232-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Loading from a File" + }, + { + "text": "(224),\n transforms.ToTensor()])\n\ntrain_data = datasets.ImageFolder(train_dir, transform=data_transform)\nval_data = datasets.ImageFolder(val_dir, transform=data_transform)\n\n# print out some data stats\nprint('Num training images: ', len(train_data))\nprint('Num validation images: ', len(val_data))\n\nPython implementation:\n# define dataloader parameters\nbatch_size = 20\nnum_workers = 0\n\n# prepare data loaders\ntrain_loader = torch.utils.data.DataLoader(train_data, batch_size=batch_size,\n num_workers=num_workers, shuffle=True)\nval_loader = torch.utils.data.DataLoader(val_data, batch_size=batch_size,\n num_workers=num_workers, shuffle=True)\n\nPython implementation:\n# Visualize some sample data\n\n# obtain one batch of training images\ndataiter = iter(train_loader)\nimages, labels = next(dataiter)\nimages = images.numpy() # convert images to numpy for display\n\n# plot the images in the batch, along with the corresponding labels\nfig = plt.figure(figsize=(25, 4))\nfor idx in np.arange(20):", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "232-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Loading from a File" + }, + { + "text": "ataiter)\nimages = images.numpy() # convert images to numpy for display\n\n# plot the images in the batch, along with the corresponding labels\nfig = plt.figure(figsize=(25, 4))\nfor idx in np.arange(20):\n ax = fig.add_subplot(2, 10, idx+1, xticks=[], yticks=[])\n plt.imshow(np.transpose(images[idx], (1, 2, 0)))\n ax.set_title(classes[labels[idx]])", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "232-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Loading from a File" + }, + { + "text": "Explanation:\n### AlexNet Implementation\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim #for gradient descent\n\n# alexnet\nimport torchvision.models\n\ntorch.manual_seed(1) # set the random seed\n\n# obtain one batch of training images\ndataiter = iter(train_loader)\nimages, labels =next(dataiter)\n\n# confirm output from AlexNet feature extraction\nalexNet = torchvision.models.alexnet(pretrained=True)\nfeatures = alexNet.features(images)\nfeatures.shape\n\nPython implementation:\n#Artifical Neural Network Architecture\nclass ANNClassifier(nn.Module):\n def __init__(self):\n super(ANNClassifier, self).__init__()\n self.fc1 = nn.Linear(256 * 6 * 6, 10)\n self.fc2 = nn.Linear(10, 5)\n\n def forward(self, x):\n x = x.view(-1, 256 * 6 * 6) #flatten feature data\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n return x\n\nPython implementation:\ndef get_accuracy(model, train=False):\n if train:\n data_loader = train_loader\n else:", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "233-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet Implementation" + }, + { + "text": "(-1, 256 * 6 * 6) #flatten feature data\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n return x\n\nPython implementation:\ndef get_accuracy(model, train=False):\n if train:\n data_loader = train_loader\n else:\n data_loader = val_loader\n\n correct = 0\n total = 0\n for imgs, labels in data_loader:\n\n imgs = alexNet.features(imgs) #SLOW\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n output = model(imgs)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total\n\nPython implementation:\ndef train(model, data, batch_size=20, num_epochs=1):\n #train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n train_loader = torch.utils.data.DataLoader(train_data, batch_size=batch_size,", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "233-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet Implementation" + }, + { + "text": "ain(model, data, batch_size=20, num_epochs=1):\n #train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n train_loader = torch.utils.data.DataLoader(train_data, batch_size=batch_size,\n num_workers=num_workers, shuffle=True)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n for epoch in range(num_epochs):\n for imgs, labels in iter(train_loader):\n\n imgs = features = alexNet.features(imgs) #SLOW\n print(n)\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "233-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet Implementation" + }, + { + "text": "model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model, train=True)) # compute training accuracy\n val_acc.append(get_accuracy(model, train=False)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "233-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet Implementation" + }, + { + "text": "lt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\nPython implementation:\nuse_cuda = False\n\nmodel = ANNClassifier()\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\n#proper model\ntrain(model, [], num_epochs=1)\n\nExplanation:\nThe previous example illustrates the core workflow of transfer learning. In practice, different pretrained architectures such as ResNet, EfficientNet, DenseNet, or Vision Transformers (ViTs) can be used depending on the target application and dataset.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "233-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "AlexNet Implementation" + }, + { + "text": "Explanation:\n# Preventing Overfitting\n\nDeep neural networks often achieve excellent performance on the training data but may fail to generalize to unseen examples. This phenomenon, known as **overfitting**, occurs when a model learns patterns that are specific to the training set rather than capturing the underlying data distribution.\n\nIn contrast, **underfitting** occurs when a model is too simple to learn meaningful relationships from the data. Modern deep learning workflows typically favor models with sufficient capacity and rely on regularization techniques to improve generalization.\n\nCommon strategies for reducing overfitting include:\n\n- Collecting additional training data\n- Using smaller or more efficient model architectures\n- Weight sharing through convolutional layers\n- Early stopping\n- Transfer learning\n- Data normalization\n- Data augmentation\n- Weight decay\n- Dropout\n\nThe following sections demonstrate several of these techniques using the MNIST handwritten digit dataset.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "234-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Preventing Overfitting" + }, + { + "text": "pping\n- Transfer learning\n- Data normalization\n- Data augmentation\n- Weight decay\n- Dropout\n\nThe following sections demonstrate several of these techniques using the MNIST handwritten digit dataset.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport matplotlib.pyplot as plt\nfrom torchvision import datasets, transforms\n\n# for reproducibility\ntorch.manual_seed(1)\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)\nmnist_train = mnist_data[:20] # 20 train images\nmnist_val = mnist_data[100:5100] # 2000 validation images\n\nExplanation:\nWe will also use the `MNISTClassifier` from the last few weeks as our base model:\n\nPython implementation:\nclass MNISTClassifier(nn.Module):\n def __init__(self):\n super(MNISTClassifier, self).__init__()\n self.layer1 = nn.Linear(28 * 28, 50)\n self.layer2 = nn.Linear(50, 20)\n self.layer3 = nn.Linear(20, 10)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "234-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Preventing Overfitting" + }, + { + "text": "class MNISTClassifier(nn.Module):\n def __init__(self):\n super(MNISTClassifier, self).__init__()\n self.layer1 = nn.Linear(28 * 28, 50)\n self.layer2 = nn.Linear(50, 20)\n self.layer3 = nn.Linear(20, 10)\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = F.relu(self.layer1(flattened))\n activation2 = F.relu(self.layer2(activation1))\n output = self.layer3(activation2)\n return output\n\nExplanation:\nAnd of course, our training code, with minor modifications that we will explain as we go along.\n\nPython implementation:\ndef train(model, train, valid, batch_size=20, num_iters=1, learn_rate=0.01, weight_decay=0):\n train_loader = torch.utils.data.DataLoader(train,\n batch_size=batch_size,\n shuffle=True) # shuffle after every epoch\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=learn_rate, momentum=0.9, weight_decay=weight_decay)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n while True:", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "234-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Preventing Overfitting" + }, + { + "text": "optim.SGD(model.parameters(), lr=learn_rate, momentum=0.9, weight_decay=weight_decay)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n\n # training\n n = 0 # the number of iterations\n while True:\n if n >= num_iters:\n break\n for imgs, labels in iter(train_loader):\n model.train() #*****************************#\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n if n % 10 == 9:\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model, train)) # compute training accuracy\n val_acc.append(get_accuracy(model, valid)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.figure(figsize=(10,4))\n plt.subplot(1,2,1)\n plt.title(\"Training Curve\")", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "234-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Preventing Overfitting" + }, + { + "text": "# compute training accuracy\n val_acc.append(get_accuracy(model, valid)) # compute validation accuracy\n n += 1\n\n # plotting\n plt.figure(figsize=(10,4))\n plt.subplot(1,2,1)\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n\n plt.subplot(1,2,2)\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\ntrain_acc_loader = torch.utils.data.DataLoader(mnist_train, batch_size=100)\nval_acc_loader = torch.utils.data.DataLoader(mnist_val, batch_size=1000)\n\ndef get_accuracy(model, data):\n correct = 0\n total = 0\n model.eval() #*********#\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n output = model(imgs)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "234-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Preventing Overfitting" + }, + { + "text": "st_val, batch_size=1000)\n\ndef get_accuracy(model, data):\n correct = 0\n total = 0\n model.eval() #*********#\n for imgs, labels in torch.utils.data.DataLoader(data, batch_size=64):\n output = model(imgs)\n pred = output.max(1, keepdim=True)[1] # get the index of the max logit\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n return correct / total\n\nExplanation:\nWithout any intervention, our model gets to about 52-53% accuracy on the validation set.\n\nPython implementation:\nmodel = MNISTClassifier()\ntrain(model, mnist_train, mnist_val, num_iters=500)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "234-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Preventing Overfitting" + }, + { + "text": "Explanation:\n## Improving Model Generalization\n\nModern deep learning models can easily memorize training data, leading to poor performance on unseen examples. Several techniques help improve model generalization by making training more stable and reducing overfitting.\n\n### Data Normalization\n\nData normalization scales input features to a consistent range, making optimization more stable and accelerating model convergence during training.\n\nFor image data, pixel intensities are naturally represented on a common scale. The PyTorch transform `transforms.ToTensor()` converts images to tensors and scales pixel values from **[0, 255]** to **[0, 1]**.\n\nAdditional normalization is often applied using the dataset's mean and standard deviation:\n\n```python\ntransforms.Normalize(mean, std)\n```\n\nThis centers each image channel around zero with unit variance, which typically improves training stability for convolutional neural networks.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "235-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Improving Model Generalization" + }, + { + "text": "Explanation:\nThis transform subtracts 0.5 from each pixel, and divides the\nresult by 0.5. So, each pixel intensity will be in the range `[-1, 1]`.\nIn general, having both positive and negative input values helps\nthe network trains quickly (because of the way weights are initialized).\nSticking with each pixel being in the range `[0, 1]` is usually fine.\n\n## Data Augmentation\n\nWhile it is often expensive to gather more data, we can often\nprogrammatically *generate* more data points from our existing\ndata set. We can make small alterations to our training set to obtain\nslightly different input data, but that is still valid.\nCommon ways of obtaining new (image) data include:\n\n- Flipping each image horizontally or vertically (won't work for digit recognition, but might for other tasks)\n- Shifting each pixel a little to the left or right\n- Rotating the images a little\n- Adding noise to the image\n\n... or even a combination of the above. For demonstration purposes, let's randomly", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "236-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Augmentation" + }, + { + "text": "sks)\n- Shifting each pixel a little to the left or right\n- Rotating the images a little\n- Adding noise to the image\n\n... or even a combination of the above. For demonstration purposes, let's randomly\nrotate our digits a little to get new training samples.\n\nHere are the 20 images in our training set:\n\nPython implementation:\ndef show20(data):\n plt.figure(figsize=(10,2))\n for n, (img, label) in enumerate(data):\n if n >= 20:\n break\n plt.subplot(2, 10, n+1)\n plt.imshow(img)\n\nmnist_imgs = datasets.MNIST('data', train=True, download=True)\nshow20(mnist_imgs)\n\nExplanation:\nHere are the 20 images in our training set, each rotated randomly, by up to 25 degrees.\n\nPython implementation:\nmnist_new = datasets.MNIST('data', train=True, download=True,\n transform=transforms.RandomRotation(25))\nshow20(mnist_new)\n\nExplanation:\nIf we apply the transformation again, we can get images with different rotations:\n\nPython implementation:", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "236-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Augmentation" + }, + { + "text": "rue, download=True,\n transform=transforms.RandomRotation(25))\nshow20(mnist_new)\n\nExplanation:\nIf we apply the transformation again, we can get images with different rotations:\n\nPython implementation:\nmnist_new = datasets.MNIST('data', train=True, download=True, transform=transforms.RandomRotation(25))\nshow20(mnist_new)\n\nExplanation:\nWe can augment our data set by, say, randomly rotating each training data point 100 times:\n\nPython implementation:\naugmented_train_data = []\n\nmy_transform = transforms.Compose([\n transforms.RandomRotation(25),\n transforms.ToTensor(),\n])\n\nfor i in range(100):\n mnist_new = datasets.MNIST('data', train=True, download=True, transform=my_transform)\n for j, item in enumerate(mnist_new):\n if j >= 20:\n break\n augmented_train_data.append(item)\n\nlen(augmented_train_data)\n\nExplanation:\nWe obtain a better validation accuracy after training on our expanded dataset.\n\nPython implementation:\nmodel = MNISTClassifier()", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "236-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Augmentation" + }, + { + "text": "ented_train_data.append(item)\n\nlen(augmented_train_data)\n\nExplanation:\nWe obtain a better validation accuracy after training on our expanded dataset.\n\nPython implementation:\nmodel = MNISTClassifier()\ntrain(model, augmented_train_data, mnist_val, num_iters=500)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "236-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Data Augmentation" + }, + { + "text": "Explanation:\n## Weight Decay\n\nA more interesting technique that prevents overfitting is the idea of weight decay.\nThe idea is to **penalize large weights**. We avoid large weights, because large weights\nmean that the prediction relies a lot on the content of one pixel, or on one unit. Intuitively,\nit does not make sense that the classification of an image should depend heavily on the\ncontent of one pixel, or even a few pixels.\n\nMathematically, we penalize large weights by adding an extra term to the loss function,\nthe term can look like the following:\n\n- $L^1$ regularization: $\\sum_k |w_k|$\n - Mathematically, this term encourages weights to be exactly 0\n- $L^2$ regularization: $\\sum_k w_k^2$\n - Mathematically, in each iteration the weight is pushed towards 0\n- Combination of $L^1$ and $L^2$ regularization: add a term $\\sum_k |w_k| + w_k^2$ to the loss function.\n\nIn PyTorch, weight decay can also be done automatically inside an optimizer. The parameter `weight_decay`", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "237-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Weight Decay" + }, + { + "text": "n of $L^1$ and $L^2$ regularization: add a term $\\sum_k |w_k| + w_k^2$ to the loss function.\n\nIn PyTorch, weight decay can also be done automatically inside an optimizer. The parameter `weight_decay`\nof `optim.SGD` and most other optimizers uses $L^2$ regularization for weight decay. The value of the\n`weight_decay` parameter is another tunable hyperparameter.\n\nPython implementation:\nmodel = MNISTClassifier()\ntrain(model, mnist_train, mnist_val, num_iters=500, weight_decay=0.001)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "237-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Weight Decay" + }, + { + "text": "Explanation:\n## Dropout\n\nYet another way to prevent overfitting is to build **many** models, then average\ntheir predictions at test time. Each model might have a different set of\ninitial weights.\n\nWe won't show an example of model averaging here. Instead, we will show another\nidea that sounds drastically different on the surface.\n\nThis idea is called **dropout**: we will randomly \"drop out\", \"zero out\", or \"remove\" a portion\nof neurons from each training iteration.\n\n![](imgs/dropout.png)\n\nIn different iterations of training, we will drop out a different set of neurons.\n\nThe technique has an effect of preventing weights from being overly dependent on\neach other: for example for one weight to be unnecessarily large to compensate for\nanother unnecessarily large weight with the opposite sign. Weights are encouraged\nto be \"more independent\" of one another.\n\nDuring test time though, we will not drop out any neurons; instead we will use\nthe entire set of weights.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "238-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Dropout" + }, + { + "text": "eight with the opposite sign. Weights are encouraged\nto be \"more independent\" of one another.\n\nDuring test time though, we will not drop out any neurons; instead we will use\nthe entire set of weights. This means that our training time and test time behaviour\nof dropout layers are *different*. In the code for the function `train` and `get_accuracy`,\nwe use `model.train()` and `model.eval()` to flag whether we want the model's training behaviour,\nor test time behaviour.\n\nWhile unintuitive, using all connections is a form\nof model averaging! We are effectively averaging over many different networks\nof various connectivity structures.\n\nPython implementation:\nclass MNISTClassifierWithDropout(nn.Module):\n def __init__(self):\n super(MNISTClassifierWithDropout, self).__init__()\n self.layer1 = nn.Linear(28 * 28, 50)\n self.layer2 = nn.Linear(50, 20)\n self.layer3 = nn.Linear(20, 10)\n self.dropout1 = nn.Dropout(0.4) # drop out layer with 20% dropped out neuron\n self.dropout2 = nn.Dropout(0.4)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "238-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Dropout" + }, + { + "text": "nn.Linear(28 * 28, 50)\n self.layer2 = nn.Linear(50, 20)\n self.layer3 = nn.Linear(20, 10)\n self.dropout1 = nn.Dropout(0.4) # drop out layer with 20% dropped out neuron\n self.dropout2 = nn.Dropout(0.4)\n self.dropout3 = nn.Dropout(0.4)\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = F.relu(self.layer1(self.dropout1(flattened)))\n activation2 = F.relu(self.layer2(self.dropout2(activation1)))\n output = self.layer3(self.dropout3(activation2))\n return output\n\nmodel = MNISTClassifierWithDropout()\ntrain(model, mnist_train, mnist_val, num_iters=500)", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "238-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Dropout" + }, + { + "text": "Explanation:\n## Summary\n\nIn this notebook, we explored transfer learning with pretrained convolutional neural networks using PyTorch. We examined feature extraction with AlexNet, implemented transfer learning for image classification, visualized learned representations, and applied practical techniques—including normalization, data augmentation, weight decay, and dropout—to improve model generalization.", + "source": "02_transfer_learning.ipynb", + "file_type": "ipynb", + "chunk_id": "239-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/02_transfer_learning.ipynb", + "section": "Summary" + }, + { + "text": "Explanation:\n# Hand Gesture Recognition with PyTorch\n\n## Overview\n\nThis project demonstrates an end-to-end deep learning workflow for recognizing American Sign Language (ASL) hand gestures using Convolutional Neural Networks (CNNs) in PyTorch. It covers the complete computer vision pipeline, from custom dataset creation and preprocessing to model training, evaluation, and transfer learning using pretrained convolutional neural networks.\n\n## Learning Objectives\n\nBy completing this project, you will learn how to:\n\n- Build a custom image dataset for hand gesture recognition\n- Preprocess and organize image data for deep learning\n- Split datasets into training, validation, and testing sets\n- Implement and train Convolutional Neural Networks (CNNs) in PyTorch\n- Apply transfer learning using pretrained computer vision models\n- Evaluate image classification models using appropriate performance metrics\n- Improve model generalization through data augmentation and regularization techniques", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "240-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Hand Gesture Recognition with PyTorch" + }, + { + "text": "g pretrained computer vision models\n- Evaluate image classification models using appropriate performance metrics\n- Improve model generalization through data augmentation and regularization techniques\n- Visualize training progress and model predictions\n\nThis notebook demonstrates practical computer vision techniques that can be applied to a wide range of image classification tasks beyond hand gesture recognition.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "240-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Hand Gesture Recognition with PyTorch" + }, + { + "text": "Explanation:\n# Part 1", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "241-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 1" + }, + { + "text": "Explanation:\n# Part A — Dataset Preparation\n\n## Overview\n\nHigh-quality datasets are essential for building reliable computer vision models. Unlike benchmark datasets such as MNIST or CIFAR, many real-world applications require collecting, organizing, and preprocessing custom image datasets before model development can begin.\n\nThis project demonstrates the complete workflow for creating a custom dataset for hand gesture recognition, including image collection, preprocessing, quality control, and dataset organization.\n\n---\n\n## American Sign Language (ASL)\n\nAmerican Sign Language (ASL) is a complete visual language that communicates through hand gestures, facial expressions, and body posture. In this project, we focus on recognizing a subset of ASL alphabet gestures using deep learning.\n\nSpecifically, the classification task consists of recognizing hand gestures corresponding to the letters:\n\n**A – I (9 gesture classes)**", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "242-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part A — Dataset Preparation" + }, + { + "text": "izing a subset of ASL alphabet gestures using deep learning.\n\nSpecifically, the classification task consists of recognizing hand gestures corresponding to the letters:\n\n**A – I (9 gesture classes)**\n\n![ASL Alphabet](https://www.disabled-world.com/pics/1/asl-alphabet.jpg)\n\n---\n\n## Dataset Collection\n\nA custom dataset is created by capturing multiple images for each ASL hand gesture.\n\n### Recommended Data Collection Guidelines\n\n- Capture multiple images for each gesture from slightly different viewpoints.\n- Ensure consistent lighting conditions.\n- Use a clean, uncluttered background whenever possible.\n- Keep the hand clearly visible and centered in the image.\n- Minimize shadows and background distractions.\n- Capture sufficient variation to improve model generalization.\n\n---\n\n## Image Preprocessing\n\nTo ensure consistent model performance, all images are standardized before training.\n\n### Preprocessing Steps\n\n- Crop the hand region\n- Resize all images to **224 × 224 pixels**", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "242-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part A — Dataset Preparation" + }, + { + "text": "Image Preprocessing\n\nTo ensure consistent model performance, all images are standardized before training.\n\n### Preprocessing Steps\n\n- Crop the hand region\n- Resize all images to **224 × 224 pixels**\n- Preserve RGB color channels\n- Center the hand within the image\n- Save images in JPEG format\n\nThese preprocessing steps reduce unnecessary variation and improve the quality of the training dataset.\n\n---\n\n## Dataset Organization\n\nThe dataset is organized into one directory for each gesture class:\n\n```text\ndataset/\n│\n├── A/\n├── B/\n├── C/\n├── D/\n├── E/\n├── F/\n├── G/\n├── H/\n└── I/\n```\n\nEach folder contains multiple images representing the corresponding ASL gesture.\n\n---\n\n## Example Dataset\n\nThe figure below illustrates examples from the hand gesture dataset used for training the convolutional neural network.\n\n![Example Dataset](https://github.com/UTNeural/APS360/blob/master/Gesture%20Images.PNG?raw=true)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "242-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part A — Dataset Preparation" + }, + { + "text": "Explanation:\n# Part B — Building a Convolutional Neural Network\n\n## Overview\n\nIn this section, a Convolutional Neural Network (CNN) is implemented from scratch using PyTorch to classify American Sign Language (ASL) hand gestures. The model is trained on the prepared dataset and evaluated on unseen images to assess its ability to generalize.\n\nThe implementation demonstrates the complete deep learning workflow, including model architecture design, training, validation, performance evaluation, and prediction.\n\n## Objectives\n\nThe goals of this section are to:\n\n- Design and implement a Convolutional Neural Network (CNN) in PyTorch\n- Train the network for multi-class image classification\n- Monitor training and validation performance\n- Evaluate the trained model on unseen test data\n- Analyze model accuracy and learning behavior\n- Explore techniques for improving generalization and reducing overfitting\n\n## Implementation Notes\n\nThe CNN is implemented using PyTorch's neural network modules.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "243-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part B — Building a Convolutional Neural Network" + }, + { + "text": "model accuracy and learning behavior\n- Explore techniques for improving generalization and reducing overfitting\n\n## Implementation Notes\n\nThe CNN is implemented using PyTorch's neural network modules. The training pipeline follows modern deep learning practices, including:\n\n- Efficient mini-batch training with `DataLoader`\n- GPU acceleration (when available)\n- Vectorized tensor operations\n- Training and validation loops\n- Loss and accuracy monitoring\n- Model evaluation on a held-out test set\n\nThroughout this section, training and validation metrics are visualized to better understand the learning process and evaluate model performance.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "243-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part B — Building a Convolutional Neural Network" + }, + { + "text": "Explanation:\n# Part 2 — Data Loading and Dataset Splitting\n\n## Overview\n\nThe dataset is loaded using PyTorch's `torchvision.datasets.ImageFolder`, which automatically assigns class labels based on the directory structure. Images are then transformed into tensors and organized into training, validation, and test datasets.\n\n### Dataset Splitting Strategy\n\nThe dataset is randomly partitioned into three non-overlapping subsets using PyTorch's `random_split` function:\n\n- **Training set:** 80%\n- **Validation set:** 10%\n- **Test set:** 10%\n\nThis strategy ensures that:\n\n- The model is trained only on the training data.\n- The validation set is used for model selection and hyperparameter tuning.\n- The test set remains completely unseen during training and validation, providing an unbiased estimate of the model's generalization performance.\n\nThe three subsets are mutually exclusive, ensuring that no image appears in more than one dataset.\n\n### Dataset Statistics\n\n| Dataset | Images |", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "244-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2 — Data Loading and Dataset Splitting" + }, + { + "text": "d estimate of the model's generalization performance.\n\nThe three subsets are mutually exclusive, ensuring that no image appears in more than one dataset.\n\n### Dataset Statistics\n\n| Dataset | Images |\n|---------|-------:|\n| Training | 1,920 |\n| Validation | 256 |\n| Test | 255 |\n| **Total** | **2,431** |\n\n### Data Loading\n\nThe dataset is loaded using PyTorch's `ImageFolder` class, which automatically reads images from class-specific folders and assigns the appropriate labels. This approach provides a simple and scalable pipeline for image classification tasks and integrates seamlessly with PyTorch's `DataLoader` for efficient mini-batch training.\n\nDataset:https://drive.google.com/file/d/1tzxyXdI3-nL2YZwUfNjtQb7sl1BJ1A-J/view?usp=sharing\n\nPython implementation:\nimport glob\nimport os\nimport zipfile\nimport time\nimport numpy as np\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.autograd import Variable\nimport torch.utils.data as data\nimport torchvision", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "244-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2 — Data Loading and Dataset Splitting" + }, + { + "text": "import zipfile\nimport time\nimport numpy as np\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.autograd import Variable\nimport torch.utils.data as data\nimport torchvision\nfrom torchvision import datasets, transforms\nimport torch.optim as optim\nfrom torch.utils.data.sampler import SubsetRandomSampler\nimport gdown\n\n# 1. Download the zip file directly via the shared link\nshared_url = 'https://drive.google.com/file/d/1tzxyXdI3-nL2YZwUfNjtQb7sl1BJ1A-J/view?usp=sharing'\nzip_output = 'Lab_3b_Gesture_Dataset.zip'\nextract_path = 'Lab_3b_Gesture_Dataset'\n\nif not os.path.exists(zip_output):\n print(\"Downloading dataset from Google Drive...\")\n gdown.download(url=shared_url, output=zip_output, quiet=False)\n\n# 2. Unzip the file to the local Colab disk\nif not os.path.exists(extract_path):\n print(\"Extracting dataset...\")\n with zipfile.ZipFile(zip_output, 'r') as zip_ref:\n zip_ref.extractall(extract_path)\n print(\"Extraction completed!\")\nelse:", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "244-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2 — Data Loading and Dataset Splitting" + }, + { + "text": "ab disk\nif not os.path.exists(extract_path):\n print(\"Extracting dataset...\")\n with zipfile.ZipFile(zip_output, 'r') as zip_ref:\n zip_ref.extractall(extract_path)\n print(\"Extraction completed!\")\nelse:\n print(\"Dataset already extracted.\")\n\n# 3. Data Transformations\nnormalizer = transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5))\n\ntransformed = transforms.Compose([\n transforms.Resize((224, 224)),\n transforms.ToTensor(),\n normalizer\n])\n\n# 4. Load Data (Fixed: root now points to the locally extracted path)\nimages_all = torchvision.datasets.ImageFolder(root=extract_path, transform=transformed)\n\n# 5. Split Dataset Dynamically (Safe from manual calculation errors)\ntorch.manual_seed(100)\ntotal_count = len(images_all)\n\n# Set up your split sizes mathematically based on your target counts\ntrain_count = 1920\nval_count = 256\ntest_count = total_count - train_count - val_count # Remainder goes to test set safely", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "244-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2 — Data Loading and Dataset Splitting" + }, + { + "text": "s_all)\n\n# Set up your split sizes mathematically based on your target counts\ntrain_count = 1920\nval_count = 256\ntest_count = total_count - train_count - val_count # Remainder goes to test set safely\n\ntrain_set, val_set, test_set = torch.utils.data.random_split(images_all, [train_count, val_count, test_count])\n\nprint(f\"Total dataset size: {total_count}\")\nprint(f\"Split sizes -> Train: {len(train_set)}, Val: {len(val_set)}, Test: {len(test_set)}\")\n\n# Note: Removed train_set = list(train_set) to preserve RAM stability.\n# You can pass train_set directly into a torch.utils.data.DataLoader now!\n\nExplanation:\n Visualizing some images \n\nPython implementation:\nimport matplotlib.pyplot as plt\nk = 0\nfor images, labels in torch.utils.data.DataLoader(train_set, batch_size=1):\n # since batch_size = 1, there is only 1 image in `images`\n image = images[0]\n # place the colour channel at the end, instead of at the beginning\n img = np.transpose(image, [1,2,0])", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "244-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2 — Data Loading and Dataset Splitting" + }, + { + "text": "t, batch_size=1):\n # since batch_size = 1, there is only 1 image in `images`\n image = images[0]\n # place the colour channel at the end, instead of at the beginning\n img = np.transpose(image, [1,2,0])\n # normalize pixel intensity values to [0, 1]\n img = img / 2 + 0.5\n plt.subplot(3, 5, k+1)\n plt.axis('off')\n plt.imshow(img)\n\n k += 1\n if k > 14:\n break", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "244-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2 — Data Loading and Dataset Splitting" + }, + { + "text": "Explanation:\n## Part 2. Model Building and Sanity Checking", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "245-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2. Model Building and Sanity Checking" + }, + { + "text": "Explanation:\n### Part 2(a) — Convolutional Neural Network Architecture\n\nA custom Convolutional Neural Network (CNN) was developed for multi-class hand gesture classification. Multiple CNN architectures were explored and compared during development, and the final model was selected based on its classification performance while maintaining a relatively small number of trainable parameters for efficient training.\n\n### Model Architecture\n\nThe final network consists of:\n\n- **Two convolutional layers** for hierarchical feature extraction.\n- **Two fully connected (linear) layers** for image classification.\n- **ReLU activation functions** after each convolutional layer to introduce non-linearity and enable efficient gradient propagation during training.\n- **Max-pooling layers** with a kernel size of **2 × 2** and stride **2** to progressively reduce the spatial dimensions while retaining the most informative features.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "246-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(a) — Convolutional Neural Network Architecture" + }, + { + "text": "nt propagation during training.\n- **Max-pooling layers** with a kernel size of **2 × 2** and stride **2** to progressively reduce the spatial dimensions while retaining the most informative features.\n- A **stride of 4** in the first convolutional layer to reduce the spatial resolution early in the network, decreasing the number of parameters required in the fully connected layers and improving computational efficiency.\n\n### Design Choices\n\n- The number of output channels is defined as a configurable hyperparameter, allowing the architecture to be easily tuned during experimentation.\n- The number of hidden units in the first fully connected layer is determined by the output size of the final convolutional layer. For example, when the second convolutional layer has **10 output channels**, the flattened feature vector contains **250 features (10 × 5 × 5)**.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "246-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(a) — Convolutional Neural Network Architecture" + }, + { + "text": "by the output size of the final convolutional layer. For example, when the second convolutional layer has **10 output channels**, the flattened feature vector contains **250 features (10 × 5 × 5)**.\n- A relatively lightweight architecture was selected to balance classification performance, computational cost, and training time while achieving good generalization on the hand gesture dataset.\n\nPython implementation:\n# Convolutional Neural Network Architecture\nclass Classifier(nn.Module):\n def __init__(self, output_conv1, output_conv2):\n super(Classifier, self).__init__()\n self.conv1 = nn.Conv2d(3, output_conv1, 10, 4) # in_channels, out_channels, kernel_size\n self.pool = nn.MaxPool2d(2, 2) # kernel_size, stride\n self.conv2 = nn.Conv2d(output_conv1, output_conv2, 7, 2) # in_channels, out_channels, kernel_size\n self.fc1 = nn.Linear(output_conv2 * 5 * 5, 32)\n self.fc2 = nn.Linear(32, 9)\n self.a = output_conv2\n\n def forward(self, x):\n x = self.pool(F.relu(self.conv1(x)))", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "246-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(a) — Convolutional Neural Network Architecture" + }, + { + "text": "n_channels, out_channels, kernel_size\n self.fc1 = nn.Linear(output_conv2 * 5 * 5, 32)\n self.fc2 = nn.Linear(32, 9)\n self.a = output_conv2\n\n def forward(self, x):\n x = self.pool(F.relu(self.conv1(x)))\n x = self.pool(F.relu(self.conv2(x)))\n x = x.view(-1, self.a * 5 * 5)\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n return x\n\nPython implementation:\nmodel = Classifier( output_conv1=10, output_conv2=20)\nN_p_large=0\nfor param in model.parameters():\n N_p_large+=param.numel()\nprint('Total number of parameters in large_net= ', N_p_large)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "246-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(a) — Convolutional Neural Network Architecture" + }, + { + "text": "Explanation:\n#### Note: two more architecture has been tried, but the accuracy was slightly lower and number of parameters higher for them. \n\nPython implementation:\n# class Classifier(nn.Module):\n # def __init__(self, output_conv1):\n # super(Classifier, self).__init__()\n # self.conv1 = nn.Conv2d(3, output_conv1, 5) #in_channels, out_chanels, kernel_size\n # self.pool = nn.MaxPool2d(2, 2) #kernel_size, stride\n # self.conv2 = nn.Conv2d(output_conv1, 10, 5) #in_channels, out_chanels, kernel_size\n # self.fc1 = nn.Linear(53*53*10, 32)\n # self.fc2 = nn.Linear(32, 10)\n\n # def forward(self, x):\n # x = self.pool(F.relu(self.conv1(x)))\n # x = self.pool(F.relu(self.conv2(x)))\n # x = x.view(-1, 53*53*10)\n # x = F.relu(self.fc1(x))\n # x = self.fc2(x)\n # return x\n\nPython implementation:\n# class Classifier(nn.Module):\n # def __init__(self):\n # super(Classifier, self).__init__()\n # self.conv1 = nn.Conv2d(3, 5, 8,3) #in_channels, out_chanels, kernel_size", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "247-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": " Note: two more architecture has been tried, but the accuracy was slightly lower and number of parameters higher for them. " + }, + { + "text": "eturn x\n\nPython implementation:\n# class Classifier(nn.Module):\n # def __init__(self):\n # super(Classifier, self).__init__()\n # self.conv1 = nn.Conv2d(3, 5, 8,3) #in_channels, out_chanels, kernel_size\n # self.pool = nn.MaxPool2d(2, 2) #kernel_size, stride\n # self.conv2 = nn.Conv2d(5, 10, 8,2) #in_channels, out_chanels, kernel_size\n # self.fc1 = nn.Linear(7*7*10, 32)\n # self.fc2 = nn.Linear(32, 10)\n\n # def forward(self, x):\n # x = self.pool(F.relu(self.conv1(x)))\n # x = self.pool(F.relu(self.conv2(x)))\n # x = x.view(-1, 7*7*10)\n # x = F.relu(self.fc1(x))\n # x = self.fc2(x)\n # return x", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "247-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": " Note: two more architecture has been tried, but the accuracy was slightly lower and number of parameters higher for them. " + }, + { + "text": "Explanation:\n### Part 2(b) — Model Training\n\nA reusable training pipeline was implemented to simplify experimentation with different neural network architectures and hyperparameter settings. The training framework allows key parameters such as the learning rate, batch size, number of epochs, and model architecture to be modified with minimal code changes.\n\n### Training Pipeline\n\nThe implementation includes:\n\n- A dedicated function for computing classification accuracy.\n- A modular training function that performs both training and validation.\n- Support for configurable hyperparameters, including:\n - Batch size\n - Learning rate\n - Number of epochs\n - Optimizer\n - Model architecture\n- Automatic model checkpointing after every training epoch, allowing intermediate models to be saved and training to be resumed if necessary.\n\n### Loss Function", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "hs\n - Optimizer\n - Model architecture\n- Automatic model checkpointing after every training epoch, allowing intermediate models to be saved and training to be resumed if necessary.\n\n### Loss Function\n\nThe model is trained using **Cross-Entropy Loss (`nn.CrossEntropyLoss`)**, which is the standard loss function for multi-class classification problems. It combines the softmax activation and negative log-likelihood into a single objective function, making it well suited for predicting one of multiple gesture classes.\n\n### Optimizer\n\n**Stochastic Gradient Descent (SGD)** is used to optimize the network parameters. Compared with standard batch gradient descent, SGD updates the model more frequently using mini-batches of data, resulting in faster training, lower memory requirements, and often improved convergence and generalization for deep neural networks.\n\nPython implementation:\ndef get_accuracy(model, criterion, batch_size, train=False, test=False):\n if test:\n data=test_set\n else:", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "ften improved convergence and generalization for deep neural networks.\n\nPython implementation:\ndef get_accuracy(model, criterion, batch_size, train=False, test=False):\n if test:\n data=test_set\n else:\n if train:\n data = train_set\n else:\n data = val_set\n\n correct = 0\n error=0\n total = 0\n total_loss = 0.0\n for i, data in enumerate(torch.utils.data.DataLoader(data, batch_size=batch_size), 0):\n imgs, labels = data\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n output = model(imgs)\n\n loss = criterion(output, labels)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n error += pred.ne(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n total_loss += loss.item()\n acc=correct/ total\n err=error/ total\n loss = float(total_loss) / (i + 1)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "s(pred)).sum().item()\n error += pred.ne(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n total_loss += loss.item()\n acc=correct/ total\n err=error/ total\n loss = float(total_loss) / (i + 1)\n return acc, loss, err\n\nPython implementation:\ndef train(model, data, output_conv1, output_conv2, batch_size=256, learning_rate=0.01, num_epochs=30):\n torch.manual_seed(1000)\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=learning_rate, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n losses_val= []\n losses_train= []\n val_err= []\n train_err= []\n\n # training\n n = 0 # the number of iterations\n start_time = time.time()\n for epoch in range(num_epochs):\n\n for imgs, labels in iter(train_loader):\n\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "for imgs, labels in iter(train_loader):\n\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=True)[0]) # compute training accuracy\n val_acc.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=False)[0]) # compute validation accuracy\n n += 1", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "atch_size=batch_size, train=True)[0]) # compute training accuracy\n val_acc.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=False)[0]) # compute validation accuracy\n n += 1\n\n losses_train.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=True)[1]) # compute validation loss\n losses_val.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=False)[1]) # compute validation loss\n train_err.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=True)[2]) # compute training accuracy\n val_err.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=False)[2]) # compute validation accuracy\n print((\"Epoch {}: Train err: {:.3f}, Train loss: {:.3f} |\"+\n \"Validation err: {:.3f}, Validation loss: {:.3f}\").format(\n epoch + 1,\n train_err[epoch],\n losses_train[epoch],\n val_err[epoch],\n losses_val[epoch]))\n\n # Save the current model (checkpoint) to a file", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "\"Validation err: {:.3f}, Validation loss: {:.3f}\").format(\n epoch + 1,\n train_err[epoch],\n losses_train[epoch],\n val_err[epoch],\n losses_val[epoch]))\n\n # Save the current model (checkpoint) to a file\n model_path =\"model_bs{0}_lr{1}_epoch{2}_conv{3}_conv{4}\".format(batch_size,\n learning_rate,\n epoch,\n output_conv1,\n output_conv2)\n torch.save(model.state_dict(), model_path)\n print('Finished Training')\n end_time = time.time()\n elapsed_time = end_time - start_time\n print(\"Total time elapsed: {:.2f} seconds\".format(elapsed_time))\n # Write the train/test loss/err into CSV file for plotting later\n epochs = np.arange(1, num_epochs + 1)\n np.savetxt(\"{}_train_err.csv\".format(model_path), train_acc)\n np.savetxt(\"{}_train_loss.csv\".format(model_path), losses)\n np.savetxt(\"{}_val_err.csv\".format(model_path), val_acc)\n # np.savetxt(\"{}_val_loss.csv\".format(model_path), val_loss)\n\n # plotting\n fig = plt.figure(figsize=(12, 10))\n plt.subplot(2, 2, 1)\n plt.title(\"Training Curve\")", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "err.csv\".format(model_path), val_acc)\n # np.savetxt(\"{}_val_loss.csv\".format(model_path), val_loss)\n\n # plotting\n fig = plt.figure(figsize=(12, 10))\n plt.subplot(2, 2, 1)\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n\n plt.subplot(2, 2, 2)\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n\n x = len(losses_train)\n plt.subplot(2, 2, 3)\n plt.title(\"Train vs Validation Loss\")\n plt.plot(range(1,x+1), losses_train, label=\"Train\")\n plt.plot(range(1,x+1), losses_val, label=\"Validation\")\n plt.xlabel(\"epoch\")\n plt.ylabel(\"Loss\")\n plt.legend(loc='best')\n\n plt.subplot(2, 2, 4)\n plt.title(\"Train vs Validation error\")\n plt.plot(range(1,x+1), train_err, label=\"Train\")\n plt.plot(range(1,x+1), val_err, label=\"Validation\")\n plt.xlabel(\"epoch\")\n plt.ylabel(\"error\")", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "bplot(2, 2, 4)\n plt.title(\"Train vs Validation error\")\n plt.plot(range(1,x+1), train_err, label=\"Train\")\n plt.plot(range(1,x+1), val_err, label=\"Validation\")\n plt.xlabel(\"epoch\")\n plt.ylabel(\"error\")\n plt.legend(loc='best')\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "248-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(b) — Model Training" + }, + { + "text": "Explanation:\n### Part 2(c) — Model Validation on a Small Dataset\n\nBefore training on the full dataset, a sanity check was performed by training the network on a small subset of images. This experiment verifies that both the model architecture and the training pipeline are implemented correctly.\n\n### Validation Strategy\n\n- A subset containing **64 training images** was selected.\n- The model was trained using a large batch size so that all images were processed together during each optimization step.\n- Training continued until the network was able to memorize the small dataset.\n\n### Results\n\nThe model successfully achieved **100% training accuracy** on the 64-image subset, demonstrating that:\n\n- The CNN architecture is capable of learning the training data.\n- The forward and backward propagation are implemented correctly.\n- The optimization process and loss function operate as expected.\n- The data loading and preprocessing pipeline function correctly.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "249-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(c) — Model Validation on a Small Dataset" + }, + { + "text": "- The forward and backward propagation are implemented correctly.\n- The optimization process and loss function operate as expected.\n- The data loading and preprocessing pipeline function correctly.\n\nSuccessfully overfitting a small dataset is a common validation step in deep learning, providing confidence that the model and training pipeline are correctly implemented before scaling to the full dataset.\n\nPython implementation:\nfrom torch.utils.data import Subset\n\n# 1. Safely slice the first 32 samples using a Subset wrapper\ndebug_indices = list(range(32))\ndebug_data = Subset(train_set, debug_indices)\n\n# 2. Setup your model and GPU config\nuse_cuda = True\n\nmodel = Classifier(output_conv1=5, output_conv2=10)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\n# 3. Train your model on the debug subset", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "249-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(c) — Model Validation on a Small Dataset" + }, + { + "text": "uda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\n# 3. Train your model on the debug subset\ntrain(model, debug_data, batch_size=512, output_conv1=5, output_conv2=10, num_epochs=200)\n\n# 4. Obtain accuracy on your 32 debug samples\n# Note: Changed your comment/batch size to 32 here since debug_data only has 32 samples.\ncorrect = 0\ntotal = 0\nfor imgs, labels in torch.utils.data.DataLoader(debug_data, batch_size=32):\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n\n output = model(imgs)\n # select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n\nprint('Accuracy on debug batch: ', correct / total)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "249-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 2(c) — Model Validation on a Small Dataset" + }, + { + "text": "Explanation:\n## Part 3. Hyperparameter Search", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "250-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 3. Hyperparameter Search" + }, + { + "text": "Explanation:\n### Part 4(a) — Hyperparameter Selection\n\nSeveral hyperparameters can significantly influence the performance of a convolutional neural network. The following hyperparameters were selected for tuning during model development:\n\n- **Learning rate** – Controls the step size used by the optimizer during parameter updates and has a significant impact on convergence speed and model performance.\n- **Batch size** – Determines the number of training samples processed in each optimization step, affecting training stability, convergence, and computational efficiency.\n- **Number of output channels in the first convolutional layer** – Controls the number of low-level features extracted from the input images.\n- **Number of output channels in the second convolutional layer** – Determines the model's capacity to learn higher-level feature representations before the fully connected layers.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "251-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(a) — Hyperparameter Selection" + }, + { + "text": "the input images.\n- **Number of output channels in the second convolutional layer** – Determines the model's capacity to learn higher-level feature representations before the fully connected layers.\n\nAmong these, the **number of output channels in the convolutional layers** is a model architecture hyperparameter, while the **learning rate** and **batch size** are optimization hyperparameters.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "251-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(a) — Hyperparameter Selection" + }, + { + "text": "Explanation:\n### Part 3(b) Tuning for convolutional out_chanels size for first and second convolution layer.\n\nPython implementation:\nuse_cuda = True\n\nmodel = Classifier(output_conv1=5, output_conv2=10)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=5, output_conv2=10)\n\nPython implementation:\nimport time\nuse_cuda = True\n\nmodel = Classifier(output_conv1=10, output_conv2=10)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=10, output_conv2=10)\n\nPython implementation:\nimport time\nuse_cuda = True\n\nmodel = Classifier(output_conv1=10, output_conv2=20)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "252-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 3(b) Tuning for convolutional out_chanels size for first and second convolution layer." + }, + { + "text": "ion:\nimport time\nuse_cuda = True\n\nmodel = Classifier(output_conv1=10, output_conv2=20)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=10, output_conv2=20)\n\nPython implementation:\nimport time\nuse_cuda = True\n\nmodel = Classifier(output_conv1=5, output_conv2=20)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=5, output_conv2=20)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "252-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 3(b) Tuning for convolutional out_chanels size for first and second convolution layer." + }, + { + "text": "Explanation:\n#### tunning batch_size \n\nPython implementation:\nuse_cuda = True\n\nmodel = Classifier(output_conv1=5, output_conv2=20)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=5, output_conv2=20, batch_size=512)\n\nPython implementation:\nuse_cuda = True\n\nmodel = Classifier(output_conv1=5, output_conv2=20)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=5, output_conv2=20, batch_size=128)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "253-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": " tunning batch_size " + }, + { + "text": "Explanation:\n#### Tuning learning rate \n\nPython implementation:\nuse_cuda = True\n\nmodel = Classifier(output_conv1=5, output_conv2=20)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=5, output_conv2=20, learning_rate=0.005)\n\nPython implementation:\nuse_cuda = True\n\nmodel = Classifier(output_conv1=5, output_conv2=20)\n\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrain(model, train_set, output_conv1=5, output_conv2=20, learning_rate=0.05)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "254-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": " Tuning learning rate " + }, + { + "text": "Explanation:\n### Part 3(c)\n\n A model with `output_conv1=5`, `output_conv2=20`, `batch_size=128`, `learning_rate=0.05`, and when `epoch=27` has been chosen as the best model ", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "255-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 3(c)" + }, + { + "text": "Explanation:\n### Part 3(d) - Test accuracy of the chosen model (using test set) was 85% and for validation set, accuracy set was 81.6%\n\nPython implementation:\nmodel = Classifier(output_conv1=5, output_conv2=20)\nmodel_path =\"model_bs256_lr0.05_epoch26_conv5_conv20\"\nstate = torch.load(model_path)\nmodel.load_state_dict(state)\n\nPython implementation:\ncriterion = nn.CrossEntropyLoss()\nmodel.cuda()\nprint(\"accuracy, loss and Error of the model for test set=\", get_accuracy(model, criterion=criterion, batch_size=256, train=True, test=True))\n\nPython implementation:\ncriterion = nn.CrossEntropyLoss()\nmodel.cuda()\nprint(\"accuracy, loss and Error of the model for train set=\", get_accuracy(model, criterion=criterion, batch_size=256, train=True))\n\nPython implementation:\ncriterion = nn.CrossEntropyLoss()\nmodel.cuda()\nprint(\"accuracy, loss and Error of the model for validation set=\", get_accuracy(model, criterion=criterion, batch_size=256))", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "256-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 3(d) - Test accuracy of the chosen model (using test set) was 85% and for validation set, accuracy set was 81.6%" + }, + { + "text": "Explanation:\n## Part 4. Transfer Learning [8 pt]\nFor many image classification tasks, it is generally not a good idea to train a very large deep neural network\nmodel from scratch due to the enormous compute requirements and lack of sufficient amounts of training\ndata.\n\nOne of the better options is to try using an existing model that performs a similar task to the one you need\nto solve. This method of utilizing a pre-trained network for other similar tasks is broadly termed **Transfer\nLearning**. In this assignment, we will use Transfer Learning to extract features from the hand gesture\nimages. Then, train a smaller network to use these features as input and classify the hand gestures.\n\nAs you have learned from the CNN lecture, convolution layers extract various features from the images which\nget utilized by the fully connected layers for correct classification. AlexNet architecture played a pivotal", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "257-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4. Transfer Learning [8 pt]" + }, + { + "text": "ed from the CNN lecture, convolution layers extract various features from the images which\nget utilized by the fully connected layers for correct classification. AlexNet architecture played a pivotal\nrole in establishing Deep Neural Nets as a go-to tool for image classification problems and we will use an\nImageNet pre-trained AlexNet model to extract features in this assignment.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "257-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4. Transfer Learning [8 pt]" + }, + { + "text": "Explanation:\n### Part 4(a) -\nHere is the code to load the AlexNet network, with pretrained weights. When you first run the code, PyTorch\nwill download the pretrained weights from the internet.\n\nPython implementation:\nimport torchvision.models\nalexnet = torchvision.models.alexnet(pretrained=True)\n\nExplanation:\nThe alexnet model is split up into two components: *alexnet.features* and *alexnet.classifier*. The\nfirst neural network component, *alexnet.features*, is used to compute convolutional features, which are\ntaken as input in *alexnet.classifier*.\n\nThe neural network alexnet.features expects an image tensor of shape Nx3x224x224 as input and it will\noutput a tensor of shape Nx256x6x6 . (N = batch size).\nIn the following cell of code a function has been defined to calculate and save alexnet features of our datasets for a specificed bach size\n\nPython implementation:\ndef save_features(batch_size):\n torch.manual_seed(1) # set the random seed", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "258-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(a) -" + }, + { + "text": "n has been defined to calculate and save alexnet features of our datasets for a specificed bach size\n\nPython implementation:\ndef save_features(batch_size):\n torch.manual_seed(1) # set the random seed\n train_load=torch.utils.data.DataLoader(train_set, batch_size=batch_size)\n val_load=torch.utils.data.DataLoader(val_set, batch_size=batch_size)\n test_load=torch.utils.data.DataLoader(test_set, batch_size=batch_size)\n\n n=0\n for images, labels in iter(train_load):\n train_features=alexnet.features(images)\n torch.save(train_features,\"train_features_save{0}\".format(n))\n n+=1\n\n n=0\n for images, labels in iter(val_load):\n val_features=alexnet.features(images)\n torch.save(val_features,\"val_features_save{0}\".format(n))\n n+=1\n\n n=0\n for images, labels in iter(test_load):\n test_features=alexnet.features(images)\n torch.save(test_features,\"test_features_save{0}\".format(n))\n\nExplanation:\n**Save the computed features**.", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "258-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(a) -" + }, + { + "text": "Explanation:\n### Part 4(b) — Transfer Learning Classifier\n\nA lightweight neural network classifier was developed to classify the feature representations extracted by the pretrained AlexNet model. Since the convolutional layers of AlexNet already learn rich visual features, only a small fully connected network is required for the final classification task.\n\n### Model Architecture\n\nThe classifier consists of:\n\n- **Two fully connected (linear) layers** for multi-class classification.\n- **ReLU activation** between the fully connected layers to introduce non-linearity and improve optimization.\n- **No convolutional or pooling layers**, as feature extraction is performed entirely by the pretrained AlexNet model.\n- A **hidden layer with 20 units**, providing sufficient model capacity while keeping the classifier lightweight and computationally efficient.\n\n### Design Choices", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "259-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(b) — Transfer Learning Classifier" + }, + { + "text": "y by the pretrained AlexNet model.\n- A **hidden layer with 20 units**, providing sufficient model capacity while keeping the classifier lightweight and computationally efficient.\n\n### Design Choices\n\n- The pretrained AlexNet serves as a fixed feature extractor, eliminating the need to train a deep convolutional network from scratch.\n- A compact classifier reduces the number of trainable parameters, lowering computational cost and decreasing the risk of overfitting on a relatively small dataset.\n- The ReLU activation function enables efficient gradient propagation and faster convergence during training.\n\nPython implementation:\n# # features = ... load precomputed alexnet.features(img) ...\n# output = model(features)\n# prob = F.softmax(output)\n\nPython implementation:\n#Artifical Neural Network Architecture\nclass Model_features(nn.Module):\n def __init__(self):\n super(Model_features, self).__init__()\n self.fc1 = nn.Linear(256 * 6 * 6, 20)\n self.fc2 = nn.Linear(20, 9)\n\n def forward(self, x):", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "259-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(b) — Transfer Learning Classifier" + }, + { + "text": "Architecture\nclass Model_features(nn.Module):\n def __init__(self):\n super(Model_features, self).__init__()\n self.fc1 = nn.Linear(256 * 6 * 6, 20)\n self.fc2 = nn.Linear(20, 9)\n\n def forward(self, x):\n x = x.view(-1, 256 * 6 * 6) #flatten feature data\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n return x", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "259-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(b) — Transfer Learning Classifier" + }, + { + "text": "Explanation:\n### Part 4 (c)\nTrain your new network, including any hyperparameter tuning. Plot and submit the training curve of your\nbest model only.\n\nNote: Depending on how you are caching (saving) your AlexNet features, PyTorch might still be tracking\nupdates to the **AlexNet weights**, which we are not tuning. One workaround is to convert your AlexNet\nfeature tensor into a numpy array, and then back into a PyTorch tensor.\n\nPython implementation:\ndef get_accuracy(model, criterion, batch_size, train=False, test=False):\n if test:\n data=test_set\n a=\"test_features_save\"\n else:\n if train:\n data = train_set\n a=\"train_features_save\"\n else:\n data = val_set\n a=\"val_features_save\"\n\n correct = 0\n error=0\n total = 0\n total_loss = 0.0\n n=0\n for i, data in enumerate(torch.utils.data.DataLoader(data, batch_size=batch_size), 0):\n imgs, labels = data\n if test:\n imgs=alexnet.features(imgs)\n else:\n imgs=torch.load(a+\"{0}\".format(n))\n n+=1\n imgs = torch.from_numpy(imgs.detach().numpy())", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": "taLoader(data, batch_size=batch_size), 0):\n imgs, labels = data\n if test:\n imgs=alexnet.features(imgs)\n else:\n imgs=torch.load(a+\"{0}\".format(n))\n n+=1\n imgs = torch.from_numpy(imgs.detach().numpy())\n #############################################\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n output = model(imgs)\n\n loss = criterion(output, labels)\n\n #select index with maximum prediction score\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n error += pred.ne(labels.view_as(pred)).sum().item()\n total += imgs.shape[0]\n total_loss += loss.item()\n acc=correct/ total\n err=error/ total\n loss = float(total_loss) / (i + 1)\n return acc, loss, err\n\nPython implementation:\ndef trainNet(model, data, batch_size=256, learning_rate=0.01, num_epochs=30):\n torch.manual_seed(1000)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": "=error/ total\n loss = float(total_loss) / (i + 1)\n return acc, loss, err\n\nPython implementation:\ndef trainNet(model, data, batch_size=256, learning_rate=0.01, num_epochs=30):\n torch.manual_seed(1000)\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n criterion = nn.CrossEntropyLoss()\n optimizer = optim.SGD(model.parameters(), lr=learning_rate, momentum=0.9)\n\n iters, losses, train_acc, val_acc = [], [], [], []\n losses_val= []\n losses_train= []\n val_err= []\n train_err= []\n\n # training\n n = 0 # the number of iterations\n start_time = time.time()\n for epoch in range(num_epochs):\n m=0\n for imgs, labels in iter(train_loader):\n\n #############################################\n imgs=torch.load(\"train_features_save{0}\".format(m))\n imgs = torch.from_numpy(imgs.detach().numpy())\n m+=1\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n out = model(imgs) # forward pass", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": ")\n m+=1\n #To Enable GPU Usage\n if use_cuda and torch.cuda.is_available():\n imgs = imgs.cuda()\n labels = labels.cuda()\n #############################################\n\n out = model(imgs) # forward pass\n loss = criterion(out, labels) # compute the total loss\n loss.backward() # backward pass (compute parameter updates)\n optimizer.step() # make the updates for each parameter\n optimizer.zero_grad() # a clean up step for PyTorch\n\n # save the current training information\n iters.append(n)\n losses.append(float(loss)/batch_size) # compute *average* loss\n train_acc.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=True)[0]) # compute training accuracy\n val_acc.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=False)[0]) # compute validation accuracy\n n += 1\n\n losses_train.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=True)[1]) # compute validation loss", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": "h_size=batch_size, train=False)[0]) # compute validation accuracy\n n += 1\n\n losses_train.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=True)[1]) # compute validation loss\n losses_val.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=False)[1]) # compute validation loss\n train_err.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=True)[2]) # compute training accuracy\n val_err.append(get_accuracy(model,criterion=criterion,batch_size=batch_size, train=False)[2]) # compute validation accuracy\n print((\"Epoch {}: Train err: {:.3f}, Train loss: {:.3f} |\"+\n \"Validation err: {:.3f}, Validation loss: {:.3f}\").format(\n epoch + 1,\n train_err[epoch],\n losses_train[epoch],\n val_err[epoch],\n losses_val[epoch]))\n\n # Save the current model (checkpoint) to a file\n model_path =\"model_2_bs{0}_lr{1}_epoch{2}\".format(batch_size,\n learning_rate,\n epoch)\n torch.save(model.state_dict(), model_path)\n print('Finished Training')", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": "the current model (checkpoint) to a file\n model_path =\"model_2_bs{0}_lr{1}_epoch{2}\".format(batch_size,\n learning_rate,\n epoch)\n torch.save(model.state_dict(), model_path)\n print('Finished Training')\n end_time = time.time()\n elapsed_time = end_time - start_time\n print(\"Total time elapsed: {:.2f} seconds\".format(elapsed_time))\n # Write the train/test loss/err into CSV file for plotting later\n epochs = np.arange(1, num_epochs + 1)\n np.savetxt(\"{}_train_err.csv\".format(model_path), train_acc)\n np.savetxt(\"{}_train_loss.csv\".format(model_path), losses)\n np.savetxt(\"{}_val_err.csv\".format(model_path), val_acc)\n # np.savetxt(\"{}_val_loss.csv\".format(model_path), val_loss)\n\n # plotting\n fig = plt.figure(figsize=(12, 10))\n plt.subplot(2, 2, 1)\n plt.title(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n\n plt.subplot(2, 2, 2)\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": "(\"Training Curve\")\n plt.plot(iters, losses, label=\"Train\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Loss\")\n\n plt.subplot(2, 2, 2)\n plt.title(\"Training Curve\")\n plt.plot(iters, train_acc, label=\"Train\")\n plt.plot(iters, val_acc, label=\"Validation\")\n plt.xlabel(\"Iterations\")\n plt.ylabel(\"Training Accuracy\")\n plt.legend(loc='best')\n\n x = len(losses_train)\n plt.subplot(2, 2, 3)\n plt.title(\"Train vs Validation Loss\")\n plt.plot(range(1,x+1), losses_train, label=\"Train\")\n plt.plot(range(1,x+1), losses_val, label=\"Validation\")\n plt.xlabel(\"epoch\")\n plt.ylabel(\"Loss\")\n plt.legend(loc='best')\n\n plt.subplot(2, 2, 4)\n plt.title(\"Train vs Validation error\")\n plt.plot(range(1,x+1), train_err, label=\"Train\")\n plt.plot(range(1,x+1), val_err, label=\"Validation\")\n plt.xlabel(\"epoch\")\n plt.ylabel(\"error\")\n plt.legend(loc='best')\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\nPython implementation:\nbatch_size=128\nuse_cuda = True", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": "plt.legend(loc='best')\n\n print(\"Final Training Accuracy: {}\".format(train_acc[-1]))\n print(\"Final Validation Accuracy: {}\".format(val_acc[-1]))\n\nPython implementation:\nbatch_size=128\nuse_cuda = True\nsave_features(batch_size=batch_size)\nmodel = Model_features()\nif use_cuda and torch.cuda.is_available():\n model.cuda()\n print('CUDA is available! Training on GPU ...')\nelse:\n print('CUDA is not available. Training on CPU ...')\n\ntrainNet(model, train_set, batch_size=batch_size, learning_rate=0.01)", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "260-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4 (c)" + }, + { + "text": "Explanation:\n### Part 4(d) Test accuracy of the model (using transfer learning) was 94.9%. By using transfer learning the test accuracy has increased by 10 %\n\nPython implementation:\nmodel = Model_features()\nmodel_path =\"model_2_bs128_lr0.01_epoch21\"\nstate = torch.load(model_path)\nmodel.load_state_dict(state)\n\nPython implementation:\ncriterion = nn.CrossEntropyLoss()\nmodel.cuda()\nprint(\"accuracy, loss and Error of the model for test set=\", get_accuracy(model, criterion=criterion, batch_size=128, train=True, test=True))", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "261-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Part 4(d) Test accuracy of the model (using transfer learning) was 94.9%. By using transfer learning the test accuracy has increased by 10 %" + }, + { + "text": "Explanation:\n# Results\n\n## Performance Summary\n\nTwo approaches were evaluated for American Sign Language (ASL) hand gesture classification:\n\n1. A custom Convolutional Neural Network (CNN) trained from scratch.\n2. A transfer learning approach using pretrained AlexNet features.\n\n### Custom CNN\n\n| Metric | Value |\n|---------|-------:|\n| Training Accuracy | **96.7%** |\n| Validation Accuracy | **81.6%** |\n| Test Accuracy | **85.9%** |\n\nThe custom CNN achieved strong training performance while demonstrating good generalization on unseen test images.\n\n---\n\n### Transfer Learning (AlexNet)\n\n| Metric | Value |\n|---------|-------:|\n| Training Accuracy | **99.0%** |\n| Validation Accuracy | **89.5%** |\n| Test Accuracy | **94.9%** |\n\nUsing pretrained AlexNet features substantially improved performance over the custom CNN, increasing the test accuracy by approximately **9 percentage points** (from **85.9%** to **94.9%**).\n\n---", + "source": "03_hand_gesture_recognition.ipynb", + "file_type": "ipynb", + "chunk_id": "262-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "computer-vision-pytorch", + "relative_path": "notebooks/03_hand_gesture_recognition.ipynb", + "section": "Results" + }, + { + "text": "# Generative Deep Learning with PyTorch\n\nA collection of practical **Generative Deep Learning** projects implemented using **PyTorch**. This repository explores representation learning, synthetic image generation, and model robustness through Autoencoders, Generative Adversarial Networks (GANs), adversarial attacks, and modern generative modeling techniques. Each notebook has been reorganized, documented, and expanded to demonstrate practical deep learning workflows for generative AI.\n\n---\n\n# 🚀 Repository Overview\n\nThis repository contains practical generative deep learning projects built with **PyTorch**, covering autoencoders, generative adversarial networks, adversarial robustness, and image generation through reproducible, well-documented notebooks.\n\n---\n\n# ✨ Project Highlights\n\n- Autoencoders and Variational Autoencoders (VAEs)\n- Convolutional Autoencoders\n- Generative Adversarial Networks (GANs)\n- Conditional GANs (cGANs)\n- Image colorization using CNNs and GANs", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "ighlights\n\n- Autoencoders and Variational Autoencoders (VAEs)\n- Convolutional Autoencoders\n- Generative Adversarial Networks (GANs)\n- Conditional GANs (cGANs)\n- Image colorization using CNNs and GANs\n- Adversarial attack techniques for evaluating model robustness\n- GPU-ready training pipelines\n- Well-documented notebooks suitable for learning and experimentation\n\n---\n\n# 📚 Topics Covered\n\n- Convolutional Autoencoder\n- Convolutional Autoencoders\n- Latent space representation learning\n- Image reconstruction\n- Generative Adversarial Networks (GANs)\n- Generator and Discriminator architectures\n- Synthetic image generation\n- Adversarial attacks\n- Model robustness\n- Deep generative models\n- PyTorch implementation\n- GPU acceleration with CUDA\n- Deep learning best practices\n- Variational Autoencoders (VAEs)\n- Conditional GANs (cGANs)\n\n---\n\n# 🛠️ Technologies\n\n- Python\n- PyTorch\n- Torchvision\n- NumPy\n- Matplotlib\n- Scikit-learn\n- Pillow (PIL)\n\n---\n\n# 📂 Repository Structure\n\n```text", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "nal Autoencoders (VAEs)\n- Conditional GANs (cGANs)\n\n---\n\n# 🛠️ Technologies\n\n- Python\n- PyTorch\n- Torchvision\n- NumPy\n- Matplotlib\n- Scikit-learn\n- Pillow (PIL)\n\n---\n\n# 📂 Repository Structure\n\n```text\ngenerative-deep-learning-pytorch/\n│\n├── README.md\n├── requirements.txt\n├── .gitignore\n│\n├── notebooks/\n│ ├── 01_autoencoder.ipynb\n│ ├── 02_gan.ipynb\n│ ├── 03_adversarial_attacks.ipynb\n│ └── 04_image_colorization.ipynb\n│\n├── Images/\n│ ├── Colorization.jpg\n│ ├── ColorizationGANs.png\n│ ├── ColorizationVAEs.png\n│ ├── Gans.jpg\n│ ├── Adversarial_Attacks.jpg\n│ └── autoencoder_overview.png\n│\n└── models/\n```\n\n---\n\n# 📖 Notebooks\n\n## Notebook 1 — Autoencoders and Variational Autoencoders\n\n\n\n\n\n\n\n
\n\n\n\n**Autoencoders**\n\n\n\n\n\n**Variational Autoencoder (VAE)**\n\n
\n\n### Overview", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "toencoders**\n\n\n\n\n\n\n\n**Variational Autoencoder (VAE)**\n\n\n\n\n\n### Overview\n\nThis notebook introduces modern autoencoder architectures for unsupervised representation learning using PyTorch. It explores fully connected autoencoders, convolutional autoencoders, and variational autoencoders (VAEs), demonstrating how neural networks learn compact latent representations, reconstruct images, and generate new samples from learned probability distributions.\n\n### Topics Covered\n\n- Fully Connected Autoencoders\n- Convolutional Autoencoders\n- Variational Autoencoders (VAEs)\n- Encoder–Decoder architectures\n- Latent space representation learning\n- Image reconstruction\n- Image generation\n- Transpose convolutions\n- KL-divergence loss\n- MNIST dataset\n- Unsupervised learning\n\n### Notebook\n\n`notebooks/01_autoencoder.ipynb`\n\n---", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "entation learning\n- Image reconstruction\n- Image generation\n- Transpose convolutions\n- KL-divergence loss\n- MNIST dataset\n- Unsupervised learning\n\n### Notebook\n\n`notebooks/01_autoencoder.ipynb`\n\n---\n\n## Notebook 2 — Generative Adversarial Networks (GANs)\n\n### Overview\n\nThis notebook explores **Generative Adversarial Networks (GANs)** for realistic image synthesis using PyTorch. It begins by examining the limitations of reconstruction-based autoencoders for image generation before introducing the adversarial learning framework. The notebook implements both standard GANs and Conditional GANs (cGANs), demonstrating how adversarial training enables the generation of high-quality synthetic images from random latent vectors and class-conditioned inputs.\n\n

\n \n

", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "dversarial training enables the generation of high-quality synthetic images from random latent vectors and class-conditioned inputs.\n\n

\n \n

\n\n* Generative Adversarial Networks consist of a Generator that synthesizes images from random latent vectors and a Discriminator that learns to distinguish generated images from real images.*\n\n### Topics Covered\n\n- Generative Adversarial Networks (GANs)\n- Conditional GANs (cGANs)\n- Generator networks\n- Discriminator networks\n- Adversarial learning\n- Latent space sampling\n- Image synthesis\n- Conditional image generation\n- GAN training challenges\n- MNIST dataset\n\n### Notebook\n\n`notebooks/02_gan.ipynb`\n\n---\n\n## Notebook 3 — Adversarial Attacks on Deep Neural Networks\n\n### Overview\n\nThis notebook investigates adversarial attacks against deep neural networks using PyTorch.", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "ok\n\n`notebooks/02_gan.ipynb`\n\n---\n\n## Notebook 3 — Adversarial Attacks on Deep Neural Networks\n\n### Overview\n\nThis notebook investigates adversarial attacks against deep neural networks using PyTorch. It demonstrates how small, carefully optimized perturbations can cause image classifiers to make incorrect predictions, highlighting important challenges in AI robustness and model security.\n\n![Adversarial Attacks](Images/Adversarial_Attacks.jpg)\n\n### Topics Covered\n\n- Adversarial examples\n- Targeted adversarial attacks\n- Gradient-based optimization\n- Adversarial perturbations\n- Neural network robustness\n- Model security\n- Visualization of adversarial examples\n- PyTorch implementation\n\n### Notebook\n\n`notebooks/03_adversarial_attacks.ipynb`\n\n---\n\n## Notebook 4 — Image Colorization with CNNs and Conditional GANs\n\n\n\n", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "NNs and Conditional GANs\n\n
\n\n\n\n**Part A:** CNN-Based Image Colorization (Encoder–Decoder & U-Net)\n\n
\n\n\n\n\n\n
\n\n\n\n**Part A:** CNN-Based Image Colorization (Encoder–Decoder & U-Net)\n\n\n\n\n\n**Part B:** Conditional GAN (cGAN) for Image Colorization\n\n
\n\n### Overview\n\nThis notebook presents an end-to-end image colorization project implemented in **PyTorch**, bringing together the concepts introduced throughout the previous notebooks. It explores two complementary deep learning approaches for predicting realistic color images from grayscale inputs.\n\n**Part A** formulates image colorization as a supervised image-to-image translation problem using convolutional encoder–decoder networks and U-Net architectures with skip connections.", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "scale inputs.\n\n**Part A** formulates image colorization as a supervised image-to-image translation problem using convolutional encoder–decoder networks and U-Net architectures with skip connections.\n\n**Part B** extends this approach by implementing a **Conditional Generative Adversarial Network (cGAN)**, where a generator learns to produce realistic color images while a discriminator encourages visually plausible outputs through adversarial training.\n\nTogether, these two approaches demonstrate both reconstruction-based and adversarial methods for image colorization.\n\n### Topics Covered\n\n- Image colorization\n- Image-to-image translation\n- Convolutional Neural Networks (CNNs)\n- Encoder–Decoder architectures\n- U-Net with skip connections\n- Image regression\n- Conditional GANs (cGANs)\n- Generator and Discriminator architectures\n- Adversarial learning\n- CIFAR-10 dataset\n- Deep learning with PyTorch\n\n### Notebook\n\n`notebooks/04_image_colorization.ipynb`\n\n---\n\n# 🎯 Learning Objectives", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "Generator and Discriminator architectures\n- Adversarial learning\n- CIFAR-10 dataset\n- Deep learning with PyTorch\n\n### Notebook\n\n`notebooks/04_image_colorization.ipynb`\n\n---\n\n# 🎯 Learning Objectives\n\nThroughout this repository, I explore how to:\n\n- Build autoencoders for unsupervised representation learning.\n- Learn compact latent representations of image data.\n- Reconstruct images using encoder-decoder architectures.\n- Train Generative Adversarial Networks (GANs) and Conditional GANs (cGANs) for image synthesis and image-to-image translation.\n- Evaluate neural network robustness using adversarial attacks.\n- Understand modern generative deep learning techniques.\n- Develop complete generative AI workflows using PyTorch.\n\n---\n\n# 📈 Key Learning Outcomes\n\nThis repository demonstrates practical implementations of:\n\n- Fully Connected Autoencoders\n- Convolutional Autoencoders\n- Variational Autoencoders (VAEs)\n- Generative Adversarial Networks (GANs)\n- Conditional GANs (cGANs)", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-9", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "trates practical implementations of:\n\n- Fully Connected Autoencoders\n- Convolutional Autoencoders\n- Variational Autoencoders (VAEs)\n- Generative Adversarial Networks (GANs)\n- Conditional GANs (cGANs)\n- Adversarial attacks on deep neural networks\n- Image colorization using CNNs and GANs\n- Image-to-image translation\n\n---\n\n# 🚀 Getting Started\n\nClone the repository:\n\n```bash\ngit clone https://github.com/Miladsaeedi70/generative-deep-learning-pytorch.git\n```\n\nNavigate to the project:\n\n```bash\ncd generative-deep-learning-pytorch\n```\n\nInstall the required packages:\n\n```bash\npip install -r requirements.txt\n```\n\nLaunch Jupyter Notebook:\n\n```bash\njupyter notebook\n```\n\nOpen the notebooks in numerical order:\n\n1. Autoencoders\n2. Generative Adversarial Networks\n3. Adversarial Attacks\n4. Generative Modeling\n\n---\n\n# ⭐ About\nThis repository showcases practical implementations of modern generative deep learning techniques using PyTorch.", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-10", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "tive Adversarial Networks\n3. Adversarial Attacks\n4. Generative Modeling\n\n---\n\n# ⭐ About\nThis repository showcases practical implementations of modern generative deep learning techniques using PyTorch. Starting from representation learning with autoencoders and variational autoencoders, it progresses through generative adversarial networks, adversarial robustness, and concludes with an end-to-end image colorization project using both encoder–decoder networks and Conditional GANs.\n\nThe notebooks have been reorganized and modernized to improve readability, reproducibility, and compatibility with recent versions of PyTorch and Python.\n\n---\n\n# 📄 License\n\nThis repository is intended for educational and portfolio purposes.", + "source": "README.md", + "file_type": "md", + "chunk_id": "263-11", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "README.md", + "section": null + }, + { + "text": "Explanation:\n# Autoencoders with PyTorch\n\n## Overview\n\nThis notebook introduces autoencoders and convolutional autoencoders using PyTorch. It demonstrates how neural networks can learn compact latent representations of image data for unsupervised representation learning and image reconstruction.\n\nThe notebook covers both fully connected and convolutional autoencoder architectures, explains transpose convolutions, and trains models to reconstruct handwritten digit images from the MNIST dataset.\n\n![Autoencoder Overview](mushroom_encoder.png)\n\n*Figure 1. An autoencoder compresses the input into a low-dimensional latent representation and reconstructs the original image through a decoder network.*\n\n## Learning Objectives\n\n- Build fully connected autoencoders\n- Design convolutional autoencoders\n- Learn compact latent representations\n- Reconstruct images from latent embeddings\n- Train autoencoders using reconstruction loss\n- Understand transpose convolutions\n- Visualize reconstructed images", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "264-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Autoencoders with PyTorch" + }, + { + "text": "- Learn compact latent representations\n- Reconstruct images from latent embeddings\n- Train autoencoders using reconstruction loss\n- Understand transpose convolutions\n- Visualize reconstructed images\n\nExplanation:\nAutoencoders are neural networks designed for **unsupervised representation learning**. They consist of two complementary components:\n\n- **Encoder:** Compresses the input into a low-dimensional latent representation.\n- **Decoder:** Reconstructs the original input from the latent representation.\n\nThe network is trained to minimize the reconstruction error between the input and the reconstructed output. Through this process, the encoder learns compact feature representations that capture the most informative characteristics of the data.\n\nIn this notebook, fully connected, convolutional, denoising, and variational autoencoders are implemented using PyTorch and applied to the MNIST handwritten digit dataset.\n\nPython implementation:\nimport torch\nimport torch.nn as nn", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "264-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Autoencoders with PyTorch" + }, + { + "text": "ed, convolutional, denoising, and variational autoencoders are implemented using PyTorch and applied to the MNIST handwritten digit dataset.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport matplotlib.pyplot as plt\nfrom torchvision import datasets, transforms\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)[:4096]", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "264-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Autoencoders with PyTorch" + }, + { + "text": "Explanation:\n#### Architecture\nThe architecture is very similar to what we have seen in the past, except now the output will be the same size as the input. Notice also that we apply a sigmoid on the output data, this is to scale the output from 0 to 1.\n\nPython implementation:\nclass Autoencoder(nn.Module):\n def __init__(self):\n super(Autoencoder, self).__init__()\n encoding_dim = 32\n # encoder\n self.fc1 = nn.Linear(28 * 28, encoding_dim)\n # decoder\n self.fc2 = nn.Linear(encoding_dim, 28*28)\n\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n x = F.relu(self.fc1(flattened))\n # sigmoid for scaling output from 0 to 1\n x = F.sigmoid(self.fc2(x))\n return x", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "265-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Architecture" + }, + { + "text": "Explanation:\n#### Training an Autoencoder\n\nHow do we train an autoencoder? How do we know what\nkind of \"encoder\" and \"decoder\" we want?\n\nOne observation is that if we pass an image through the encoder,\nthen pass the result through the decoder, we should get\nroughly the same image back. Ideally, reducing the\ndimensionality and then generating the image should\ngive us the same result.\n\nThis observation provides us a training strategy: we will\nminimize the reconstruction error of the autoencoder\nacross our training data.\nWe use a loss function called 'MSELoss', which\ncomputes the square error at every pixel.\n\nBeyond using a different loss function, the training\nscheme is roughly the same. Note that in the code below,\nwe are using a the optimizer called 'Adam'.\n\nWe switched to this optimizer not because it is specifically\nused for autoencoders, but because this is the optimizer that\npeople tend to use in practice. Feel free to use Adam for your\nother neural networks.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "266-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Training an Autoencoder" + }, + { + "text": "d to this optimizer not because it is specifically\nused for autoencoders, but because this is the optimizer that\npeople tend to use in practice. Feel free to use Adam for your\nother neural networks.\n\nPython implementation:\ndef train(model, num_epochs=5, batch_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss() # mean square error loss\n optimizer = torch.optim.Adam(model.parameters(),\n lr=learning_rate,\n weight_decay=1e-5) # <--\n train_loader = torch.utils.data.DataLoader(mnist_data,\n batch_size=batch_size,\n shuffle=True)\n outputs = []\n for epoch in range(num_epochs):\n for data in train_loader:\n img, _ = data\n recon = model(img)\n img = img.view(-1, 28 * 28)\n loss = criterion(recon, img)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img, recon),)\n return outputs\n\nPython implementation:\nmodel = Autoencoder()\nmax_epochs = 20", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "266-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Training an Autoencoder" + }, + { + "text": "optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img, recon),)\n return outputs\n\nPython implementation:\nmodel = Autoencoder()\nmax_epochs = 20\noutputs = train(model, num_epochs=max_epochs)\n\nExplanation:\nJust like with our ANN we can have additional layers to make a deep autoencoder, also known as a stacked autoencoder.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "266-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Training an Autoencoder" + }, + { + "text": "Explanation:\n## Convolutional Autoencoder\nWhen working with image data it is often better to use a convolutional neural network and take advantage of the spatial relationships. The architecture for the encoder stage of a convolutional autoencoder will consist of standard convolutional layers that we have seen in our previous architectures. The decoder step will be a bit more tricky since we need a way to increase the resolution.\n\nWe need something akin to convolution, but that goes in the *opposite* direction. We will use something called a **transpose convolution**. Transpose convolutions were first called *deconvolutions*, since it is the ``inverse'' of a convolution operation. However, the terminology was confusing since it has nothing to do with the mathematical notion of deconvolution.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "267-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Convolutional Autoencoder" + }, + { + "text": "Explanation:\n### Convolution Transpose\n\nFirst, let's illustrate how convolution transposes can be 'inverses' of convolution layers.\nWe begin by creating a convolutional layer in PyTorch. This is the convolution that we will\ntry to find an ``inverse'' for.\n\nPython implementation:\nconv = nn.Conv2d(in_channels=8,\n out_channels=8,\n kernel_size=5)\n\nExplanation:\nTo illustrate how convolutional layers work, we'll create a random tensor\nand see how the convolution acts on that tensor:\n\nPython implementation:\nx = torch.randn(2, 8, 64, 64)\ny = conv(x)\ny.shape\n\nExplanation:\nA convolution transpose layer with the exact same specifications as above\nwould have the ``reverse'' effect on the shape.\n\nPython implementation:\nconvt = nn.ConvTranspose2d(in_channels=8,\n out_channels=8,\n kernel_size=5)\nconvt(y).shape # should be same as x.shape\n\nExplanation:\nAnd it does! Notice that the weights of this convolution transpose layer are all\nrandom, and are unrelated to the weights of the original `Conv2d`.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "268-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Convolution Transpose" + }, + { + "text": "t(y).shape # should be same as x.shape\n\nExplanation:\nAnd it does! Notice that the weights of this convolution transpose layer are all\nrandom, and are unrelated to the weights of the original `Conv2d`. So, the layer\n`convt` is not the mathematical inverse of the layer `conv`. However, with training,\nthe convolution transpose has the potential to learn to act as an approximate\ninverse to `conv`.\n\nHere is another example of `convt` in action:\n\nPython implementation:\nx = torch.randn(32, 8, 64, 64)\ny = convt(x)\ny.shape\n\nExplanation:\nNotice that the width and height of `y` is `68x68`, because the `kernel_size` is 5\nand we have not added any padding. You can verify that if we start with a tensor\nwith resolution `68x68` and applied a `5x5` convolution, we would end up with\na tensor with resolution `64x64`.\n\nPython implementation:\nconv = nn.Conv2d(in_channels=8,\n out_channels=16,\n kernel_size=5)\ny = torch.randn(32, 8, 68, 68)\nx = conv(y)\nx.shape\n\nExplanation:", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "268-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Convolution Transpose" + }, + { + "text": "nd up with\na tensor with resolution `64x64`.\n\nPython implementation:\nconv = nn.Conv2d(in_channels=8,\n out_channels=16,\n kernel_size=5)\ny = torch.randn(32, 8, 68, 68)\nx = conv(y)\nx.shape\n\nExplanation:\nAs before, we can add a padding to our convolution transpose, just like we added\npadding to our convolution operations:\n\nPython implementation:\nconvt = nn.ConvTranspose2d(in_channels=16,\n out_channels=8,\n kernel_size=5,\n padding=2)\nx = torch.randn(32, 16, 64, 64)\ny = convt(x)\ny.shape\n\nExplanation:\nMore interestingly, we can add a stride to the convolution to increase our resolution!\n\nPython implementation:\nconvt = nn.ConvTranspose2d(in_channels=16,\n out_channels=8,\n kernel_size=5,\n stride=2,\n output_padding=1, # needed because stride=2\n padding=0)\nx = torch.randn(32, 16, 64, 64)\ny = convt(x)\ny.shape\n\nExplanation:\nOur resolution has doubled.\n\nBut what is actually happening? Essentially, we are adding a padding of zeros\nin between every row and every column of `x`.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "268-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Convolution Transpose" + }, + { + "text": "Explanation:\n### Implementation of a Convolutional Autoencoder\n\nTo demonstrate the use of convolution transpose operations,\nwe will build a **convolutional autoencoder**. Below is an example of a *convolutional* autoencoder that uses solely convolutional layers:\n\nPython implementation:\nclass Autoencoder(nn.Module):\n def __init__(self):\n super(Autoencoder, self).__init__()\n self.encoder = nn.Sequential( # like the Composition layer you built\n nn.Conv2d(1, 16, 3, stride=2, padding=1),\n nn.ReLU(),\n nn.Conv2d(16, 32, 3, stride=2, padding=1),\n nn.ReLU(),\n nn.Conv2d(32, 64, 7)\n )\n self.decoder = nn.Sequential(\n nn.ConvTranspose2d(64, 32, 7),\n nn.ReLU(),\n nn.ConvTranspose2d(32, 16, 3, stride=2, padding=1, output_padding=1),\n nn.ReLU(),\n nn.ConvTranspose2d(16, 1, 3, stride=2, padding=1, output_padding=1),\n nn.Sigmoid()\n )\n\n def forward(self, x):\n x = self.encoder(x)\n x = self.decoder(x)\n return x\n\nPython implementation:\n#encoder\n nn.Conv2d(1, 16, 3, stride=2, padding=1),", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "269-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Implementation of a Convolutional Autoencoder" + }, + { + "text": "padding=1, output_padding=1),\n nn.Sigmoid()\n )\n\n def forward(self, x):\n x = self.encoder(x)\n x = self.decoder(x)\n return x\n\nPython implementation:\n#encoder\n nn.Conv2d(1, 16, 3, stride=2, padding=1),\n nn.Conv2d(16, 32, 3, stride=2, padding=1),\n nn.Conv2d(32, 64, 7)\n\n #decoder\n nn.ConvTranspose2d(64, 32, 7),\n nn.ConvTranspose2d(32, 16, 3, stride=2, padding=1, output_padding=1),\n nn.ConvTranspose2d(16, 1, 3, stride=2, padding=1, output_padding=1),", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "269-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Implementation of a Convolutional Autoencoder" + }, + { + "text": "Explanation:\n#### Training a Convolutional Autoencoder\n\nThe training of the convolutional autoencoder will be the same as with the fully-connected autoencoder architecture we introduced in the beginning.\n\nPython implementation:\ndef train(model, num_epochs=5, batch_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss() # mean square error loss\n optimizer = torch.optim.Adam(model.parameters(),\n lr=learning_rate,\n weight_decay=1e-5) # <--\n train_loader = torch.utils.data.DataLoader(mnist_data,\n batch_size=batch_size,\n shuffle=True)\n outputs = []\n for epoch in range(num_epochs):\n for data in train_loader:\n img, _ = data\n recon = model(img)\n loss = criterion(recon, img)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img, recon),)\n return outputs\n\nExplanation:\nNow, we can train this network.\n\nPython implementation:\nmodel = Autoencoder()\nmax_epochs = 20", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "270-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Training a Convolutional Autoencoder" + }, + { + "text": "{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img, recon),)\n return outputs\n\nExplanation:\nNow, we can train this network.\n\nPython implementation:\nmodel = Autoencoder()\nmax_epochs = 20\noutputs = train(model, num_epochs=max_epochs)\n\nExplanation:\nThe loss goes down as we train, meaning that our reconstructed images look more\nand more like the actual images!\n\nLet's look at the training progression: that is, the reconstructed images at\nvarious points of training:\n\nPython implementation:\nfor k in range(0, max_epochs, 5):\n plt.figure(figsize=(9, 2))\n imgs = outputs[k][1].detach().numpy()\n recon = outputs[k][2].detach().numpy()\n for i, item in enumerate(imgs):\n if i >= 9: break\n plt.subplot(2, 9, i+1)\n plt.imshow(item[0])\n\n for i, item in enumerate(recon):\n if i >= 9: break\n plt.subplot(2, 9, 9+i+1)\n plt.imshow(item[0])\n\nExplanation:", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "270-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Training a Convolutional Autoencoder" + }, + { + "text": "em in enumerate(imgs):\n if i >= 9: break\n plt.subplot(2, 9, i+1)\n plt.imshow(item[0])\n\n for i, item in enumerate(recon):\n if i >= 9: break\n plt.subplot(2, 9, 9+i+1)\n plt.imshow(item[0])\n\nExplanation:\nDuring the early stages of training, the reconstructed images contain little recognizable structure because the network has not yet learned meaningful feature representations. As optimization progresses, reconstruction quality improves as the encoder learns increasingly informative latent embeddings.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "270-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Training a Convolutional Autoencoder" + }, + { + "text": "Explanation:\n## Denoising Autoencoder\nWe can add noise to our data and see if we can train an autoencoder to clean out the noise added to our images.\n\nPython implementation:\ndef train(model, num_epochs=5, batch_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss() # mean square error loss\n optimizer = torch.optim.Adam(model.parameters(),\n lr=learning_rate,\n weight_decay=1e-5)\n train_loader = torch.utils.data.DataLoader(mnist_data,\n batch_size=batch_size,\n shuffle=True)\n noise = 0.5\n outputs = []\n for epoch in range(num_epochs):\n for data in train_loader:\n img, _ = data\n\n img_noisy = img + noise * torch.randn(*img.shape)\n img_noisy = np.clip(img_noisy, 0., 1.)\n\n recon = model(img_noisy)\n #img = img.view(-1, 28 * 28)\n loss = criterion(recon, img)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img_noisy, recon),)\n return outputs\n\nPython implementation:", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "271-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Denoising Autoencoder" + }, + { + "text": "s.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img_noisy, recon),)\n return outputs\n\nPython implementation:\nimport numpy as np\n\n# train denoising autoencoder\nmodel = Autoencoder()\nmax_epochs = 20\noutputs = train(model, num_epochs=max_epochs)\n\nPython implementation:\n# reconstructed images at various parts of training\nfor k in range(0, max_epochs, 5):\n plt.figure(figsize=(9, 2))\n imgs = outputs[k][1].detach().numpy()\n recon = outputs[k][2].detach().numpy()\n for i, item in enumerate(imgs):\n if i >= 9: break\n plt.subplot(2, 9, i+1)\n plt.imshow(item[0])\n\n for i, item in enumerate(recon):\n if i >= 9: break\n plt.subplot(2, 9, 9+i+1)\n plt.imshow(item[0])\n\nPython implementation:\ndef train(model, num_epochs=5, batch_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss() # mean square error loss\n optimizer = torch.optim.Adam(model.parameters(),\n lr=learning_rate,", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "271-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Denoising Autoencoder" + }, + { + "text": "model, num_epochs=5, batch_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss() # mean square error loss\n optimizer = torch.optim.Adam(model.parameters(),\n lr=learning_rate,\n weight_decay=1e-5)\n train_loader = torch.utils.data.DataLoader(mnist_data,\n batch_size=batch_size,\n shuffle=True)\n noise = 0.5\n outputs = []\n for epoch in range(num_epochs):\n for data in train_loader:\n img, _ = data\n\n img_noisy = img + noise * torch.randn(*img.shape)\n img_noisy = np.clip(img_noisy, 0., 1.)\n\n recon = model(img_noisy)\n #img = img.view(-1, 28 * 28)\n loss = criterion(recon, img)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img_noisy, recon),)\n return outputs", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "271-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Denoising Autoencoder" + }, + { + "text": "Explanation:\n### Testing on new images\n\nPython implementation:\nbatch_size = 64\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)[4096:4160]\n\ntest_loader = torch.utils.data.DataLoader(mnist_data,\n batch_size=batch_size,\n shuffle=True)\n\n# obtain one batch of test images\ndataiter = iter(test_loader)\nimg, labels = dataiter.next()\n\n# add noise to the test images\nnoise = 0.5\nimg_noisy = img + noise * torch.randn(*img.shape)\nimg_noisy = np.clip(img_noisy, 0., 1.)\n\n# get sample outputs\nrecon = model(img_noisy)\n# prep images for display\nimg_noisy = img_noisy.numpy()\nrecon = recon.detach().numpy()\n\n# reconstructed images at various parts of training\nfor k in range(1):\n plt.figure(figsize=(9, 2))\n\n for i, item in enumerate(img_noisy):\n if i >= 9: break\n plt.subplot(2, 9, i+1)\n plt.imshow(item[0])\n\n for i, item in enumerate(recon):\n if i >= 9: break\n plt.subplot(2, 9, 9+i+1)\n plt.imshow(item[0])\n\nExplanation:", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "272-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Testing on new images" + }, + { + "text": "enumerate(img_noisy):\n if i >= 9: break\n plt.subplot(2, 9, i+1)\n plt.imshow(item[0])\n\n for i, item in enumerate(recon):\n if i >= 9: break\n plt.subplot(2, 9, 9+i+1)\n plt.imshow(item[0])\n\nExplanation:\nAutoencoders are well suited for extracting compressed representations of images and can correct things that don't match the expectation. This approach can be extended to other applications such as handling object occlusion, or filling in missing segments of an image.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "272-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Testing on new images" + }, + { + "text": "Explanation:\n## Structure in the Embeddings\n\nSince we are drastically reducing the dimensionality of the image, there has to be\nsome kind of structure in the embedding space. That is, the network should be able\nto \"save\" space by mapping similar images to similar embeddings.\n\nWe will demonstrate the structure of the embedding space by having\nsome fun with our autoencoders. Let's begin with two images in our training set.\nFor now, we'll choose images of the same digit.\n\nExplanation:\nFirst load pre-denoising autoencoder architecture and training although you could also do this with the denosing autoencoder.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport matplotlib.pyplot as plt\nfrom torchvision import datasets, transforms\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)[:4096]\n\nclass Autoencoder(nn.Module):\n def __init__(self):", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "273-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Structure in the Embeddings" + }, + { + "text": "s, transforms\n\nmnist_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\nmnist_data = list(mnist_data)[:4096]\n\nclass Autoencoder(nn.Module):\n def __init__(self):\n super(Autoencoder, self).__init__()\n self.encoder = nn.Sequential(\n nn.Conv2d(1, 16, 3, stride=2, padding=1),\n nn.ReLU(),\n nn.Conv2d(16, 32, 3, stride=2, padding=1),\n nn.ReLU(),\n nn.Conv2d(32, 64, 7)\n )\n self.decoder = nn.Sequential(\n nn.ConvTranspose2d(64, 32, 7),\n nn.ReLU(),\n nn.ConvTranspose2d(32, 16, 3, stride=2, padding=1, output_padding=1),\n nn.ReLU(),\n nn.ConvTranspose2d(16, 1, 3, stride=2, padding=1, output_padding=1),\n nn.Sigmoid()\n )\n\n def forward(self, x):\n x = self.encoder(x)\n x = self.decoder(x)\n return x\n\ndef train(model, num_epochs=5, batch_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss() # mean square error loss\n optimizer = torch.optim.Adam(model.parameters(),\n lr=learning_rate,\n weight_decay=1e-5) # <--", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "273-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Structure in the Embeddings" + }, + { + "text": "_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss() # mean square error loss\n optimizer = torch.optim.Adam(model.parameters(),\n lr=learning_rate,\n weight_decay=1e-5) # <--\n train_loader = torch.utils.data.DataLoader(mnist_data,\n batch_size=batch_size,\n shuffle=True)\n outputs = []\n for epoch in range(num_epochs):\n for data in train_loader:\n img, _ = data\n recon = model(img)\n loss = criterion(recon, img)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img, recon),)\n return outputs\n\nPython implementation:\nmodel = Autoencoder()\nmax_epochs = 20\noutputs = train(model, num_epochs=max_epochs)\n\nExplanation:\nOutput two sample images\n\nPython implementation:\nimgs = outputs[max_epochs-1][1].detach().numpy()\nplt.subplot(1, 2, 1)\nplt.imshow(imgs[0][0])\nplt.subplot(1, 2, 2)\nplt.imshow(imgs[8][0])\n\nExplanation:", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "273-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Structure in the Embeddings" + }, + { + "text": "Output two sample images\n\nPython implementation:\nimgs = outputs[max_epochs-1][1].detach().numpy()\nplt.subplot(1, 2, 1)\nplt.imshow(imgs[0][0])\nplt.subplot(1, 2, 2)\nplt.imshow(imgs[8][0])\n\nExplanation:\nWe will then compute the **low-dimensional embeddings** of both images,\nby applying the **encoder**:\n\nPython implementation:\nx1 = outputs[max_epochs-1][1][0,:,:,:] # first image\nx2 = outputs[max_epochs-1][1][8,:,:,:] # second image\nx = torch.stack([x1,x2]) # stack them together so we only call `encoder` once\nembedding = model.encoder(x)\ne1 = embedding[0] # embedding of first image\ne2 = embedding[1] # embedding of second image\n\nExplanation:\nNow we will do something interesting. Not only are we goign to run the\ndecoder on those two embeddings `e1` and `e2`, we are also going to **interpolate**\nbetween the two embeddings and decode those as well!\n\nPython implementation:\nembedding_values = []\nfor i in range(0, 10):\n e = e1 * (i/10) + e2 * (10-i)/10\n embedding_values.append(e)", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "273-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Structure in the Embeddings" + }, + { + "text": "**interpolate**\nbetween the two embeddings and decode those as well!\n\nPython implementation:\nembedding_values = []\nfor i in range(0, 10):\n e = e1 * (i/10) + e2 * (10-i)/10\n embedding_values.append(e)\nembedding_values = torch.stack(embedding_values)\n\nrecons = model.decoder(embedding_values)\n\nExplanation:\nLet's plot the reconstructions of each interpolated values.\nThe original images are shown below too:\n\nPython implementation:\nplt.figure(figsize=(10, 2))\nfor i, recon in enumerate(recons.detach().numpy()):\n plt.subplot(2,10,i+1)\n plt.imshow(recon[0])\nplt.subplot(2,10,11)\nplt.imshow(imgs[8][0])\nplt.subplot(2,10,20)\nplt.imshow(imgs[0][0])\n\nExplanation:\nNotice that there is a smooth transition between the two images!\nThe middle images are likely new, in that there are no training images\nthat are exactly like any of the generated images.\n\nAs promised, we can do the same thing with two images containing\ndifferent digits. There should be a smooth transition between\nthe two digits.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "273-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Structure in the Embeddings" + }, + { + "text": "ges\nthat are exactly like any of the generated images.\n\nAs promised, we can do the same thing with two images containing\ndifferent digits. There should be a smooth transition between\nthe two digits.\n\nPython implementation:\ndef interpolate(index1, index2):\n x1 = mnist_data[index1][0]\n x2 = mnist_data[index2][0]\n x = torch.stack([x1,x2])\n embedding = model.encoder(x)\n e1 = embedding[0] # embedding of first image\n e2 = embedding[1] # embedding of second image\n\n embedding_values = []\n for i in range(0, 10):\n e = e1 * (i/10) + e2 * (10-i)/10\n embedding_values.append(e)\n embedding_values = torch.stack(embedding_values)\n\n recons = model.decoder(embedding_values)\n\n plt.figure(figsize=(10, 2))\n for i, recon in enumerate(recons.detach().numpy()):\n plt.subplot(2,10,i+1)\n plt.imshow(recon[0])\n plt.subplot(2,10,11)\n plt.imshow(x2[0])\n plt.subplot(2,10,20)\n plt.imshow(x1[0])\n\ninterpolate(0, 1)\n\nPython implementation:\ndef interpolate_pixel(index1, index2):\n x1 = mnist_data[index1][0]", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "273-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Structure in the Embeddings" + }, + { + "text": "con[0])\n plt.subplot(2,10,11)\n plt.imshow(x2[0])\n plt.subplot(2,10,20)\n plt.imshow(x1[0])\n\ninterpolate(0, 1)\n\nPython implementation:\ndef interpolate_pixel(index1, index2):\n x1 = mnist_data[index1][0]\n x2 = mnist_data[index2][0]\n\n interpolated_values = []\n for i in range(0, 10):\n e = x1 * (i/10) + x2 * (10-i)/10\n interpolated_values.append(e)\n\n plt.figure(figsize=(10, 2))\n for i, recon in enumerate(interpolated_values):\n plt.subplot(2,10,i+1)\n plt.imshow(recon[0])\n plt.subplot(2,10,11)\n plt.imshow(x2[0])\n plt.subplot(2,10,20)\n plt.imshow(x1[0])\n\nExplanation:\nWhat happens if we randomly initialize in the embedding space?\n\nExplanation:\nThe variational autoencoder will allow us to randomly initialize in the embedding space to generate new MNIST-like samples. Provided below is sample code showing how to do that.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "273-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Structure in the Embeddings" + }, + { + "text": "Explanation:\n## Variational Autoencoder\n\n![alt text](http://kvfrans.com/content/images/2016/08/vae.jpg)\n\nTo allow us to sample from the embedding space and generate new images we add a constraint on the encoding network, that forces it to generate latent vectors that roughly follow a unit Gaussian distribution. This constraint is what separates a variational autoencoder from the ones we've seen up to now.\n\nNow generating new images requires that we sample a latent vector from the unit gaussian and pass it into the decoder.\n\nAs shown in the figure, we will have encoding and decoding networks similar to what we used before, whether fully-connected or convolutional layers. Then we add two additional linear layers to hold the mean and standard deviation vectors of the embedding space. We will need some way to generate a sampled latent space which will act as input to the decoding network.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "274-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Variational Autoencoder" + }, + { + "text": "tional linear layers to hold the mean and standard deviation vectors of the embedding space. We will need some way to generate a sampled latent space which will act as input to the decoding network.\n\nWe will also need to update our loss function to use Kullback-Leibler divergence to constrain the embedding space to follow a unit Gaussian distribution. You will not be required to know the math behind this.\n\nA demonstration of the variational autoencoder is provided below.\n\nPython implementation:\n#Setup\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nfrom torchvision import datasets, transforms\n\n# train data\ntrain_loader = torch.utils.data.DataLoader(\n datasets.MNIST('data', train=True, download=True,\n transform=transforms.ToTensor()),\n batch_size=64, shuffle=True)\n\n# test data\ntest_loader = torch.utils.data.DataLoader(\n datasets.MNIST('data', train=False, transform=transforms.ToTensor()),\n batch_size=64, shuffle=True)", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "274-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Variational Autoencoder" + }, + { + "text": "oTensor()),\n batch_size=64, shuffle=True)\n\n# test data\ntest_loader = torch.utils.data.DataLoader(\n datasets.MNIST('data', train=False, transform=transforms.ToTensor()),\n batch_size=64, shuffle=True)\n\nPython implementation:\n# dimensions of latent space\nzdim = 25\n\n# Variational Autoencoder\nclass Autoencoder(nn.Module):\n def __init__(self):\n super(Autoencoder, self).__init__()\n\n # encoder\n self.fc1 = nn.Linear(28 * 28, 350)\n self.relu = nn.ReLU()\n self.fc2m = nn.Linear(350, zdim) # mu layer\n self.fc2s = nn.Linear(350, zdim) # sd layer\n\n # decoder\n self.fc3 = nn.Linear(zdim, 350)\n self.fc4 = nn.Linear(350, 28 * 28)\n self.sigmoid = nn.Sigmoid()\n\n def encode(self, x):\n h1 = self.relu(self.fc1(x))\n return self.fc2m(h1), self.fc2s(h1)\n\n # reparameterize\n def reparameterize(self, mu, logvar):\n if self.training:\n std = logvar.mul(0.5).exp_()\n eps = std.data.new(std.size()).normal_()\n return eps.mul(std).add_(mu)\n else:\n return mu\n\n def decode(self, z):\n h3 = self.relu(self.fc3(z))", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "274-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Variational Autoencoder" + }, + { + "text": "logvar):\n if self.training:\n std = logvar.mul(0.5).exp_()\n eps = std.data.new(std.size()).normal_()\n return eps.mul(std).add_(mu)\n else:\n return mu\n\n def decode(self, z):\n h3 = self.relu(self.fc3(z))\n return self.sigmoid(self.fc4(h3))\n\n def forward(self, x):\n mu, logvar = self.encode(x.view(-1, 28 * 28))\n z = self.reparameterize(mu, logvar)\n return self.decode(z), mu, logvar\n\nPython implementation:\n# loss function for VAE are unique and use Kullback-Leibler\n# divergence measure to force distribution to match unit Gaussian\ndef loss_function(recon_x, x, mu, logvar):\n bce = F.binary_cross_entropy(recon_x, x.view(-1, 28 * 28))\n kld = -0.5 * torch.sum(1 + logvar - mu.pow(2) - logvar.exp())\n kld /= batch_size * 28 * 28\n return bce + kld\n\nPython implementation:\ndef train(model, num_epochs = 1, batch_size = 64, learning_rate = 1e-3):\n model.train() #train mode\n torch.manual_seed(42)\n\n train_loader = torch.utils.data.DataLoader(datasets.MNIST('data',", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "274-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Variational Autoencoder" + }, + { + "text": "ntation:\ndef train(model, num_epochs = 1, batch_size = 64, learning_rate = 1e-3):\n model.train() #train mode\n torch.manual_seed(42)\n\n train_loader = torch.utils.data.DataLoader(datasets.MNIST('data',\n train=True, download=True, transform=transforms.ToTensor()),\n batch_size = batch_size, shuffle = True)\n\n optimizer = optim.Adam(model.parameters(), learning_rate)\n\n for epoch in range(num_epochs):\n for data in train_loader: # load batch\n img, _ = data\n\n recon, mu, logvar = model(img)\n loss = loss_function(recon, img, mu, logvar) # calculate loss\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n\nPython implementation:\nbatch_size = 64\n\nmodel = Autoencoder()\ntrain(model, num_epochs = 30, batch_size = batch_size)\n\nPython implementation:\n# generate random samples in latent space\nmodel.eval()\nsample = torch.randn(64, zdim)\nsample = model.decode(sample)\n\nimport matplotlib.pyplot as plt", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "274-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Variational Autoencoder" + }, + { + "text": ", batch_size = batch_size)\n\nPython implementation:\n# generate random samples in latent space\nmodel.eval()\nsample = torch.randn(64, zdim)\nsample = model.decode(sample)\n\nimport matplotlib.pyplot as plt\nimgs = sample.data.view(64, 28, 28).numpy()\nplt.imshow(imgs[4])\n\nPython implementation:\n# display images\nfor k in range(1):\n plt.figure(figsize=(8, 8))\n\n for i, item in enumerate(imgs):\n plt.subplot(8, 8, i+1)\n plt.imshow(item)\n\nExplanation:\nIn summary we have learned about several different autoencoders. The different architectures we explored with stacked, convolutional, denoising, and variational autoencoders can be combined to take advantage of their strengths and weaknesses to develop an architecture best suited for your problem.\n\nIn this tutorial we didn't go over **semi-supervised learning** which is another very practical application of autoencoders.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "274-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Variational Autoencoder" + }, + { + "text": "hs and weaknesses to develop an architecture best suited for your problem.\n\nIn this tutorial we didn't go over **semi-supervised learning** which is another very practical application of autoencoders. By learning embeddings from unlabeled data, autoencoders can improve model performance in situations where labeled data may be scarce.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "274-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Variational Autoencoder" + }, + { + "text": "Explanation:\n# Results\n\n## Summary\n\nThis notebook demonstrated several autoencoder architectures for unsupervised representation learning using the MNIST dataset.\n\n### Implemented Models\n\n- Fully Connected Autoencoder\n- Convolutional Autoencoder\n- Denoising Autoencoder\n- Variational Autoencoder (VAE)\n\n### Key Outcomes\n\n- Learned compact latent representations of handwritten digits.\n- Successfully reconstructed input images with both fully connected and convolutional architectures.\n- Demonstrated image denoising using a denoising autoencoder.\n- Explored interpolation within the latent space to generate smooth transitions between handwritten digits.\n- Introduced variational autoencoders for probabilistic generative modeling.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "275-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Results" + }, + { + "text": "Explanation:\n# Conclusion\n\nAutoencoders provide an effective framework for learning compressed feature representations without requiring labeled data. Throughout this notebook, progressively more advanced architectures—including convolutional, denoising, and variational autoencoders—were implemented to demonstrate reconstruction, representation learning, image denoising, and generative modeling. These techniques form the foundation for many modern deep generative models and self-supervised learning approaches.", + "source": "01_autoencoder.ipynb", + "file_type": "ipynb", + "chunk_id": "276-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/01_autoencoder.ipynb", + "section": "Conclusion" + }, + { + "text": "Explanation:\n# Generative Adversarial Networks (GANs) with PyTorch\n\n## Overview\n\nThis notebook introduces **Generative Adversarial Networks (GANs)** for image generation using PyTorch. It begins by reviewing the limitations of traditional autoencoders for generative tasks and motivates the adversarial learning framework used by GANs.\n\nThe notebook implements both standard GANs and Conditional GANs (cGANs), demonstrating how a generator and discriminator are trained simultaneously to synthesize realistic handwritten digit images. It also explores adversarial optimization, latent space sampling, and conditional image generation.\n\n## Learning Objectives\n\nBy completing this notebook, you will learn how to:\n\n- Understand the limitations of autoencoders for image generation\n- Build Generator and Discriminator networks\n- Train a Generative Adversarial Network (GAN)\n- Understand adversarial loss functions\n- Generate realistic images from random latent vectors", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "277-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Generative Adversarial Networks (GANs) with PyTorch" + }, + { + "text": "e generation\n- Build Generator and Discriminator networks\n- Train a Generative Adversarial Network (GAN)\n- Understand adversarial loss functions\n- Generate realistic images from random latent vectors\n- Implement Conditional GANs (cGANs)\n- Generate class-conditioned images\n- Visualize synthetic image generation", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "277-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Generative Adversarial Networks (GANs) with PyTorch" + }, + { + "text": "Explanation:\n## From Autoencoders to GANs\n\nBefore introducing GANs, we briefly review autoencoders and discuss why reconstruction-based models often produce blurry outputs. This limitation motivates the adversarial learning framework used by Generative Adversarial Networks.\n\nHere is the code that we wrote back in the autoencoder lecture.\nThe autoencoder model consists of an **encoder** that maps\nimages to a vector embedding, and a **decoder** that reconstructs\nimages from an embedding.\n\nPython implementation:\nclass Autoencoder(nn.Module):\n def __init__(self):\n super(Autoencoder, self).__init__()\n self.encoder = nn.Sequential(\n nn.Conv2d(1, 16, 3, stride=2, padding=1),\n nn.ReLU(),\n nn.Conv2d(16, 32, 3, stride=2, padding=1),\n nn.ReLU(),\n nn.Conv2d(32, 64, 7)\n )\n self.decoder = nn.Sequential(\n nn.ConvTranspose2d(64, 32, 7),\n nn.ReLU(),\n nn.ConvTranspose2d(32, 16, 3, stride=2, padding=1, output_padding=1),\n nn.ReLU(),\n nn.ConvTranspose2d(16, 1, 3, stride=2, padding=1, output_padding=1),", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "278-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "From Autoencoders to GANs" + }, + { + "text": "l(\n nn.ConvTranspose2d(64, 32, 7),\n nn.ReLU(),\n nn.ConvTranspose2d(32, 16, 3, stride=2, padding=1, output_padding=1),\n nn.ReLU(),\n nn.ConvTranspose2d(16, 1, 3, stride=2, padding=1, output_padding=1),\n nn.Sigmoid()\n )\n\n def forward(self, x):\n x = self.encoder(x)\n x = self.decoder(x)\n return x\n\nExplanation:\nWe trained an autoencoder model on the reconstruction loss: the difference in\npixel intensities between a real image and its reconstruction.\n\nPython implementation:\ndef train(model, num_epochs=5, batch_size=64, learning_rate=1e-3):\n torch.manual_seed(42)\n criterion = nn.MSELoss()\n optimizer = torch.optim.Adam(model.parameters(), lr=learning_rate, weight_decay=1e-5)\n train_loader = torch.utils.data.DataLoader(mnist_data, batch_size=batch_size, shuffle=True)\n outputs = []\n for epoch in range(num_epochs):\n for data in train_loader:\n img, label = data\n recon = model(img)\n loss = criterion(recon, img)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "278-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "From Autoencoders to GANs" + }, + { + "text": "utputs = []\n for epoch in range(num_epochs):\n for data in train_loader:\n img, label = data\n recon = model(img)\n loss = criterion(recon, img)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n print('Epoch:{}, Loss:{:.4f}'.format(epoch+1, float(loss)))\n outputs.append((epoch, img, recon),)\n torch.save(model.state_dict(), \"autoencoder%d.pt\" % epoch)\n return outputs\n\nmodel = Autoencoder()\noutputs = train(model, num_epochs=5)\n\nPython implementation:\n# Choose a model to load -- after 2 epochs of training\nckpt = torch.load(\"autoencoder1.pt\")\nmodel.load_state_dict(ckpt)\n\nExplanation:\nLet's take a look at one MNIST image from training,\nand its autoencoder reconstruction:\n\nPython implementation:\noriginal = mnist_data[0][0].unsqueeze(0)\nemb = model.encoder(original)\nrecon_img = model.decoder(emb).detach().numpy()[0,0,:,:]\n\n# plot the original image\nplt.subplot(1,2,1)\nplt.title(\"original\")\nplt.imshow(original[0][0], cmap='gray')\n\n# plot the reconstructed\nplt.subplot(1,2,2)", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "278-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "From Autoencoders to GANs" + }, + { + "text": "= model.decoder(emb).detach().numpy()[0,0,:,:]\n\n# plot the original image\nplt.subplot(1,2,1)\nplt.title(\"original\")\nplt.imshow(original[0][0], cmap='gray')\n\n# plot the reconstructed\nplt.subplot(1,2,2)\nplt.title(\"reconstruction\")\nplt.imshow(recon_img, cmap='gray')\n\nExplanation:\nThe reconstruction is reasonable, but notice that the reconstruction\nis blurrier than\nthe original image. If we perturb the embedding to generate a new image, we\nstill should see this blurriness:\n\nPython implementation:\n# Run this a few times\nx = emb + 10 * torch.randn(1, 64, 1, 1) # add a random perturbation\n\n# reconstruct image and plot\nimg = model.decoder(x)[0,0,:,:]\nimg = img.detach().numpy()\nplt.title(\"perturbed reconstruction\")\nplt.imshow(img, cmap='gray')", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "278-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "From Autoencoders to GANs" + }, + { + "text": "Explanation:\n## Limitations of Autoencoders\n\nAutoencoders are trained using reconstruction losses such as Mean Squared Error (MSE). While effective for representation learning, these losses encourage the network to predict average pixel values, often producing blurry reconstructions. Since the objective measures pixel-wise similarity rather than perceptual realism, generated images may appear smooth and lack fine details.\n\nGenerative Adversarial Networks address this limitation by replacing the handcrafted reconstruction objective with a learned discriminator that evaluates the realism of generated images.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass Discriminator(nn.Module):\n\n def __init__(self):\n super(Discriminator, self).__init__()\n self.fc1 = nn.Linear(28*28, 128)\n self.fc2 = nn.Linear(128, 64)\n self.fc3 = nn.Linear(64, 32)\n self.fc4 = nn.Linear(32, 1)\n self.dropout = nn.Dropout(0.3)\n\n def forward(self, x):", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "279-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Limitations of Autoencoders" + }, + { + "text": "r, self).__init__()\n self.fc1 = nn.Linear(28*28, 128)\n self.fc2 = nn.Linear(128, 64)\n self.fc3 = nn.Linear(64, 32)\n self.fc4 = nn.Linear(32, 1)\n self.dropout = nn.Dropout(0.3)\n\n def forward(self, x):\n x = x.view(-1, 28*28) # flatten image\n x = F.leaky_relu(self.fc1(x), 0.2)\n x = self.dropout(x)\n x = F.leaky_relu(self.fc2(x), 0.2)\n x = self.dropout(x)\n x = F.leaky_relu(self.fc3(x), 0.2)\n x = self.dropout(x)\n out = self.fc4(x)\n return out\n\nclass Generator(nn.Module):\n\n def __init__(self):\n super(Generator, self).__init__()\n self.fc1 = nn.Linear(100, 32)\n self.fc2 = nn.Linear(32, 64)\n self.fc3 = nn.Linear(64, 128)\n self.fc4 = nn.Linear(128, 28*28)\n self.dropout = nn.Dropout(0.3)\n\n def forward(self, x):\n x = F.leaky_relu(self.fc1(x), 0.2)\n x = self.dropout(x)\n x = F.leaky_relu(self.fc2(x), 0.2)\n x = self.dropout(x)\n x = F.leaky_relu(self.fc3(x), 0.2)\n x = self.dropout(x)\n out = F.tanh(self.fc4(x))\n return out\n\nD = Discriminator()\nG = Generator()", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "279-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Limitations of Autoencoders" + }, + { + "text": "Explanation:\n## GAN Architecture\n\nA Generative Adversarial Network consists of two neural networks trained simultaneously:\n\n- **Generator (G):** Learns to transform random latent vectors into realistic synthetic images.\n- **Discriminator (D):** Learns to distinguish between real images and images generated by the generator.\n\nDuring training, the generator attempts to fool the discriminator, while the discriminator continuously improves its ability to detect generated samples. This adversarial process enables the generator to produce increasingly realistic images.\n\nExplanation:\nLet's try training the network.\n\nPython implementation:\nfrom torchvision import datasets\nimport torchvision.transforms as transforms\nimport numpy as np\nimport torch\nimport matplotlib.pyplot as plt\nimport torch.optim as optim\nimport torch.utils.data\n\ndef train(G, D, lr=0.002, batch_size=64, num_epochs=20):\n\n rand_size = 100;\n\n # optimizers for generator and discriminator", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "280-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "GAN Architecture" + }, + { + "text": "atplotlib.pyplot as plt\nimport torch.optim as optim\nimport torch.utils.data\n\ndef train(G, D, lr=0.002, batch_size=64, num_epochs=20):\n\n rand_size = 100;\n\n # optimizers for generator and discriminator\n d_optimizer = optim.Adam(D.parameters(), lr)\n g_optimizer = optim.Adam(G.parameters(), lr)\n\n # define loss function\n criterion = nn.BCEWithLogitsLoss()\n\n # get the training datasets\n train_data = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\n\n # prepare data loader\n train_loader = torch.utils.data.DataLoader(train_data, batch_size=batch_size, shuffle=True)\n\n # keep track of loss and generated, \"fake\" samples\n samples = []\n losses = []\n\n # fixed data for testing\n sample_size=16\n test_noise = np.random.uniform(-1, 1, size=(sample_size, rand_size))\n test_noise = torch.from_numpy(test_noise).float()\n\n for epoch in range(num_epochs):\n D.train()\n G.train()\n\n for batch_i, (real_images, _) in enumerate(train_loader):\n\n batch_size = real_images.size(0)", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "280-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "GAN Architecture" + }, + { + "text": "st_noise = torch.from_numpy(test_noise).float()\n\n for epoch in range(num_epochs):\n D.train()\n G.train()\n\n for batch_i, (real_images, _) in enumerate(train_loader):\n\n batch_size = real_images.size(0)\n\n # rescale images to range -1 to 1\n real_images = real_images*2 - 1\n\n # === Train the Discriminator ===\n\n d_optimizer.zero_grad()\n\n # discriminator losses on real images\n D_real = D(real_images)\n labels = torch.zeros(batch_size)\n d_real_loss = criterion(D_real.squeeze(), labels)\n\n # discriminator losses on fake images\n z = np.random.uniform(-1, 1, size=(batch_size, rand_size))\n z = torch.from_numpy(z).float()\n fake_images = G(z)\n\n D_fake = D(fake_images)\n labels = torch.ones(batch_size) # fake labels = 1\n d_fake_loss = criterion(D_fake.squeeze(), labels)\n\n # add up losses and update parameters\n d_loss = d_real_loss + d_fake_loss\n d_loss.backward()\n d_optimizer.step()\n\n # === Train the Generator ===\n g_optimizer.zero_grad()\n\n # generator losses on fake images", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "280-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "GAN Architecture" + }, + { + "text": "up losses and update parameters\n d_loss = d_real_loss + d_fake_loss\n d_loss.backward()\n d_optimizer.step()\n\n # === Train the Generator ===\n g_optimizer.zero_grad()\n\n # generator losses on fake images\n z = np.random.uniform(-1, 1, size=(batch_size, rand_size))\n z = torch.from_numpy(z).float()\n fake_images = G(z)\n\n D_fake = D(fake_images)\n labels = torch.zeros(batch_size) #flipped labels\n\n # compute loss and update parameters\n g_loss = criterion(D_fake.squeeze(), labels)\n g_loss.backward()\n g_optimizer.step()\n\n # print loss\n print('Epoch [%d/%d], d_loss: %.4f, g_loss: %.4f, '\n % (epoch + 1, num_epochs, d_loss.item(), g_loss.item()))\n\n # append discriminator loss and generator loss\n losses.append((d_loss.item(), g_loss.item()))\n\n # plot images\n G.eval()\n D.eval()\n test_images = G(test_noise)\n\n plt.figure(figsize=(9, 3))\n for k in range(16):\n plt.subplot(2, 8, k+1)\n plt.imshow(test_images[k,:].data.numpy().reshape(28, 28), cmap='Greys')\n plt.show()\n\n return losses\n\nPython implementation:", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "280-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "GAN Architecture" + }, + { + "text": "plt.figure(figsize=(9, 3))\n for k in range(16):\n plt.subplot(2, 8, k+1)\n plt.imshow(test_images[k,:].data.numpy().reshape(28, 28), cmap='Greys')\n plt.show()\n\n return losses\n\nPython implementation:\nfig, ax = plt.subplots()\nlosses = np.array(losses)\nplt.plot(losses.T[0], label='Discriminator')\nplt.plot(losses.T[1], label='Generator')\nplt.title(\"Training Losses\")\nplt.legend()", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "280-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "GAN Architecture" + }, + { + "text": "Explanation:\n## Training Challenges\n\nTraining GANs is significantly more challenging than training conventional supervised neural networks. Because the generator and discriminator continuously compete against one another, their losses often fluctuate throughout training rather than decreasing monotonically.\n\nSuccessful GAN training typically requires careful tuning of learning rates, network architectures, and optimization strategies to maintain a balance between the two networks.", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "281-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Training Challenges" + }, + { + "text": "Explanation:\n## Conditional Generative Adversarial Networks (cGANs)\n\nStandard GANs generate images without user control over the output. Conditional GANs (cGANs) address this limitation by conditioning both the generator and discriminator on additional information, such as class labels. This enables the model to generate images belonging to a specified category while retaining the adversarial training framework.", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "282-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Conditional Generative Adversarial Networks (cGANs)" + }, + { + "text": "Explanation:\n## Example Conditional GAN\n\nThis kernel is a PyTorch implementation of Conditional GAN, which is a GAN that allows you to choose the label of the generated image. The generator and the discriminator are going to be simple feedforward networks, so I guess the images won't be as good as in this nice kernel by Sergio Gámez. I used this implementation by eriklindernoren as inspiration.\n\nPython implementation:\n%matplotlib inline\nimport torch\nimport torch.nn as nn\nimport pandas as pd\nimport numpy as np\nfrom torchvision import transforms\nfrom torchvision import datasets\n\nfrom torch.utils.data import Dataset, DataLoader\nfrom PIL import Image\nfrom torch import autograd\nfrom torch.autograd import Variable\nfrom torchvision.utils import make_grid\nimport matplotlib.pyplot as plt\n\nExplanation:\nLet's start by defining a Dataset class:\n\n- Data Loading and Processing Tutorial on PyTorch's documentation\n- torchvision has a built-in class for Fashion MNIST\n\nPython implementation:", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": "Explanation:\nLet's start by defining a Dataset class:\n\n- Data Loading and Processing Tutorial on PyTorch's documentation\n- torchvision has a built-in class for Fashion MNIST\n\nPython implementation:\n#load from google drive\nfrom google.colab import drive\ndrive.mount('/content/gdrive')\n\n# location on Google Drive\nmaster_path = '/content/gdrive/My Drive/1_0 Teaching Related/MIE1517 - Deep Learning/Project 3 - GAN/cGAN_MNIST/'\n#master_path = '/content/gdrive/My Drive/'\n\nPython implementation:\n# transform = transforms.Compose([\n# transforms.ToTensor(),\n# transforms.Normalize(mean=(0.5, 0.5, 0.5), std=(0.5, 0.5, 0.5))\n# ])\n\ntransform = transforms.Compose([\n transforms.ToTensor(),\n transforms.Normalize(mean=(0.5,), std=(0.5,))\n])\n\n# get the training datasets\ndataset = datasets.MNIST('data', train=True, download=True, transform=transform)\n\n# prepare data loader\ndata_loader = torch.utils.data.DataLoader(dataset, batch_size=64, shuffle=True)\n\nExplanation:", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": "s\ndataset = datasets.MNIST('data', train=True, download=True, transform=transform)\n\n# prepare data loader\ndata_loader = torch.utils.data.DataLoader(dataset, batch_size=64, shuffle=True)\n\nExplanation:\nNow let's define the generator and the discriminator, which are simple MLPs. I'm going to use an embedding layer for the label:\n\nPython implementation:\nclass Discriminator(nn.Module):\n def __init__(self):\n super().__init__()\n\n self.label_emb = nn.Embedding(10, 10)\n\n self.model = nn.Sequential(\n nn.Linear(794, 1024),\n nn.LeakyReLU(0.2, inplace=True),\n nn.Dropout(0.3),\n nn.Linear(1024, 512),\n nn.LeakyReLU(0.2, inplace=True),\n nn.Dropout(0.3),\n nn.Linear(512, 256),\n nn.LeakyReLU(0.2, inplace=True),\n nn.Dropout(0.3),\n nn.Linear(256, 1),\n nn.Sigmoid()\n )\n\n def forward(self, x, labels):\n x = x.view(x.size(0), 784)\n c = self.label_emb(labels)\n x = torch.cat([x, c], 1)\n out = self.model(x)\n return out.squeeze()\n\nPython implementation:\nclass Generator(nn.Module):\n def __init__(self):", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": ":\n x = x.view(x.size(0), 784)\n c = self.label_emb(labels)\n x = torch.cat([x, c], 1)\n out = self.model(x)\n return out.squeeze()\n\nPython implementation:\nclass Generator(nn.Module):\n def __init__(self):\n super().__init__()\n\n self.label_emb = nn.Embedding(10, 10)\n\n self.model = nn.Sequential(\n nn.Linear(110, 256),\n nn.LeakyReLU(0.2, inplace=True),\n nn.Linear(256, 512),\n nn.LeakyReLU(0.2, inplace=True),\n nn.Linear(512, 1024),\n nn.LeakyReLU(0.2, inplace=True),\n nn.Linear(1024, 784),\n nn.Tanh()\n )\n\n def forward(self, z, labels):\n z = z.view(z.size(0), 100)\n c = self.label_emb(labels)\n x = torch.cat([z, c], 1)\n out = self.model(x)\n return out.view(x.size(0), 28, 28)\n\nPython implementation:\ngenerator = Generator().cuda()\ndiscriminator = Discriminator().cuda()\n\nPython implementation:\ncriterion = nn.BCELoss()\nd_optimizer = torch.optim.Adam(discriminator.parameters(), lr=1e-4)\ng_optimizer = torch.optim.Adam(generator.parameters(), lr=1e-4)\n\nPython implementation:", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": "on implementation:\ncriterion = nn.BCELoss()\nd_optimizer = torch.optim.Adam(discriminator.parameters(), lr=1e-4)\ng_optimizer = torch.optim.Adam(generator.parameters(), lr=1e-4)\n\nPython implementation:\ndef generator_train_step(batch_size, discriminator, generator, g_optimizer, criterion):\n g_optimizer.zero_grad()\n z = Variable(torch.randn(batch_size, 100)).cuda()\n fake_labels = Variable(torch.LongTensor(np.random.randint(0, 10, batch_size))).cuda()\n fake_images = generator(z, fake_labels)\n validity = discriminator(fake_images, fake_labels)\n g_loss = criterion(validity, Variable(torch.ones(batch_size)).cuda())\n g_loss.backward()\n g_optimizer.step()\n # return g_loss.data[0]\n return g_loss.item()\n\nPython implementation:\ndef discriminator_train_step(batch_size, discriminator, generator, d_optimizer, criterion, real_images, labels):\n d_optimizer.zero_grad()\n\n # train with real images\n real_validity = discriminator(real_images, labels)", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": "_train_step(batch_size, discriminator, generator, d_optimizer, criterion, real_images, labels):\n d_optimizer.zero_grad()\n\n # train with real images\n real_validity = discriminator(real_images, labels)\n real_loss = criterion(real_validity, Variable(torch.ones(batch_size)).cuda())\n\n # train with fake images\n z = Variable(torch.randn(batch_size, 100)).cuda()\n fake_labels = Variable(torch.LongTensor(np.random.randint(0, 10, batch_size))).cuda()\n fake_images = generator(z, fake_labels)\n fake_validity = discriminator(fake_images, fake_labels)\n fake_loss = criterion(fake_validity, Variable(torch.zeros(batch_size)).cuda())\n\n d_loss = real_loss + fake_loss\n d_loss.backward()\n d_optimizer.step()\n # return d_loss.data[0]\n return d_loss.item()\n\nPython implementation:\nnum_epochs = 30\nn_critic = 5\ndisplay_step = 300\nfor epoch in range(num_epochs):\n print('Starting epoch {}...'.format(epoch))\n for i, (images, labels) in enumerate(data_loader):\n real_images = Variable(images).cuda()", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": "n_critic = 5\ndisplay_step = 300\nfor epoch in range(num_epochs):\n print('Starting epoch {}...'.format(epoch))\n for i, (images, labels) in enumerate(data_loader):\n real_images = Variable(images).cuda()\n labels = Variable(labels).cuda()\n generator.train()\n batch_size = real_images.size(0)\n d_loss = discriminator_train_step(len(real_images), discriminator,\n generator, d_optimizer, criterion,\n real_images, labels)\n\n g_loss = generator_train_step(batch_size, discriminator, generator, g_optimizer, criterion)\n\n generator.eval()\n print('g_loss: {}, d_loss: {}'.format(g_loss, d_loss))\n z = Variable(torch.randn(9, 100)).cuda()\n labels = Variable(torch.LongTensor(np.arange(9))).cuda()\n sample_images = generator(z, labels).unsqueeze(1).data.cpu()\n grid = make_grid(sample_images, nrow=3, normalize=True).permute(1,2,0).numpy()\n plt.imshow(grid)\n plt.show()\n\nPython implementation:\nz = Variable(torch.randn(100, 100)).cuda()", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": "nsqueeze(1).data.cpu()\n grid = make_grid(sample_images, nrow=3, normalize=True).permute(1,2,0).numpy()\n plt.imshow(grid)\n plt.show()\n\nPython implementation:\nz = Variable(torch.randn(100, 100)).cuda()\nlabels = Variable(torch.LongTensor([i for _ in range(10) for i in range(10)])).cuda()\nsample_images = generator(z, labels).unsqueeze(1).data.cpu()\ngrid = make_grid(sample_images, nrow=10, normalize=True).permute(1,2,0).numpy()\nfig, ax = plt.subplots(figsize=(15,15))\nax.imshow(grid)\n_ = plt.yticks([])\n_ = plt.xticks(np.arange(15, 300, 30), ['Zero', 'One', 'Two', 'Three', 'Four', 'Five', 'Six', 'Seven', 'Eight', 'Nine'], rotation=45, fontsize=20)", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "283-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Example Conditional GAN" + }, + { + "text": "Explanation:\n# Results\n\n## Summary\n\nThis notebook implemented Generative Adversarial Networks for image synthesis using the MNIST dataset.\n\n### Implemented Models\n\n- Fully Connected GAN\n- Conditional GAN (cGAN)\n\n### Key Outcomes\n\n- Learned to generate realistic handwritten digit images from random latent vectors.\n- Demonstrated adversarial training between generator and discriminator networks.\n- Explored conditional image generation using class labels.\n- Compared GAN-based image generation with autoencoder reconstruction.\n- Highlighted common GAN training challenges, including instability and hyperparameter sensitivity.\n\nThese experiments demonstrate how adversarial learning enables neural networks to generate realistic synthetic images and forms the foundation for many modern generative AI models.", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "284-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Results" + }, + { + "text": "Explanation:\n# Conclusion\n\nGenerative Adversarial Networks provide a powerful framework for learning complex data distributions through adversarial optimization. Unlike reconstruction-based autoencoders, GANs learn to generate visually realistic samples by competing against a discriminator network. This notebook introduced both standard GANs and Conditional GANs, illustrating how adversarial learning enables controllable, high-quality image generation and serves as a foundation for modern generative AI techniques.", + "source": "02_gan.ipynb", + "file_type": "ipynb", + "chunk_id": "285-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/02_gan.ipynb", + "section": "Conclusion" + }, + { + "text": "Explanation:\n# Adversarial Attacks on Deep Neural Networks with PyTorch\n\n## Overview\n\nThis notebook explores **adversarial attacks** on deep neural networks using PyTorch. It demonstrates how carefully crafted perturbations, often imperceptible to humans, can cause image classification models to make incorrect predictions. By optimizing these perturbations through gradient-based methods, the notebook highlights vulnerabilities in deep learning models and introduces techniques for evaluating model robustness.\n\nThe notebook implements targeted adversarial attacks, visualizes adversarial examples, and analyzes the impact of adversarial perturbations on image classification performance.\n\n## Learning Objectives\n\nBy completing this notebook, you will learn how to:\n\n- Understand adversarial examples\n- Generate targeted adversarial attacks\n- Optimize perturbations using gradient descent\n- Evaluate neural network robustness\n- Visualize adversarial perturbations", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "286-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Adversarial Attacks on Deep Neural Networks with PyTorch" + }, + { + "text": ":\n\n- Understand adversarial examples\n- Generate targeted adversarial attacks\n- Optimize perturbations using gradient descent\n- Evaluate neural network robustness\n- Visualize adversarial perturbations\n- Compare original and adversarial images\n- Analyze classifier vulnerabilities\n- Understand the importance of robust deep learning", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "286-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Adversarial Attacks on Deep Neural Networks with PyTorch" + }, + { + "text": "Explanation:\n## Introduction\n\nDeep neural networks achieve remarkable performance across many computer vision tasks. However, despite their high accuracy, they can be surprisingly vulnerable to **adversarial examples**—inputs that have been intentionally modified with small, carefully designed perturbations that are often imperceptible to humans but cause incorrect model predictions.\n\nAdversarial attacks provide valuable insights into the limitations of modern deep learning models and have become an important area of research in AI security, model robustness, and trustworthy machine learning.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\n\nfrom torchvision import datasets, transforms\nimport matplotlib.pyplot as plt\n\nmnist_images = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\n\nclass FCNet(nn.Module):\n def __init__(self):\n super(FCNet, self).__init__()", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "287-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Introduction" + }, + { + "text": "atplotlib.pyplot as plt\n\nmnist_images = datasets.MNIST('data', train=True, download=True, transform=transforms.ToTensor())\n\nclass FCNet(nn.Module):\n def __init__(self):\n super(FCNet, self).__init__()\n self.layer1 = nn.Linear(28 * 28, 50)\n self.layer2 = nn.Linear(50, 20)\n self.layer3 = nn.Linear(20, 10)\n def forward(self, img):\n flattened = img.view(-1, 28 * 28)\n activation1 = F.relu(self.layer1(flattened))\n activation2 = F.relu(self.layer2(activation1))\n output = self.layer3(activation2)\n return output\n\nclass ConvNet(nn.Module):\n def __init__(self):\n super(ConvNet, self).__init__()\n self.conv1 = nn.Conv2d(1, 5, 5, padding=2)\n self.pool = nn.MaxPool2d(2, 2)\n self.conv2 = nn.Conv2d(5, 10, 5, padding=2)\n self.fc1 = nn.Linear(10 * 7 * 7, 32)\n self.fc2 = nn.Linear(32, 10)\n\n def forward(self, x):\n x = self.pool(F.relu(self.conv1(x)))\n x = self.pool(F.relu(self.conv2(x)))\n x = x.view(-1, 10 * 7 * 7)\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n x = x.squeeze(1)\n return x\n\nPython implementation:", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "287-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Introduction" + }, + { + "text": ":\n x = self.pool(F.relu(self.conv1(x)))\n x = self.pool(F.relu(self.conv2(x)))\n x = x.view(-1, 10 * 7 * 7)\n x = F.relu(self.fc1(x))\n x = self.fc2(x)\n x = x.squeeze(1)\n return x\n\nPython implementation:\ndef train(model, data, batch_size=64, lr=0.001, num_iters=1000, print_every=100):\n train_loader = torch.utils.data.DataLoader(data, batch_size=batch_size)\n optimizer = optim.Adam(model.parameters(), lr=lr)\n criterion = nn.CrossEntropyLoss()\n\n total_loss = 0\n n = 0\n\n while True:\n for imgs, labels in iter(train_loader):\n out = model(imgs)\n loss = criterion(out, labels)\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n total_loss += loss.item()\n n += 1\n\n if n % print_every == 0:\n print(\"Iter %d. Avg.Loss: %f\" % (n, total_loss/print_every))\n total_loss = 0\n if n > num_iters:\n return\n\nPython implementation:\nfc_model = FCNet()\ntrain(fc_model, mnist_images, num_iters=1000)\n\nPython implementation:\ncnn_model = ConvNet()\ntrain(cnn_model, mnist_images, num_iters=1000)", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "287-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Introduction" + }, + { + "text": "Explanation:\n## Targetted Adversarial Attack\n\nThe purpose of an adversarial attack is to perturb an input\n(usually an image $x$) so that a neural network $f$ misclassifies\nthe perturbed image $x + \\epsilon$. In a targeted attack, we\nwant the network $f$ to misclassify the perturbed image into\na class of our choosing.\n\nLet's begin with this image. We will perturb the image so our model\nthinks that the image is of the digit 3, when in fact it is of the\ndigit 5.\n\nPython implementation:\nimage = mnist_images[0][0]\ntarget_label = 3\nmodel = fc_model\n\nplt.imshow(image[0])\n\nExplanation:\nOur approach is as follows:\n\n- We will create a random noise $\\epsilon$ that is the same size\n as the image.\n- We will use an optimizer to tune the values of $\\epsilon$ to\n make the neural network misclassify $x + \\epsilon$ to our target class\n\nThe second step might sound a little mysterious, but is actually\nvery similar to tuning the weights of a neural network!\n\nFirst, let's create some noise values.", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "288-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Targetted Adversarial Attack" + }, + { + "text": "ify $x + \\epsilon$ to our target class\n\nThe second step might sound a little mysterious, but is actually\nvery similar to tuning the weights of a neural network!\n\nFirst, let's create some noise values. In order for PyTorch to be\nable to tune these values using an optimizer, we need to\nset `requires_grad=True`:\n\nPython implementation:\nnoise = torch.randn(1, 28, 28) * 0.01\nnoise.requires_grad = True\n\nExplanation:\nNow, we will tune the noise:\n\nPython implementation:\noptimizer = optim.Adam([noise], lr=0.01, weight_decay=1)\ncriterion = nn.CrossEntropyLoss()\n\nfor i in range(1000):\n adv_image = torch.clamp(image + noise, 0, 1)\n out = model(adv_image.unsqueeze(0))\n loss = criterion(out, torch.Tensor([target_label]).long())\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\nExplanation:\nTo keep the pixel values in `noise` small,\nwe use a fairly large `weight_decay`. We use the `CrossEntropyLoss`,\nbut maximize the neural network prediction of our `target_label`.", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "288-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Targetted Adversarial Attack" + }, + { + "text": "grad()\n\nExplanation:\nTo keep the pixel values in `noise` small,\nwe use a fairly large `weight_decay`. We use the `CrossEntropyLoss`,\nbut maximize the neural network prediction of our `target_label`.\n\nNotice also that the `adv_image` is clamped so that the pixel values\nare kept in the range [0, 1].\n\nNow, let's see the resulting image:\n\nPython implementation:\nadv_image = torch.clamp(image + noise, 0, 1)\nadv_label = torch.argmax(model(adv_image), dim=1).item()\nadv_percent = torch.softmax(model(adv_image), dim=1)[0,target_label].item()\n\nplt.subplot(1, 3, 1)\nplt.title(\"Original\")\nplt.imshow(image[0])\n\nplt.subplot(1, 3, 2)\nplt.title(\"Noise\")\nplt.imshow(noise[0].detach().numpy())\n\nplt.subplot(1, 3, 3)\nplt.title(\"Label=%d P(label=%d)=%.2f\" % (adv_label, target_label, adv_percent))\nplt.imshow(adv_image.detach().numpy()[0])\n\nExplanation:\nThe image on the right still looks like a \"5\" to a human.\nHowever, the neural network misclassifies the image.", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "288-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Targetted Adversarial Attack" + }, + { + "text": "el, target_label, adv_percent))\nplt.imshow(adv_image.detach().numpy()[0])\n\nExplanation:\nThe image on the right still looks like a \"5\" to a human.\nHowever, the neural network misclassifies the image.\n\nThe same steps can be used to create an adversarial attack\nfor other images, and for other model architectures.\n\nPython implementation:\ndef create_adversarial_example(model, image, target_label):\n noise = torch.randn(1, 28, 28)\n noise.requires_grad = True\n\n optimizer = optim.Adam([noise], lr=0.01, weight_decay=1)\n criterion = nn.CrossEntropyLoss()\n\n for i in range(1000):\n adv_image = torch.clamp(image + noise, 0, 1)\n out = model(adv_image.unsqueeze(0))\n loss = criterion(out, torch.Tensor([target_label]).long())\n loss.backward()\n optimizer.step()\n optimizer.zero_grad()\n\n adv_image = torch.clamp(image + noise, 0, 1)\n adv_label = torch.argmax(model(adv_image.unsqueeze(0)), dim=1).item()\n adv_percent = torch.softmax(model(adv_image.unsqueeze(0)), dim=1)[0,target_label].item()", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "288-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Targetted Adversarial Attack" + }, + { + "text": "ge = torch.clamp(image + noise, 0, 1)\n adv_label = torch.argmax(model(adv_image.unsqueeze(0)), dim=1).item()\n adv_percent = torch.softmax(model(adv_image.unsqueeze(0)), dim=1)[0,target_label].item()\n\n plt.subplot(1, 3, 1)\n plt.title(\"Original\")\n plt.imshow(image[0])\n\n plt.subplot(1, 3, 2)\n plt.title(\"Noise\")\n plt.imshow(noise[0].detach().numpy())\n\n plt.subplot(1, 3, 3)\n plt.title(\"Label=%d P(label=%d)=%.2f\" % (adv_label, target_label, adv_percent))\n plt.imshow(adv_image.detach().numpy()[0])", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "288-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Targetted Adversarial Attack" + }, + { + "text": "Explanation:\n# Results\n\n## Summary\n\nThis notebook demonstrated gradient-based adversarial attacks against image classification models.\n\n### Implemented Techniques\n\n- Targeted adversarial attacks\n- Gradient-based perturbation optimization\n- Adversarial image generation\n- Robustness evaluation\n\n### Key Outcomes\n\n- Successfully generated adversarial examples capable of changing model predictions.\n- Visualized perturbations that remain nearly imperceptible to humans.\n- Demonstrated how small input modifications can significantly affect neural network predictions.\n- Explored optimization strategies for generating targeted adversarial examples.\n\nThese experiments highlight the importance of evaluating model robustness alongside predictive accuracy when developing deep learning systems.", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "289-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Results" + }, + { + "text": "Explanation:\n# Conclusion\n\nAdversarial attacks reveal important limitations of deep neural networks by demonstrating that carefully optimized perturbations can cause significant prediction errors while remaining visually indistinguishable from the original images. Understanding these vulnerabilities is essential for developing more robust, reliable, and secure AI systems. This notebook provides a practical introduction to adversarial machine learning and serves as a foundation for studying adversarial defenses and robust model training.", + "source": "03_adversarial_attacks.ipynb", + "file_type": "ipynb", + "chunk_id": "290-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/03_adversarial_attacks.ipynb", + "section": "Conclusion" + }, + { + "text": "Explanation:\n# Image Colorization with CNNs and Conditional GANs\n\n## Overview\n\nThis notebook demonstrates an end-to-end image colorization project implemented in PyTorch. The objective is to predict realistic color images from grayscale inputs using deep learning.\n\nThe notebook is organized into two complementary parts.\n\n**Part A** investigates image colorization as a supervised regression problem using convolutional encoder–decoder networks and U-Net architectures with skip connections.\n\n**Part B** builds upon these models by implementing a Conditional Generative Adversarial Network (cGAN), where a generator learns to produce realistic color images while a discriminator encourages visually plausible outputs through adversarial training.\n\nTogether, these approaches illustrate both reconstruction-based and adversarial methods for image-to-image translation.\n\nSteps:\n\n1. Clean and process the dataset and create greyscale images.\n2. Implement and modify an autoencoder architecture.\n3.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "291-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Image Colorization with CNNs and Conditional GANs" + }, + { + "text": "onstruction-based and adversarial methods for image-to-image translation.\n\nSteps:\n\n1. Clean and process the dataset and create greyscale images.\n2. Implement and modify an autoencoder architecture.\n3. Tune the hyperparameters of an autoencoder.\n4. Implement skip connections and other techniques to improve performance.\n5. Implement a cGAN and compare with an autoencoder.\n6. Improve on the cGAN by trying one of several techniques to enhance training.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "291-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Image Colorization with CNNs and Conditional GANs" + }, + { + "text": "Explanation:\n# Part A — CNN-Based Image Colorization\n\nIn this part we will construct and compare different autoencoder models for the image colourization task.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "292-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part A — CNN-Based Image Colorization" + }, + { + "text": "Explanation:\n#### Helper code\n\nProvided are some helper functions for loading and preparing the data. Note that you will need to use the Colab GPU for this assignment.\n\nPython implementation:\n\"\"\"\nColourization of CIFAR-10 Horses via classification.\n\"\"\"\nimport argparse\nimport math\nimport time\n\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport numpy.random as npr\nimport scipy.misc\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.autograd import Variable\n\nPython implementation:\n######################################################################\n# Setup working directory\n######################################################################\n%mkdir -p /content/a3/\n%cd /content/a3\n\nPython implementation:\n######################################################################\n# Helper functions for loading data\n######################################################################\n# adapted from", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "293-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Helper code" + }, + { + "text": "ation:\n######################################################################\n# Helper functions for loading data\n######################################################################\n# adapted from\n# https://github.com/fchollet/keras/blob/master/keras/datasets/cifar10.py\n\nimport os\nimport pickle\nimport sys\nimport tarfile\n\nimport numpy as np\nfrom PIL import Image\nfrom six.moves.urllib.request import urlretrieve\n\ndef get_file(fname, origin, untar=False, extract=False, archive_format=\"auto\", cache_dir=\"data\"):\n datadir = os.path.join(cache_dir)\n if not os.path.exists(datadir):\n os.makedirs(datadir)\n\n if untar:\n untar_fpath = os.path.join(datadir, fname)\n fpath = untar_fpath + \".tar.gz\"\n else:\n fpath = os.path.join(datadir, fname)\n\n print(\"File path: %s\" % fpath)\n if not os.path.exists(fpath):\n print(\"Downloading data from\", origin)\n\n error_msg = \"URL fetch failure on {}: {} -- {}\"\n try:\n try:\n urlretrieve(origin, fpath)\n except URLError as e:", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "293-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Helper code" + }, + { + "text": "h: %s\" % fpath)\n if not os.path.exists(fpath):\n print(\"Downloading data from\", origin)\n\n error_msg = \"URL fetch failure on {}: {} -- {}\"\n try:\n try:\n urlretrieve(origin, fpath)\n except URLError as e:\n raise Exception(error_msg.format(origin, e.errno, e.reason))\n except HTTPError as e:\n raise Exception(error_msg.format(origin, e.code, e.msg))\n except (Exception, KeyboardInterrupt) as e:\n if os.path.exists(fpath):\n os.remove(fpath)\n raise\n\n if untar:\n if not os.path.exists(untar_fpath):\n print(\"Extracting file.\")\n with tarfile.open(fpath) as archive:\n archive.extractall(datadir)\n return untar_fpath\n\n if extract:\n _extract_archive(fpath, datadir, archive_format)\n\n return fpath\n\ndef load_batch(fpath, label_key=\"labels\"):\n \"\"\"Internal utility for parsing CIFAR data.\n # Arguments\n fpath: path the file to parse.\n label_key: key for label data in the retrieve\n dictionary.\n # Returns\n A tuple `(data, labels)`.\n \"\"\"\n f = open(fpath, \"rb\")\n if sys.version_info < (3,):\n d = pickle.load(f)\n else:", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "293-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Helper code" + }, + { + "text": "he file to parse.\n label_key: key for label data in the retrieve\n dictionary.\n # Returns\n A tuple `(data, labels)`.\n \"\"\"\n f = open(fpath, \"rb\")\n if sys.version_info < (3,):\n d = pickle.load(f)\n else:\n d = pickle.load(f, encoding=\"bytes\")\n # decode utf8\n d_decoded = {}\n for k, v in d.items():\n d_decoded[k.decode(\"utf8\")] = v\n d = d_decoded\n f.close()\n data = d[\"data\"]\n labels = d[label_key]\n\n data = data.reshape(data.shape[0], 3, 32, 32)\n return data, labels\n\ndef load_cifar10(transpose=False):\n \"\"\"Loads CIFAR10 dataset.\n # Returns\n Tuple of Numpy arrays: `(x_train, y_train), (x_test, y_test)`.\n \"\"\"\n dirname = \"cifar-10-batches-py\"\n origin = \"http://www.cs.toronto.edu/~kriz/cifar-10-python.tar.gz\"\n path = get_file(dirname, origin=origin, untar=True)\n\n num_train_samples = 50000\n\n x_train = np.zeros((num_train_samples, 3, 32, 32), dtype=\"uint8\")\n y_train = np.zeros((num_train_samples,), dtype=\"uint8\")\n\n for i in range(1, 6):\n fpath = os.path.join(path, \"data_batch_\" + str(i))", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "293-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Helper code" + }, + { + "text": "x_train = np.zeros((num_train_samples, 3, 32, 32), dtype=\"uint8\")\n y_train = np.zeros((num_train_samples,), dtype=\"uint8\")\n\n for i in range(1, 6):\n fpath = os.path.join(path, \"data_batch_\" + str(i))\n data, labels = load_batch(fpath)\n x_train[(i - 1) * 10000 : i * 10000, :, :, :] = data\n y_train[(i - 1) * 10000 : i * 10000] = labels\n\n fpath = os.path.join(path, \"test_batch\")\n x_test, y_test = load_batch(fpath)\n\n y_train = np.reshape(y_train, (len(y_train), 1))\n y_test = np.reshape(y_test, (len(y_test), 1))\n\n if transpose:\n x_train = x_train.transpose(0, 2, 3, 1)\n x_test = x_test.transpose(0, 2, 3, 1)\n return (x_train, y_train), (x_test, y_test)\n\nPython implementation:\n# Download CIFAR dataset\nm = load_cifar10()", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "293-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Helper code" + }, + { + "text": "Explanation:\n## Part 1. Data Preparation\n\nTo start off run the above code to load the CIFAR dataset and then work through the following questions/tasks.\n\n### Part (a)\nVerify that the dataset has loaded correctly. How many samples do we have? How is the data organized?\n\n$\\color{blue}{\\text{Answer: }}$\n\n- The data has been loaded correctly. The data has been loaded in a tuple with two groups: train and test set. \n- train and test set are tuple with lenght of two: one for images with their pixels and one for labels \n\n- We have 50000 images with their labels in train set. We have 10 groups of images\n\n- We have 10000 images with their labels in test set.\n\n- We have 50000 RBG images with 32*32 pixels and label of every image has been loaded as numpy.ndarray.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "294-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 1. Data Preparation" + }, + { + "text": "0 images with their labels in test set.\n\n- We have 50000 RBG images with 32*32 pixels and label of every image has been loaded as numpy.ndarray.\n\n- We have 10000 RBG images with 32*32 pixels and label of every image has been loaded.\n\nPython implementation:\nprint('type of the data:',type(m))\nprint('length of data with tuple type=',len(m))\ntrain, test= m\nprint('type of the train set:', type(train))\nprint('length of train set with tuple type', len(train))\n\nprint('type of the test set:', type(test))\nprint('length of test set with tuple type', len(test))\nx_train, train_label=train\nx_test, test_label=test\nprint('type of the x_train set:', type(x_train))\nprint('Shape of the training images as numpy.ndarray:', x_train.shape)\nprint('Shape of the training images as numpy.ndarray:', x_test.shape)\n\nprint('Number of images in train set=', len(x_train))\nprint('Number of labels in train set=', len(train_label))", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "294-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 1. Data Preparation" + }, + { + "text": ", x_train.shape)\nprint('Shape of the training images as numpy.ndarray:', x_test.shape)\n\nprint('Number of images in train set=', len(x_train))\nprint('Number of labels in train set=', len(train_label))\nprint('Number of images in test set=', len(x_test))\nprint('Number of labels in test set=', len(test_label))\n\nlen(np.unique(test_label, axis=0))", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "294-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 1. Data Preparation" + }, + { + "text": "Explanation:\n### Part (b)\nPreprocess the data to select only images of horses. Learning to generate only hourse images will make our task easier. Your function will also convert the colour images to greyscale to create our input data.\n\nPython implementation:\n# select a single category.\nHORSE_CATEGORY = 7\n\n# convert colour images into greyscale\ndef process(xs, ys, max_pixel=256.0, downsize_input=False):\n \"\"\"\n Pre-process CIFAR10 images by taking only the horse category,\n shuffling, and have colour values be bound between 0 and 1\n\n Args:\n xs: the colour RGB pixel values\n ys: the category labels\n max_pixel: maximum pixel value in the original data\n Returns:\n xs: value normalized and shuffled colour images\n grey: greyscale images, also normalized so values are between 0 and 1\n \"\"\"\n xs = xs / max_pixel\n xs = xs[np.where(ys == HORSE_CATEGORY)[0], :, :, :]\n npr.shuffle(xs)\n\n grey = np.mean(xs, axis=1, keepdims=True)\n\n if downsize_input:\n downsize_module = nn.Sequential(\n nn.AvgPool2d(2),", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "295-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "/ max_pixel\n xs = xs[np.where(ys == HORSE_CATEGORY)[0], :, :, :]\n npr.shuffle(xs)\n\n grey = np.mean(xs, axis=1, keepdims=True)\n\n if downsize_input:\n downsize_module = nn.Sequential(\n nn.AvgPool2d(2),\n nn.AvgPool2d(2),\n nn.Upsample(scale_factor=2),\n nn.Upsample(scale_factor=2),\n )\n xs_downsized = downsize_module.forward(torch.from_numpy(xs).float())\n xs_downsized = xs_downsized.data.numpy()\n return (xs, xs_downsized)\n else:\n return (xs, grey)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "295-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\nCreate a dataloader (or function) to batch the samples.\n\nPython implementation:\n# dataloader for batching samples\n\ndef get_batch(x, y, batch_size):\n \"\"\"\n Generated that yields batches of data\n\n Args:\n x: input values\n y: output values\n batch_size: size of each batch\n Yields:\n batch_x: a batch of inputs of size at most batch_size\n batch_y: a batch of outputs of size at most batch_size\n \"\"\"\n N = np.shape(x)[0]\n assert N == np.shape(y)[0]\n for i in range(0, N, batch_size):\n batch_x = x[i : i + batch_size, :, :, :]\n batch_y = y[i : i + batch_size, :, :, :]\n yield (batch_x, batch_y)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "296-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (e)\nVerify and visualize that we are able to generate different batches of data.\n\nPython implementation:\n# code to load different batches of horse dataset\n\nprint(\"Loading data...\")\n(x_train, y_train), (x_test, y_test) = load_cifar10()\n\nprint(\"Transforming data...\")\ntrain_rgb, train_grey = process(x_train, y_train)\ntest_rgb, test_grey = process(x_test, y_test)\n\nPython implementation:\n# shape of training data\nprint('Training Data: ', train_rgb.shape, train_grey.shape)\n# shape of testing data\nprint('Testing Data: ', test_rgb.shape, test_grey.shape)\n\nExplanation:\nLoad Batches\n\nPython implementation:\n# obtain batches of images\nxs, ys = next(iter(get_batch(train_grey, train_rgb, 10)))\nprint(xs.shape, ys.shape)\n\nExplanation:\nVisualization\n\n$\\color{blue}{\\text{ }}$\n\n- 5 images in train set with corresponding gray images has been ploted ", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "297-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (e)" + }, + { + "text": "rgb, 10)))\nprint(xs.shape, ys.shape)\n\nExplanation:\nVisualization\n\n$\\color{blue}{\\text{ }}$\n\n- 5 images in train set with corresponding gray images has been ploted \n\n- 5 images in test set with corresponding gray images has been ploted \n\nPython implementation:\n# visualize 5 train/test images\nimport matplotlib.pyplot as plt\n\nk = 0\nfor images in train_rgb:\n image = images\n # place the colour channel at the end, instead of at the beginning\n img = np.transpose(image, [1,2,0])\n plt.subplot(2, 5, k+1)\n plt.axis('off')\n plt.imshow(img)\n\n k += 1\n if k > 4:\n break\n\nfor images in train_grey:\n image = images\n # place the colour channel at the end, instead of at the beginning\n img = np.transpose(image, [1,2,0])\n plt.subplot(2, 5, k+1)\n plt.axis('off')\n plt.imshow(img[:,:,0], cmap='gray')\n\n k += 1\n if k > 9:\n break\n\nPython implementation:\nk = 0\nfor images in test_rgb:\n image = images", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "297-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (e)" + }, + { + "text": "ranspose(image, [1,2,0])\n plt.subplot(2, 5, k+1)\n plt.axis('off')\n plt.imshow(img[:,:,0], cmap='gray')\n\n k += 1\n if k > 9:\n break\n\nPython implementation:\nk = 0\nfor images in test_rgb:\n image = images\n # place the colour channel at the end, instead of at the beginning\n img = np.transpose(image, [1,2,0])\n plt.subplot(2, 5, k+1)\n plt.axis('off')\n plt.imshow(img)\n\n k += 1\n if k > 4:\n break\n\nfor images in test_grey:\n image = images\n # place the colour channel at the end, instead of at the beginning\n img = np.transpose(image, [1,2,0])\n plt.subplot(2, 5, k+1)\n plt.axis('off')\n plt.imshow(img[:,:,0],cmap='gray')\n\n k += 1\n if k > 9:\n break", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "297-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (e)" + }, + { + "text": "Explanation:\n## Part 2. Colourization as Regression\n\nThere are many ways to frame the problem of image colourization as a machine learning problem. One naive approach is to frame it as a regression problem, where we build a model to predict the RGB intensities at each pixel given the greyscale input. In this case, the outputs are continuous, and so squared error can be used to train the model.\n\nIn this section, you will get familar with training neural networks using cloud GPUs. Run the helper code and answer the questions that follow.\n\n#### Helper Code\n\nExplanation:\nRegression Architecture\n\nPython implementation:\nclass RegressionCNN(nn.Module):\n def __init__(self, kernel, num_filters):\n # first call parent's initialization function\n super().__init__()\n padding = kernel // 2\n\n self.downconv1 = nn.Sequential(\n nn.Conv2d(1, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "l // 2\n\n self.downconv1 = nn.Sequential(\n nn.Conv2d(1, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(\n nn.Conv2d(num_filters, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n\n self.rfconv = nn.Sequential(\n nn.Conv2d(num_filters*2, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU())\n\n self.upconv1 = nn.Sequential(\n nn.Conv2d(num_filters*2, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.upconv2 = nn.Sequential(\n nn.Conv2d(num_filters, 3, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(3),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.finalconv = nn.Conv2d(3, 3, kernel_size=kernel, padding=padding)\n\n def forward(self, x):\n out = self.downconv1(x)\n out = self.downconv2(out)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "orm2d(3),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.finalconv = nn.Conv2d(3, 3, kernel_size=kernel, padding=padding)\n\n def forward(self, x):\n out = self.downconv1(x)\n out = self.downconv2(out)\n out = self.rfconv(out)\n out = self.upconv1(out)\n out = self.upconv2(out)\n out = self.finalconv(out)\n return out\n\nExplanation:\n Compute loss \n\nPython implementation:\ndef get_loss(gen, batch_size,):\n criterion = nn.MSELoss()\n losses = []\n for i, (xs, ys) in enumerate(get_batch(test_grey, test_rgb, args.batch_size)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n outputs = gen(images)\n loss = criterion(outputs, labels)\n losses.append(loss.data.item())\n Test_loss=np.mean(losses)\n return Test_loss\n\nExplanation:\nTraining code\n\nPython implementation:\nclass AttrDict(dict):\n def __init__(self, *args, **kwargs):\n super(AttrDict, self).__init__(*args, **kwargs)\n self.__dict__ = self\n\ndef get_torch_vars(xs, ys, gpu=False):\n \"\"\"", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "Python implementation:\nclass AttrDict(dict):\n def __init__(self, *args, **kwargs):\n super(AttrDict, self).__init__(*args, **kwargs)\n self.__dict__ = self\n\ndef get_torch_vars(xs, ys, gpu=False):\n \"\"\"\n Helper function to convert numpy arrays to pytorch tensors.\n If GPU is used, move the tensors to GPU.\n\n Args:\n xs (float numpy tenosor): greyscale input\n ys (int numpy tenosor): rgb as labels\n gpu (bool): whether to move pytorch tensor to GPU\n Returns:\n Variable(xs), Variable(ys)\n \"\"\"\n xs = torch.from_numpy(xs).float()\n ys = torch.from_numpy(ys).float()\n if gpu:\n xs = xs.cuda()\n ys = ys.cuda()\n return Variable(xs), Variable(ys)\n\ndef train(args, gen=None):\n\n # Numpy random seed\n npr.seed(args.seed)\n\n # Save directory\n save_dir = \"outputs/\" + args.experiment_name\n\n # LOAD THE MODEL\n if gen is None:\n Net = globals()[args.model]\n gen = Net(args.kernel, args.num_filters)\n\n # LOSS FUNCTION\n criterion = nn.MSELoss()\n optimizer = torch.optim.Adam(gen.parameters(), lr=args.learn_rate)\n\n # DATA", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "s None:\n Net = globals()[args.model]\n gen = Net(args.kernel, args.num_filters)\n\n # LOSS FUNCTION\n criterion = nn.MSELoss()\n optimizer = torch.optim.Adam(gen.parameters(), lr=args.learn_rate)\n\n # DATA\n print(\"Loading data...\")\n (x_train, y_train), (x_test, y_test) = load_cifar10()\n\n print(\"Transforming data...\")\n train_rgb, train_grey = process(x_train, y_train, downsize_input=args.downsize_input)\n test_rgb, test_grey = process(x_test, y_test, downsize_input=args.downsize_input)\n\n # Create the outputs folder if not created already\n if not os.path.exists(save_dir):\n os.makedirs(save_dir)\n\n print(\"Beginning training ...\")\n if args.gpu:\n gen.cuda()\n start = time.time()\n\n train_losses = []\n valid_losses = []\n valid_accs = []\n for epoch in range(args.epochs):\n # Train the Model\n gen.train() # Change model to 'train' mode\n losses = []\n for i, (xs, ys) in enumerate(get_batch(train_grey, train_rgb, args.batch_size)):\n images, labels = get_torch_vars(xs, ys, args.gpu)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "the Model\n gen.train() # Change model to 'train' mode\n losses = []\n for i, (xs, ys) in enumerate(get_batch(train_grey, train_rgb, args.batch_size)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n # Forward + Backward + Optimize\n optimizer.zero_grad()\n outputs = gen(images)\n\n loss = criterion(outputs, labels)\n loss.backward()\n optimizer.step()\n losses.append(loss.data.item())\n\n print(epoch, loss.cpu().detach())\n\n train_losses.append(np.mean(losses)) # compute Train loss\n valid_losses.append(get_loss(gen, args.batch_size)) # compute Test loss\n if args.plot:\n visual(images, labels, outputs, args.gpu, 1)\n\n print(\"final train loss=\", np.mean(losses) )\n print(\"final test loss=\", get_loss(gen, args.batch_size))\n x = len(valid_losses)\n plt.title(\"Train vs Test Loss\")\n plt.plot(range(1,x+1), train_losses, label=\"Train\")\n plt.plot(range(1,x+1), valid_losses, label=\"Test\")\n plt.xlabel(\"epoch\")\n plt.ylabel(\"Loss\")\n plt.legend(loc='best')\n return gen\n\nExplanation:\nTraining visualization code", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "rain_losses, label=\"Train\")\n plt.plot(range(1,x+1), valid_losses, label=\"Test\")\n plt.xlabel(\"epoch\")\n plt.ylabel(\"Loss\")\n plt.legend(loc='best')\n return gen\n\nExplanation:\nTraining visualization code\n\nPython implementation:\n# visualize 5 train/test images\ndef visual(img_grey, img_real, img_fake, gpu = 0, flag_torch = 0):\n\n if gpu:\n img_grey = img_grey.cpu().detach()\n img_real = img_real.cpu().detach()\n img_fake = img_fake.cpu().detach()\n\n if flag_torch:\n img_grey = img_grey.numpy()\n img_real = img_real.numpy()\n img_fake = img_fake.numpy()\n\n if flag_torch == 2:\n img_real = np.transpose(img_real[:, :, :, :, :], [0, 4, 2, 3, 1]).squeeze()\n img_fake = np.transpose(img_fake[:, :, :, :, :], [0, 4, 2, 3, 1]).squeeze()\n\n #correct image structure\n img_grey = np.transpose(img_grey[:5, :, :, :], [0, 2, 3, 1]).squeeze()\n img_real = np.transpose(img_real[:5, :, :, :], [0, 2, 3, 1])\n img_fake = np.transpose(img_fake[:5, :, :, :], [0, 2, 3, 1])\n\n for i in range(5):\n ax = plt.subplot(3, 5, i + 1)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "[0, 2, 3, 1]).squeeze()\n img_real = np.transpose(img_real[:5, :, :, :], [0, 2, 3, 1])\n img_fake = np.transpose(img_fake[:5, :, :, :], [0, 2, 3, 1])\n\n for i in range(5):\n ax = plt.subplot(3, 5, i + 1)\n ax.imshow(img_grey[i], cmap='gray')\n ax.axis(\"off\")\n ax = plt.subplot(3, 5, i + 1 + 5)\n ax.imshow(img_real[i])\n ax.axis(\"off\")\n ax = plt.subplot(3, 5, i + 1 + 10)\n ax.imshow(img_fake[i])\n ax.axis(\"off\")\n plt.show()\n\nExplanation:\nMain training loop for regression CNN\n\nPython implementation:\n#Main training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"RegressionCNN\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\nprint(\"Show 5 pics from test set:\")\nk = 0", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": ",\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\nprint(\"Show 5 pics from test set:\")\nk = 0\nfor i, (xs, ys) in enumerate(get_batch(test_grey, test_rgb, args.batch_size)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n outputs = cnn(images)\n visual(images, labels, outputs, args.gpu, 1)\n k += 1\n if k > 0:\n break", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "298-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Colourization as Regression" + }, + { + "text": "Explanation:\n### Part (a)\nDescribe the model RegressionCNN. How many convolution layers does it have? What are the filter sizes and number of filters at each layer? Construct a table or draw a diagram.\n\nExplanation:\n$\\color{blue}{\\text{}}$\n\n- We have 6 convolution layers in the architecture \n\n\n\n |Layer # | number of input channel | Number of filters| filter size| description |\n:-------------------|:---------------:|--------------------:|--------------------:|--------------------:\nFirst layer| 1|32|3*3| There is a BatchNorm2d - Relu - and maxpool2d after the convolution layer |\nSecond layer| 32|64|3*3| There is a BatchNorm2d - Relu - and maxpool2d after the convolution layer |\nThird layer| 64|64|3*3| There is a BatchNorm2d - and Relu after the convolution layer |\nFourth layer| 64|32|3*3| There is a BatchNorm2d - Relu - and Upsample after the convolution layer |", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "299-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": "ion layer |\nThird layer| 64|64|3*3| There is a BatchNorm2d - and Relu after the convolution layer |\nFourth layer| 64|32|3*3| There is a BatchNorm2d - Relu - and Upsample after the convolution layer |\nFifth layer| 32|3|3*3| There is a BatchNorm2d - Relu - and Upsample after the convolution layer |\nSixth layer| 3|3|3*3| - |\n\n", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "299-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\nRun the regression training code (should run without errors). This will generate some images. How many epochs are we training the CNN model in the given setting?\n\nExplanation:\n$\\color{Blue}{\\text{ }}$\n\n- We are training the CNN model in the given setting for 25 epochs ", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "300-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\nRe-train a couple of new models using a different number of training epochs. You may train each new models in a new code cell by copying and modifying the code from the last notebook cell. Comment on how the results (output images, training loss) change as we increase or decrease the number of epochs.\n\nExplanation:\n$\\color{blue}{\\text{ }}$\n\n- By increasing epoch the quality and predicted color of images will increase slightly and color of predicted images would be closer to the images' color. Moreover, the training loss would decrease. For example, when epoch is 100 training loss=0.0055 that is less than 0.0088 for epoch=25. \n\n- By decreasing epoch the quality and predicted color of images will decrease. Moreover, the training loss would increase. For example, when epoch is 12 training loss=0.01 that is more than 0.0088 when epoch=25.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "301-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "ing epoch the quality and predicted color of images will decrease. Moreover, the training loss would increase. For example, when epoch is 12 training loss=0.01 that is more than 0.0088 when epoch=25.\nwhen epoch is 3 the color would not be predicted at all and training loss is so high.\n\n\nPython implementation:\n#Main training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"RegressionCNN\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 50,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n#Main training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"RegressionCNN\",\n \"kernel\": 3,", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "301-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"RegressionCNN\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 100,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n#Main training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"RegressionCNN\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 12,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n#Main training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "301-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "tion_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n#Main training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"RegressionCNN\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 8,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n#Main training loop for CNN\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"RegressionCNN\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 3,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "301-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "l\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 3,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\n\nargs.update(args_dict)\ncnn = train(args)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "301-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n## Part 3. Skip Connections\nA skip connection in a neural network is a connection which skips one or more layer and connects to a later layer. We will introduce skip connections.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "302-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 3. Skip Connections" + }, + { + "text": "Explanation:\n### Part (a)\nAdd a skip connection from the first layer to the last, second layer to the second last, etc.\nThat is, the final convolution should have both the output of the previous layer and the initial greyscale input as input. This type of skip-connection is introduced by [3], and is called a \"UNet\". Following the CNN class that you have completed, complete the __init__ and forward methods of the UNet class.\nHint: You will need to use the function torch.cat.\n\nPython implementation:\n#complete the code\n\nclass UNet(nn.Module):\n def __init__(self, kernel, num_filters, num_colours=3, num_in_channels=1):\n super().__init__()\n\n # Useful parameters\n stride = 2\n padding = kernel // 2\n output_padding = 1\n\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n self.downconv1 = nn.Sequential(\n nn.Conv2d(num_in_channels, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "303-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": "##########################\n self.downconv1 = nn.Sequential(\n nn.Conv2d(num_in_channels, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(\n nn.Conv2d(num_filters, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n\n self.rfconv = nn.Sequential(\n nn.Conv2d(num_filters*2, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU())\n\n self.upconv1 = nn.Sequential(\n nn.Conv2d(num_filters*4, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.upconv2 = nn.Sequential(\n nn.Conv2d(num_filters*2, num_colours, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_colours),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.finalconv = nn.Conv2d(num_colours+num_in_channels, 3, kernel_size=kernel, padding=padding)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "303-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": "ze=kernel, padding=padding),\n nn.BatchNorm2d(num_colours),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.finalconv = nn.Conv2d(num_colours+num_in_channels, 3, kernel_size=kernel, padding=padding)\n\n def forward(self, x):\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n L1 = self.downconv1(x)\n L2 = self.downconv2(L1)\n L3 = self.rfconv(L2)\n L4 = self.upconv1( torch.cat((L3, L2), dim=1))\n L5 = self.upconv2(torch.cat((L4, L1), dim=1))\n L6 = self.finalconv(torch.cat((L5, x), dim=1))\n\n return L6", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "303-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\nTrain the \"UNet\" model for the same amount of epochs as the previous CNN and plot the training curve using a batch size of 100. How does the result compare to the previous model? Did skip connections improve the validation loss and accuracy? Did the skip connections improve the output qualitatively? How? Give at least two reasons why skip connections might improve the performance of our CNN models.\n\n$\\color{blue}{\\text{ }}$\n\n- The resuts of this model is better than the results of the previous CNN. \n\n- The results has been improved because of skip connection. \n\n- Training and test loss have been decreased: \n\n- The model for this part didn't overfit. \n\n\n\n | | The model with skip connection| The previous CNN|\n:-------------------|:---------------:|--------------------:\ntrain loss| 0.0055| 0.00559|", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "304-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "didn't overfit. \n\n\n\n | | The model with skip connection| The previous CNN|\n:-------------------|:---------------:|--------------------:\ntrain loss| 0.0055| 0.00559|\ntest loss| 0.00557|0.00655|\n\n\n$\\color{blue}{\\text{two reasons : }}$\n\n- The quality of images have been increased because of skip connection as it help with vanishing gradient problem.\n\n- keeping some information from previous layer will help to use information that has been forgotten. So, it makes sure that information is not lost. Moreover, by using skip connection the value of loss will decrease.\n\nPython implementation:\n# Main training loop for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 25,", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "304-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "\"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 100,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "304-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\nRe-train a few more \"UNet\" models using different mini batch sizes with a fixed number of epochs. Describe the effect of batch sizes on the training/validation loss, and the final image output.\n\n$\\color{blue}{\\text{}}$\n\n- By increasing batch size (`batch size=200`) The quality of images will decrease \n\n- By decreasing batch size (`batch size=25`) the quality of predicted images increase and loss value will decrease.\n\n- By too much decreasing batch size (`batch size=5`) the model overfit. The quality of images in train set would be so good, but the quality of images in test set will not increase. \n\nPython implementation:\n# complete the code\n\n# Main training loop for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "305-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "p for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 200,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n# complete the code\n\n# Main training loop for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 400,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n# Main training loop for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "305-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "ion_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n# Main training loop for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 50,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n# Main training loop for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 25,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "305-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "32,\n 'learn_rate':0.001,\n \"batch_size\": 25,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\n# Main training loop for UNet\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"UNet\",\n \"kernel\": 3,\n \"num_filters\": 32,\n 'learn_rate':0.001,\n \"batch_size\": 5,\n \"epochs\": 25,\n \"seed\": 0,\n \"plot\": True,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)\n\nPython implementation:\nprint(\"Show 5 pics from test set:\")\nk = 0\nfor i, (xs, ys) in enumerate(get_batch(test_grey, test_rgb, args.batch_size)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n outputs = cnn(images)\n visual(images, labels, outputs, args.gpu, 1)\n k += 1\n if k > 0:\n break", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "305-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n# PART B - Conditional GAN\n\nIn this second half of the assignment we will construct a conditional generative adversarial network for our image colourization task.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "306-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "PART B - Conditional GAN" + }, + { + "text": "Explanation:\n## Part 1. Conditional GAN\nTo start we will be modifying the previous sample code to construct and train a conditional GAN. We will exploring the different architectures to identify and select our best image colourization model.\n\nNote: This second half of the assignment should be started after the lecture on generative adversarial networks (GANs).", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "307-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 1. Conditional GAN" + }, + { + "text": "Explanation:\n### Part (a)\nModify the provided training code to implement a generator. Then test to verify it works on the desired input (Hint: you can reuse some of your earlier autoencoder models here to act as a generator)\n\n$\\color{blue}{\\text{ }}$\n\n- I will use the archtecture of Skip connection that we had. \n\nPython implementation:\nclass Generator(nn.Module):\n def __init__(self, kernel, num_filters, num_colours=3, num_in_channels=1):\n super().__init__()\n\n # Useful parameters\n stride = 2\n padding = kernel // 2\n output_padding = 1\n\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n self.downconv1 = nn.Sequential(\n nn.Conv2d(num_in_channels, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(\n nn.Conv2d(num_filters, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "308-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": ".BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(\n nn.Conv2d(num_filters, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n\n self.rfconv = nn.Sequential(\n nn.Conv2d(num_filters*2, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU())\n\n self.upconv1 = nn.Sequential(\n nn.Conv2d(num_filters*4, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.upconv2 = nn.Sequential(\n nn.Conv2d(num_filters*2, num_colours, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_colours),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.finalconv = nn.Conv2d(num_colours+num_in_channels, 3, kernel_size=kernel, padding=padding)\n\n def forward(self, x):\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n L1 = self.downconv1(x)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "308-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": "nels, 3, kernel_size=kernel, padding=padding)\n\n def forward(self, x):\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n L1 = self.downconv1(x)\n L2 = self.downconv2(L1)\n L3 = self.rfconv(L2)\n L4 = self.upconv1( torch.cat((L3, L2), dim=1))\n L5 = self.upconv2(torch.cat((L4, L1), dim=1))\n out = self.finalconv(torch.cat((L5, x), dim=1))\n #out=torch.cat((x, L6), dim=1)\n\n return out\n\nPython implementation:\n#test generator architecture\n\nmodel=Generator(3,32)\nmodel.cuda()\n(x_train, y_train), (x_test, y_test) = load_cifar10()\ntrain_rgb, train_grey = process(x_train, y_train, downsize_input=False)\n(xs, ys) = next(iter(get_batch(train_grey, train_rgb, batch_size=100)))\nimages, labels = get_torch_vars(xs, ys, gpu=True)\noutputs = model(images)\nprint(outputs.shape)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "308-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\nModify the provided training code to implement a discriminator. Then test to verify it works on the desired input.\n\n$\\color{blue}{\\text{ }}$\n\n- I will use encode parts of the skip connection with several leanir layer as discriminator. \n\nPython implementation:\n# discriminator code\n\nclass Discriminator(nn.Module):\n def __init__(self, kernel, num_filters, num_colours=3, num_in_channels=1):\n super().__init__()\n\n # Useful parameters\n stride = 2\n padding = kernel // 2\n output_padding = 1\n\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n\n self.downconv1 = nn.Sequential(\n nn.Conv2d(3, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(\n nn.Conv2d(num_filters, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU(),\n nn.MaxPool2d(2),)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "309-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "eLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(\n nn.Conv2d(num_filters, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n\n self.rfconv = nn.Sequential(\n nn.Conv2d(num_filters*2, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU())\n self.fc1=nn.Linear(3072, 1024)\n self.act2=nn.LeakyReLU(0.2, inplace=True)\n self.dropout=nn.Dropout(0.3)\n self.fc2=nn.Linear(1024, 64)\n self.fc3=nn.Linear(64, 32)\n self.fc4=nn.Linear(32, 1)\n self.act3=nn.Sigmoid()\n\n def forward(self, x, img_greyscale): ### def forward(self, x, img_greyscale):\n\n ############### YOUR CODE GOES HERE ###############\n # ###################################################\n L1 = self.downconv1(x)\n L2 = self.downconv2(L1)\n x = self.rfconv(L2)\n\n x = x.view(x.size(0), -1)\n # print(x.shape)\n c = img_greyscale.view(img_greyscale.size(0),-1)\n x = torch.cat([x, c], 1)\n x=self.fc1(x)\n x=self.act2(x)\n x=self.dropout(x)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "309-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "v2(L1)\n x = self.rfconv(L2)\n\n x = x.view(x.size(0), -1)\n # print(x.shape)\n c = img_greyscale.view(img_greyscale.size(0),-1)\n x = torch.cat([x, c], 1)\n x=self.fc1(x)\n x=self.act2(x)\n x=self.dropout(x)\n x=self.fc2(x)\n x=self.act2(x)\n x=self.dropout(x)\n x=self.fc3(x)\n x=self.act2(x)\n x=self.dropout(x)\n x=self.fc4(x)\n out=self.act3(x)\n\n return out.squeeze() # Flatten to [batch_size]\n\nPython implementation:\n# test discriminator architecture\nimport torch.nn as nn\n\nmodel=Discriminator(3,16)\nmodel.cuda()\n(x_train, y_train), (x_test, y_test) = load_cifar10()\ntrain_rgb, train_grey = process(x_train, y_train, downsize_input=False)\n(xs, ys) = next(iter(get_batch(train_grey, train_rgb, batch_size=50)))\nimages, labels = get_torch_vars(xs, ys, gpu=True)\noutputs = model(labels, images)\nprint(outputs.shape)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "309-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\nModify the provided training code to implement a conditional GAN.\n\n$\\color{blue}{\\text{}}$\n\n- has been done \n\nPython implementation:\nclass AttrDict(dict):\n def __init__(self, *args, **kwargs):\n super(AttrDict, self).__init__(*args, **kwargs)\n self.__dict__ = self\n\ndef get_torch_vars(xs, ys, gpu=False):\n \"\"\"\n Helper function to convert numpy arrays to pytorch tensors.\n If GPU is used, move the tensors to GPU.\n\n Args:\n xs (float numpy tenosor): greyscale input\n ys (int numpy tenosor): categorical labels\n gpu (bool): whether to move pytorch tensor to GPU\n Returns:\n Variable(xs), Variable(ys)\n \"\"\"\n xs = torch.from_numpy(xs).float()\n ys = torch.from_numpy(ys).float() #--> ADDED for cGAN\n if gpu:\n xs = xs.cuda()\n ys = ys.cuda()\n return Variable(xs), Variable(ys)\n\ndef train(args, cnn=None):\n # Set the maximum number of threads to prevent crash in Teaching Labs\n # TODO: necessary?\n torch.set_num_threads(5)\n # Numpy random seed", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "310-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "return Variable(xs), Variable(ys)\n\ndef train(args, cnn=None):\n # Set the maximum number of threads to prevent crash in Teaching Labs\n # TODO: necessary?\n torch.set_num_threads(5)\n # Numpy random seed\n npr.seed(args.seed)\n\n # Save directory\n save_dir = \"outputs/\" + args.experiment_name\n\n # LOAD THE COLOURS CATEGORIES\n\n # INPUT CHANNEL\n num_in_channels = 1 if not args.downsize_input else 3\n # LOAD THE MODEL\n if cnn is None:\n Net = globals()[args.model]\n cnn =Generator(args.kernel,args.num_filters)\n discriminator = Discriminator (args.kernel,args.num_filters)\n\n # LOSS FUNCTION\n\n criterion = nn.BCELoss()\n g_optimizer = torch.optim.Adam(cnn.parameters(), lr=1e-4)\n d_optimizer = torch.optim.Adam(discriminator.parameters(), lr=1e-4)\n\n # DATA\n print(\"Loading data...\")\n (x_train, y_train), (x_test, y_test) = load_cifar10()\n\n print(\"Transforming data...\")\n train_rgb, train_grey = process(x_train, y_train, downsize_input=args.downsize_input)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "310-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "print(\"Loading data...\")\n (x_train, y_train), (x_test, y_test) = load_cifar10()\n\n print(\"Transforming data...\")\n train_rgb, train_grey = process(x_train, y_train, downsize_input=args.downsize_input)\n test_rgb, test_grey = process(x_test, y_test, downsize_input=args.downsize_input)\n\n # Create the outputs folder if not created already\n if not os.path.exists(save_dir):\n os.makedirs(save_dir)\n\n print(\"Beginning training ...\")\n if args.gpu:\n cnn.cuda()\n discriminator.cuda()\n start = time.time()\n\n train_losses = []\n valid_losses = []\n valid_accs = []\n for epoch in range(args.epochs):\n # Train the Model\n cnn.train()\n discriminator.train()\n losses = []\n\n for i, (xs, ys) in enumerate(get_batch(train_grey, train_rgb, args.batch_size)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n\n #--->ADDED 5\n img_grey = images\n img_real = labels\n batch_size = args.batch_size\n\n #discriminator training\n d_optimizer.zero_grad()\n real_prob = discriminator(img_real,img_grey)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "310-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "s(xs, ys, args.gpu)\n\n #--->ADDED 5\n img_grey = images\n img_real = labels\n batch_size = args.batch_size\n\n #discriminator training\n d_optimizer.zero_grad()\n real_prob = discriminator(img_real,img_grey)\n real_labels=Variable(torch.ones(batch_size)).cuda()\n real_loss = criterion(real_prob, real_labels)\n fake_images = cnn(img_grey)\n fake_prob = discriminator(fake_images, img_grey)\n fake_labels=Variable(torch.zeros(batch_size)).cuda()\n fake_loss = criterion(fake_prob, fake_labels)\n d_loss = real_loss + fake_loss\n d_loss.backward()\n d_optimizer.step()\n\n # generator training\n g_optimizer.zero_grad()\n fake_images = cnn(img_grey)\n make_real = discriminator(fake_images, img_grey)\n g_loss= criterion(make_real, real_labels)\n g_loss.backward()\n g_optimizer.step()\n\n # print and visualize\n print(epoch, g_loss.cpu().detach(), d_loss.cpu().detach())\n visual(images, labels, fake_images, args.gpu, 1)\n\n return cnn", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "310-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (d)\nTrain a conditional GAN for image colourization.\n\nPython implementation:\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"Generator\",\n \"kernel\": 3,\n \"num_filters\": 16,\n 'learn_rate':0.001,\n \"batch_size\": 50,\n \"epochs\": 125,\n \"seed\": 0,\n \"plot\": False,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)\n\n#batch size of 50 with 100 epochs seamed to work", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "311-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n### Part (e)\nHow does the performance of the cGAN compare with the autoencoder models that you tested in the first half of this assignment?\n\nExplanation:\n$\\color{blue}{\\text{ }}$\n\n- The cGAN outperforms the autoencoder models we previously evaluated. This method works better because we have a generator(skip conection architecture and a discriminator that improve the preformance of the generator by optimizing its loss function.The quality of images have increased and colors are closer to real ones. ", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "312-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (e)" + }, + { + "text": "Explanation:\n### Part (f)\nA colour space is a choice of mapping of colours into three-dimensional coordinates. Some colours could be close together in one colour space, but further apart in others. The RGB colour space is probably the most familiar to you, the model used in in our regression colourization example computes squared error in RGB colour space. But, most state of the art colourization models\ndo not use RGB colour space. How could using the RGB colour space be problematic? Your answer should relate how human perception of colour is different than the squared distance. You may use the Wikipedia article on colour space to help you answer the question.\n\nExplanation:\n$\\color{blue}{\\text{ }}$\n\n- Color perception is handled by cones and rods, which are photoreceptors. Rods are responsible for vision at low light levels (scotopic vision). They do not mediate color vision, and have a low spatial acuity.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "313-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (f)" + }, + { + "text": "rception is handled by cones and rods, which are photoreceptors. Rods are responsible for vision at low light levels (scotopic vision). They do not mediate color vision, and have a low spatial acuity. Cones are active at higher light levels (photopic vision), are capable of color vision and are responsible for high spatial acuity.\n- We are not equally sensitive to all of those colours; for example, green light is more sensitive than blue or red light. So, we lose a corresponding sensitivity to the shades of green. This means that if we were to implement RGB colour space, this could be problematic because we are only accounting for the kind of light that is emitted, and so the squared distance between changing the hue of green with a variation of x for example, versus changing hue of red with a same variation of x would give us the same squared distance, but could be preceived by humans differently.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "313-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (f)" + }, + { + "text": "hanging the hue of green with a variation of x for example, versus changing hue of red with a same variation of x would give us the same squared distance, but could be preceived by humans differently. Also, RGB could be problematic because it does not take into account the lightness of an image.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "313-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part (f)" + }, + { + "text": "Explanation:\n## Part 2. Exploration\n\nAt this point we have trained a few different generative models for our image colourization task with varying results. What makes this work exciting is that there many other approaches we could take. In this part of the assignment you will be exploring at least one of several approaches towards improving our performance on the image colourization task. Some well known approaches you can consider include:\n\n- lab colour space representation instead of RBG which simplifies the problem and requires you to predict two output channels instead of three\n- k-means to represent RBG colourspace by 'k' distinct colours, this effectively changes the problem from regression to classification.\n\nOther interesting approaches include:\n- combining L1 loss along with the discriminator-based loss\n- starting with a pretrained generator (i.e. Resnet)\n- patch discriminator trained on local regions", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "Other interesting approaches include:\n- combining L1 loss along with the discriminator-based loss\n- starting with a pretrained generator (i.e. Resnet)\n- patch discriminator trained on local regions\n\nA great example of some of these different approaches can be found in a blog post by Moein Shariatnia.\n\nNote you are only required to pick one of the suggested modifications.\n\n$\\color{blue}{\\text{Answer: }}$\n\n- The first approach has been choosen for improving the quality of images. After using this image quality of image increased and their color are closer to the color of real images\n\nPython implementation:\n# provide your code here\nclass Generator_two(nn.Module):\n def __init__(self, kernel, num_filters, num_colours=2, num_in_channels=1):\n super().__init__()\n\n # Useful parameters\n stride = 2\n padding = kernel // 2", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "your code here\nclass Generator_two(nn.Module):\n def __init__(self, kernel, num_filters, num_colours=2, num_in_channels=1):\n super().__init__()\n\n # Useful parameters\n stride = 2\n padding = kernel // 2\n output_padding = 1\n\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n self.downconv1 = nn.Sequential(\n nn.Conv2d(num_in_channels, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n self.downconv2 = nn.Sequential(\n nn.Conv2d(num_filters, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU(),\n nn.MaxPool2d(2),)\n\n self.rfconv = nn.Sequential(\n nn.Conv2d(num_filters*2, num_filters*2, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU())\n\n self.upconv1 = nn.Sequential(\n nn.Conv2d(num_filters*4, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "ng),\n nn.BatchNorm2d(num_filters*2),\n nn.ReLU())\n\n self.upconv1 = nn.Sequential(\n nn.Conv2d(num_filters*4, num_filters, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_filters),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.upconv2 = nn.Sequential(\n nn.Conv2d(num_filters*2, num_colours, kernel_size=kernel, padding=padding),\n nn.BatchNorm2d(num_colours),\n nn.ReLU(),\n nn.Upsample(scale_factor=2),)\n self.finalconv = nn.Conv2d(num_colours+num_in_channels, 2, kernel_size=kernel, padding=padding)\n\n def forward(self, x):\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n L1 = self.downconv1(x)\n L2 = self.downconv2(L1)\n L3 = self.rfconv(L2)\n L4 = self.upconv1( torch.cat((L3, L2), dim=1))\n L5 = self.upconv2(torch.cat((L4, L1), dim=1))\n out = self.finalconv(torch.cat((L5, x), dim=1))\n #out=torch.cat((x, L6), dim=1)\n\n return out\n\nPython implementation:\n#test generator architecture\n\nmodel=Generator_two(3,32)\nmodel.cuda()", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "L1), dim=1))\n out = self.finalconv(torch.cat((L5, x), dim=1))\n #out=torch.cat((x, L6), dim=1)\n\n return out\n\nPython implementation:\n#test generator architecture\n\nmodel=Generator_two(3,32)\nmodel.cuda()\n(x_train, y_train), (x_test, y_test) = load_cifar10()\ntrain_rgb, train_grey = process(x_train, y_train, downsize_input=False)\n(xs, ys) = next(iter(get_batch(train_grey, train_rgb, batch_size=16)))\nimages, labels = get_torch_vars(xs, ys, gpu=True)\noutputs = model(images)\nprint(outputs.shape)\n\nPython implementation:\n# discriminator code\n\nclass Discriminator_two(nn.Module):\n def __init__(self, kernel, num_filters, num_colours=3, num_in_channels=1):\n super().__init__()\n\n # Useful parameters\n stride = 2\n padding = kernel // 2\n output_padding = 1\n\n ############### YOUR CODE GOES HERE ###############\n ###################################################\n self.fc1=nn.Linear(3072, 1024)\n self.act2=nn.LeakyReLU(0.2, inplace=True)\n self.dropout=nn.Dropout(0.3)\n self.fc2=nn.Linear(1024, 64)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "###########\n ###################################################\n self.fc1=nn.Linear(3072, 1024)\n self.act2=nn.LeakyReLU(0.2, inplace=True)\n self.dropout=nn.Dropout(0.3)\n self.fc2=nn.Linear(1024, 64)\n self.fc3=nn.Linear(64, 32)\n self.fc4=nn.Linear(32, 1)\n self.act3=nn.Sigmoid()\n\n def forward(self, x, img_greyscale): ### def forward(self, x, img_greyscale):\n\n # ############### YOUR CODE GOES HERE ###############\n # # # ###################################################\n\n x = x.reshape(x.size(0), -1)\n # print(x.shape)\n c = img_greyscale.reshape(img_greyscale.size(0),-1)\n x = torch.cat([x, c], 1)\n x=self.fc1(x)\n x=self.act2(x)\n x=self.dropout(x)\n x=self.fc2(x)\n x=self.act2(x)\n x=self.dropout(x)\n x=self.fc3(x)\n x=self.act2(x)\n x=self.dropout(x)\n x=self.fc4(x)\n out=self.act3(x)\n\n return out.squeeze() # Flatten to [batch_size]\n\nPython implementation:\n# test discriminator architecture\nimport torch.nn as nn\n\nmodel=Discriminator_two(3,16)\nmodel.cuda()", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "f.fc4(x)\n out=self.act3(x)\n\n return out.squeeze() # Flatten to [batch_size]\n\nPython implementation:\n# test discriminator architecture\nimport torch.nn as nn\n\nmodel=Discriminator_two(3,16)\nmodel.cuda()\n(x_train, y_train), (x_test, y_test) = load_cifar10()\ntrain_rgb, train_grey = process(x_train, y_train, downsize_input=False)\n(xs, ys) = next(iter(get_batch(train_grey, train_rgb, batch_size=16)))\nimages, labels = get_torch_vars(xs, ys, gpu=True)\nprint(labels[:,1:3,:,:].shape)\nprint(images.shape)\noutputs = model(labels[:,1:3,:,:], images)\nprint(outputs.shape)\n\nPython implementation:\nfrom skimage.color import rgb2lab, lab2rgb\n\ndef labtorgb(L1,ab1): #Function to convert from lab 2 rgb\n L1=L1.detach().cpu().numpy()\n ab1=ab1.detach().cpu().numpy()\n ab1= ab1 * 110 #Unnormalizing\n L1=(L1+1.)*50 #Unnormalizing\n\n trial=np.concatenate((L1,ab1),axis=1)\n\n trial=np.transpose(trial, (0,2,3,1)) #Makes the dimensions appropriate for lab2rgb\n trial_lab=[]\n\n for i in trial:\n img=lab2rgb(i)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "=(L1+1.)*50 #Unnormalizing\n\n trial=np.concatenate((L1,ab1),axis=1)\n\n trial=np.transpose(trial, (0,2,3,1)) #Makes the dimensions appropriate for lab2rgb\n trial_lab=[]\n\n for i in trial:\n img=lab2rgb(i)\n trial_lab.append(img*255.) #scales back to RGB\n\n trial_lab=np.array(trial_lab)\n trial_lab=trial_lab.astype(np.uint8)\n\n return np.stack(trial_lab,axis=0), L1\n\nPython implementation:\nclass AttrDict(dict):\n def __init__(self, *args, **kwargs):\n super(AttrDict, self).__init__(*args, **kwargs)\n self.__dict__ = self\n\ndef get_torch_vars(xs, ys, gpu=False):\n \"\"\"\n Helper function to convert numpy arrays to pytorch tensors.\n If GPU is used, move the tensors to GPU.\n\n Args:\n xs (float numpy tenosor): greyscale input\n ys (int numpy tenosor): categorical labels\n gpu (bool): whether to move pytorch tensor to GPU\n Returns:\n Variable(xs), Variable(ys)\n \"\"\"\n xs = torch.from_numpy(xs).float()\n ys = torch.from_numpy(ys).float() #--> ADDED for cGAN\n if gpu:\n xs = xs.cuda()\n ys = ys.cuda()", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "ove pytorch tensor to GPU\n Returns:\n Variable(xs), Variable(ys)\n \"\"\"\n xs = torch.from_numpy(xs).float()\n ys = torch.from_numpy(ys).float() #--> ADDED for cGAN\n if gpu:\n xs = xs.cuda()\n ys = ys.cuda()\n return Variable(xs), Variable(ys)\n\ndef train(args, cnn=None):\n # Set the maximum number of threads to prevent crash in Teaching Labs\n # TODO: necessary?\n torch.set_num_threads(5)\n # Numpy random seed\n npr.seed(args.seed)\n\n # Save directory\n save_dir = \"outputs/\" + args.experiment_name\n\n # LOAD THE COLOURS CATEGORIES\n\n # INPUT CHANNEL\n num_in_channels = 1 if not args.downsize_input else 3\n # LOAD THE MODEL\n if cnn is None:\n Net = globals()[args.model]\n cnn =Generator_two(args.kernel,args.num_filters)\n discriminator = Discriminator_two(args.kernel,args.num_filters)\n\n # LOSS FUNCTION\n\n criterion = nn.BCELoss()\n g_optimizer = torch.optim.Adam(cnn.parameters(), lr=1e-4)\n d_optimizer = torch.optim.Adam(discriminator.parameters(), lr=1e-4)\n\n # DATA\n print(\"Loading data...\")", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "FUNCTION\n\n criterion = nn.BCELoss()\n g_optimizer = torch.optim.Adam(cnn.parameters(), lr=1e-4)\n d_optimizer = torch.optim.Adam(discriminator.parameters(), lr=1e-4)\n\n # DATA\n print(\"Loading data...\")\n (x_train, y_train), (x_test, y_test) = load_cifar10(transpose=True)\n\n train_rgb, train_grey = process(x_train, y_train, downsize_input=False)\n test_rgb, test_grey = process(x_test, y_test, downsize_input=False)\n\n train_lab=rgb2lab(train_rgb).astype(\"float32\")\n\n train_lab[:,:,:,:1]=train_lab[:,:,:,:1] / 50. - 1. #normalizing\n train_lab[:,:,:,1:]= train_lab[:,:,:,1:] / 110. #normalizing\n train_lab=np.transpose(train_lab,(0,3,1,2))\n train_L=train_lab[:,:1,:,:]\n train_ab=train_lab[:,1:,:,:]\n ###################################################################################\n\n # Create the outputs folder if not created already\n if not os.path.exists(save_dir):\n os.makedirs(save_dir)\n\n print(\"Beginning training ...\")\n if args.gpu:\n cnn.cuda()\n discriminator.cuda()\n start = time.time()", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-9", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "outputs folder if not created already\n if not os.path.exists(save_dir):\n os.makedirs(save_dir)\n\n print(\"Beginning training ...\")\n if args.gpu:\n cnn.cuda()\n discriminator.cuda()\n start = time.time()\n\n train_losses = []\n valid_losses = []\n valid_accs = []\n for epoch in range(args.epochs):\n # Train the Model\n cnn.train()\n discriminator.train()\n losses = []\n\n for i, (xs, ys) in enumerate(get_batch(train_L, train_ab, args.batch_size)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n\n #--->ADDED 5\n img_grey = images\n img_real = labels\n batch_size = args.batch_size\n\n #discriminator training\n d_optimizer.zero_grad()\n real_prob = discriminator(img_real,img_grey)\n real_labels=Variable(torch.ones(batch_size)).cuda()\n real_loss = criterion(real_prob, real_labels)\n fake_images = cnn(img_grey)\n fake_prob = discriminator(fake_images, img_grey)\n fake_labels=Variable(torch.zeros(batch_size)).cuda()\n fake_loss = criterion(fake_prob, fake_labels)\n d_loss = real_loss + fake_loss\n d_loss.backward()", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-10", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "ake_prob = discriminator(fake_images, img_grey)\n fake_labels=Variable(torch.zeros(batch_size)).cuda()\n fake_loss = criterion(fake_prob, fake_labels)\n d_loss = real_loss + fake_loss\n d_loss.backward()\n d_optimizer.step()\n\n # generator training\n g_optimizer.zero_grad()\n fake_images = cnn(img_grey)\n make_real = discriminator(fake_images, img_grey)\n g_loss= criterion(make_real, real_labels)\n g_loss.backward()\n g_optimizer.step()\n\n img_real_final, gray_im=labtorgb(img_grey,img_real)\n print(gray_im.shape)\n print(type(img_real))\n\n img_fake_final, gray_im=labtorgb(img_grey,fake_images)\n gray_im=torch.tensor(gray_im)\n\n img_real=torch.tensor(img_real_final).permute(0,3,1,2)\n fake_images=torch.tensor(img_fake_final).permute(0,3,1,2)\n\n print(epoch, g_loss.cpu().detach(), d_loss.cpu().detach())\n print(torch.max(fake_images), torch.min(fake_images))\n print(torch.max(gray_im), torch.min(gray_im))\n print(torch.max(img_real), torch.min(img_real))\n visual(gray_im, img_real, fake_images, args.gpu, 1)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-11", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "t(torch.max(fake_images), torch.min(fake_images))\n print(torch.max(gray_im), torch.min(gray_im))\n print(torch.max(img_real), torch.min(img_real))\n visual(gray_im, img_real, fake_images, args.gpu, 1)\n\n return cnn\n\nPython implementation:\nargs = AttrDict()\nargs_dict = {\n \"gpu\": True,\n \"valid\": False,\n \"checkpoint\": \"\",\n \"colours\": \"./data/colours/colour_kmeans24_cat7.npy\",\n \"model\": \"Generator_two\",\n \"kernel\": 3,\n \"num_filters\": 16,\n 'learn_rate':0.5,\n \"batch_size\": 20,\n \"epochs\": 100,\n \"seed\": 0,\n \"plot\": False,\n \"experiment_name\": \"colourization_cnn\",\n \"visualize\": False,\n \"downsize_input\": False,\n}\nargs.update(args_dict)\ncnn = train(args)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "314-12", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 2. Exploration" + }, + { + "text": "Explanation:\n## Part 3. New Data\nRetrieve sample pictures from online and demonstrate how well your best model performs. Provide all your code.\n\n$\\color{blue}{\\text{ }}$\n\n- 25 sample images have been used for evaluating the last model( the best one). The quality and colors of images are close to the real ones. However they are not the same.Also, by using a pretrained model for generator, we can improve the model significantly \n\nPython implementation:\n# provide your code here\nfrom google.colab import drive\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torchvision\nimport torchvision.transforms as transforms\nfrom torchvision import datasets\n\ndrive.mount('/content/drive')\n\nImg_path='/content/drive/MyDrive/Colab Notebooks/ASP360/Lab3'\n\ntransformations = transforms.Compose([transforms.Resize((32,32)),transforms.ToTensor()])\n\nfull_data = torchvision.datasets.ImageFolder(Img_path,transform=transformations)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "315-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 3. New Data" + }, + { + "text": "olab Notebooks/ASP360/Lab3'\n\ntransformations = transforms.Compose([transforms.Resize((32,32)),transforms.ToTensor()])\n\nfull_data = torchvision.datasets.ImageFolder(Img_path,transform=transformations)\nhorse_loader = torch.utils.data.DataLoader(full_data,batch_size=1,num_workers=0, shuffle=True)\n\nPython implementation:\ntest_img=[]\nfor i, data in enumerate(horse_loader,0):\n inputs,labels=data\n inputs=inputs.squeeze(0)\n inputs=inputs.numpy()\n test_img.append(inputs)\n horse_test=np.array(test_img)\n\nPython implementation:\ntest_lab=rgb2lab(horse_test.transpose(0, 2, 3, 1)).astype(\"float32\")\ntest_lab[:,:,:,:1]=test_lab[:,:,:,:1] / 50. - 1. #normalizing\ntest_lab[:,:,:,1:]= test_lab[:,:,:,1:] / 110. #normalizing\ntest_lab=np.transpose(test_lab,(0,3,1,2))\ntest_L=test_lab[:,:1,:,:]\ntest_ab=test_lab[:,1:,:,:]\n\nfor i, (xs, ys) in enumerate(get_batch(test_L, test_ab, batch_size=5)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n\n #--->ADDED 5\n img_grey = images\n img_real = labels\n batch_size = 5", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "315-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 3. New Data" + }, + { + "text": ":,:,:]\n\nfor i, (xs, ys) in enumerate(get_batch(test_L, test_ab, batch_size=5)):\n images, labels = get_torch_vars(xs, ys, args.gpu)\n\n #--->ADDED 5\n img_grey = images\n img_real = labels\n batch_size = 5\n fake_images = cnn(img_grey)\n\n img_real_final, gray_im=labtorgb(img_grey,img_real)\n print(gray_im.shape)\n print(type(img_real))\n\n img_fake_final, gray_im=labtorgb(img_grey,fake_images)\n gray_im=torch.tensor(gray_im)\n\n img_real=torch.tensor(img_real_final).permute(0,3,1,2)\n fake_images=torch.tensor(img_fake_final).permute(0,3,1,2)\n visual(gray_im, img_real, fake_images, args.gpu, 1)", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "315-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Part 3. New Data" + }, + { + "text": "Explanation:\n# Results\n\n## Summary\n\nThis notebook applied multiple deep learning architectures to the image colorization problem.\n\n### Implemented Models\n\n**Part A**\n\n- Regression CNN\n- U-Net\n- Skip Connections\n\n**Part B**\n\n- Conditional GAN (cGAN)\n- Generator\n- Discriminator\n\n### Key Outcomes\n\n- Successfully formulated image colorization as a supervised learning task.\n- Compared regression-based CNNs with U-Net architectures.\n- Demonstrated the benefits of skip connections for preserving image details.\n- Implemented a Conditional GAN for image-to-image translation.\n- Compared reconstruction-based and adversarial approaches for realistic image colorization.\n- Evaluated model performance on unseen images.\n\nThe experiments demonstrate how adversarial learning can generate more realistic colorizations than regression-based methods while building upon the encoder–decoder architectures introduced earlier in the repository.", + "source": "04_image_colorization.ipynb", + "file_type": "ipynb", + "chunk_id": "316-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "generative-deep-learning", + "relative_path": "notebooks/04_image_colorization.ipynb", + "section": "Results" + }, + { + "text": "# Module 1 – Introduction to Transformers\n\n![Transformer Architecture](transformers_architecture.png)\n\nThis module introduces the fundamentals of Transformer-based models using the Hugging Face Transformers library.\n\n## Topics Covered\n\n- Transformer pipelines\n- Sentiment analysis\n- Text generation\n- Zero-shot classification\n- Named Entity Recognition\n- Question Answering\n\n## Notebook\n\n| Notebook | Description |\n|----------|-------------|\n| `01_transformers_intro.ipynb` | Introduction to Transformer pipelines and inference with pretrained models. |\n\n## Skills Demonstrated\n\n- Hugging Face `pipeline`\n- Pretrained Transformer models\n- NLP inference\n- Python\n- PyTorch", + "source": "README.md", + "file_type": "md", + "chunk_id": "317-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/README.md", + "section": null + }, + { + "text": "Explanation:\n# Introduction to Transformers\n\nThis notebook contains my implementation of the introductory practical exercise from the Hugging Face Large Language Model Course.\n\n## Learning Objectives\n\n- Understand the Transformer architecture\n- Use Hugging Face `pipeline`\n- Perform inference with pretrained models\n- Explore common NLP tasks\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n\n---", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "318-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Introduction to Transformers" + }, + { + "text": "Explanation:\n## Text Classification\n\nText classification assigns one or more labels to a piece of text.\nIn this example, we use a pretrained sentiment analysis model through the Hugging Face `pipeline` API.\n\nPython implementation:\nfrom transformers import pipeline\n\nclassifier = pipeline(\"sentiment-analysis\")\nclassifier(\"I've been waiting for a HuggingFace course my whole life.\")\n\nPython implementation:\nfrom transformers import pipeline\n\nclassifier = pipeline(\"sentiment-analysis\",\n device=\"mps\")\nclassifier(\"I've been waiting for a HuggingFace course my whole life.\")\n\nPython implementation:\nclassifier(\n [\"I've been waiting for a HuggingFace course my whole life.\", \"I hate this so much!\"]\n)\n\nPython implementation:\nfrom transformers import pipeline\n\nclassifier = pipeline(\"zero-shot-classification\")\nclassifier(\n \"This is a course about the Transformers library\",\n candidate_labels=[\"education\", \"politics\", \"business\"],\n)", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "319-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Text Classification" + }, + { + "text": "Explanation:\n## Text Generation\n\nText generation predicts the next tokens in a sequence to generate coherent natural language.\n\nPython implementation:\nfrom transformers import pipeline\ngenerator = pipeline(\"text-generation\",\n model=\"distilgpt2\",\n device='mps')\ngenerator(\"In this course, we will teach you how to\")\n\nPython implementation:\nfrom transformers import pipeline\n\ngenerator = pipeline(\"text-generation\", model=\"distilgpt2\")\ngenerator(\n \"In this course, we will teach you how to\",\n max_length=30,\n num_return_sequences=2,\n)", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "320-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Text Generation" + }, + { + "text": "Explanation:\n## Masked Language Modeling (Mask Filling)\n\nMasked language modeling predicts missing words within a sentence. Here, we use a pretrained masked language model to infer the most probable token for a masked position in the input text.\n\nPython implementation:\nfrom transformers import pipeline\n\nunmasker = pipeline(\"fill-mask\")\nunmasker(\"This course will teach you all about models.\", top_k=2)", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "321-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Masked Language Modeling (Mask Filling)" + }, + { + "text": "Explanation:\n## Named Entity Recognition (NER)\n\nNamed Entity Recognition identifies and classifies entities such as people, organizations, locations, and dates within text. This example demonstrates how to extract structured information using a pretrained NER model.\n\nPython implementation:\nfrom transformers import pipeline\n\nner = pipeline(\"ner\", aggregation_strategy=\"simple\")\nner(\"My name is Sylvain and I work at Hugging Face in Brooklyn.\")", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "322-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Named Entity Recognition (NER)" + }, + { + "text": "Explanation:\n## Question Answering\n\nQuestion answering locates the most relevant answer to a question within a given context. In this example, we use a pretrained extractive question answering model to identify the answer span from the provided passage.\n\nPython implementation:\nfrom transformers import pipeline\n\nquestion_answerer = pipeline(\"question-answering\")\nquestion_answerer(\n question=\"Where do I work?\",\n context=\"My name is Sylvain and I work at Hugging Face in Brooklyn\",\n)", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "323-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Question Answering" + }, + { + "text": "Explanation:\n## Summarization\n\nText summarization generates a concise version of a longer document while preserving its key information. Here, we use a pretrained sequence-to-sequence Transformer model to automatically summarize text.\n\nPython implementation:\nfrom transformers import pipeline\n\nsummarizer = pipeline(\"summarization\")\nsummarizer(\n \"\"\"\n America has changed dramatically during recent years. Not only has the number of\n graduates in traditional engineering disciplines such as mechanical, civil,\n electrical, chemical, and aeronautical engineering declined, but in most of\n the premier American universities engineering curricula now concentrate on\n and encourage largely the study of engineering science. As a result, there\n are declining offerings in engineering subjects dealing with infrastructure,\n the environment, and related issues, and greater concentration on high\n technology subjects, largely supporting increasingly complex scientific\n developments.", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "324-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Summarization" + }, + { + "text": "ng subjects dealing with infrastructure,\n the environment, and related issues, and greater concentration on high\n technology subjects, largely supporting increasingly complex scientific\n developments. While the latter is important, it should not be at the expense\n of more traditional engineering.\n\n Rapidly developing economies such as China and India, as well as other\n industrial countries in Europe and Asia, continue to encourage and advance\n the teaching of engineering. Both China and India, respectively, graduate\n six and eight times as many traditional engineers as does the United States.\n Other industrial countries at minimum maintain their output, while America\n suffers an increasingly serious decline in the number of engineering graduates\n and a lack of well-educated engineers.\n\"\"\"\n)", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "324-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Summarization" + }, + { + "text": "Explanation:\n## Machine Translation\n\nMachine translation converts text from one language to another using pretrained sequence-to-sequence Transformer models.\n\nPython implementation:\nfrom transformers import pipeline\n\ntranslator = pipeline(\"translation\", model=\"Helsinki-NLP/opus-mt-fr-en\")\ntranslator(\"Ce cours est produit par Hugging Face.\")", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "325-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Machine Translation" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nIn this notebook I learned how to:\n\n- Load pretrained Transformer models using Hugging Face.\n- Use the `pipeline` API for inference.\n- Apply Transformer models to common NLP tasks.\n- Understand the workflow for running pretrained language models.\n\nThis notebook serves as the foundation for the remaining practical exercises in the Hugging Face LLM Course.", + "source": "transformers_intro.ipynb", + "file_type": "ipynb", + "chunk_id": "326-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "01-transformers/transformers_intro.ipynb", + "section": "Key Takeaways" + }, + { + "text": "# Module 2 – Understanding the Transformer Pipeline\n\n![Hugging Face Pipeline Workflow](full_nlp_pipeline.png)\n\n*Figure 1. End-to-end inference workflow in the Hugging Face `pipeline`: raw text is tokenized, processed by a Transformer model, and converted into human-readable predictions.*\n\nThis module explains how the Hugging Face `pipeline` API works internally by examining its core components: tokenizers, Transformer models, and inference workflows.\n\n## Topics Covered\n\n- Behind the pipeline\n- Transformer models\n- Tokenization\n- Processing multiple sequences\n- Building the complete inference pipeline\n\n## Notebook\n\n| Notebook | Description |\n|----------|-------------|\n| `02_pipeline_internals.ipynb` | Understanding the internal workflow of the Hugging Face Transformers pipeline. |\n\n## Skills Demonstrated\n\n- Hugging Face Transformers\n- PyTorch\n- Tokenization\n- Model inference\n- NLP preprocessing\n- Pipeline internals\n\n---\n\n## Summary", + "source": "README.md", + "file_type": "md", + "chunk_id": "327-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/README.md", + "section": null + }, + { + "text": "of the Hugging Face Transformers pipeline. |\n\n## Skills Demonstrated\n\n- Hugging Face Transformers\n- PyTorch\n- Tokenization\n- Model inference\n- NLP preprocessing\n- Pipeline internals\n\n---\n\n## Summary\n\nThis module provides a detailed look at the internal workflow of the Hugging Face `pipeline`, illustrating how tokenization, model inference, and post-processing work together to perform NLP tasks efficiently.", + "source": "README.md", + "file_type": "md", + "chunk_id": "327-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/README.md", + "section": null + }, + { + "text": "Explanation:\n# Understanding the Transformer Pipeline\n\nThis notebook explores the internal components of the Hugging Face `pipeline` API by examining how Transformer models, tokenizers, and inference workflows interact to process natural language.\n\n## Learning Objectives\n\n- Understand the internal workflow of the Hugging Face pipeline\n- Load and inspect Transformer models\n- Explore how tokenizers convert text into model inputs\n- Process single and multiple text sequences\n- Build the complete inference pipeline manually\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n\nPython implementation:\nfrom transformers import pipeline\n\nclassifier = pipeline(\"sentiment-analysis\")\nclassifier(\n [\n \"I've been waiting for a HuggingFace course my whole life.\",\n \"I hate this so much!\",\n ]\n)\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "328-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Understanding the Transformer Pipeline" + }, + { + "text": "s so much!\",\n ]\n)\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\n\nPython implementation:\nraw_inputs = [\n \"I've been waiting for a HuggingFace course my whole life.\",\n \"I hate this so much!\",\n]\ninputs = tokenizer(raw_inputs, padding=True, truncation=True, return_tensors=\"pt\")\nprint(inputs)\n\nPython implementation:\nfrom transformers import AutoModel\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\nmodel = AutoModel.from_pretrained(checkpoint)\n\nPython implementation:\nfrom transformers import AutoModelForSequenceClassification\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint)\noutputs = model(**inputs)\n\nPython implementation:\nimport torch\n\npredictions = torch.nn.functional.softmax(outputs.logits, dim=-1)\nprint(predictions)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "328-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Understanding the Transformer Pipeline" + }, + { + "text": "Explanation:\n## Transformer Models\n\nTransformer models are responsible for learning contextual representations of text and performing inference. In this section, we explore how pretrained models are loaded and how they generate predictions from tokenized inputs.\n\nPython implementation:\nfrom transformers import BertConfig, BertModel\n\n# Building the config\nconfig = BertConfig()\n\n# Building the model from the config\nmodel = BertModel(config)\n\nPython implementation:\nfrom transformers import BertConfig, BertModel\n\nconfig = BertConfig()\nmodel = BertModel(config)\n\n# Model is randomly initialized!\n\nPython implementation:\nfrom transformers import BertModel\n\nmodel = BertModel.from_pretrained(\"bert-base-cased\")\n\nPython implementation:\nencoded_sequences = [\n [101, 7592, 999, 102],\n [101, 4658, 1012, 102],\n [101, 3835, 999, 102],\n]\n\nPython implementation:\nimport torch\n\nmodel_inputs = torch.tensor(encoded_sequences)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "329-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Transformer Models" + }, + { + "text": "Explanation:\n## Tokenization\n\nBefore text can be processed by a Transformer model, it must be converted into numerical tokens. This section demonstrates how tokenizers prepare raw text by splitting, encoding, and formatting inputs for the model.\n\nPython implementation:\nfrom transformers import BertTokenizer\n\ntokenizer = BertTokenizer.from_pretrained(\"bert-base-cased\")\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ntokenizer = AutoTokenizer.from_pretrained(\"bert-base-cased\")\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ntokenizer = AutoTokenizer.from_pretrained(\"bert-base-cased\")\n\nsequence = \"Using a Transformer network is simple\"\ntokens = tokenizer.tokenize(sequence)\n\nprint(tokens)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "330-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Tokenization" + }, + { + "text": "Explanation:\n## Processing Multiple Sequences\n\nTransformer models can process sentence pairs or batches of text. This section demonstrates how multiple sequences are tokenized and represented using attention masks and token type IDs.\n\nPython implementation:\nimport torch\nfrom transformers import AutoTokenizer, AutoModelForSequenceClassification\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint)\n\nsequence = \"I've been waiting for a HuggingFace course my whole life.\"\n\ntokens = tokenizer.tokenize(sequence)\nids = tokenizer.convert_tokens_to_ids(tokens)\ninput_ids = torch.tensor(ids)\n# This line will fail.\nmodel(input_ids)\n\nPython implementation:\nimport torch\nfrom transformers import AutoTokenizer, AutoModelForSequenceClassification\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "331-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Processing Multiple Sequences" + }, + { + "text": "torch\nfrom transformers import AutoTokenizer, AutoModelForSequenceClassification\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint)\n\nsequence = \"I've been waiting for a HuggingFace course my whole life.\"\n\ntokens = tokenizer.tokenize(sequence)\nids = tokenizer.convert_tokens_to_ids(tokens)\n\ninput_ids = torch.tensor([ids])\nprint(\"Input IDs:\", input_ids)\n\noutput = model(input_ids)\nprint(\"Logits:\", output.logits)\n\nPython implementation:\nbatched_ids = [\n [200, 200, 200],\n [200, 200]\n]\n\nPython implementation:\npadding_id = 100\n\nbatched_ids = [\n [200, 200, 200],\n [200, 200, padding_id],\n]\n\nPython implementation:\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint)\n\nsequence1_ids = [[200, 200, 200]]\nsequence2_ids = [[200, 200]]\nbatched_ids = [\n [200, 200, 200],\n [200, 200, tokenizer.pad_token_id],\n]", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "331-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Processing Multiple Sequences" + }, + { + "text": "utoModelForSequenceClassification.from_pretrained(checkpoint)\n\nsequence1_ids = [[200, 200, 200]]\nsequence2_ids = [[200, 200]]\nbatched_ids = [\n [200, 200, 200],\n [200, 200, tokenizer.pad_token_id],\n]\n\nprint(model(torch.tensor(sequence1_ids)).logits)\nprint(model(torch.tensor(sequence2_ids)).logits)\nprint(model(torch.tensor(batched_ids)).logits)\n\nPython implementation:\nbatched_ids = [\n [200, 200, 200],\n [200, 200, tokenizer.pad_token_id],\n]\n\nattention_mask = [\n [1, 1, 1],\n [1, 1, 0],\n]\n\noutputs = model(torch.tensor(batched_ids), attention_mask=torch.tensor(attention_mask))\nprint(outputs.logits)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "331-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Processing Multiple Sequences" + }, + { + "text": "Explanation:\n## Building the Complete Inference Pipeline\n\nThis section combines tokenization, model inference, and prediction post-processing to illustrate how the Hugging Face `pipeline` API works internally.\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\n\nsequence = \"I've been waiting for a HuggingFace course my whole life.\"\n\nmodel_inputs = tokenizer(sequence)\n\nPython implementation:\nsequence = \"I've been waiting for a HuggingFace course my whole life.\"\n\nmodel_inputs = tokenizer(sequence)\n\nPython implementation:\nsequences = [\"I've been waiting for a HuggingFace course my whole life.\", \"So have I!\"]\n\nmodel_inputs = tokenizer(sequences)\n\nPython implementation:\n# Will pad the sequences up to the maximum sequence length\nmodel_inputs = tokenizer(sequences, padding=\"longest\")\n\n# Will pad the sequences up to the model max length", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "332-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Building the Complete Inference Pipeline" + }, + { + "text": "quences)\n\nPython implementation:\n# Will pad the sequences up to the maximum sequence length\nmodel_inputs = tokenizer(sequences, padding=\"longest\")\n\n# Will pad the sequences up to the model max length\n# (512 for BERT or DistilBERT)\nmodel_inputs = tokenizer(sequences, padding=\"max_length\")\n\n# Will pad the sequences up to the specified max length\nmodel_inputs = tokenizer(sequences, padding=\"max_length\", max_length=8)\n\nPython implementation:\nsequences = [\"I've been waiting for a HuggingFace course my whole life.\", \"So have I!\"]\n\n# Will truncate the sequences that are longer than the model max length\n# (512 for BERT or DistilBERT)\nmodel_inputs = tokenizer(sequences, truncation=True)\n\n# Will truncate the sequences that are longer than the specified max length\nmodel_inputs = tokenizer(sequences, max_length=8, truncation=True)\n\nPython implementation:\nsequences = [\"I've been waiting for a HuggingFace course my whole life.\", \"So have I!\"]\n\n# Returns PyTorch tensors", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "332-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Building the Complete Inference Pipeline" + }, + { + "text": "inputs = tokenizer(sequences, max_length=8, truncation=True)\n\nPython implementation:\nsequences = [\"I've been waiting for a HuggingFace course my whole life.\", \"So have I!\"]\n\n# Returns PyTorch tensors\nmodel_inputs = tokenizer(sequences, padding=True, return_tensors=\"pt\")\n\n# Returns TensorFlow tensors\nmodel_inputs = tokenizer(sequences, padding=True, return_tensors=\"tf\")\n\n# Returns NumPy arrays\nmodel_inputs = tokenizer(sequences, padding=True, return_tensors=\"np\")\n\nPython implementation:\nsequence = \"I've been waiting for a HuggingFace course my whole life.\"\n\nmodel_inputs = tokenizer(sequence)\nprint(model_inputs[\"input_ids\"])\n\ntokens = tokenizer.tokenize(sequence)\nids = tokenizer.convert_tokens_to_ids(tokens)\nprint(ids)\n\nPython implementation:\nimport torch\nfrom transformers import AutoTokenizer, AutoModelForSequenceClassification\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "332-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Building the Complete Inference Pipeline" + }, + { + "text": "torch\nfrom transformers import AutoTokenizer, AutoModelForSequenceClassification\n\ncheckpoint = \"distilbert-base-uncased-finetuned-sst-2-english\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint)\nsequences = [\"I've been waiting for a HuggingFace course my whole life.\", \"So have I!\"]\n\ntokens = tokenizer(sequences, padding=True, truncation=True, return_tensors=\"pt\")\noutput = model(**tokens)", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "332-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Building the Complete Inference Pipeline" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nThis notebook helps to:\n\n- Understand the internal workflow of the Hugging Face `pipeline`.\n- Tokenize raw text for Transformer models.\n- Run inference using pretrained Transformer models.\n- Process multiple input sequences.\n- Reconstruct the complete NLP inference pipeline manually.", + "source": "pipeline_internals.ipynb", + "file_type": "ipynb", + "chunk_id": "333-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "02-transformer-workflow/pipeline_internals.ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Fine-Tuning Transformer Models\n\nThis notebook demonstrates how to prepare datasets and fine-tune pretrained Transformer models using the Hugging Face Transformers library and the Trainer API.\n\n## Learning Objectives\n\n- Prepare datasets for model training\n- Tokenize and preprocess text efficiently\n- Configure training hyperparameters\n- Fine-tune pretrained Transformer models\n- Evaluate model performance\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Evaluate", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "334-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Fine-Tuning Transformer Models" + }, + { + "text": "Explanation:\n## Performing a Single Training Step\n\nBefore training a model on an entire dataset, it is useful to understand what happens during a single optimization step.\n\nIn this example, we fine-tune a pretrained Transformer model using **one batch** of labeled examples. The workflow includes:\n\n1. Tokenizing the input text.\n2. Performing a forward pass through the model.\n3. Computing the training loss.\n4. Calculating gradients using backpropagation.\n5. Updating the model parameters with the AdamW optimizer.\n\nThis single-batch example illustrates the fundamental operations that are repeated throughout the full training process.\n\nPython implementation:\nimport torch\nfrom torch.optim import AdamW\nfrom transformers import AutoTokenizer, AutoModelForSequenceClassification\n\n# Load the pretrained tokenizer and sequence classification model\ncheckpoint = \"bert-base-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "335-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Performing a Single Training Step" + }, + { + "text": "Tokenizer, AutoModelForSequenceClassification\n\n# Load the pretrained tokenizer and sequence classification model\ncheckpoint = \"bert-base-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint)\n\n# Example training samples\nsequences = [\n \"I've been waiting for a Hugging Face course my whole life.\",\n \"This course is amazing!\",\n]\n\n# Tokenize the input text and convert it into PyTorch tensors\nbatch = tokenizer(\n sequences,\n padding=True,\n truncation=True,\n return_tensors=\"pt\",\n)\n\n# Ground-truth sentiment labels (1 = positive)\nbatch[\"labels\"] = torch.tensor([1, 1])\n\n# Initialize the AdamW optimizer\noptimizer = AdamW(model.parameters())\n\n# Clear any previously accumulated gradients\noptimizer.zero_grad()\n\n# Perform a forward pass and compute the training loss\nloss = model(**batch).loss\n\n# Compute gradients using backpropagation\nloss.backward()\n\n# Update the model parameters\noptimizer.step()", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "335-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Performing a Single Training Step" + }, + { + "text": "o_grad()\n\n# Perform a forward pass and compute the training loss\nloss = model(**batch).loss\n\n# Compute gradients using backpropagation\nloss.backward()\n\n# Update the model parameters\noptimizer.step()\n\n# Display the training loss\nprint(f\"Training loss: {loss.item():.4f}\")", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "335-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Performing a Single Training Step" + }, + { + "text": "Explanation:\n### Loading a Dataset\n\nBefore training a Transformer model, we need a labeled dataset. The Hugging Face `datasets` library provides access to a large collection of benchmark datasets and utilities for downloading, processing, and managing them efficiently.\n\nIn this example, we load a sentiment analysis dataset that will be used throughout the remainder of the notebook.\n\nPython implementation:\nfrom datasets import load_dataset\n\nraw_datasets = load_dataset(\"glue\", \"mrpc\")\nraw_datasets\n\nPython implementation:\nraw_train_dataset = raw_datasets[\"train\"]\nraw_train_dataset[0]", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "336-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Loading a Dataset" + }, + { + "text": "Explanation:\n### Preprocessing the Dataset\n\nRaw text cannot be fed directly into a Transformer model. It must first be tokenized and converted into numerical representations that the model can process.\n\nThis preprocessing step applies the tokenizer to every example in the dataset while generating the necessary input tensors required for training.\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ncheckpoint = \"bert-base-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\ntokenized_sentences_1 = tokenizer(raw_datasets[\"train\"][\"sentence1\"])\ntokenized_sentences_2 = tokenizer(raw_datasets[\"train\"][\"sentence2\"])\n\nPython implementation:\ninputs = tokenizer(\"This is the first sentence.\", \"This is the second one.\")\ninputs\n\nPython implementation:\ntokenized_dataset = tokenizer(\n raw_datasets[\"train\"][\"sentence1\"],\n raw_datasets[\"train\"][\"sentence2\"],\n padding=True,\n truncation=True,\n)\n\nPython implementation:\ndef tokenize_function(example):", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "337-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Preprocessing the Dataset" + }, + { + "text": "on:\ntokenized_dataset = tokenizer(\n raw_datasets[\"train\"][\"sentence1\"],\n raw_datasets[\"train\"][\"sentence2\"],\n padding=True,\n truncation=True,\n)\n\nPython implementation:\ndef tokenize_function(example):\n return tokenizer(example[\"sentence1\"], example[\"sentence2\"], truncation=True)\n\nPython implementation:\ntokenized_datasets = raw_datasets.map(tokenize_function, batched=True)\ntokenized_datasets\n\nPython implementation:\nfrom transformers import DataCollatorWithPadding\n\ndata_collator = DataCollatorWithPadding(tokenizer=tokenizer)\n\nPython implementation:\nsamples = tokenized_datasets[\"train\"][:8]\nsamples = {k: v for k, v in samples.items() if k not in [\"idx\", \"sentence1\", \"sentence2\"]}\n[len(x) for x in samples[\"input_ids\"]]\n\nPython implementation:\nbatch = data_collator(samples)\n{k: v.shape for k, v in batch.items()}", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "337-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Preprocessing the Dataset" + }, + { + "text": "Explanation:\n## Fine-Tuning with the Hugging Face Trainer API\n\nThe Hugging Face `Trainer` API provides a high-level interface for fine-tuning Transformer models. It automates many aspects of the training process, including optimization, evaluation, checkpointing, and logging, allowing developers to focus on model development instead of boilerplate training code.\n\nPython implementation:\nfrom datasets import load_dataset\nfrom transformers import AutoTokenizer, DataCollatorWithPadding\n\nraw_datasets = load_dataset(\"glue\", \"mrpc\")\ncheckpoint = \"bert-base-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\n\ndef tokenize_function(example):\n return tokenizer(example[\"sentence1\"], example[\"sentence2\"], truncation=True)\n\ntokenized_datasets = raw_datasets.map(tokenize_function, batched=True)\ndata_collator = DataCollatorWithPadding(tokenizer=tokenizer)\n\nPython implementation:\nfrom transformers import TrainingArguments\n\ntraining_args = TrainingArguments(\"test-trainer\")", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "338-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Fine-Tuning with the Hugging Face Trainer API" + }, + { + "text": "tion, batched=True)\ndata_collator = DataCollatorWithPadding(tokenizer=tokenizer)\n\nPython implementation:\nfrom transformers import TrainingArguments\n\ntraining_args = TrainingArguments(\"test-trainer\")\n\nPython implementation:\nfrom transformers import AutoModelForSequenceClassification\n\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)\n\nPython implementation:\nimport torch\ndevice = torch.device(\"mps\" if torch.backends.mps.is_available() else \"cpu\")\n\nmodel = model.to(device)\n\nPython implementation:\ntokenized_datasets[\"train\"] = (\n tokenized_datasets[\"train\"]\n .shuffle(seed=42)\n .select(range(1000))\n)\n\ntokenized_datasets[\"validation\"] = (\n tokenized_datasets[\"validation\"]\n .shuffle(seed=42)\n .select(range(200))\n)\n\nPython implementation:\nfrom transformers import Trainer\n\ntrainer = Trainer(\n model,\n training_args,\n train_dataset=tokenized_datasets[\"train\"],\n eval_dataset=tokenized_datasets[\"validation\"],\n data_collator=data_collator,", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "338-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Fine-Tuning with the Hugging Face Trainer API" + }, + { + "text": "from transformers import Trainer\n\ntrainer = Trainer(\n model,\n training_args,\n train_dataset=tokenized_datasets[\"train\"],\n eval_dataset=tokenized_datasets[\"validation\"],\n data_collator=data_collator,\n processing_class=tokenizer,\n)\n\nPython implementation:\nimport numpy as np\n\npreds = np.argmax(predictions.predictions, axis=-1)\n\nPython implementation:\nimport evaluate\n\nmetric = evaluate.load(\"glue\", \"mrpc\")\nmetric.compute(predictions=preds, references=predictions.label_ids)\n\nPython implementation:\ndef compute_metrics(eval_preds):\n metric = evaluate.load(\"glue\", \"mrpc\")\n logits, labels = eval_preds\n predictions = np.argmax(logits, axis=-1)\n return metric.compute(predictions=predictions, references=labels)\n\nPython implementation:\ntraining_args = TrainingArguments(\"test-trainer\", eval_strategy=\"epoch\")\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)\n\ntrainer = Trainer(\n model,\n training_args,\n train_dataset=tokenized_datasets[\"train\"],", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "338-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Fine-Tuning with the Hugging Face Trainer API" + }, + { + "text": "r\", eval_strategy=\"epoch\")\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)\n\ntrainer = Trainer(\n model,\n training_args,\n train_dataset=tokenized_datasets[\"train\"],\n eval_dataset=tokenized_datasets[\"validation\"],\n data_collator=data_collator,\n processing_class=tokenizer,\n compute_metrics=compute_metrics,\n)", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "338-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Fine-Tuning with the Hugging Face Trainer API" + }, + { + "text": "Explanation:\n## End-to-End Fine-Tuning Workflow\n\nThe following example demonstrates the complete supervised fine-tuning workflow. It combines dataset loading, preprocessing, model initialization, training configuration, optimization, and evaluation into a single pipeline using the Hugging Face ecosystem.\n\nPython implementation:\nfrom datasets import load_dataset\nfrom transformers import AutoTokenizer, DataCollatorWithPadding\n\nraw_datasets = load_dataset(\"glue\", \"mrpc\")\ncheckpoint = \"bert-base-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(checkpoint)\n\ndef tokenize_function(example):\n return tokenizer(example[\"sentence1\"], example[\"sentence2\"], truncation=True)\n\ntokenized_datasets = raw_datasets.map(tokenize_function, batched=True)\ndata_collator = DataCollatorWithPadding(tokenizer=tokenizer)\n\nPython implementation:\ntokenized_datasets = tokenized_datasets.remove_columns([\"sentence1\", \"sentence2\", \"idx\"])\ntokenized_datasets = tokenized_datasets.rename_column(\"label\", \"labels\")", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "339-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "End-to-End Fine-Tuning Workflow" + }, + { + "text": "r=tokenizer)\n\nPython implementation:\ntokenized_datasets = tokenized_datasets.remove_columns([\"sentence1\", \"sentence2\", \"idx\"])\ntokenized_datasets = tokenized_datasets.rename_column(\"label\", \"labels\")\ntokenized_datasets.set_format(\"torch\")\ntokenized_datasets[\"train\"].column_names\n\nPython implementation:\ntokenized_datasets[\"train\"] = (\n tokenized_datasets[\"train\"]\n .shuffle(seed=42)\n .select(range(500))\n)\n\ntokenized_datasets[\"validation\"] = (\n tokenized_datasets[\"validation\"]\n .shuffle(seed=42)\n .select(range(200))\n)\n\nPython implementation:\nfrom torch.utils.data import DataLoader\n\ntrain_dataloader = DataLoader(\n tokenized_datasets[\"train\"], shuffle=True, batch_size=8, collate_fn=data_collator\n)\neval_dataloader = DataLoader(\n tokenized_datasets[\"validation\"], batch_size=8, collate_fn=data_collator\n)\n\nPython implementation:\nfor batch in train_dataloader:\n break\n{k: v.shape for k, v in batch.items()}\n\nPython implementation:\nfrom transformers import AutoModelForSequenceClassification", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "339-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "End-to-End Fine-Tuning Workflow" + }, + { + "text": "ta_collator\n)\n\nPython implementation:\nfor batch in train_dataloader:\n break\n{k: v.shape for k, v in batch.items()}\n\nPython implementation:\nfrom transformers import AutoModelForSequenceClassification\n\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)\n\nPython implementation:\nfrom torch.optim import AdamW\n\noptimizer = AdamW(model.parameters(), lr=5e-5)\n\nPython implementation:\nfrom transformers import get_scheduler\n\nnum_epochs = 3\nnum_training_steps = num_epochs * len(train_dataloader)\nlr_scheduler = get_scheduler(\n \"linear\",\n optimizer=optimizer,\n num_warmup_steps=0,\n num_training_steps=num_training_steps,\n)\nprint(num_training_steps)\n\nPython implementation:\nimport torch\n\ndevice = torch.device(\"cuda\") if torch.cuda.is_available() else torch.device(\"cpu\")\nmodel.to(device)\ndevice\n\nPython implementation:\nfrom tqdm.auto import tqdm\n\nprogress_bar = tqdm(range(num_training_steps))\n\nmodel.train()\nfor epoch in range(num_epochs):\n for batch in train_dataloader:", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "339-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "End-to-End Fine-Tuning Workflow" + }, + { + "text": "del.to(device)\ndevice\n\nPython implementation:\nfrom tqdm.auto import tqdm\n\nprogress_bar = tqdm(range(num_training_steps))\n\nmodel.train()\nfor epoch in range(num_epochs):\n for batch in train_dataloader:\n batch = {k: v.to(device) for k, v in batch.items()}\n outputs = model(**batch)\n loss = outputs.loss\n loss.backward()\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()\n progress_bar.update(1)\n\nPython implementation:\nfrom transformers import AutoModelForSequenceClassification, get_scheduler\nfrom torch.optim import AdamW\n\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)\noptimizer = AdamW(model.parameters(), lr=3e-5)\n\ndevice = torch.device(\"cuda\") if torch.cuda.is_available() else torch.device(\"cpu\")\nmodel.to(device)\n\nnum_epochs = 3\nnum_training_steps = num_epochs * len(train_dataloader)\nlr_scheduler = get_scheduler(\n \"linear\",\n optimizer=optimizer,\n num_warmup_steps=0,\n num_training_steps=num_training_steps,\n)", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "339-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "End-to-End Fine-Tuning Workflow" + }, + { + "text": "num_epochs = 3\nnum_training_steps = num_epochs * len(train_dataloader)\nlr_scheduler = get_scheduler(\n \"linear\",\n optimizer=optimizer,\n num_warmup_steps=0,\n num_training_steps=num_training_steps,\n)\n\nprogress_bar = tqdm(range(num_training_steps))\n\nmodel.train()\nfor epoch in range(num_epochs):\n for batch in train_dataloader:\n batch = {k: v.to(device) for k, v in batch.items()}\n outputs = model(**batch)\n loss = outputs.loss\n loss.backward()\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()\n progress_bar.update(1)\n\nPython implementation:\nfrom accelerate import Accelerator\nfrom transformers import AutoModelForSequenceClassification, get_scheduler\nfrom torch.optim import AdamW\n\naccelerator = Accelerator()\n\nmodel = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)\noptimizer = AdamW(model.parameters(), lr=3e-5)\n\ntrain_dl, eval_dl, model, optimizer = accelerator.prepare(\n train_dataloader, eval_dataloader, model, optimizer\n)\n\nnum_epochs = 3", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "339-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "End-to-End Fine-Tuning Workflow" + }, + { + "text": "kpoint, num_labels=2)\noptimizer = AdamW(model.parameters(), lr=3e-5)\n\ntrain_dl, eval_dl, model, optimizer = accelerator.prepare(\n train_dataloader, eval_dataloader, model, optimizer\n)\n\nnum_epochs = 3\nnum_training_steps = num_epochs * len(train_dl)\nlr_scheduler = get_scheduler(\n \"linear\",\n optimizer=optimizer,\n num_warmup_steps=0,\n num_training_steps=num_training_steps,\n)\n\nprogress_bar = tqdm(range(num_training_steps))\n\nmodel.train()\nfor epoch in range(num_epochs):\n for batch in train_dl:\n outputs = model(**batch)\n loss = outputs.loss\n accelerator.backward(loss)\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()\n progress_bar.update(1)", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "339-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "End-to-End Fine-Tuning Workflow" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Prepare datasets for Transformer training.\n- Tokenize and preprocess text efficiently.\n- Configure training with the Hugging Face Trainer API.\n- Fine-tune pretrained Transformer models.\n- Evaluate model performance on downstream NLP tasks.", + "source": "03_model_training.ipynb", + "file_type": "ipynb", + "chunk_id": "340-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/03_model_training.ipynb", + "section": "Key Takeaways" + }, + { + "text": "# Module 3 – Fine-Tuning Transformer Models\n\nThis module demonstrates how to prepare datasets and fine-tune pretrained Transformer models using the Hugging Face Transformers library and the `Trainer` API.\n\nUnlike the later LLM-specific modules (e.g., SFT, LoRA, and GRPO), this notebook focuses on the standard supervised fine-tuning workflow for Transformer models on downstream NLP tasks.\n\n---\n\n## Learning Objectives\n\n- Load and inspect NLP datasets\n- Tokenize and preprocess text\n- Perform a single optimization step\n- Configure training with `TrainingArguments`\n- Fine-tune pretrained Transformer models using the `Trainer` API\n- Evaluate model performance\n\n---\n\n## Topics Covered\n\n- Dataset loading\n- Dataset preprocessing\n- Tokenization\n- Dynamic padding with data collators\n- Single-batch training\n- Hugging Face Trainer API\n- Training configuration\n- Model evaluation\n\n---\n\n## Notebook\n\n| Notebook | Description |\n|----------|-------------|", + "source": "README.md", + "file_type": "md", + "chunk_id": "341-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/README.md", + "section": null + }, + { + "text": "mic padding with data collators\n- Single-batch training\n- Hugging Face Trainer API\n- Training configuration\n- Model evaluation\n\n---\n\n## Notebook\n\n| Notebook | Description |\n|----------|-------------|\n| `03_model_training.ipynb` | End-to-end supervised fine-tuning of a pretrained Transformer model using the Hugging Face Trainer API. |\n\n---\n\n## Skills Demonstrated\n\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Evaluate\n- Hugging Face Trainer API\n- Dataset preprocessing\n- Tokenization\n- Model fine-tuning\n- Model evaluation\n\n---\n\n## Requirements\n\nInstall the required packages from the repository root:\n\n```bash\npip install -r ../requirements.txt\n```\n\n---\n\n## Repository Structure\n\n```\n03-model-training/\n│\n├── README.md\n└── 03_model_training.ipynb\n```\n\n---\n\n## Key Takeaways\n\n- Prepare datasets for Transformer training.\n- Build an end-to-end supervised fine-tuning pipeline.\n- Configure and train models using the Hugging Face `Trainer`.", + "source": "README.md", + "file_type": "md", + "chunk_id": "341-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/README.md", + "section": null + }, + { + "text": "ng.ipynb\n```\n\n---\n\n## Key Takeaways\n\n- Prepare datasets for Transformer training.\n- Build an end-to-end supervised fine-tuning pipeline.\n- Configure and train models using the Hugging Face `Trainer`.\n- Evaluate fine-tuned models on downstream NLP tasks.", + "source": "README.md", + "file_type": "md", + "chunk_id": "341-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "03-model-training/README.md", + "section": null + }, + { + "text": "# Module 4 – Tokenization\n\nThis module explores how Hugging Face tokenizers transform raw text into numerical representations that can be processed by Transformer models. It covers the complete tokenization pipeline, from text normalization and pre-tokenization to building custom tokenizers using different tokenization algorithms.\n\n## Learning Objectives\n\n- Understand the architecture of modern tokenizers\n- Explore fast tokenizer capabilities\n- Learn token alignment for Question Answering and Named Entity Recognition\n- Understand text normalization and pre-tokenization\n- Build custom tokenizers from scratch\n- Compare WordPiece, Byte-Pair Encoding (BPE), and Unigram tokenization\n\n## Topics Covered\n\n- Fast tokenizers\n- Token-to-word and character offset mappings\n- Question Answering tokenization\n- Named Entity Recognition tokenization\n- Text normalization\n- Pre-tokenization\n- WordPiece tokenization\n- Byte-Pair Encoding (BPE)\n- Byte-Level BPE\n- Unigram tokenization", + "source": "README.md", + "file_type": "md", + "chunk_id": "342-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/README.md", + "section": null + }, + { + "text": "tion Answering tokenization\n- Named Entity Recognition tokenization\n- Text normalization\n- Pre-tokenization\n- WordPiece tokenization\n- Byte-Pair Encoding (BPE)\n- Byte-Level BPE\n- Unigram tokenization\n- Training tokenizer vocabularies\n- Tokenizer post-processing\n- Saving and reloading custom tokenizers\n\n## Notebook\n\n| Notebook | Description |\n|----------|-------------|\n| `04_tokenizers.ipynb` | Exploring Hugging Face fast tokenizers and building custom tokenizers from scratch. |\n\n## Skills Demonstrated\n\n- Hugging Face Tokenizers\n- Hugging Face Transformers\n- Text preprocessing\n- Offset mappings\n- Question Answering preprocessing\n- Named Entity Recognition preprocessing\n- Tokenization pipelines\n- Vocabulary training\n- Custom tokenizer development\n\n## Requirements\n\nInstall the required dependencies from the repository root:\n\n```bash\npip install -r ../requirements.txt\n```\n\n## Key Takeaways\n\nAfter completing this module, you will understand how to:", + "source": "README.md", + "file_type": "md", + "chunk_id": "342-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/README.md", + "section": null + }, + { + "text": "Requirements\n\nInstall the required dependencies from the repository root:\n\n```bash\npip install -r ../requirements.txt\n```\n\n## Key Takeaways\n\nAfter completing this module, you will understand how to:\n\n- Transform raw text into model-ready token IDs.\n- Use fast tokenizers for NLP tasks such as Question Answering and Named Entity Recognition.\n- Build and train custom tokenizers using WordPiece, Byte-Pair Encoding (BPE), Byte-Level BPE, and Unigram algorithms.\n- Understand each stage of the tokenization pipeline, from normalization to post-processing.", + "source": "README.md", + "file_type": "md", + "chunk_id": "342-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/README.md", + "section": null + }, + { + "text": "Explanation:\n# Understanding Fast Tokenizers\n\nThis notebook explores the internal components of Hugging Face fast tokenizers and demonstrates how text is transformed into numerical representations for Transformer models.\n\n## Learning Objectives\n\n- Understand the advantages of fast tokenizers\n- Explore tokenization in Question Answering pipelines\n- Learn normalization and pre-tokenization\n- Build a tokenizer from its individual components\n\n### Tokenization Pipeline\n\nThe tokenization process transforms raw text into numerical token IDs that can be processed by Transformer models.\n\n```text\nRaw Text\n │\n ▼\nNormalization\n │\n ▼\nPre-tokenization\n │\n ▼\nTokenizer Model\n(BPE / WordPiece / Unigram)\n │\n ▼\nPost-processing\n │\n ▼\nToken IDs\n```\n\nEach stage has a specific purpose:\n\n- **Normalization:** Standardizes the input text (e.g., lowercasing or Unicode normalization).\n- **Pre-tokenization:** Splits the text into preliminary word-like units.", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "343-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Understanding Fast Tokenizers" + }, + { + "text": "stage has a specific purpose:\n\n- **Normalization:** Standardizes the input text (e.g., lowercasing or Unicode normalization).\n- **Pre-tokenization:** Splits the text into preliminary word-like units.\n- **Tokenizer Model:** Applies a tokenization algorithm such as **WordPiece**, **Byte-Pair Encoding (BPE)**, or **Unigram** to generate subword tokens.\n- **Post-processing:** Inserts special tokens (e.g., `[CLS]`, `[SEP]`) and formats the sequence for the target model.\n- **Token IDs:** Converts the final tokens into numerical IDs that serve as input to Transformer models.", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "343-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Understanding Fast Tokenizers" + }, + { + "text": "Explanation:\n## Fast Tokenizers\n\nFast tokenizers are implemented in Rust and provide significant performance improvements over Python-based tokenizers. More importantly, they expose detailed alignment information that links tokens to their original words and character positions.\n\nIn this section, we explore how fast tokenizers:\n\n- Generate tokens from raw text.\n- Map tokens back to their corresponding words.\n- Retrieve character offsets.\n- Support token-level NLP tasks such as Named Entity Recognition (NER).\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ntokenizer = AutoTokenizer.from_pretrained(\"bert-base-cased\")\nexample = \"My name is Sylvain and I work at Hugging Face in Brooklyn.\"\nencoding = tokenizer(example)\nprint(type(encoding))\n\nPython implementation:\nstart, end = encoding.word_to_chars(3)\nexample[start:end]\n\nPython implementation:\nfrom transformers import pipeline\n\ntoken_classifier = pipeline(\"token-classification\")", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "344-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers" + }, + { + "text": "ding))\n\nPython implementation:\nstart, end = encoding.word_to_chars(3)\nexample[start:end]\n\nPython implementation:\nfrom transformers import pipeline\n\ntoken_classifier = pipeline(\"token-classification\")\ntoken_classifier(\"My name is Sylvain and I work at Hugging Face in Brooklyn.\")\n\nPython implementation:\nfrom transformers import pipeline\n\ntoken_classifier = pipeline(\"token-classification\", aggregation_strategy=\"simple\")\ntoken_classifier(\"My name is Sylvain and I work at Hugging Face in Brooklyn.\")\n\nPython implementation:\nfrom transformers import AutoTokenizer, AutoModelForTokenClassification\n\nmodel_checkpoint = \"dbmdz/bert-large-cased-finetuned-conll03-english\"\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint)\nmodel = AutoModelForTokenClassification.from_pretrained(model_checkpoint)\n\nexample = \"My name is Sylvain and I work at Hugging Face in Brooklyn.\"\ninputs = tokenizer(example, return_tensors=\"pt\")\noutputs = model(**inputs)\n\nPython implementation:\nimport torch", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "344-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers" + }, + { + "text": "el_checkpoint)\n\nexample = \"My name is Sylvain and I work at Hugging Face in Brooklyn.\"\ninputs = tokenizer(example, return_tensors=\"pt\")\noutputs = model(**inputs)\n\nPython implementation:\nimport torch\n\nprobabilities = torch.nn.functional.softmax(outputs.logits, dim=-1)[0].tolist()\npredictions = outputs.logits.argmax(dim=-1)[0].tolist()\nprint(predictions)\n\nPython implementation:\nresults = []\ntokens = inputs.tokens()\n\nfor idx, pred in enumerate(predictions):\n label = model.config.id2label[pred]\n if label != \"O\":\n results.append(\n {\"entity\": label, \"score\": probabilities[idx][pred], \"word\": tokens[idx]}\n )\n\nprint(results)\n\nPython implementation:\ninputs_with_offsets = tokenizer(example, return_offsets_mapping=True)\ninputs_with_offsets[\"offset_mapping\"]\n\nPython implementation:\nresults = []\ninputs_with_offsets = tokenizer(example, return_offsets_mapping=True)\ntokens = inputs_with_offsets.tokens()\noffsets = inputs_with_offsets[\"offset_mapping\"]\n\nfor idx, pred in enumerate(predictions):", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "344-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers" + }, + { + "text": "]\ninputs_with_offsets = tokenizer(example, return_offsets_mapping=True)\ntokens = inputs_with_offsets.tokens()\noffsets = inputs_with_offsets[\"offset_mapping\"]\n\nfor idx, pred in enumerate(predictions):\n label = model.config.id2label[pred]\n if label != \"O\":\n start, end = offsets[idx]\n results.append(\n {\n \"entity\": label,\n \"score\": probabilities[idx][pred],\n \"word\": tokens[idx],\n \"start\": start,\n \"end\": end,\n }\n )\n\nprint(results)\n\nPython implementation:\nimport numpy as np\n\nresults = []\ninputs_with_offsets = tokenizer(example, return_offsets_mapping=True)\ntokens = inputs_with_offsets.tokens()\noffsets = inputs_with_offsets[\"offset_mapping\"]\n\nidx = 0\nwhile idx < len(predictions):\n pred = predictions[idx]\n label = model.config.id2label[pred]\n if label != \"O\":\n # Remove the B- or I-\n label = label[2:]\n start, _ = offsets[idx]\n\n # Grab all the tokens labeled with I-label\n all_scores = []\n while (\n idx < len(predictions)\n and model.config.id2label[predictions[idx]] == f\"I-{label}\"\n ):", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "344-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers" + }, + { + "text": "el = label[2:]\n start, _ = offsets[idx]\n\n # Grab all the tokens labeled with I-label\n all_scores = []\n while (\n idx < len(predictions)\n and model.config.id2label[predictions[idx]] == f\"I-{label}\"\n ):\n all_scores.append(probabilities[idx][pred])\n _, end = offsets[idx]\n idx += 1\n\n # The score is the mean of all the scores of the tokens in that grouped entity\n score = np.mean(all_scores).item()\n word = example[start:end]\n results.append(\n {\n \"entity_group\": label,\n \"score\": score,\n \"word\": word,\n \"start\": start,\n \"end\": end,\n }\n )\n idx += 1\n\nprint(results)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "344-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers" + }, + { + "text": "Explanation:\n## Fast Tokenizers for Question Answering\n\nQuestion Answering models predict the start and end positions of an answer within a passage of text.\n\nFast tokenizers provide character offset mappings that enable these predicted token positions to be converted back into readable text spans. They also support sliding-window tokenization, allowing models to process documents that exceed the maximum input length.\n\nPython implementation:\nfrom transformers import pipeline\n\nquestion_answerer = pipeline(\"question-answering\")\ncontext = \"\"\"\n🤗 Transformers is backed by the three most popular deep learning libraries — Jax, PyTorch, and TensorFlow — with a seamless integration\nbetween them. It's straightforward to train your models with one before loading them for inference with the other.\n\"\"\"\nquestion = \"Which deep learning libraries back 🤗 Transformers?\"\nquestion_answerer(question=question, context=context)\n\nPython implementation:\nlong_context = \"\"\"\n🤗 Transformers: State of the Art NLP", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "question = \"Which deep learning libraries back 🤗 Transformers?\"\nquestion_answerer(question=question, context=context)\n\nPython implementation:\nlong_context = \"\"\"\n🤗 Transformers: State of the Art NLP\n\n🤗 Transformers provides thousands of pretrained models to perform tasks on texts such as classification, information extraction,\nquestion answering, summarization, translation, text generation and more in over 100 languages.\nIts aim is to make cutting-edge NLP easier to use for everyone.\n\n🤗 Transformers provides APIs to quickly download and use those pretrained models on a given text, fine-tune them on your own datasets and\nthen share them with the community on our model hub. At the same time, each python module defining an architecture is fully standalone and\ncan be modified to enable quick research experiments.\n\nWhy should I use transformers?\n\n1. Easy-to-use state-of-the-art models:\n - High performance on NLU and NLG tasks.\n - Low barrier to entry for educators and practitioners.", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "quick research experiments.\n\nWhy should I use transformers?\n\n1. Easy-to-use state-of-the-art models:\n - High performance on NLU and NLG tasks.\n - Low barrier to entry for educators and practitioners.\n - Few user-facing abstractions with just three classes to learn.\n - A unified API for using all our pretrained models.\n - Lower compute costs, smaller carbon footprint:\n\n2. Researchers can share trained models instead of always retraining.\n - Practitioners can reduce compute time and production costs.\n - Dozens of architectures with over 10,000 pretrained models, some in more than 100 languages.\n\n3. Choose the right framework for every part of a model's lifetime:\n - Train state-of-the-art models in 3 lines of code.\n - Move a single model between TF2.0/PyTorch frameworks at will.\n - Seamlessly pick the right framework for training, evaluation and production.\n\n4. Easily customize a model or an example to your needs:", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "Move a single model between TF2.0/PyTorch frameworks at will.\n - Seamlessly pick the right framework for training, evaluation and production.\n\n4. Easily customize a model or an example to your needs:\n - We provide examples for each architecture to reproduce the results published by its original authors.\n - Model internals are exposed as consistently as possible.\n - Model files can be used independently of the library for quick experiments.\n\n🤗 Transformers is backed by the three most popular deep learning libraries — Jax, PyTorch and TensorFlow — with a seamless integration\nbetween them. It's straightforward to train your models with one before loading them for inference with the other.\n\"\"\"\nquestion_answerer(question=question, context=long_context)\n\nPython implementation:\nfrom transformers import AutoTokenizer, AutoModelForQuestionAnswering\n\nmodel_checkpoint = \"distilbert-base-cased-distilled-squad\"\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "entation:\nfrom transformers import AutoTokenizer, AutoModelForQuestionAnswering\n\nmodel_checkpoint = \"distilbert-base-cased-distilled-squad\"\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint)\nmodel = AutoModelForQuestionAnswering.from_pretrained(model_checkpoint)\n\ninputs = tokenizer(question, context, return_tensors=\"pt\")\noutputs = model(**inputs)\n\nPython implementation:\nstart_logits = outputs.start_logits\nend_logits = outputs.end_logits\nprint(start_logits.shape, end_logits.shape)\n\nPython implementation:\nimport torch\n\nsequence_ids = inputs.sequence_ids()\n# Mask everything apart from the tokens of the context\nmask = [i != 1 for i in sequence_ids]\n# Unmask the [CLS] token\nmask[0] = False\nmask = torch.tensor(mask)[None]\n\nstart_logits[mask] = -10000\nend_logits[mask] = -10000\n\nPython implementation:\nstart_probabilities = torch.nn.functional.softmax(start_logits, dim=-1)[0]\nend_probabilities = torch.nn.functional.softmax(end_logits, dim=-1)[0]\n\nPython implementation:", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "10000\n\nPython implementation:\nstart_probabilities = torch.nn.functional.softmax(start_logits, dim=-1)[0]\nend_probabilities = torch.nn.functional.softmax(end_logits, dim=-1)[0]\n\nPython implementation:\nmax_index = scores.argmax().item()\nstart_index = max_index // scores.shape[1]\nend_index = max_index % scores.shape[1]\nprint(scores[start_index, end_index])\n\nPython implementation:\ninputs_with_offsets = tokenizer(question, context, return_offsets_mapping=True)\noffsets = inputs_with_offsets[\"offset_mapping\"]\n\nstart_char, _ = offsets[start_index]\n_, end_char = offsets[end_index]\nanswer = context[start_char:end_char]\n\nPython implementation:\nresult = {\n \"answer\": answer,\n \"start\": start_char,\n \"end\": end_char,\n \"score\": scores[start_index, end_index],\n}\nprint(result)\n\nPython implementation:\nsentence = \"This sentence is not too long but we are going to split it anyway.\"\ninputs = tokenizer(\n sentence, truncation=True, return_overflowing_tokens=True, max_length=6, stride=2\n)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "plementation:\nsentence = \"This sentence is not too long but we are going to split it anyway.\"\ninputs = tokenizer(\n sentence, truncation=True, return_overflowing_tokens=True, max_length=6, stride=2\n)\n\nfor ids in inputs[\"input_ids\"]:\n print(tokenizer.decode(ids))\n\nPython implementation:\nsentences = [\n \"This sentence is not too long but we are going to split it anyway.\",\n \"This sentence is shorter but will still get split.\",\n]\ninputs = tokenizer(\n sentences, truncation=True, return_overflowing_tokens=True, max_length=6, stride=2\n)\n\nprint(inputs[\"overflow_to_sample_mapping\"])\n\nPython implementation:\ninputs = tokenizer(\n question,\n long_context,\n stride=128,\n max_length=384,\n padding=\"longest\",\n truncation=\"only_second\",\n return_overflowing_tokens=True,\n return_offsets_mapping=True,\n)\n\nPython implementation:\n_ = inputs.pop(\"overflow_to_sample_mapping\")\noffsets = inputs.pop(\"offset_mapping\")\n\ninputs = inputs.convert_to_tensors(\"pt\")\nprint(inputs[\"input_ids\"].shape)\n\nPython implementation:", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "implementation:\n_ = inputs.pop(\"overflow_to_sample_mapping\")\noffsets = inputs.pop(\"offset_mapping\")\n\ninputs = inputs.convert_to_tensors(\"pt\")\nprint(inputs[\"input_ids\"].shape)\n\nPython implementation:\noutputs = model(**inputs)\n\nstart_logits = outputs.start_logits\nend_logits = outputs.end_logits\nprint(start_logits.shape, end_logits.shape)\n\nPython implementation:\nsequence_ids = inputs.sequence_ids()\n# Mask everything apart from the tokens of the context\nmask = [i != 1 for i in sequence_ids]\n# Unmask the [CLS] token\nmask[0] = False\n# Mask all the [PAD] tokens\nmask = torch.logical_or(torch.tensor(mask)[None], (inputs[\"attention_mask\"] == 0))\n\nstart_logits[mask] = -10000\nend_logits[mask] = -10000\n\nPython implementation:\nstart_probabilities = torch.nn.functional.softmax(start_logits, dim=-1)\nend_probabilities = torch.nn.functional.softmax(end_logits, dim=-1)\n\nPython implementation:\ncandidates = []\nfor start_probs, end_probs in zip(start_probabilities, end_probabilities):", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "_logits, dim=-1)\nend_probabilities = torch.nn.functional.softmax(end_logits, dim=-1)\n\nPython implementation:\ncandidates = []\nfor start_probs, end_probs in zip(start_probabilities, end_probabilities):\n scores = start_probs[:, None] * end_probs[None, :]\n idx = torch.triu(scores).argmax().item()\n\n start_idx = idx // scores.shape[1]\n end_idx = idx % scores.shape[1]\n score = scores[start_idx, end_idx].item()\n candidates.append((start_idx, end_idx, score))\n\nprint(candidates)\n\nPython implementation:\nfor candidate, offset in zip(candidates, offsets):\n start_token, end_token, score = candidate\n start_char, _ = offset[start_token]\n _, end_char = offset[end_token]\n answer = long_context[start_char:end_char]\n result = {\"answer\": answer, \"start\": start_char, \"end\": end_char, \"score\": score}\n print(result)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "345-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Fast Tokenizers for Question Answering" + }, + { + "text": "Explanation:\n## Text Normalization and Pre-Tokenization\n\nBefore text is converted into tokens, it passes through several preprocessing stages.\n\nIn this section, we examine how different Transformer models normalize text, split it into preliminary word units, and prepare it for subword tokenization.\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ntokenizer = AutoTokenizer.from_pretrained(\"bert-base-uncased\")\nprint(type(tokenizer.backend_tokenizer))\n\nPython implementation:\ntokenizer = AutoTokenizer.from_pretrained(\"gpt2\")\ntokenizer.backend_tokenizer.pre_tokenizer.pre_tokenize_str(\"Hello, how are you?\")\n\nPython implementation:\ntokenizer = AutoTokenizer.from_pretrained(\"t5-small\")\ntokenizer.backend_tokenizer.pre_tokenizer.pre_tokenize_str(\"Hello, how are you?\")", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "346-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Text Normalization and Pre-Tokenization" + }, + { + "text": "Explanation:\n## Building a Tokenizer from Scratch\n\nA tokenizer is composed of several modular components that transform raw text into model-ready tokens.\n\nIn this section, we construct custom tokenizers by combining normalization, pre-tokenization, vocabulary training, post-processing, and decoding components.", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "347-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Building a Tokenizer from Scratch" + }, + { + "text": "Explanation:\n### Preparing the Training Corpus\n\nBefore training a tokenizer, we first create a text corpus that will be used to learn the vocabulary.\n\nPython implementation:\nfrom datasets import load_dataset\n\ndataset = load_dataset(\"wikitext\", name=\"wikitext-2-raw-v1\", split=\"train\")\n\ndef get_training_corpus():\n for i in range(0, len(dataset), 1000):\n yield dataset[i : i + 1000][\"text\"]\n\nPython implementation:\nwith open(\"wikitext-2.txt\", \"w\", encoding=\"utf-8\") as f:\n for i in range(len(dataset)):\n f.write(dataset[i][\"text\"] + \"\\n\")", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "348-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Preparing the Training Corpus" + }, + { + "text": "Explanation:\n### Building a WordPiece Tokenizer\n\nWordPiece is the tokenization algorithm used by BERT. It constructs a vocabulary of subword units that balance vocabulary size with the ability to represent rare words.\n\nPython implementation:\nfrom tokenizers import (\n decoders,\n models,\n normalizers,\n pre_tokenizers,\n processors,\n trainers,\n Tokenizer,\n)\n\ntokenizer = Tokenizer(models.WordPiece(unk_token=\"[UNK]\"))\n\nPython implementation:\ntokenizer.normalizer = normalizers.Sequence(\n [normalizers.NFD(), normalizers.Lowercase(), normalizers.StripAccents()]\n)\n\nPython implementation:\npre_tokenizer = pre_tokenizers.WhitespaceSplit()\npre_tokenizer.pre_tokenize_str(\"Let's test my pre-tokenizer.\")\n\nPython implementation:\npre_tokenizer = pre_tokenizers.Sequence(\n [pre_tokenizers.WhitespaceSplit(), pre_tokenizers.Punctuation()]\n)\npre_tokenizer.pre_tokenize_str(\"Let's test my pre-tokenizer.\")\n\nPython implementation:\nspecial_tokens = [\"[UNK]\", \"[PAD]\", \"[CLS]\", \"[SEP]\", \"[MASK]\"]", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "349-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Building a WordPiece Tokenizer" + }, + { + "text": "itespaceSplit(), pre_tokenizers.Punctuation()]\n)\npre_tokenizer.pre_tokenize_str(\"Let's test my pre-tokenizer.\")\n\nPython implementation:\nspecial_tokens = [\"[UNK]\", \"[PAD]\", \"[CLS]\", \"[SEP]\", \"[MASK]\"]\ntrainer = trainers.WordPieceTrainer(vocab_size=25000, special_tokens=special_tokens)\n\nPython implementation:\ntokenizer.model = models.WordPiece(unk_token=\"[UNK]\")\ntokenizer.train([\"wikitext-2.txt\"], trainer=trainer)\n\nPython implementation:\ncls_token_id = tokenizer.token_to_id(\"[CLS]\")\nsep_token_id = tokenizer.token_to_id(\"[SEP]\")\nprint(cls_token_id, sep_token_id)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "349-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Building a WordPiece Tokenizer" + }, + { + "text": "Explanation:\n### Post-Processing\n\nAfter tokenization, special tokens such as `[CLS]` and `[SEP]` are automatically inserted to create model-ready inputs.\n\nPython implementation:\ntokenizer.post_processor = processors.TemplateProcessing(\n single=f\"[CLS]:0 $A:0 [SEP]:0\",\n pair=f\"[CLS]:0 $A:0 [SEP]:0 $B:1 [SEP]:1\",\n special_tokens=[(\"[CLS]\", cls_token_id), (\"[SEP]\", sep_token_id)],\n)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "350-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Post-Processing" + }, + { + "text": "Explanation:\n### Saving and Reloading the Tokenizer\n\nOnce trained, a tokenizer can be saved to disk and reloaded for future training or inference.\n\nPython implementation:\nfrom transformers import PreTrainedTokenizerFast\n\nwrapped_tokenizer = PreTrainedTokenizerFast(\n tokenizer_object=tokenizer,\n # tokenizer_file=\"tokenizer.json\", # You can load from the tokenizer file, alternatively\n unk_token=\"[UNK]\",\n pad_token=\"[PAD]\",\n cls_token=\"[CLS]\",\n sep_token=\"[SEP]\",\n mask_token=\"[MASK]\",\n)\n\nPython implementation:\nfrom transformers import BertTokenizerFast\n\nwrapped_tokenizer = BertTokenizerFast(tokenizer_object=tokenizer)\n\nPython implementation:\ntrainer = trainers.BpeTrainer(vocab_size=25000, special_tokens=[\"<|endoftext|>\"])\ntokenizer.train_from_iterator(get_training_corpus(), trainer=trainer)\n\nPython implementation:\ntokenizer.model = models.BPE()\ntokenizer.train([\"wikitext-2.txt\"], trainer=trainer)\n\nPython implementation:\nsentence = \"Let's test this tokenizer.\"", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "351-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Saving and Reloading the Tokenizer" + }, + { + "text": "_corpus(), trainer=trainer)\n\nPython implementation:\ntokenizer.model = models.BPE()\ntokenizer.train([\"wikitext-2.txt\"], trainer=trainer)\n\nPython implementation:\nsentence = \"Let's test this tokenizer.\"\nencoding = tokenizer.encode(sentence)\nstart, end = encoding.offsets[4]\nsentence[start:end]\n\nPython implementation:\nfrom transformers import PreTrainedTokenizerFast\n\nwrapped_tokenizer = PreTrainedTokenizerFast(\n tokenizer_object=tokenizer,\n bos_token=\"<|endoftext|>\",\n eos_token=\"<|endoftext|>\",\n)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "351-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Saving and Reloading the Tokenizer" + }, + { + "text": "Explanation:\n### Building a Byte-Level BPE Tokenizer\n\nGPT-2 uses Byte-Level Byte Pair Encoding (BPE), which represents text at the byte level before learning subword merges. This approach allows it to tokenize virtually any Unicode text without requiring an unknown token.\n\nPython implementation:\nfrom transformers import GPT2TokenizerFast\n\nwrapped_tokenizer = GPT2TokenizerFast(tokenizer_object=tokenizer)\n\nPython implementation:\nfrom tokenizers import Regex\n\ntokenizer.normalizer = normalizers.Sequence(\n [\n normalizers.Replace(\"``\", '\"'),\n normalizers.Replace(\"''\", '\"'),\n normalizers.NFKD(),\n normalizers.StripAccents(),\n normalizers.Replace(Regex(\" {2,}\"), \" \"),\n ]\n)\n\nPython implementation:\nspecial_tokens = [\"\", \"\", \"\", \"\", \"\", \"\", \"\"]\ntrainer = trainers.UnigramTrainer(\n vocab_size=25000, special_tokens=special_tokens, unk_token=\"\"\n)\ntokenizer.train_from_iterator(get_training_corpus(), trainer=trainer)\n\nPython implementation:", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "352-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Building a Byte-Level BPE Tokenizer" + }, + { + "text": "trainer = trainers.UnigramTrainer(\n vocab_size=25000, special_tokens=special_tokens, unk_token=\"\"\n)\ntokenizer.train_from_iterator(get_training_corpus(), trainer=trainer)\n\nPython implementation:\ntokenizer.model = models.Unigram()\ntokenizer.train([\"wikitext-2.txt\"], trainer=trainer)\n\nPython implementation:\ncls_token_id = tokenizer.token_to_id(\"\")\nsep_token_id = tokenizer.token_to_id(\"\")\nprint(cls_token_id, sep_token_id)\n\nPython implementation:\ntokenizer.post_processor = processors.TemplateProcessing(\n single=\"$A:0 :0 :2\",\n pair=\"$A:0 :0 $B:1 :1 :2\",\n special_tokens=[(\"\", sep_token_id), (\"\", cls_token_id)],\n)\n\nPython implementation:\nfrom transformers import PreTrainedTokenizerFast\n\nwrapped_tokenizer = PreTrainedTokenizerFast(\n tokenizer_object=tokenizer,\n bos_token=\"\",\n eos_token=\"\",\n unk_token=\"\",\n pad_token=\"\",\n cls_token=\"\",\n sep_token=\"\",\n mask_token=\"\",\n padding_side=\"left\",\n)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "352-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Building a Byte-Level BPE Tokenizer" + }, + { + "text": "Explanation:\n### Building a Unigram Tokenizer\n\nXLNet and SentencePiece use the Unigram algorithm, which begins with a large vocabulary and iteratively removes tokens while maximizing the likelihood of the training corpus.\n\nPython implementation:\nfrom transformers import XLNetTokenizerFast\n\nwrapped_tokenizer = XLNetTokenizerFast(tokenizer_object=tokenizer)", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "353-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Building a Unigram Tokenizer" + }, + { + "text": "Explanation:\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Use Hugging Face fast tokenizers.\n- Retrieve token-level metadata and character offsets.\n- Apply tokenization to Question Answering and Named Entity Recognition.\n- Understand normalization and pre-tokenization.\n- Build custom WordPiece, Byte-Level BPE, and Unigram tokenizers.\n- Train tokenizer vocabularies from raw text corpora.", + "source": "Tokenizers.ipynb", + "file_type": "ipynb", + "chunk_id": "354-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "04-tokenization/Tokenizers.ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Token Classification with Transformers\n\nThis notebook demonstrates how to build an end-to-end token classification pipeline using the Hugging Face Transformers library. Token classification assigns a label to each token in a sentence and is commonly used for tasks such as Named Entity Recognition (NER).\n\n## Learning Objectives\n\n- Understand token classification tasks\n- Load and explore NER datasets\n- Tokenize text while preserving label alignment\n- Build data collators for token classification\n- Fine-tune Transformer models\n- Evaluate token classification models\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Evaluate\n- Accelerate\n\nExplanation:\nYou will need to setup git, adapt your email and name in the following cell.\n\nPython implementation:\n!git config --global user.email \"milad.saeedi@mail.utoronto.ca\"\n!git config --global user.name \"Milad Saeedi\"\n\nExplanation:", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "355-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Token Classification with Transformers" + }, + { + "text": "t, adapt your email and name in the following cell.\n\nPython implementation:\n!git config --global user.email \"milad.saeedi@mail.utoronto.ca\"\n!git config --global user.name \"Milad Saeedi\"\n\nExplanation:\nYou will also need to be logged in to the Hugging Face Hub. Execute the following and enter your credentials.", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "355-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Token Classification with Transformers" + }, + { + "text": "Explanation:\n## Loading the Dataset\n\nThe first step is to load a labeled Named Entity Recognition (NER) dataset. Unlike sentence classification tasks, token classification datasets contain a sequence of words together with a label assigned to each individual token.\n\nPython implementation:\nfrom datasets import load_dataset\n\nraw_datasets = load_dataset(\n \"conll2003\",\n trust_remote_code=True\n)\n\nPython implementation:\nraw_datasets[\"train\"] = (\n raw_datasets[\"train\"]\n .shuffle(seed=42)\n .select(range(1000))\n)\n\nraw_datasets[\"validation\"] = (\n raw_datasets[\"validation\"]\n .shuffle(seed=42)\n .select(range(200))\n)\n\nraw_datasets[\"test\"] = (\n raw_datasets[\"test\"]\n .shuffle(seed=42)\n .select(range(200))\n)", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "356-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Loading the Dataset" + }, + { + "text": "Explanation:\n### Exploring the Dataset\n\nBefore training a model, it is useful to inspect the dataset structure. In this section, we examine the input tokens and their corresponding NER labels to better understand how the data is organized.\n\nPython implementation:\nner_feature = raw_datasets[\"train\"].features[\"ner_tags\"]\nner_feature\n\nPython implementation:\nlabel_names = ner_feature.feature.names\nlabel_names\n\nPython implementation:\nwords = raw_datasets[\"train\"][0][\"tokens\"]\nlabels = raw_datasets[\"train\"][0][\"ner_tags\"]\nline1 = \"\"\nline2 = \"\"\nfor word, label in zip(words, labels):\n full_label = label_names[label]\n max_length = max(len(word), len(full_label))\n line1 += word + \" \" * (max_length - len(word) + 1)\n line2 += full_label + \" \" * (max_length - len(full_label) + 1)\n\nprint(line1)\nprint(line2)", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "357-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Exploring the Dataset" + }, + { + "text": "Explanation:\n## Tokenizing the Dataset\n\nTransformer models operate on subword tokens rather than complete words. This section demonstrates how to tokenize the input text while preserving the relationship between the original words and their corresponding labels.\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\nmodel_checkpoint = \"bert-base-cased\"\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint)\n\nPython implementation:\ninputs = tokenizer(raw_datasets[\"train\"][0][\"tokens\"], is_split_into_words=True)\ninputs.tokens()", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "358-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Tokenizing the Dataset" + }, + { + "text": "Explanation:\n### Aligning Labels with Subword Tokens\n\nA single word may be split into multiple subword tokens during tokenization. Since the original dataset provides one label per word, we must align these labels with the generated subword tokens before training.\n\nPython implementation:\ndef align_labels_with_tokens(labels, word_ids):\n new_labels = []\n current_word = None\n for word_id in word_ids:\n if word_id != current_word:\n # Start of a new word!\n current_word = word_id\n label = -100 if word_id is None else labels[word_id]\n new_labels.append(label)\n elif word_id is None:\n # Special token\n new_labels.append(-100)\n else:\n # Same word as previous token\n label = labels[word_id]\n # If the label is B-XXX we change it to I-XXX\n if label % 2 == 1:\n label += 1\n new_labels.append(label)\n\n return new_labels", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "359-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Aligning Labels with Subword Tokens" + }, + { + "text": "Explanation:\n### Preprocessing the Dataset\n\nAfter defining the alignment procedure, we apply it to every example in the dataset. This produces tokenized inputs together with correctly aligned labels that are ready for model training.\n\nPython implementation:\nlabels = raw_datasets[\"train\"][0][\"ner_tags\"]\nword_ids = inputs.word_ids()\nprint(labels)\nprint(align_labels_with_tokens(labels, word_ids))\n\nPython implementation:\ndef tokenize_and_align_labels(examples):\n tokenized_inputs = tokenizer(\n examples[\"tokens\"], truncation=True, is_split_into_words=True\n )\n all_labels = examples[\"ner_tags\"]\n new_labels = []\n for i, labels in enumerate(all_labels):\n word_ids = tokenized_inputs.word_ids(i)\n new_labels.append(align_labels_with_tokens(labels, word_ids))\n\n tokenized_inputs[\"labels\"] = new_labels\n return tokenized_inputs\n\nPython implementation:\ntokenized_datasets = raw_datasets.map(\n tokenize_and_align_labels,\n batched=True,\n remove_columns=raw_datasets[\"train\"].column_names,\n)", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "360-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Preprocessing the Dataset" + }, + { + "text": "Explanation:\n## Dynamic Padding\n\nTraining batches often contain sequences of different lengths. The data collator dynamically pads each batch while ensuring that token labels remain correctly aligned with the input tokens.\n\nPython implementation:\nfrom transformers import DataCollatorForTokenClassification\n\ndata_collator = DataCollatorForTokenClassification(tokenizer=tokenizer)\n\nPython implementation:\nbatch = data_collator([tokenized_datasets[\"train\"][i] for i in range(2)])\nbatch[\"labels\"]", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "361-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Dynamic Padding" + }, + { + "text": "Explanation:\n## Evaluation Metrics\n\nToken classification models are commonly evaluated using precision, recall, and F1-score. In this notebook, we use the `seqeval` library to compute these metrics for Named Entity Recognition.\n\nPython implementation:\nimport evaluate\n\nmetric = evaluate.load(\"seqeval\")\n\nPython implementation:\nlabels = raw_datasets[\"train\"][0][\"ner_tags\"]\nlabels = [label_names[i] for i in labels]\nlabels\n\nPython implementation:\npredictions = labels.copy()\npredictions[2] = \"O\"\nmetric.compute(predictions=[predictions], references=[labels])\n\nPython implementation:\nimport numpy as np\n\ndef compute_metrics(eval_preds):\n logits, labels = eval_preds\n predictions = np.argmax(logits, axis=-1)\n\n # Remove ignored index (special tokens) and convert to labels\n true_labels = [[label_names[l] for l in label if l != -100] for label in labels]\n true_predictions = [\n [label_names[p] for (p, l) in zip(prediction, label) if l != -100]\n for prediction, label in zip(predictions, labels)\n ]", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "362-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Evaluation Metrics" + }, + { + "text": "es[l] for l in label if l != -100] for label in labels]\n true_predictions = [\n [label_names[p] for (p, l) in zip(prediction, label) if l != -100]\n for prediction, label in zip(predictions, labels)\n ]\n all_metrics = metric.compute(predictions=true_predictions, references=true_labels)\n return {\n \"precision\": all_metrics[\"overall_precision\"],\n \"recall\": all_metrics[\"overall_recall\"],\n \"f1\": all_metrics[\"overall_f1\"],\n \"accuracy\": all_metrics[\"overall_accuracy\"],\n }\n\nPython implementation:\nid2label = {i: label for i, label in enumerate(label_names)}\nlabel2id = {v: k for k, v in id2label.items()}\n\nPython implementation:\nfrom transformers import AutoModelForTokenClassification\n\nmodel = AutoModelForTokenClassification.from_pretrained(\n model_checkpoint,\n id2label=id2label,\n label2id=label2id,\n)", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "362-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Evaluation Metrics" + }, + { + "text": "Explanation:\n## Configuring Model Training\n\nThe `TrainingArguments` class defines the training configuration, including learning rate, batch size, evaluation strategy, checkpointing, and logging behavior.\n\nPython implementation:\nfrom transformers import TrainingArguments\n\nargs = TrainingArguments(\n \"bert-finetuned-ner-tokenclass\",\n eval_strategy=\"epoch\",\n save_strategy=\"epoch\",\n learning_rate=2e-5,\n num_train_epochs=3,\n weight_decay=0.01,\n push_to_hub=True,\n)", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "363-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Configuring Model Training" + }, + { + "text": "Explanation:\n### Fine-Tuning with the Trainer API\n\nThe Hugging Face `Trainer` API manages the complete training workflow, including optimization, evaluation, checkpointing, and metric computation.\n\nPython implementation:\nfrom transformers import Trainer\n\ntrainer = Trainer(\n model=model,\n args=args,\n train_dataset=tokenized_datasets[\"train\"],\n eval_dataset=tokenized_datasets[\"validation\"],\n data_collator=data_collator,\n compute_metrics=compute_metrics,\n tokenizer=tokenizer,\n)\ntrainer.train()", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "364-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Fine-Tuning with the Trainer API" + }, + { + "text": "Explanation:\n## Manual Training Loop with Accelerate\n\nIn addition to the high-level `Trainer` API, Hugging Face provides the `Accelerate` library for implementing custom training loops. This approach offers greater flexibility while simplifying multi-device training.\n\nPython implementation:\nfrom torch.utils.data import DataLoader\n\ntrain_dataloader = DataLoader(\n tokenized_datasets[\"train\"],\n shuffle=True,\n collate_fn=data_collator,\n batch_size=8,\n)\neval_dataloader = DataLoader(\n tokenized_datasets[\"validation\"], collate_fn=data_collator, batch_size=8\n)\n\nPython implementation:\nmodel = AutoModelForTokenClassification.from_pretrained(\n model_checkpoint,\n id2label=id2label,\n label2id=label2id,\n)\n\nPython implementation:\nfrom torch.optim import AdamW\n\noptimizer = AdamW(model.parameters(), lr=2e-5)\n\nPython implementation:\nfrom accelerate import Accelerator\n\naccelerator = Accelerator()\nmodel, optimizer, train_dataloader, eval_dataloader = accelerator.prepare(", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "365-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Manual Training Loop with Accelerate" + }, + { + "text": "= AdamW(model.parameters(), lr=2e-5)\n\nPython implementation:\nfrom accelerate import Accelerator\n\naccelerator = Accelerator()\nmodel, optimizer, train_dataloader, eval_dataloader = accelerator.prepare(\n model, optimizer, train_dataloader, eval_dataloader\n)\n\nPython implementation:\nfrom transformers import get_scheduler\n\nnum_train_epochs = 3\nnum_update_steps_per_epoch = len(train_dataloader)\nnum_training_steps = num_train_epochs * num_update_steps_per_epoch\n\nlr_scheduler = get_scheduler(\n \"linear\",\n optimizer=optimizer,\n num_warmup_steps=0,\n num_training_steps=num_training_steps,\n)\n\nPython implementation:\nfrom huggingface_hub import Repository, get_full_repo_name\n\nmodel_name = \"bert-finetuned-ner-accelerate\"\nrepo_name = get_full_repo_name(model_name)\nrepo_name\n\nPython implementation:\noutput_dir = \"bert-finetuned-ner-accelerate\"\nrepo = Repository(output_dir, clone_from=repo_name)\n\nPython implementation:\ndef postprocess(predictions, labels):", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "365-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Manual Training Loop with Accelerate" + }, + { + "text": "_name)\nrepo_name\n\nPython implementation:\noutput_dir = \"bert-finetuned-ner-accelerate\"\nrepo = Repository(output_dir, clone_from=repo_name)\n\nPython implementation:\ndef postprocess(predictions, labels):\n predictions = predictions.detach().cpu().clone().numpy()\n labels = labels.detach().cpu().clone().numpy()\n\n # Remove ignored index (special tokens) and convert to labels\n true_labels = [[label_names[l] for l in label if l != -100] for label in labels]\n true_predictions = [\n [label_names[p] for (p, l) in zip(prediction, label) if l != -100]\n for prediction, label in zip(predictions, labels)\n ]\n return true_labels, true_predictions\n\nPython implementation:\nfrom tqdm.auto import tqdm\nimport torch\n\nprogress_bar = tqdm(range(num_training_steps))\n\nfor epoch in range(num_train_epochs):\n # Training\n model.train()\n for batch in train_dataloader:\n outputs = model(**batch)\n loss = outputs.loss\n accelerator.backward(loss)\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "365-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Manual Training Loop with Accelerate" + }, + { + "text": "):\n # Training\n model.train()\n for batch in train_dataloader:\n outputs = model(**batch)\n loss = outputs.loss\n accelerator.backward(loss)\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()\n progress_bar.update(1)\n\n # Evaluation\n model.eval()\n for batch in eval_dataloader:\n with torch.no_grad():\n outputs = model(**batch)\n\n predictions = outputs.logits.argmax(dim=-1)\n labels = batch[\"labels\"]\n\n # Necessary to pad predictions and labels for being gathered\n predictions = accelerator.pad_across_processes(predictions, dim=1, pad_index=-100)\n labels = accelerator.pad_across_processes(labels, dim=1, pad_index=-100)\n\n predictions_gathered = accelerator.gather(predictions)\n labels_gathered = accelerator.gather(labels)\n\n true_predictions, true_labels = postprocess(predictions_gathered, labels_gathered)\n metric.add_batch(predictions=true_predictions, references=true_labels)\n\n results = metric.compute()\n print(\n f\"epoch {epoch}:\",\n {\n key: results[f\"overall_{key}\"]", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "365-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Manual Training Loop with Accelerate" + }, + { + "text": "dictions_gathered, labels_gathered)\n metric.add_batch(predictions=true_predictions, references=true_labels)\n\n results = metric.compute()\n print(\n f\"epoch {epoch}:\",\n {\n key: results[f\"overall_{key}\"]\n for key in [\"precision\", \"recall\", \"f1\", \"accuracy\"]\n },\n )\n\nPython implementation:\naccelerator.wait_for_everyone()\nunwrapped_model = accelerator.unwrap_model(model)\nunwrapped_model.save_pretrained(output_dir, save_function=accelerator.save)\n\nPython implementation:\nfrom transformers import pipeline\n\n# Replace this with your own checkpoint\nmodel_checkpoint = \"huggingface-course/bert-finetuned-ner\"\ntoken_classifier = pipeline(\n \"token-classification\", model=model_checkpoint, aggregation_strategy=\"simple\"\n)\ntoken_classifier(\"My name is Sylvain and I work at Hugging Face in Brooklyn.\")", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "365-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Manual Training Loop with Accelerate" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nIn this notebook:\n\n- Load and explore Named Entity Recognition datasets.\n- Tokenize text while preserving word-label alignment.\n- Handle subword tokenization for token classification.\n- Dynamically pad batches for efficient training.\n- Fine-tune pretrained Transformer models for token classification.\n- Evaluate NER models using precision, recall, and F1-score.\n- Implement both high-level (`Trainer`) and custom (`Accelerate`) training workflows.", + "source": "5.1.Token_classification_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "366-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.1.Token_classification_(PyTorch).ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Fine-Tuning a Masked Language Model\n\nThis notebook demonstrates how to fine-tune a pretrained masked language model using the Hugging Face Transformers library. It covers masked token prediction, dataset preparation, dynamic masking, supervised fine-tuning, and custom training with the Accelerate library.\n\n## Learning Objectives\n\n- Understand masked language modeling (MLM)\n- Prepare text datasets for MLM training\n- Apply dynamic token masking\n- Fine-tune pretrained language models\n- Evaluate language models\n- Train using both the Trainer API and Accelerate\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Evaluate\n- Accelerate\n\nExplanation:\nYou will need to setup git, adapt your email and name in the following cell.\n\nPython implementation:\n!git config --global user.email \"milad.saeedi@mail.utoronto.ca\"\n!git config --global user.name \"Milad Saeedi\"\n\nExplanation:", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "367-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Fine-Tuning a Masked Language Model" + }, + { + "text": "t, adapt your email and name in the following cell.\n\nPython implementation:\n!git config --global user.email \"milad.saeedi@mail.utoronto.ca\"\n!git config --global user.name \"Milad Saeedi\"\n\nExplanation:\nYou will also need to be logged in to the Hugging Face Hub. Execute the following and enter your credentials.\n\nPython implementation:\nfrom huggingface_hub import notebook_login\n\nnotebook_login()", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "367-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Fine-Tuning a Masked Language Model" + }, + { + "text": "Explanation:\n## Understanding Masked Language Models\n\nMasked Language Models (MLMs) are trained to predict intentionally hidden words within a sentence. During pretraining, selected tokens are replaced with the special `[MASK]` token, and the model learns to recover the original words using the surrounding context.\n\nIn this section, we load a pretrained DistilBERT model, perform masked token prediction, and explore how tokenization converts raw text into numerical inputs suitable for Transformer models.\n\nPython implementation:\nfrom transformers import AutoModelForMaskedLM\n\nmodel_checkpoint = \"distilbert-base-uncased\"\nmodel = AutoModelForMaskedLM.from_pretrained(model_checkpoint)\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint)\n\nPython implementation:\nimport torch\n\ninputs = tokenizer(text, return_tensors=\"pt\")\ntoken_logits = model(**inputs).logits\n# Find the location of [MASK] and extract its logits", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "368-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Understanding Masked Language Models" + }, + { + "text": "rained(model_checkpoint)\n\nPython implementation:\nimport torch\n\ninputs = tokenizer(text, return_tensors=\"pt\")\ntoken_logits = model(**inputs).logits\n# Find the location of [MASK] and extract its logits\nmask_token_index = torch.where(inputs[\"input_ids\"] == tokenizer.mask_token_id)[1]\nmask_token_logits = token_logits[0, mask_token_index, :]\n# Pick the [MASK] candidates with the highest logits\ntop_5_tokens = torch.topk(mask_token_logits, 5, dim=1).indices[0].tolist()\n\nfor token in top_5_tokens:\n print(f\"'>>> {text.replace(tokenizer.mask_token, tokenizer.decode([token]))}'\")", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "368-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Understanding Masked Language Models" + }, + { + "text": "Explanation:\n## Preparing the Dataset\n\nBefore fine-tuning, raw text must be transformed into fixed-length training examples. This section demonstrates how to load a text dataset, tokenize the input, concatenate multiple documents, and divide them into uniform sequences that can be processed efficiently during training.\n\nPython implementation:\nfrom datasets import load_dataset\n\nimdb_dataset = load_dataset(\"imdb\")\nimdb_dataset\n\nPython implementation:\nsample = imdb_dataset[\"train\"].shuffle(seed=42).select(range(3))\n\nfor row in sample:\n print(f\"\\n'>>> Review: {row['text']}'\")\n print(f\"'>>> Label: {row['label']}'\")\n\nPython implementation:\ndef tokenize_function(examples):\n result = tokenizer(examples[\"text\"])\n if tokenizer.is_fast:\n result[\"word_ids\"] = [result.word_ids(i) for i in range(len(result[\"input_ids\"]))]\n return result\n\n# Use batched=True to activate fast multithreading!\ntokenized_datasets = imdb_dataset.map(\n tokenize_function, batched=True, remove_columns=[\"text\", \"label\"]\n)", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "369-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Preparing the Dataset" + }, + { + "text": "(result[\"input_ids\"]))]\n return result\n\n# Use batched=True to activate fast multithreading!\ntokenized_datasets = imdb_dataset.map(\n tokenize_function, batched=True, remove_columns=[\"text\", \"label\"]\n)\ntokenized_datasets\n\nPython implementation:\n# Slicing produces a list of lists for each feature\ntokenized_samples = tokenized_datasets[\"train\"][:3]\n\nfor idx, sample in enumerate(tokenized_samples[\"input_ids\"]):\n print(f\"'>>> Review {idx} length: {len(sample)}'\")\n\nPython implementation:\nconcatenated_examples = {\n k: sum(tokenized_samples[k], []) for k in tokenized_samples.keys()\n}\ntotal_length = len(concatenated_examples[\"input_ids\"])\nprint(f\"'>>> Concatenated reviews length: {total_length}'\")\n\nPython implementation:\nchunks = {\n k: [t[i : i + chunk_size] for i in range(0, total_length, chunk_size)]\n for k, t in concatenated_examples.items()\n}\n\nfor chunk in chunks[\"input_ids\"]:\n print(f\"'>>> Chunk length: {len(chunk)}'\")\n\nPython implementation:\ndef group_texts(examples):", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "369-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Preparing the Dataset" + }, + { + "text": "tal_length, chunk_size)]\n for k, t in concatenated_examples.items()\n}\n\nfor chunk in chunks[\"input_ids\"]:\n print(f\"'>>> Chunk length: {len(chunk)}'\")\n\nPython implementation:\ndef group_texts(examples):\n # Concatenate all texts\n concatenated_examples = {k: sum(examples[k], []) for k in examples.keys()}\n # Compute length of concatenated texts\n total_length = len(concatenated_examples[list(examples.keys())[0]])\n # We drop the last chunk if it's smaller than chunk_size\n total_length = (total_length // chunk_size) * chunk_size\n # Split by chunks of max_len\n result = {\n k: [t[i : i + chunk_size] for i in range(0, total_length, chunk_size)]\n for k, t in concatenated_examples.items()\n }\n # Create a new labels column\n result[\"labels\"] = result[\"input_ids\"].copy()\n return result\n\nPython implementation:\nlm_datasets = tokenized_datasets.map(group_texts, batched=True)\nlm_datasets\n\nPython implementation:\nfrom transformers import DataCollatorForLanguageModeling", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "369-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Preparing the Dataset" + }, + { + "text": "()\n return result\n\nPython implementation:\nlm_datasets = tokenized_datasets.map(group_texts, batched=True)\nlm_datasets\n\nPython implementation:\nfrom transformers import DataCollatorForLanguageModeling\n\ndata_collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm_probability=0.15)\n\nPython implementation:\nsamples = [lm_datasets[\"train\"][i] for i in range(2)]\nfor sample in samples:\n _ = sample.pop(\"word_ids\")\n\nfor chunk in data_collator(samples)[\"input_ids\"]:\n print(f\"\\n'>>> {tokenizer.decode(chunk)}'\")", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "369-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Preparing the Dataset" + }, + { + "text": "Explanation:\n## Dynamic Token Masking\n\nRather than masking the same words every time, masked language models typically apply dynamic masking during training. At each training step, different tokens are randomly selected and replaced with the `[MASK]` token, exposing the model to a wider variety of prediction tasks and improving generalization.\n\nPython implementation:\nimport collections\nimport numpy as np\n\nfrom transformers import default_data_collator\n\nwwm_probability = 0.2\n\ndef whole_word_masking_data_collator(features):\n for feature in features:\n word_ids = feature.pop(\"word_ids\")\n\n # Create a map between words and corresponding token indices\n mapping = collections.defaultdict(list)\n current_word_index = -1\n current_word = None\n for idx, word_id in enumerate(word_ids):\n if word_id is not None:\n if word_id != current_word:\n current_word = word_id\n current_word_index += 1\n mapping[current_word_index].append(idx)\n\n # Randomly mask words", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "370-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Dynamic Token Masking" + }, + { + "text": "word_id in enumerate(word_ids):\n if word_id is not None:\n if word_id != current_word:\n current_word = word_id\n current_word_index += 1\n mapping[current_word_index].append(idx)\n\n # Randomly mask words\n mask = np.random.binomial(1, wwm_probability, (len(mapping),))\n input_ids = feature[\"input_ids\"]\n labels = feature[\"labels\"]\n new_labels = [-100] * len(labels)\n for word_id in np.where(mask)[0]:\n word_id = word_id.item()\n for idx in mapping[word_id]:\n new_labels[idx] = labels[idx]\n input_ids[idx] = tokenizer.mask_token_id\n feature[\"labels\"] = new_labels\n\n return default_data_collator(features)\n\nPython implementation:\nsamples = [lm_datasets[\"train\"][i] for i in range(2)]\nbatch = whole_word_masking_data_collator(samples)\n\nfor chunk in batch[\"input_ids\"]:\n print(f\"\\n'>>> {tokenizer.decode(chunk)}'\")\n\nPython implementation:\ntrain_size = 5_000\ntest_size = int(0.1 * train_size)\n\ndownsampled_dataset = lm_datasets[\"train\"].train_test_split(\n train_size=train_size, test_size=test_size, seed=42\n)", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "370-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Dynamic Token Masking" + }, + { + "text": ")}'\")\n\nPython implementation:\ntrain_size = 5_000\ntest_size = int(0.1 * train_size)\n\ndownsampled_dataset = lm_datasets[\"train\"].train_test_split(\n train_size=train_size, test_size=test_size, seed=42\n)\ndownsampled_dataset", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "370-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Dynamic Token Masking" + }, + { + "text": "Explanation:\n## Fine-Tuning with the Trainer API\n\nThe Hugging Face `Trainer` API simplifies supervised fine-tuning by managing optimization, evaluation, checkpointing, and logging. After training, the model is evaluated using perplexity, a common metric for language modeling, and can be shared through the Hugging Face Hub.\n\nPython implementation:\nfrom transformers import TrainingArguments\n\nbatch_size = 64\n# Show the training loss with every epoch\nlogging_steps = len(downsampled_dataset[\"train\"]) // batch_size\nmodel_name = model_checkpoint.split(\"/\")[-1]\n\ntraining_args = TrainingArguments(\n output_dir=f\"{model_name}-finetuned-imdb\",\n overwrite_output_dir=True,\n eval_strategy=\"epoch\",\n learning_rate=2e-5,\n weight_decay=0.01,\n per_device_train_batch_size=batch_size,\n per_device_eval_batch_size=batch_size,\n push_to_hub=True,\n fp16=True,\n logging_steps=logging_steps,\n)\n\nPython implementation:\nfrom transformers import Trainer\n\ntrainer = Trainer(\n model=model,\n args=training_args,", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "371-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Fine-Tuning with the Trainer API" + }, + { + "text": "l_batch_size=batch_size,\n push_to_hub=True,\n fp16=True,\n logging_steps=logging_steps,\n)\n\nPython implementation:\nfrom transformers import Trainer\n\ntrainer = Trainer(\n model=model,\n args=training_args,\n train_dataset=downsampled_dataset[\"train\"],\n eval_dataset=downsampled_dataset[\"test\"],\n data_collator=data_collator,\n tokenizer=tokenizer,\n)\n\nPython implementation:\nimport math\n\neval_results = trainer.evaluate()\nprint(f\">>> Perplexity: {math.exp(eval_results['eval_loss']):.2f}\")", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "371-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Fine-Tuning with the Trainer API" + }, + { + "text": "Explanation:\n## Custom Training with Accelerate\n\nFor greater flexibility, Hugging Face provides the `Accelerate` library, which enables custom training loops while handling device placement, mixed precision, and distributed training. In this section, we build the training loop manually, optimize the model, upload the fine-tuned model to the Hugging Face Hub, and evaluate it on masked token prediction.\n\nPython implementation:\ndef insert_random_mask(batch):\n features = [dict(zip(batch, t)) for t in zip(*batch.values())]\n masked_inputs = data_collator(features)\n # Create a new \"masked\" column for each column in the dataset\n return {\"masked_\" + k: v.numpy() for k, v in masked_inputs.items()}\n\nPython implementation:\ndownsampled_dataset = downsampled_dataset.remove_columns([\"word_ids\"])\neval_dataset = downsampled_dataset[\"test\"].map(\n insert_random_mask,\n batched=True,\n remove_columns=downsampled_dataset[\"test\"].column_names,\n)\neval_dataset = eval_dataset.rename_columns(\n {", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "372-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "d_ids\"])\neval_dataset = downsampled_dataset[\"test\"].map(\n insert_random_mask,\n batched=True,\n remove_columns=downsampled_dataset[\"test\"].column_names,\n)\neval_dataset = eval_dataset.rename_columns(\n {\n \"masked_input_ids\": \"input_ids\",\n \"masked_attention_mask\": \"attention_mask\",\n \"masked_labels\": \"labels\",\n }\n)\n\nPython implementation:\nfrom torch.utils.data import DataLoader\nfrom transformers import default_data_collator\n\nbatch_size = 64\ntrain_dataloader = DataLoader(\n downsampled_dataset[\"train\"],\n shuffle=True,\n batch_size=batch_size,\n collate_fn=data_collator,\n)\neval_dataloader = DataLoader(\n eval_dataset, batch_size=batch_size, collate_fn=default_data_collator\n)\n\nPython implementation:\nfrom torch.optim import AdamW\n\noptimizer = AdamW(model.parameters(), lr=5e-5)\n\nPython implementation:\nfrom accelerate import Accelerator\n\naccelerator = Accelerator()\nmodel, optimizer, train_dataloader, eval_dataloader = accelerator.prepare(\n model, optimizer, train_dataloader, eval_dataloader\n)", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "372-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "ion:\nfrom accelerate import Accelerator\n\naccelerator = Accelerator()\nmodel, optimizer, train_dataloader, eval_dataloader = accelerator.prepare(\n model, optimizer, train_dataloader, eval_dataloader\n)\n\nPython implementation:\nfrom transformers import get_scheduler\n\nnum_train_epochs = 3\nnum_update_steps_per_epoch = len(train_dataloader)\nnum_training_steps = num_train_epochs * num_update_steps_per_epoch\n\nlr_scheduler = get_scheduler(\n \"linear\",\n optimizer=optimizer,\n num_warmup_steps=0,\n num_training_steps=num_training_steps,\n)\n\nPython implementation:\nfrom huggingface_hub import get_full_repo_name\n\nmodel_name = \"distilbert-base-uncased-finetuned-imdb \"\nrepo_name = get_full_repo_name(model_name)\nrepo_name\n\nPython implementation:\nfrom huggingface_hub import Repository\n\noutput_dir = model_name\nrepo = Repository(output_dir, clone_from=repo_name)\n\nPython implementation:\nfrom tqdm.auto import tqdm\nimport torch\nimport math\n\nprogress_bar = tqdm(range(num_training_steps))", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "372-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "utput_dir = model_name\nrepo = Repository(output_dir, clone_from=repo_name)\n\nPython implementation:\nfrom tqdm.auto import tqdm\nimport torch\nimport math\n\nprogress_bar = tqdm(range(num_training_steps))\n\nfor epoch in range(num_train_epochs):\n # Training\n model.train()\n for batch in train_dataloader:\n outputs = model(**batch)\n loss = outputs.loss\n accelerator.backward(loss)\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()\n progress_bar.update(1)\n\n # Evaluation\n model.eval()\n losses = []\n for step, batch in enumerate(eval_dataloader):\n with torch.no_grad():\n outputs = model(**batch)\n\n loss = outputs.loss\n losses.append(accelerator.gather(loss.repeat(batch_size)))\n\n losses = torch.cat(losses)\n losses = losses[: len(eval_dataset)]\n try:\n perplexity = math.exp(torch.mean(losses))\n except OverflowError:\n perplexity = float(\"inf\")\n\n print(f\">>> Epoch {epoch}: Perplexity: {perplexity}\")\n\n # Save and upload\n accelerator.wait_for_everyone()", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "372-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "erplexity = math.exp(torch.mean(losses))\n except OverflowError:\n perplexity = float(\"inf\")\n\n print(f\">>> Epoch {epoch}: Perplexity: {perplexity}\")\n\n # Save and upload\n accelerator.wait_for_everyone()\n unwrapped_model = accelerator.unwrap_model(model)\n unwrapped_model.save_pretrained(output_dir, save_function=accelerator.save)\n if accelerator.is_main_process:\n tokenizer.save_pretrained(output_dir)\n repo.push_to_hub(\n commit_message=f\"Training in progress epoch {epoch}\", blocking=False\n )\n\nPython implementation:\nfrom transformers import pipeline\n\nmask_filler = pipeline(\n \"fill-mask\", model=\"huggingface-course/distilbert-base-uncased-finetuned-imdb\"\n)\n\nPython implementation:\npreds = mask_filler(text)\n\nfor pred in preds:\n print(f\">>> {pred['sequence']}\")", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "372-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\n- Perform masked token prediction using pretrained Transformer models.\n- Prepare text datasets for masked language modeling.\n- Apply dynamic masking during training.\n- Fine-tune pretrained language models using the Hugging Face `Trainer` API.\n- Implement custom training loops with the `Accelerate` library.\n- Evaluate masked language models using perplexity and inference examples.", + "source": "5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "373-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.2.Fine_tuning_a_masked_language_model_(PyTorch).ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Neural Machine Translation with Transformers\n\nThis notebook demonstrates how to build an end-to-end neural machine translation pipeline using the Hugging Face Transformers library. It covers dataset preparation, tokenization, preprocessing, evaluation, and fine-tuning of a pretrained sequence-to-sequence Transformer model.\n\n## Learning Objectives\n\n- Load and explore a machine translation dataset\n- Understand sequence-to-sequence (Seq2Seq) models\n- Tokenize source and target languages\n- Prepare translation datasets for training\n- Fine-tune pretrained translation models\n- Evaluate translation quality using BLEU\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Evaluate\n- Hugging Face Hub\n\nPython implementation:\n!git config --global user.email \"milad.saeedi@mail.utoronto.ca\"\n!git config --global user.name \"Milad Saeedi\"", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "374-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Neural Machine Translation with Transformers" + }, + { + "text": "Explanation:\n## Loading the Translation Dataset\n\nThe first step is to load a bilingual dataset containing aligned sentence pairs. Each example consists of a source sentence and its corresponding translation in the target language.\n\nPython implementation:\nfrom datasets import load_dataset\n\nraw_datasets = load_dataset(\"kde4\", lang1=\"en\", lang2=\"fr\")", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "375-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Loading the Translation Dataset" + }, + { + "text": "Explanation:\n### Exploring the Dataset\n\nBefore preprocessing, we inspect the dataset structure to understand how the source and target translations are organized.\n\nPython implementation:\nfrom datasets import DatasetDict\nraw_datasets = DatasetDict({\n \"train\": raw_datasets[\"train\"]\n .shuffle(seed=42)\n .select(range(10000))\n})\n\nraw_datasets", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "376-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Exploring the Dataset" + }, + { + "text": "Explanation:\n## Preparing the Dataset\n\nThe dataset is divided into training and validation subsets. The training set is used to optimize the model parameters, while the validation set measures how well the model generalizes to unseen translation examples.\n\nPython implementation:\nsplit_datasets = raw_datasets[\"train\"].train_test_split(train_size=0.9, seed=20)\nsplit_datasets\n\nPython implementation:\nfrom transformers import pipeline\n\nmodel_checkpoint = \"Helsinki-NLP/opus-mt-en-fr\"\ntranslator = pipeline(\"translation\", model=model_checkpoint)\ntranslator(\"Default to expanded threads\")\n\nPython implementation:\ntranslator(\n \"Unable to import %1 using the OFX importer plugin. This file is not the correct format.\"\n)", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "377-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Preparing the Dataset" + }, + { + "text": "Explanation:\n## Tokenizing Source and Target Sentences\n\nSequence-to-sequence models require both the source language and the target language to be tokenized. This section demonstrates how the tokenizer prepares each language for training.\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\nmodel_checkpoint = \"Helsinki-NLP/opus-mt-en-fr\"\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint, return_tensors=\"pt\")\n\nPython implementation:\nen_sentence = split_datasets[\"train\"][1][\"translation\"][\"en\"]\nfr_sentence = split_datasets[\"train\"][1][\"translation\"][\"fr\"]\n\ninputs = tokenizer(en_sentence, text_target=fr_sentence)\ninputs", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "378-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Tokenizing Source and Target Sentences" + }, + { + "text": "Explanation:\n### Tokenizing the Target Language\n\nUnlike text classification tasks, translation requires tokenizing both the input sentence and its expected translation. The target tokens serve as labels during supervised training.\n\nPython implementation:\nmax_length = 128\n\ndef preprocess_function(examples):\n inputs = [ex[\"en\"] for ex in examples[\"translation\"]]\n targets = [ex[\"fr\"] for ex in examples[\"translation\"]]\n model_inputs = tokenizer(\n inputs, text_target=targets, max_length=max_length, truncation=True\n )\n return model_inputs", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "379-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Tokenizing the Target Language" + }, + { + "text": "Explanation:\n### Preprocessing the Dataset\n\nThe tokenizer is applied to every sentence pair, producing numerical representations that can be used for sequence-to-sequence model training.\n\nPython implementation:\ntokenized_datasets = split_datasets.map(\n preprocess_function,\n batched=True,\n remove_columns=split_datasets[\"train\"].column_names,\n)", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "380-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Preprocessing the Dataset" + }, + { + "text": "Explanation:\n## Loading a Sequence-to-Sequence Model\n\nMachine translation models use an encoder-decoder architecture. The encoder processes the source sentence, while the decoder generates the translated sentence one token at a time.\n\nPython implementation:\nfrom transformers import AutoModelForSeq2SeqLM\n\nmodel = AutoModelForSeq2SeqLM.from_pretrained(model_checkpoint)", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "381-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Loading a Sequence-to-Sequence Model" + }, + { + "text": "Explanation:\n### Dynamic Padding for Seq2Seq Models\n\nTranslation datasets contain sentences with varying lengths. The data collator dynamically pads each batch while ensuring that decoder inputs and labels remain correctly aligned.\n\nPython implementation:\nfrom transformers import DataCollatorForSeq2Seq\n\ndata_collator = DataCollatorForSeq2Seq(tokenizer, model=model)\n\nPython implementation:\nbatch = data_collator([tokenized_datasets[\"train\"][i] for i in range(1, 3)])\nbatch.keys()", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "382-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Dynamic Padding for Seq2Seq Models" + }, + { + "text": "Explanation:\n## Evaluating Translation Quality\n\nMachine translation models are commonly evaluated using the BLEU metric, which measures the similarity between generated translations and reference translations.\n\nIn this notebook, we use the `sacrebleu` implementation provided through the Hugging Face `evaluate` library.\n\nPython implementation:\nimport evaluate\n\nmetric = evaluate.load(\"sacrebleu\")\n\nPython implementation:\npredictions = [\n \"This plugin lets you translate web pages between several languages automatically.\"\n]\nreferences = [\n [\n \"This plugin allows you to automatically translate web pages between several languages.\"\n ]\n]\nmetric.compute(predictions=predictions, references=references)\n\nPython implementation:\npredictions = [\"This This This This\"]\nreferences = [\n [\n \"This plugin allows you to automatically translate web pages between several languages.\"\n ]\n]\nmetric.compute(predictions=predictions, references=references)\n\nPython implementation:\npredictions = [\"This plugin\"]", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "383-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Evaluating Translation Quality" + }, + { + "text": "n allows you to automatically translate web pages between several languages.\"\n ]\n]\nmetric.compute(predictions=predictions, references=references)\n\nPython implementation:\npredictions = [\"This plugin\"]\nreferences = [\n [\n \"This plugin allows you to automatically translate web pages between several languages.\"\n ]\n]\nmetric.compute(predictions=predictions, references=references)\n\nPython implementation:\nimport numpy as np\n\ndef compute_metrics(eval_preds):\n preds, labels = eval_preds\n # In case the model returns more than the prediction logits\n if isinstance(preds, tuple):\n preds = preds[0]\n\n decoded_preds = tokenizer.batch_decode(preds, skip_special_tokens=True)\n\n # Replace -100s in the labels as we can't decode them\n labels = np.where(labels != -100, labels, tokenizer.pad_token_id)\n decoded_labels = tokenizer.batch_decode(labels, skip_special_tokens=True)\n\n # Some simple post-processing\n decoded_preds = [pred.strip() for pred in decoded_preds]", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "383-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Evaluating Translation Quality" + }, + { + "text": "0, labels, tokenizer.pad_token_id)\n decoded_labels = tokenizer.batch_decode(labels, skip_special_tokens=True)\n\n # Some simple post-processing\n decoded_preds = [pred.strip() for pred in decoded_preds]\n decoded_labels = [[label.strip()] for label in decoded_labels]\n\n result = metric.compute(predictions=decoded_preds, references=decoded_labels)\n return {\"bleu\": result[\"score\"]}", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "383-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Evaluating Translation Quality" + }, + { + "text": "Explanation:\n## Configuring Model Training\n\nThe `Seq2SeqTrainingArguments` class defines the training configuration, including batch size, learning rate, evaluation strategy, generation parameters, and checkpointing.\n\nPython implementation:\nfrom transformers import Seq2SeqTrainingArguments\n\nargs = Seq2SeqTrainingArguments(\n f\"marian-finetuned-kde4-en-to-fr_translate\",\n eval_strategy=\"no\",\n save_strategy=\"epoch\",\n learning_rate=2e-5,\n per_device_train_batch_size=32,\n per_device_eval_batch_size=64,\n weight_decay=0.01,\n save_total_limit=3,\n num_train_epochs=3,\n predict_with_generate=True,\n fp16=True,\n push_to_hub=True,\n)\n\nPython implementation:\nfrom transformers import Seq2SeqTrainer\n\ntrainer = Seq2SeqTrainer(\n model,\n args,\n train_dataset=tokenized_datasets[\"train\"],\n eval_dataset=tokenized_datasets[\"validation\"],\n data_collator=data_collator,\n tokenizer=tokenizer,\n compute_metrics=compute_metrics,\n)", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "384-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Configuring Model Training" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Load and preprocess bilingual translation datasets.\n- Tokenize source and target languages for sequence-to-sequence learning.\n- Prepare translation datasets for supervised training.\n- Fine-tune pretrained encoder-decoder Transformer models.\n- Evaluate translation quality using BLEU.\n- Share fine-tuned translation models on the Hugging Face Hub.", + "source": "5.3.Translation_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "385-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.3.Translation_(PyTorch).ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Text Summarization with Transformers\n\nThis notebook demonstrates how to fine-tune a pretrained sequence-to-sequence (Seq2Seq) Transformer model for abstractive text summarization. It covers dataset preparation, preprocessing, evaluation with ROUGE, supervised fine-tuning, and custom training using the Accelerate library.\n\n## Learning Objectives\n\n- Understand abstractive text summarization\n- Prepare summarization datasets\n- Tokenize source documents and summaries\n- Evaluate summarization models using ROUGE\n- Fine-tune pretrained Seq2Seq models\n- Train using both the Trainer API and Accelerate\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Evaluate\n- Accelerate", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "386-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Text Summarization with Transformers" + }, + { + "text": "Explanation:\n## Preparing the Summarization Dataset\n\nThe first step is to prepare a dataset containing document-summary pairs. In this section, we load the dataset, inspect its structure, tokenize both the source documents and target summaries, and convert them into numerical representations suitable for sequence-to-sequence training.\n\nPython implementation:\nfrom huggingface_hub import notebook_login\n\nnotebook_login()\n\nExplanation:\nYou will also need to be logged in to the Hugging Face Hub. Execute the following and enter your credentials.\n\nPython implementation:\nfrom datasets import load_dataset\n\ndataset = load_dataset(\"cnn_dailymail\", \"3.0.0\")\ndataset\n\nPython implementation:\nfrom datasets import DatasetDict\n\nsmall_dataset = DatasetDict()\n\nfor split in dataset.keys():\n n = int(0.01 * len(dataset[split]))\n small_dataset[split] = dataset[split].select(range(n))\n\nsmall_dataset\n\nPython implementation:\ndef show_samples(dataset, num_samples=3, seed=42):", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "387-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Preparing the Summarization Dataset" + }, + { + "text": "dataset.keys():\n n = int(0.01 * len(dataset[split]))\n small_dataset[split] = dataset[split].select(range(n))\n\nsmall_dataset\n\nPython implementation:\ndef show_samples(dataset, num_samples=3, seed=42):\n sample = dataset[\"train\"].shuffle(seed=seed).select(range(num_samples))\n for example in sample:\n print(f\"\\n'>> highlights: {example['highlights']}'\")\n print(f\"'>> article: {example['article']}'\")\n\nshow_samples(small_dataset)\n\nPython implementation:\nsmall_dataset = small_dataset.filter(lambda x: len(x[\"highlights\"].split()) > 2)\nsmall_dataset\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\nmodel_checkpoint = \"google/mt5-small\"\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint)\n\nPython implementation:\ninputs = tokenizer(\"I loved reading the Hunger Games!\")\ninputs\n\nPython implementation:\nmax_input_length = 512\nmax_target_length = 30\n\ndef preprocess_function(examples):\n model_inputs = tokenizer(\n examples[\"article\"],\n max_length=max_input_length,\n truncation=True,\n )", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "387-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Preparing the Summarization Dataset" + }, + { + "text": "n implementation:\nmax_input_length = 512\nmax_target_length = 30\n\ndef preprocess_function(examples):\n model_inputs = tokenizer(\n examples[\"article\"],\n max_length=max_input_length,\n truncation=True,\n )\n labels = tokenizer(\n examples[\"highlights\"], max_length=max_target_length, truncation=True\n )\n model_inputs[\"labels\"] = labels[\"input_ids\"]\n return model_inputs\n\nPython implementation:\ngenerated_summary = \"I absolutely loved reading the Hunger Games\"\nreference_summary = \"I loved reading the Hunger Games\"", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "387-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Preparing the Summarization Dataset" + }, + { + "text": "Explanation:\n## Evaluating Summarization Quality\n\nUnlike classification tasks, summarization models are evaluated by comparing generated summaries with reference summaries. In this notebook, we use the ROUGE metric to measure the overlap between generated and reference text and establish a baseline before fine-tuning.\n\nPython implementation:\nimport evaluate\n\nrouge_score = evaluate.load(\"rouge\")\n\nPython implementation:\nscores = rouge_score.compute(\n predictions=[generated_summary], references=[reference_summary]\n)\nscores\n\nPython implementation:\nimport nltk\n\nnltk.download(\"punkt\")\n\nPython implementation:\nimport nltk\n\nnltk.download(\"punkt\")\nnltk.download(\"punkt_tab\")\n\nPython implementation:\nfrom nltk.tokenize import sent_tokenize\n\ndef three_sentence_summary(text):\n return \"\\n\".join(sent_tokenize(text)[:3])\n\nprint(three_sentence_summary(small_dataset[\"train\"][1][\"article\"]))\n\nPython implementation:\ndef evaluate_baseline(dataset, metric):", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "388-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Evaluating Summarization Quality" + }, + { + "text": "sentence_summary(text):\n return \"\\n\".join(sent_tokenize(text)[:3])\n\nprint(three_sentence_summary(small_dataset[\"train\"][1][\"article\"]))\n\nPython implementation:\ndef evaluate_baseline(dataset, metric):\n summaries = [three_sentence_summary(text) for text in dataset[\"article\"]]\n return metric.compute(predictions=summaries, references=dataset[\"highlights\"])\n\nPython implementation:\nimport pandas as pd\n\nscore = evaluate_baseline(small_dataset[\"validation\"], rouge_score)\nrouge_names = [\"rouge1\", \"rouge2\", \"rougeL\", \"rougeLsum\"]\nrouge_dict = dict((rn, round(score[rn] * 100, 2)) for rn in rouge_names)\nrouge_dict", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "388-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Evaluating Summarization Quality" + }, + { + "text": "Explanation:\n## Fine-Tuning the Summarization Model\n\nWe initialize a pretrained sequence-to-sequence Transformer model, prepare dynamic batching for variable-length documents, configure the training process, and fine-tune the model using the Hugging Face `Seq2SeqTrainer`. After training, the model is evaluated and can be shared through the Hugging Face Hub.\n\nPython implementation:\nfrom transformers import AutoModelForSeq2SeqLM\n\nmodel = AutoModelForSeq2SeqLM.from_pretrained(model_checkpoint)\n\nPython implementation:\nfrom transformers import Seq2SeqTrainingArguments\n\nbatch_size = 8\nnum_train_epochs = 8\n# Show the training loss with every epoch\nlogging_steps = len(tokenized_datasets[\"train\"]) // batch_size\nmodel_name = model_checkpoint.split(\"/\")[-1]\n\nargs = Seq2SeqTrainingArguments(\n output_dir=f\"{model_name}-finetuned-amazon-en-es\",\n eval_strategy=\"epoch\",\n learning_rate=5.6e-5,\n per_device_train_batch_size=batch_size,\n per_device_eval_batch_size=batch_size,\n weight_decay=0.01,", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "389-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Fine-Tuning the Summarization Model" + }, + { + "text": "output_dir=f\"{model_name}-finetuned-amazon-en-es\",\n eval_strategy=\"epoch\",\n learning_rate=5.6e-5,\n per_device_train_batch_size=batch_size,\n per_device_eval_batch_size=batch_size,\n weight_decay=0.01,\n save_total_limit=3,\n num_train_epochs=num_train_epochs,\n predict_with_generate=True,\n logging_steps=logging_steps,\n push_to_hub=True,\n)\n\nPython implementation:\nimport numpy as np\n\ndef compute_metrics(eval_pred):\n predictions, labels = eval_pred\n # Decode generated summaries into text\n decoded_preds = tokenizer.batch_decode(predictions, skip_special_tokens=True)\n # Replace -100 in the labels as we can't decode them\n labels = np.where(labels != -100, labels, tokenizer.pad_token_id)\n # Decode reference summaries into text\n decoded_labels = tokenizer.batch_decode(labels, skip_special_tokens=True)\n # ROUGE expects a newline after each sentence\n decoded_preds = [\"\\n\".join(sent_tokenize(pred.strip())) for pred in decoded_preds]", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "389-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Fine-Tuning the Summarization Model" + }, + { + "text": "ed_labels = tokenizer.batch_decode(labels, skip_special_tokens=True)\n # ROUGE expects a newline after each sentence\n decoded_preds = [\"\\n\".join(sent_tokenize(pred.strip())) for pred in decoded_preds]\n decoded_labels = [\"\\n\".join(sent_tokenize(label.strip())) for label in decoded_labels]\n # Compute ROUGE scores\n result = rouge_score.compute(\n predictions=decoded_preds, references=decoded_labels, use_stemmer=True\n )\n # Extract the median scores\n result = {key: value* 100 for key, value in result.items()}\n return {k: round(v, 4) for k, v in result.items()}\n\nPython implementation:\nfrom transformers import DataCollatorForSeq2Seq\n\ndata_collator = DataCollatorForSeq2Seq(tokenizer, model=model)\n\nPython implementation:\ntokenized_datasets = tokenized_datasets.remove_columns(\n small_dataset[\"train\"].column_names\n)\n\nPython implementation:\nfeatures = [tokenized_datasets[\"train\"][i] for i in range(2)]\ndata_collator(features)\n\nPython implementation:\nfrom transformers import Seq2SeqTrainer", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "389-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Fine-Tuning the Summarization Model" + }, + { + "text": "[\"train\"].column_names\n)\n\nPython implementation:\nfeatures = [tokenized_datasets[\"train\"][i] for i in range(2)]\ndata_collator(features)\n\nPython implementation:\nfrom transformers import Seq2SeqTrainer\n\ntrainer = Seq2SeqTrainer(\n model,\n args,\n train_dataset=tokenized_datasets[\"train\"],\n eval_dataset=tokenized_datasets[\"validation\"],\n data_collator=data_collator,\n tokenizer=tokenizer,\n compute_metrics=compute_metrics,\n)", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "389-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Fine-Tuning the Summarization Model" + }, + { + "text": "Explanation:\n## Custom Training with Accelerate\n\nFor greater flexibility, Hugging Face provides the `Accelerate` library, which simplifies custom training loops while supporting efficient execution on CPUs, GPUs, and distributed environments. In this section, we manually configure the training pipeline, optimizer, scheduler, and evaluation workflow.\n\nPython implementation:\nfrom torch.utils.data import DataLoader\n\nbatch_size = 8\ntrain_dataloader = DataLoader(\n tokenized_datasets[\"train\"],\n shuffle=True,\n collate_fn=data_collator,\n batch_size=batch_size,\n)\neval_dataloader = DataLoader(\n tokenized_datasets[\"validation\"], collate_fn=data_collator, batch_size=batch_size\n)\n\nPython implementation:\nfrom torch.optim import AdamW\n\noptimizer = AdamW(model.parameters(), lr=2e-5)\n\nPython implementation:\nfrom accelerate import Accelerator\n\naccelerator = Accelerator()\nmodel, optimizer, train_dataloader, eval_dataloader = accelerator.prepare(\n model, optimizer, train_dataloader, eval_dataloader\n)", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "390-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "ion:\nfrom accelerate import Accelerator\n\naccelerator = Accelerator()\nmodel, optimizer, train_dataloader, eval_dataloader = accelerator.prepare(\n model, optimizer, train_dataloader, eval_dataloader\n)\n\nPython implementation:\nfrom transformers import get_scheduler\n\nnum_train_epochs = 10\nnum_update_steps_per_epoch = len(train_dataloader)\nnum_training_steps = num_train_epochs * num_update_steps_per_epoch\n\nlr_scheduler = get_scheduler(\n \"linear\",\n optimizer=optimizer,\n num_warmup_steps=0,\n num_training_steps=num_training_steps,\n)\n\nPython implementation:\ndef postprocess_text(preds, labels):\n preds = [pred.strip() for pred in preds]\n labels = [label.strip() for label in labels]\n\n # ROUGE expects a newline after each sentence\n preds = [\"\\n\".join(nltk.sent_tokenize(pred)) for pred in preds]\n labels = [\"\\n\".join(nltk.sent_tokenize(label)) for label in labels]\n\n return preds, labels\n\nPython implementation:\nfrom huggingface_hub import get_full_repo_name", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "390-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "_tokenize(pred)) for pred in preds]\n labels = [\"\\n\".join(nltk.sent_tokenize(label)) for label in labels]\n\n return preds, labels\n\nPython implementation:\nfrom huggingface_hub import get_full_repo_name\n\nmodel_name = \"test-bert-finetuned-squad-accelerate\"\nrepo_name = get_full_repo_name(model_name)\nrepo_name\n\nPython implementation:\nfrom huggingface_hub import Repository\n\noutput_dir = \"results-mt5-finetuned-squad-accelerate\"\nrepo = Repository(output_dir, clone_from=repo_name)\n\nPython implementation:\nfrom tqdm.auto import tqdm\nimport torch\nimport numpy as np\n\nprogress_bar = tqdm(range(num_training_steps))\n\nfor epoch in range(num_train_epochs):\n # Training\n model.train()\n for step, batch in enumerate(train_dataloader):\n outputs = model(**batch)\n loss = outputs.loss\n accelerator.backward(loss)\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()\n progress_bar.update(1)\n\n # Evaluation\n model.eval()\n for step, batch in enumerate(eval_dataloader):\n with torch.no_grad():", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "390-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "ckward(loss)\n\n optimizer.step()\n lr_scheduler.step()\n optimizer.zero_grad()\n progress_bar.update(1)\n\n # Evaluation\n model.eval()\n for step, batch in enumerate(eval_dataloader):\n with torch.no_grad():\n generated_tokens = accelerator.unwrap_model(model).generate(\n batch[\"input_ids\"],\n attention_mask=batch[\"attention_mask\"],\n )\n\n generated_tokens = accelerator.pad_across_processes(\n generated_tokens, dim=1, pad_index=tokenizer.pad_token_id\n )\n labels = batch[\"labels\"]\n\n # If we did not pad to max length, we need to pad the labels too\n labels = accelerator.pad_across_processes(\n batch[\"labels\"], dim=1, pad_index=tokenizer.pad_token_id\n )\n\n generated_tokens = accelerator.gather(generated_tokens).cpu().numpy()\n labels = accelerator.gather(labels).cpu().numpy()\n\n # Replace -100 in the labels as we can't decode them\n labels = np.where(labels != -100, labels, tokenizer.pad_token_id)\n if isinstance(generated_tokens, tuple):\n generated_tokens = generated_tokens[0]", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "390-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "Replace -100 in the labels as we can't decode them\n labels = np.where(labels != -100, labels, tokenizer.pad_token_id)\n if isinstance(generated_tokens, tuple):\n generated_tokens = generated_tokens[0]\n decoded_preds = tokenizer.batch_decode(\n generated_tokens, skip_special_tokens=True\n )\n decoded_labels = tokenizer.batch_decode(labels, skip_special_tokens=True)\n\n decoded_preds, decoded_labels = postprocess_text(\n decoded_preds, decoded_labels\n )\n\n rouge_score.add_batch(predictions=decoded_preds, references=decoded_labels)\n\n # Compute metrics\n result = rouge_score.compute()\n # Extract the median ROUGE scores\n result = {key: value.mid.fmeasure * 100 for key, value in result.items()}\n result = {k: round(v, 4) for k, v in result.items()}\n print(f\"Epoch {epoch}:\", result)\n\n # Save and upload\n accelerator.wait_for_everyone()\n unwrapped_model = accelerator.unwrap_model(model)\n unwrapped_model.save_pretrained(output_dir, save_function=accelerator.save)\n if accelerator.is_main_process:", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "390-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "pload\n accelerator.wait_for_everyone()\n unwrapped_model = accelerator.unwrap_model(model)\n unwrapped_model.save_pretrained(output_dir, save_function=accelerator.save)\n if accelerator.is_main_process:\n tokenizer.save_pretrained(output_dir)\n repo.push_to_hub(\n commit_message=f\"Training in progress epoch {epoch}\", blocking=False\n )\n\nPython implementation:\nfrom transformers import pipeline\n\nhub_model_id = \"huggingface-course/mt5-small-finetuned-amazon-en-es\"\nsummarizer = pipeline(\"summarization\", model=hub_model_id)", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "390-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Custom Training with Accelerate" + }, + { + "text": "Explanation:\n## Generating Summaries\n\nAfter fine-tuning, we use the summarization pipeline to generate summaries for unseen documents and compare the generated outputs with the reference summaries to assess model performance qualitatively.\n\nPython implementation:\ndef print_summary(idx):\n review = books_dataset[\"test\"][idx][\"review_body\"]\n title = books_dataset[\"test\"][idx][\"review_title\"]\n summary = summarizer(books_dataset[\"test\"][idx][\"review_body\"])[0][\"summary_text\"]\n print(f\"'>>> Review: {review}'\")\n print(f\"\\n'>>> Title: {title}'\")\n print(f\"\\n'>>> Summary: {summary}'\")", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "391-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Generating Summaries" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Prepare document-summary datasets for sequence-to-sequence learning.\n- Tokenize source documents and target summaries.\n- Evaluate summarization models using the ROUGE metric.\n- Fine-tune pretrained Transformer models for abstractive summarization.\n- Implement both high-level (`Seq2SeqTrainer`) and custom (`Accelerate`) training workflows.\n- Generate summaries for previously unseen documents.", + "source": "5.4.Summarization_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "392-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.4.Summarization_(PyTorch).ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Training a Causal Language Model from Scratch\n\nThis notebook demonstrates how to train a GPT-style causal language model from scratch using the Hugging Face Transformers library. It covers dataset preparation, tokenizer training, model initialization, language model training, and model publishing.\n\n## Learning Objectives\n\n- Understand causal language modeling (CLM)\n- Prepare text datasets for autoregressive training\n- Train a GPT-style Transformer model from scratch\n- Fine-tune and evaluate language models\n- Publish trained models to the Hugging Face Hub\n\n## Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Hub\n\nPython implementation:\n!git config --global user.email \"milad.saeedi@mail.utoronto.ca\"\n!git config --global user.name \"Milad Saeedi\"", + "source": "5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "393-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "section": "Training a Causal Language Model from Scratch" + }, + { + "text": "Explanation:\n## Preparing the Training Dataset\n\nThe first step in training a causal language model is preparing a large text corpus. In this section, we load the dataset, inspect its contents, initialize a tokenizer, and convert raw text into fixed-length token sequences that can be used for autoregressive language modeling.\n\nPython implementation:\nfrom datasets import load_dataset, DatasetDict\n\nds_train = load_dataset(\"huggingface-course/codeparrot-ds-train\", split=\"train\")\nds_valid = load_dataset(\"huggingface-course/codeparrot-ds-valid\", split=\"validation\")\n\nraw_datasets = DatasetDict(\n {\n \"train\": ds_train.shuffle().select(range(1000)),\n \"valid\": ds_valid.shuffle().select(range(50))\n }\n)\n\nraw_datasets\n\nPython implementation:\nfrom transformers import AutoTokenizer\n\ncontext_length = 128\ntokenizer = AutoTokenizer.from_pretrained(\"huggingface-course/code-search-net-tokenizer\")\n\noutputs = tokenizer(\n raw_datasets[\"train\"][:2][\"content\"],\n truncation=True,\n max_length=context_length,", + "source": "5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "394-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "section": "Preparing the Training Dataset" + }, + { + "text": "28\ntokenizer = AutoTokenizer.from_pretrained(\"huggingface-course/code-search-net-tokenizer\")\n\noutputs = tokenizer(\n raw_datasets[\"train\"][:2][\"content\"],\n truncation=True,\n max_length=context_length,\n return_overflowing_tokens=True,\n return_length=True,\n)\n\nprint(f\"Input IDs length: {len(outputs['input_ids'])}\")\nprint(f\"Input chunk lengths: {(outputs['length'])}\")\nprint(f\"Chunk mapping: {outputs['overflow_to_sample_mapping']}\")\n\nPython implementation:\ndef tokenize(element):\n outputs = tokenizer(\n element[\"content\"],\n truncation=True,\n max_length=context_length,\n return_overflowing_tokens=True,\n return_length=True,\n )\n input_batch = []\n for length, input_ids in zip(outputs[\"length\"], outputs[\"input_ids\"]):\n if length == context_length:\n input_batch.append(input_ids)\n return {\"input_ids\": input_batch}\n\ntokenized_datasets = raw_datasets.map(\n tokenize, batched=True, remove_columns=raw_datasets[\"train\"].column_names\n)\ntokenized_datasets", + "source": "5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "394-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "section": "Preparing the Training Dataset" + }, + { + "text": "Explanation:\n## Building a GPT-Style Language Model\n\nUnlike fine-tuning a pretrained model, this notebook initializes a Transformer model from a configuration rather than loading pretrained weights. We define the model architecture, estimate its size, and prepare dynamically padded training batches for causal language modeling.\n\nPython implementation:\nfrom transformers import AutoTokenizer, GPT2LMHeadModel, AutoConfig\n\nconfig = AutoConfig.from_pretrained(\n \"gpt2\",\n vocab_size=len(tokenizer),\n n_ctx=context_length,\n bos_token_id=tokenizer.bos_token_id,\n eos_token_id=tokenizer.eos_token_id,\n)\n\nPython implementation:\nmodel = GPT2LMHeadModel(config)\nmodel_size = sum(t.numel() for t in model.parameters())\nprint(f\"GPT-2 size: {model_size/1000**2:.1f}M parameters\")\n\nPython implementation:\nfrom transformers import DataCollatorForLanguageModeling\n\ntokenizer.pad_token = tokenizer.eos_token\ndata_collator = DataCollatorForLanguageModeling(tokenizer, mlm=False)\n\nPython implementation:", + "source": "5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "395-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "section": "Building a GPT-Style Language Model" + }, + { + "text": "tion:\nfrom transformers import DataCollatorForLanguageModeling\n\ntokenizer.pad_token = tokenizer.eos_token\ndata_collator = DataCollatorForLanguageModeling(tokenizer, mlm=False)\n\nPython implementation:\nout = data_collator([tokenized_datasets[\"train\"][i] for i in range(5)])\nfor key in out:\n print(f\"{key} shape: {out[key].shape}\")", + "source": "5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "395-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "section": "Building a GPT-Style Language Model" + }, + { + "text": "Explanation:\n## Publishing the Trained Model\n\nOnce training is complete, the model can be uploaded to the Hugging Face Hub, allowing it to be versioned, shared, and reused for future inference or fine-tuning tasks.\n\nPython implementation:\nfrom transformers import Trainer, TrainingArguments\n\nargs = TrainingArguments(\n output_dir=\"codeparrot-ds\",\n per_device_train_batch_size=16,\n per_device_eval_batch_size=16,\n eval_strategy=\"steps\",\n eval_steps=5_000,\n logging_steps=5_000,\n gradient_accumulation_steps=8,\n num_train_epochs=1,\n weight_decay=0.1,\n warmup_steps=1_000,\n lr_scheduler_type=\"cosine\",\n learning_rate=5e-4,\n save_steps=5_000,\n fp16=True,\n push_to_hub=True,\n)\n\ntrainer = Trainer(\n model=model,\n tokenizer=tokenizer,\n args=args,\n data_collator=data_collator,\n train_dataset=tokenized_datasets[\"train\"],\n eval_dataset=tokenized_datasets[\"valid\"],\n)", + "source": "5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "396-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "section": "Publishing the Trained Model" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Prepare text datasets for causal language modeling.\n- Tokenize large text corpora for autoregressive training.\n- Initialize a GPT-style Transformer model from scratch.\n- Train a causal language model using the Hugging Face `Trainer` API.\n- Publish trained language models to the Hugging Face Hub.", + "source": "5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "file_type": "ipynb", + "chunk_id": "397-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/5.5.Training_a_causal_language_model_from_scratch_(PyTorch).ipynb", + "section": "Key Takeaways" + }, + { + "text": "# Module 5 – NLP Applications\n\nThis module demonstrates practical Natural Language Processing (NLP) tasks using pretrained Transformer models from the Hugging Face ecosystem.\n\nEach notebook focuses on a specific NLP application, illustrating the complete workflow from preprocessing and model loading to inference or fine-tuning.\n\n---\n\n## Topics Covered\n\n- Token Classification\n- Masked Language Modeling\n- Machine Translation\n- Text Summarization\n- Causal Language Modeling\n\n---\n\n## Notebooks\n\n| Notebook | Description |\n|----------|-------------|\n| Token Classification | Named Entity Recognition and token-level prediction |\n| Masked Language Modeling | Fine-tuning BERT-style masked language models |\n| Machine Translation | Neural machine translation with encoder-decoder models |\n| Text Summarization | Abstractive summarization using Transformer models |\n| Causal Language Modeling | Training autoregressive language models from scratch |\n\n---\n\n## Skills Demonstrated", + "source": "README.md", + "file_type": "md", + "chunk_id": "398-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/README.md", + "section": null + }, + { + "text": "models |\n| Text Summarization | Abstractive summarization using Transformer models |\n| Causal Language Modeling | Training autoregressive language models from scratch |\n\n---\n\n## Skills Demonstrated\n\n- Hugging Face Transformers\n- Token Classification\n- Named Entity Recognition\n- Machine Translation\n- Text Summarization\n- Masked Language Modeling\n- Causal Language Modeling\n- Dataset preprocessing\n- Model evaluation\n\n---\n\n## Requirements\n\n```bash\npip install -r ../requirements.txt", + "source": "README.md", + "file_type": "md", + "chunk_id": "398-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "05-nlp-tasks/README.md", + "section": null + }, + { + "text": "Explanation:\n# Exploring Chat Templates with SmolLM2\n\nThis notebook demonstrates how chat templates are used to format conversations for instruction-tuned Large Language Models (LLMs). It explores the SmolLM2 chat template, shows how conversational messages are converted into model-ready prompts, and demonstrates how chat templates can be applied to real-world conversational datasets.\n\n## Learning Objectives\n\n- Understand the purpose of chat templates\n- Format conversations for instruction-tuned LLMs\n- Convert structured messages into model-ready prompts\n- Apply chat templates to conversational datasets\n- Prepare datasets for supervised fine-tuning (SFT)\n\n## Technologies\n\n- Python\n- Hugging Face Transformers\n- Hugging Face Datasets\n- SmolLM2\n\nPython implementation:\n# Import necessary libraries\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\nfrom trl import setup_chat_format\nimport torch", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "399-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Exploring Chat Templates with SmolLM2" + }, + { + "text": "Explanation:\n## Loading the Chat Model\n\nInstruction-tuned language models expect conversations to follow a specific format. Chat templates define this format by organizing user, assistant, and system messages into a structured prompt that the model can understand.\n\nIn this section, we load the SmolLM2 tokenizer and model, then explore its built-in chat template.\n\nPython implementation:\n# Dynamically set the device\ndevice = (\n \"cuda\"\n if torch.cuda.is_available()\n else \"mps\" if torch.backends.mps.is_available() else \"cpu\"\n)\n\nmodel_name = \"HuggingFaceTB/SmolLM2-135M\"\nmodel = AutoModelForCausalLM.from_pretrained(\n pretrained_model_name_or_path=model_name\n).to(device)\ntokenizer = AutoTokenizer.from_pretrained(pretrained_model_name_or_path=model_name)\nmodel, tokenizer = setup_chat_format(model=model, tokenizer=tokenizer)\n\nPython implementation:\n# Define messages for SmolLM2\nmessages = [\n {\"role\": \"user\", \"content\": \"Hello, how are you?\"},\n {\n \"role\": \"assistant\",", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "400-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Loading the Chat Model" + }, + { + "text": "= setup_chat_format(model=model, tokenizer=tokenizer)\n\nPython implementation:\n# Define messages for SmolLM2\nmessages = [\n {\"role\": \"user\", \"content\": \"Hello, how are you?\"},\n {\n \"role\": \"assistant\",\n \"content\": \"I'm doing well, thank you! How can I assist you today?\",\n },\n]", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "400-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Loading the Chat Model" + }, + { + "text": "Explanation:\n## Understanding Chat Templates\n\nChat templates convert structured conversations into a single text prompt that follows the format expected by a conversational language model. Rather than manually writing special tokens or role markers, the tokenizer automatically applies the correct template for the selected model.\n\nIn this section, we compare the raw conversation, the formatted prompt, and the tokenized representation used during inference.", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "401-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Understanding Chat Templates" + }, + { + "text": "Explanation:\n## Applying Chat Templates to Conversational Datasets\n\nLarge Language Models are typically fine-tuned using datasets that contain multi-turn conversations. This section demonstrates how chat templates can be applied automatically to every example in a dataset, converting structured conversations into training-ready text.\n\nPython implementation:\ninput_text = tokenizer.apply_chat_template(\n messages, tokenize=True, add_generation_prompt=True\n)\n\nprint(\"Conversation decoded:\", tokenizer.decode(token_ids=input_text))", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "402-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Applying Chat Templates to Conversational Datasets" + }, + { + "text": "Explanation:\n# Tokenize the conversation\n\nOf course, the tokenizer also tokenizes the conversation and special token as ids that relate to the model's vocabulary.\n\nExplanation:\n
\n

Exercise: Process a dataset for SFT

\n

Take a dataset from the Hugging Face hub and process it for SFT.

\n

Difficulty Levels

\n

🐢 Convert the `HuggingFaceTB/smoltalk` dataset into chatml format.

\n

🐕 Convert the `openai/gsm8k` dataset into chatml format.

\n
\n\nPython implementation:\nfrom IPython.core.display import display, HTML\n\ndisplay(\n HTML(\n \"\"\"\n\"\"\"\n )\n)\n\nPython implementation:\nfrom datasets import load_dataset", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "403-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Tokenize the conversation" + }, + { + "text": "gingface.co/datasets/HuggingFaceTB/smoltalk/embed/viewer/all/train?row=0\"\n frameborder=\"0\"\n width=\"100%\"\n height=\"360px\"\n>\n\"\"\"\n )\n)\n\nPython implementation:\nfrom datasets import load_dataset\n\nds = load_dataset(\"HuggingFaceTB/smoltalk\", \"everyday-conversations\")\nds\n\nPython implementation:\nmessages=ds['train'][0]['messages']\nmessages\n\nPython implementation:\ninput_text = tokenizer.apply_chat_template(\n messages, tokenize=True, add_generation_prompt=True\n)\n\nprint(\"Conversation decoded:\", tokenizer.decode(token_ids=input_text))\n\nPython implementation:\ndef process_dataset(sample):\n sample[\"text\"] = tokenizer.apply_chat_template(\n sample[\"messages\"],\n tokenize=False\n )\n return sample\n\nds = ds.map(process_dataset)\nds", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "403-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Tokenize the conversation" + }, + { + "text": "Explanation:\n## Preparing Instruction-Following Datasets\n\nDifferent datasets use different formats for representing questions and answers. In this section, we transform an instruction-following dataset into the conversational message format expected by chat templates, making it suitable for supervised fine-tuning.\n\nPython implementation:\ndisplay(\n HTML(\n \"\"\"\n\"\"\"\n )\n)\n\nPython implementation:\nds = load_dataset(\"openai/gsm8k\", \"main\")\nds\n\nPython implementation:\ndef process_dataset(sample):\n # 1. Create a chat conversation\n messages = [\n {\n \"role\": \"user\",\n \"content\": sample[\"question\"],\n },\n {\n \"role\": \"assistant\",\n \"content\": sample[\"answer\"],\n },\n ]\n\n # 2. Apply the chat template\n sample[\"text\"] = tokenizer.apply_chat_template(\n messages,\n tokenize=False,\n )\n\n return sample\n\nds = ds.map(process_dataset)\nds", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "404-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Preparing Instruction-Following Datasets" + }, + { + "text": "Explanation:\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Format conversations using chat templates.\n- Convert structured messages into model-ready prompts.\n- Understand the relationship between conversations and tokenized inputs.\n- Apply chat templates to conversational datasets.\n- Prepare instruction-following datasets for supervised fine-tuning of Large Language Models.", + "source": "6.1.chat-templates.ipynb", + "file_type": "ipynb", + "chunk_id": "405-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.1.chat-templates.ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# Supervised Fine-Tuning with SFTTrainer\n\nThis notebook demonstrates how to fine-tune the `HuggingFaceTB/SmolLM2-135M` model using the `SFTTrainer` from the `trl` library. The notebook cells run and will finetune the model. You can select your difficulty by trying out different datasets.\n\n
\n

Exercise: Fine-Tuning SmolLM2 with SFTTrainer

\n

Take a dataset from the Hugging Face hub and finetune a model on it.

\n

Difficulty Levels

\n

🐢 Use the `HuggingFaceTB/smoltalk` dataset

\n

🐕 Try out the `bigcode/the-stack-smol` dataset and finetune a code generation model on a specific subset `data/python`.

\n

🦁 Select a dataset that relates to a real world use case your interested in

\n
\n\n## Supervised Fine-Tuning with SFTTrainer", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "406-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Supervised Fine-Tuning with SFTTrainer" + }, + { + "text": "a code generation model on a specific subset `data/python`.

\n

🦁 Select a dataset that relates to a real world use case your interested in

\n\n\n## Supervised Fine-Tuning with SFTTrainer\n\nThis notebook demonstrates how to perform Supervised Fine-Tuning (SFT) of a pretrained Large Language Model (LLM) using the Hugging Face TRL library. It covers model preparation, conversational dataset processing, training with the `SFTTrainer`, and evaluating the fine-tuned model.\n\n### Learning Objectives\n\n- Understand Supervised Fine-Tuning (SFT)\n- Prepare conversational datasets for SFT\n- Configure the `SFTTrainer`\n- Fine-tune instruction-following language models\n- Compare model behavior before and after fine-tuning\n- Publish fine-tuned models to the Hugging Face Hub\n\n### Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- TRL\n- Hugging Face Hub\n\nPython implementation:\nfrom huggingface_hub import notebook_login\n\nnotebook_login()", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "406-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Supervised Fine-Tuning with SFTTrainer" + }, + { + "text": "Explanation:\n## Loading the Base Model\n\nBefore fine-tuning, we load a pretrained instruction-following language model and evaluate its responses on a sample prompt. This baseline provides a reference for comparing the model's behavior before and after supervised fine-tuning.\n\nPython implementation:\n# Import necessary libraries\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\nfrom datasets import load_dataset\nfrom trl import SFTConfig, SFTTrainer, setup_chat_format\nimport torch\n\ndevice = (\n \"cuda\"\n if torch.cuda.is_available()\n else \"mps\" if torch.backends.mps.is_available() else \"cpu\"\n)\n\n# Load the model and tokenizer\nmodel_name = \"HuggingFaceTB/SmolLM2-135M\"\nmodel = AutoModelForCausalLM.from_pretrained(\n pretrained_model_name_or_path=model_name\n).to(device)\ntokenizer = AutoTokenizer.from_pretrained(pretrained_model_name_or_path=model_name)\n\n# Set up the chat format\nmodel, tokenizer = setup_chat_format(model=model, tokenizer=tokenizer)", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "407-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Loading the Base Model" + }, + { + "text": "name\n).to(device)\ntokenizer = AutoTokenizer.from_pretrained(pretrained_model_name_or_path=model_name)\n\n# Set up the chat format\nmodel, tokenizer = setup_chat_format(model=model, tokenizer=tokenizer)\n\n# Set our name for the finetune to be saved &/ uploaded to\nfinetune_name = \"SmolLM2-FT-MyDataset\"\nfinetune_tags = [\"smol-course\", \"module_1\"]", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "407-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Loading the Base Model" + }, + { + "text": "Explanation:\n### Generate with the base model\n\nHere we will try out the base model which does not have a chat template.\n\nPython implementation:\n# Let's test the base model before training\nprompt = \"Write a haiku about programming\"\n\n# Format with template\nmessages = [{\"role\": \"user\", \"content\": prompt}]\nformatted_prompt = tokenizer.apply_chat_template(messages, tokenize=False)\n\n# Generate response\ninputs = tokenizer(formatted_prompt, return_tensors=\"pt\").to(device)\noutputs = model.generate(**inputs, max_new_tokens=100)\nprint(\"Before training:\")\nprint(tokenizer.decode(outputs[0], skip_special_tokens=True))", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "408-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Generate with the base model" + }, + { + "text": "Explanation:\n## Preparing the Training Dataset\n\nSupervised Fine-Tuning requires examples of prompts paired with high-quality responses. In this section, we load a conversational dataset, inspect its structure, and prepare it for training using the format expected by the `SFTTrainer`.\n\n### Dataset Preparation\n\nWe will load a sample dataset and format it for training. The dataset should be structured with input-output pairs, where each input is a prompt and the output is the expected response from the model.\n\n**TRL will format input messages based on the model's chat templates.** They need to be represented as a list of dictionaries with the keys: `role` and `content`,.\n\nPython implementation:\n# Load a sample dataset\nfrom datasets import load_dataset\n\n# TODO: define your dataset and config using the path and name parameters\nds = load_dataset(path=\"HuggingFaceTB/smoltalk\", name=\"everyday-conversations\")\nds\n\nPython implementation:", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "409-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Preparing the Training Dataset" + }, + { + "text": "port load_dataset\n\n# TODO: define your dataset and config using the path and name parameters\nds = load_dataset(path=\"HuggingFaceTB/smoltalk\", name=\"everyday-conversations\")\nds\n\nPython implementation:\n# TODO: 🦁 If your dataset is not in a format that TRL can convert to the chat template, you will need to process it. Refer to the [module](../chat_templates.md)\n\ndef process_dataset(sample):\n sample[\"text\"] = tokenizer.apply_chat_template(\n sample[\"messages\"],\n tokenize=False\n )\n return sample\n\nds = ds.map(process_dataset)\nds", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "409-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Preparing the Training Dataset" + }, + { + "text": "Explanation:\n## Fine-Tuning with SFTTrainer\n\nThe TRL `SFTTrainer` simplifies supervised fine-tuning by integrating dataset processing, optimization, checkpointing, and evaluation into a single training workflow. After configuring the training parameters, we fine-tune the model on the prepared conversational dataset and upload the resulting model to the Hugging Face Hub.\n\nThe `SFTTrainer` is configured with various parameters that control the training process. These include the number of training steps, batch size, learning rate, and evaluation strategy. Adjust these parameters based on your specific requirements and computational resources.\n\nPython implementation:\n# Configure the SFTTrainer\nsft_config = SFTConfig(\n output_dir=\"./sft_output\",\n max_steps=1000, # Adjust based on dataset size and desired training duration\n per_device_train_batch_size=4, # Set according to your GPU memory capacity\n learning_rate=5e-5, # Common starting point for fine-tuning", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "410-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Fine-Tuning with SFTTrainer" + }, + { + "text": ", # Adjust based on dataset size and desired training duration\n per_device_train_batch_size=4, # Set according to your GPU memory capacity\n learning_rate=5e-5, # Common starting point for fine-tuning\n logging_steps=10, # Frequency of logging training metrics\n save_steps=100, # Frequency of saving model checkpoints\n eval_strategy=\"steps\", # Evaluate the model at regular intervals\n eval_steps=50, # Frequency of evaluation\n use_mps_device=(\n True if device == \"mps\" else False\n ), # Use MPS for mixed precision training\n hub_model_id=finetune_name, # Set a unique name for your model\n dataset_text_field=\"text\",\n)\n\n# Initialize the SFTTrainer\ntrainer = SFTTrainer(\n model=model,\n args=sft_config,\n train_dataset=ds[\"train\"],\n processing_class=tokenizer,\n eval_dataset=ds[\"test\"],\n)\n\n# TODO: 🦁 🐕 align the SFTTrainer params with your chosen dataset. For example, if you are using the `bigcode/the-stack-smol` dataset, you will need to choose the `content` column`", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "410-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Fine-Tuning with SFTTrainer" + }, + { + "text": "Explanation:\n## Training the Model\n\nWith the trainer configured, we can now proceed to train the model. The training process will involve iterating over the dataset, computing the loss, and updating the model's parameters to minimize this loss.\n\nPython implementation:\n# Train the model\ntrainer.train()\n\n# Save the model\ntrainer.save_model(f\"./{finetune_name}\")\n\nExplanation:\n
\n

Bonus Exercise: Generate with fine-tuned model

\n

🐕 Use the fine-tuned to model generate a response, just like with the base example..

\n
", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "411-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Training the Model" + }, + { + "text": "Explanation:\n## Evaluating the Fine-Tuned Model\n\nAfter training, we generate responses using the same prompt employed before fine-tuning. Comparing the outputs illustrates how supervised fine-tuning improves the model's ability to follow instructions and produce task-specific responses.\n\nPython implementation:\n# Test the fine-tuned model on the same prompt\n\n# Let's test the base model before training\nprompt = \"Write a haiku about programming\"\n\n# Format with template\nmessages = [{\"role\": \"user\", \"content\": prompt}]\nformatted_prompt = tokenizer.apply_chat_template(messages, tokenize=False)\n\n# Generate response\ninputs = tokenizer(formatted_prompt, return_tensors=\"pt\").to(device)\n\noutputs = trainer.model.generate(\n **inputs,\n max_new_tokens=100,\n do_sample=True,\n temperature=0.7,\n top_p=0.9,\n )\n\nresponse = tokenizer.decode(outputs[0], skip_special_tokens=True)\n\nprint(response)", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "412-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Evaluating the Fine-Tuned Model" + }, + { + "text": "Explanation:\n## 💐 You're done!\n\nThis notebook provided a step-by-step guide to fine-tuning the `HuggingFaceTB/SmolLM2-135M` model using the `SFTTrainer`. By following these steps, you can adapt the model to perform specific tasks more effectively. If you want to carry on working on this course, here are steps you could try out:\n\n- Try this notebook on a harder difficulty\n- Review a colleagues PR\n- Improve the course material via an Issue or PR.", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "413-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "💐 You're done!" + }, + { + "text": "Explanation:\n---\n\n# Key Takeaways\n\nIn this notebook, I learned how to:\n\n- Prepare conversational datasets for supervised fine-tuning.\n- Configure the TRL `SFTTrainer`.\n- Fine-tune a pretrained instruction-following language model.\n- Compare model behavior before and after fine-tuning.\n- Publish fine-tuned models to the Hugging Face Hub.", + "source": "6.2.supervised-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "414-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.2.supervised-fine-tuning.ipynb", + "section": "Key Takeaways" + }, + { + "text": "Explanation:\n# How to Fine-Tune LLMs with LoRA Adapters using Hugging Face TRL\n\nThis notebook demonstrates how to efficiently fine-tune large language models using LoRA (Low-Rank Adaptation) adapters. LoRA is a parameter-efficient fine-tuning technique that:\n- Freezes the pre-trained model weights\n- Adds small trainable rank decomposition matrices to attention layers\n- Typically reduces trainable parameters by ~90%\n- Maintains model performance while being memory efficient\n\nWe'll cover:\n1. Setup development environment and LoRA configuration\n2. Create and prepare the dataset for adapter training\n3. Fine-tune using `trl` and `SFTTrainer` with LoRA adapters\n4. Test the model and merge adapters (optional)", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "415-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "How to Fine-Tune LLMs with LoRA Adapters using Hugging Face TRL" + }, + { + "text": "Explanation:\n## 1. Setup development environment\n\nOur first step is to install Hugging Face Libraries and Pytorch, including trl, transformers and datasets. If you haven't heard of trl yet, don't worry. It is a new library on top of transformers and datasets, which makes it easier to fine-tune, rlhf, align open LLMs.\n\nPython implementation:\n!pip install -q \\\n transformers==4.56.1 \\\n datasets==3.6.0 \\\n evaluate==0.4.6 \\\n huggingface_hub==0.36.2\\\n trl==0.17.0\n\nPython implementation:\n# Install the requirements in Google Colab\n# !pip install transformers datasets trl huggingface_hub\n\n# Authenticate to Hugging Face\n\nfrom huggingface_hub import login\n\nlogin()\n\n# for convenience you can create an environment variable containing your hub token as HF_TOKEN", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "416-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "1. Setup development environment" + }, + { + "text": "Explanation:\n## 2. Load the dataset\n\nPython implementation:\n# Load a sample dataset\nfrom datasets import load_dataset\n\n# TODO: define your dataset and config using the path and name parameters\ndataset = load_dataset(path=\"HuggingFaceTB/smoltalk\", name=\"everyday-conversations\")\ndataset", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "417-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "2. Load the dataset" + }, + { + "text": "Explanation:\n## 3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA\n\nThe [SFTTrainer](https://huggingface.co/docs/trl/sft_trainer) from `trl` provides integration with LoRA adapters through the [PEFT](https://huggingface.co/docs/peft/en/index) library. Key advantages of this setup include:\n\n1. **Memory Efficiency**:\n - Only adapter parameters are stored in GPU memory\n - Base model weights remain frozen and can be loaded in lower precision\n - Enables fine-tuning of large models on consumer GPUs\n\n2. **Training Features**:\n - Native PEFT/LoRA integration with minimal setup\n - Support for QLoRA (Quantized LoRA) for even better memory efficiency\n\n3. **Adapter Management**:\n - Adapter weight saving during checkpoints\n - Features to merge adapters back into base model\n\nWe'll use LoRA in our example, which combines LoRA with 4-bit quantization to further reduce memory usage without sacrificing performance. The setup requires just a few configuration steps:\n1.", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "model\n\nWe'll use LoRA in our example, which combines LoRA with 4-bit quantization to further reduce memory usage without sacrificing performance. The setup requires just a few configuration steps:\n1. Define the LoRA configuration (rank, alpha, dropout)\n2. Create the SFTTrainer with PEFT config\n3. Train and save the adapter weights\n\nPython implementation:\n# Import necessary libraries\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\nfrom datasets import load_dataset\nfrom trl import SFTConfig, SFTTrainer\nimport torch\n\ndevice = (\n \"cuda\"\n if torch.cuda.is_available()\n else \"mps\" if torch.backends.mps.is_available() else \"cpu\"\n)\n\n# Load the model and tokenizer\nmodel_name = \"HuggingFaceTB/SmolLM2-135M-Instruct\"\n\nmodel = AutoModelForCausalLM.from_pretrained(\n pretrained_model_name_or_path=model_name\n).to(device)\ntokenizer = AutoTokenizer.from_pretrained(pretrained_model_name_or_path=model_name)\n\n# Set our name for the finetune to be saved &/ uploaded to", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "pretrained_model_name_or_path=model_name\n).to(device)\ntokenizer = AutoTokenizer.from_pretrained(pretrained_model_name_or_path=model_name)\n\n# Set our name for the finetune to be saved &/ uploaded to\nfinetune_name = \"SmolLM2-SFT-MyDataset\"\nfinetune_tags = [\"smol-course\", \"module_1\"]\n\nExplanation:\nThe `SFTTrainer`  supports a native integration with `peft`, which makes it super easy to efficiently tune LLMs using, e.g. LoRA. We only need to create our `LoraConfig` and provide it to the trainer.\n\n
\n

Exercise: Define LoRA parameters for finetuning

\n

Take a dataset from the Hugging Face hub and finetune a model on it.

\n

Difficulty Levels

\n

🐢 Use the general parameters for an abitrary finetune

\n

🐕 Adjust the parameters and review in weights & biases.

", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "ace hub and finetune a model on it.

\n

Difficulty Levels

\n

🐢 Use the general parameters for an abitrary finetune

\n

🐕 Adjust the parameters and review in weights & biases.

\n

🦁 Adjust the parameters and show change in inference results.

\n
\n\nPython implementation:\nfrom peft import LoraConfig\n\n# TODO: Configure LoRA parameters\n# r: rank dimension for LoRA update matrices (smaller = more compression)\nrank_dimension = 6\n# lora_alpha: scaling factor for LoRA layers (higher = stronger adaptation)\nlora_alpha = 8\n# lora_dropout: dropout probability for LoRA layers (helps prevent overfitting)\nlora_dropout = 0.05\n\npeft_config = LoraConfig(\n r=rank_dimension, # Rank dimension - typically between 4-32\n lora_alpha=lora_alpha, # LoRA scaling factor - typically 2x rank\n lora_dropout=lora_dropout, # Dropout probability for LoRA layers\n bias=\"none\", # Bias type for LoRA. the corresponding biases will be updated during training.", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "LoRA scaling factor - typically 2x rank\n lora_dropout=lora_dropout, # Dropout probability for LoRA layers\n bias=\"none\", # Bias type for LoRA. the corresponding biases will be updated during training.\n target_modules=\"all-linear\", # Which modules to apply LoRA to\n task_type=\"CAUSAL_LM\", # Task type for model architecture\n)\n\nExplanation:\nBefore we can start our training we need to define the hyperparameters (`TrainingArguments`) we want to use.\n\nPython implementation:\nmax_seq_length = 1512 # max sequence length for model and packing of the dataset\n\n# Training configuration\n# Hyperparameters based on QLoRA paper recommendations\nargs = SFTConfig(\n # Output settings\n output_dir=finetune_name, # Directory to save model checkpoints\n # Training duration\n num_train_epochs=1, # Number of training epochs\n # Batch size settings\n per_device_train_batch_size=2, # Batch size per GPU\n gradient_accumulation_steps=2, # Accumulate gradients for larger effective batch\n # Memory optimization", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "of training epochs\n # Batch size settings\n per_device_train_batch_size=2, # Batch size per GPU\n gradient_accumulation_steps=2, # Accumulate gradients for larger effective batch\n # Memory optimization\n gradient_checkpointing=True, # Trade compute for memory savings\n # Optimizer settings\n optim=\"adamw_torch_fused\", # Use fused AdamW for efficiency\n learning_rate=2e-4, # Learning rate (QLoRA paper)\n max_grad_norm=0.3, # Gradient clipping threshold\n # Learning rate schedule\n warmup_ratio=0.03, # Portion of steps for warmup\n lr_scheduler_type=\"constant\", # Keep learning rate constant after warmup\n # Logging and saving\n logging_steps=10, # Log metrics every N steps\n save_strategy=\"epoch\", # Save checkpoint every epoch\n # Precision settings\n bf16=True, # Use bfloat16 precision\n # Integration settings\n push_to_hub=False, # Don't push to HuggingFace Hub\n report_to=\"none\", # Disable external logging\n max_length=max_seq_length,\n packing=True, # Enable input packing for efficiency", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "Integration settings\n push_to_hub=False, # Don't push to HuggingFace Hub\n report_to=\"none\", # Disable external logging\n max_length=max_seq_length,\n packing=True, # Enable input packing for efficiency\n dataset_kwargs={\n \"add_special_tokens\": False, # Special tokens handled by template\n \"append_concat_token\": False, # No additional separator needed\n },\n)\n\nExplanation:\nWe now have every building block we need to create our `SFTTrainer` to start then training our model.\n\nPython implementation:\n# Create SFTTrainer with LoRA configuration\ntrainer = SFTTrainer(\n model=model,\n args=args,\n train_dataset=dataset[\"train\"],\n peft_config=peft_config, # LoRA configuration\n processing_class=tokenizer,\n)\n\nExplanation:\nStart training our model by calling the `train()` method on our `Trainer` instance. This will start the training loop and train our model for 3 epochs. Since we are using a PEFT method, we will only save the adapted model weights and not the full model.\n\nPython implementation:", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "ance. This will start the training loop and train our model for 3 epochs. Since we are using a PEFT method, we will only save the adapted model weights and not the full model.\n\nPython implementation:\n# start training, the model will be automatically saved to the hub and the output directory\ntrainer.train()\n\nPython implementation:\n# save model\ntrainer.save_model()\n\nExplanation:\nThe training with Flash Attention for 3 epochs with a dataset of 15k samples took 4:14:36 on a `g5.2xlarge`. The instance costs `1.21$/h` which brings us to a total cost of only ~`5.3$`.", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "418-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Fine-tune LLM using `trl` and the `SFTTrainer` with LoRA" + }, + { + "text": "Explanation:\n### Merge LoRA Adapter into the Original Model\n\nWhen using LoRA, we only train adapter weights while keeping the base model frozen. During training, we save only these lightweight adapter weights (~2-10MB) rather than a full model copy. However, for deployment, you might want to merge the adapters back into the base model for:\n\n1. **Simplified Deployment**: Single model file instead of base model + adapters\n2. **Inference Speed**: No adapter computation overhead\n3. **Framework Compatibility**: Better compatibility with serving frameworks\n\nPython implementation:\nfrom peft import AutoPeftModelForCausalLM\n\n# Load PEFT model on CPU\nmodel = AutoPeftModelForCausalLM.from_pretrained(\n pretrained_model_name_or_path=args.output_dir,\n torch_dtype=torch.float16,\n low_cpu_mem_usage=True,\n)\n\n# Merge LoRA and base model and save\nmerged_model = model.merge_and_unload()\nmerged_model.save_pretrained(\n args.output_dir, safe_serialization=True, max_shard_size=\"2GB\"\n)", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "419-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "Merge LoRA Adapter into the Original Model" + }, + { + "text": "Explanation:\n## 3. Test Model and run Inference\n\nAfter the training is done we want to test our model. We will load different samples from the original dataset and evaluate the model on those samples, using a simple loop and accuracy as our metric.\n\nExplanation:\n
\n

Bonus Exercise: Load LoRA Adapter

\n

Use what you learnt from the ecample note book to load your trained LoRA adapter for inference.

\n
\n\nPython implementation:\n# free the memory again\ndel model\ndel trainer\ntorch.cuda.empty_cache()\n\nPython implementation:\nimport torch\nfrom peft import AutoPeftModelForCausalLM\nfrom transformers import AutoTokenizer, pipeline\n\n# Load Model with PEFT adapter\ntokenizer = AutoTokenizer.from_pretrained(finetune_name)\nmodel = AutoPeftModelForCausalLM.from_pretrained(\n finetune_name, device_map=\"auto\", torch_dtype=torch.float16\n)\npipe = pipeline(", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "420-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Test Model and run Inference" + }, + { + "text": "PEFT adapter\ntokenizer = AutoTokenizer.from_pretrained(finetune_name)\nmodel = AutoPeftModelForCausalLM.from_pretrained(\n finetune_name, device_map=\"auto\", torch_dtype=torch.float16\n)\npipe = pipeline(\n \"text-generation\", model=merged_model, tokenizer=tokenizer, device=device\n)\n\nExplanation:\nLets test some prompt samples and see how the model performs.\n\nPython implementation:\nprompts = [\n \"What is the capital of Germany? Explain why thats the case and if it was different in the past?\",\n \"Write a Python function to calculate the factorial of a number.\",\n \"A rectangular garden has a length of 25 feet and a width of 15 feet. If you want to build a fence around the entire garden, how many feet of fencing will you need?\",\n \"What is the difference between a fruit and a vegetable? Give examples of each.\",\n]\n\ndef test_inference(prompt):\n prompt = pipe.tokenizer.apply_chat_template(\n [{\"role\": \"user\", \"content\": prompt}],\n tokenize=False,\n add_generation_prompt=True,\n )\n outputs = pipe(\n prompt,", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "420-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Test Model and run Inference" + }, + { + "text": "each.\",\n]\n\ndef test_inference(prompt):\n prompt = pipe.tokenizer.apply_chat_template(\n [{\"role\": \"user\", \"content\": prompt}],\n tokenize=False,\n add_generation_prompt=True,\n )\n outputs = pipe(\n prompt,\n )\n return outputs[0][\"generated_text\"][len(prompt) :].strip()\n\nfor prompt in prompts:\n print(f\" prompt:\\n{prompt}\")\n print(f\" response:\\n{test_inference(prompt)}\")\n print(\"-\" * 50)", + "source": "6.3.lora-fine-tuning.ipynb", + "file_type": "ipynb", + "chunk_id": "420-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/6.3.lora-fine-tuning.ipynb", + "section": "3. Test Model and run Inference" + }, + { + "text": "# Module 6 – Fine-Tuning Large Language Models\n\nThis module explores modern techniques for adapting pretrained Large Language Models (LLMs) to downstream tasks using the Hugging Face ecosystem. It covers chat templates, supervised fine-tuning (SFT), and parameter-efficient fine-tuning with Low-Rank Adaptation (LoRA).\n\n## Learning Objectives\n\n- Understand chat templates for conversational LLMs\n- Prepare datasets for supervised fine-tuning\n- Fine-tune pretrained language models using TRL\n- Apply LoRA for parameter-efficient fine-tuning\n- Compare full fine-tuning and LoRA-based approaches\n- Publish fine-tuned models to the Hugging Face Hub\n\n---\n\n## Topics Covered\n\n- Chat Templates\n- Conversational datasets\n- Supervised Fine-Tuning (SFT)\n- TRL SFTTrainer\n- PEFT\n- LoRA\n- Model inference\n- Hugging Face Hub\n\n---\n\n## Notebooks\n\n| Notebook | Description |\n|----------|-------------|\n| `01-chat-templates.ipynb` | Formatting conversational data using chat templates |", + "source": "README.md", + "file_type": "md", + "chunk_id": "421-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/README.md", + "section": null + }, + { + "text": "LoRA\n- Model inference\n- Hugging Face Hub\n\n---\n\n## Notebooks\n\n| Notebook | Description |\n|----------|-------------|\n| `01-chat-templates.ipynb` | Formatting conversational data using chat templates |\n| `02-supervised-fine-tuning.ipynb` | Fine-tuning LLMs using the TRL `SFTTrainer` |\n| `03-lora-fine-tuning.ipynb` | Parameter-efficient fine-tuning using LoRA adapters |\n\n---\n\n## Skills Demonstrated\n\n- Hugging Face Transformers\n- TRL\n- PEFT\n- LoRA\n- Chat Templates\n- Supervised Fine-Tuning\n- Large Language Models\n- Hugging Face Hub\n\n---\n\n## Requirements\n\n```bash\npip install -r ../requirements.txt\n```", + "source": "README.md", + "file_type": "md", + "chunk_id": "421-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "06-fine-tuningLLMs/README.md", + "section": null + }, + { + "text": "# Reasoning Large Language Models with GRPO (Group Relative Policy Optimization)\n\nThis project demonstrates how **Reasoning Large Language Models (Reasoning LLMs)** can be fine-tuned using **Group Relative Policy Optimization (GRPO)**, a reinforcement learning algorithm introduced by **DeepSeek R1**.\n\nUnlike traditional supervised fine-tuning, GRPO teaches an LLM to improve its reasoning by generating **multiple candidate solutions**, evaluating them with **reward functions**, and reinforcing better reasoning paths through reinforcement learning. This notebook follows the practical workflow introduced in the Hugging Face **Open R1** course and implements GRPO using the **TRL** library. :contentReference[oaicite:0]{index=0}\n\n---\n\n# Motivation\n\nPretrained Large Language Models generate fluent text because they are optimized to predict the next token. However, they often struggle with problems requiring:\n\n- Multi-step mathematical reasoning\n- Logical reasoning\n- Planning", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "ls generate fluent text because they are optimized to predict the next token. However, they often struggle with problems requiring:\n\n- Multi-step mathematical reasoning\n- Logical reasoning\n- Planning\n- Self-correction\n- Complex coding tasks\n\nRecent reasoning models such as **DeepSeek R1** demonstrated that Reinforcement Learning can significantly improve reasoning capabilities after pretraining.\n\nInstead of learning from labeled examples only, the model learns through **trial and error**, receiving rewards for producing better reasoning processes.\n\n---\n\n# What is Reinforcement Learning?\n\nReinforcement Learning (RL) trains an agent through interaction with an environment.\n\nThe basic RL loop consists of:\n\n1. Observe the environment\n2. Perform an action\n3. Receive a reward\n4. Improve the policy\n5. Repeat\n\nFor language models:\n\n- **Agent:** LLM\n- **Environment:** Prompt and reward function\n- **Action:** Generated response\n- **Reward:** Quality score\n- **Policy:** Model parameters", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "the policy\n5. Repeat\n\nFor language models:\n\n- **Agent:** LLM\n- **Environment:** Prompt and reward function\n- **Action:** Generated response\n- **Reward:** Quality score\n- **Policy:** Model parameters\n\nRather than predicting only the next token, the model gradually learns to generate responses that maximize rewards.\n\n---\n\n# Reinforcement Learning from Human Feedback (RLHF)\n\nOne of the most successful approaches for aligning LLMs is **RLHF**.\n\nThe overall process is illustrated below.\n\n

\n\n

\n\nThe workflow consists of four stages:\n\n1. Human evaluators compare multiple responses.\n2. Their preferences become reward signals.\n3. A reward model (or reward function) learns these preferences.\n4. Reinforcement learning updates the LLM to produce responses humans prefer.\n\nThis allows the model to become:\n\n- More helpful\n- More accurate\n- Better aligned with human expectations\n- Less likely to generate undesirable responses\n\n---\n\n# Why GRPO?", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "responses humans prefer.\n\nThis allows the model to become:\n\n- More helpful\n- More accurate\n- Better aligned with human expectations\n- Less likely to generate undesirable responses\n\n---\n\n# Why GRPO?\n\nSeveral reinforcement learning algorithms have been proposed for LLM alignment.\n\n| Method | Description |\n|---------|-------------|\n| PPO | Uses a policy model and a separate value model (critic). |\n| DPO | Learns directly from preference pairs without reinforcement learning. |\n| **GRPO** | Compares multiple responses within the same group and optimizes using relative rewards. |\n\nGRPO offers several advantages:\n\n- No separate value (critic) model\n- Lower GPU memory requirements\n- Stable optimization\n- Better reasoning performance\n- Flexible reward functions\n- Simple implementation using Hugging Face TRL\n\n---\n\n# How GRPO Works\n\nInstead of generating one answer, the model generates **multiple candidate solutions** for every prompt.\n\n

\n\n

", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "ng Face TRL\n\n---\n\n# How GRPO Works\n\nInstead of generating one answer, the model generates **multiple candidate solutions** for every prompt.\n\n

\n\n

\n\nThe training process is:\n\n### Step 1 — Generate Multiple Responses\n\nGiven a prompt:\n\n> Solve 24 × 17\n\nThe model may generate several different reasoning paths.\n\nExample:\n\n- Step-by-step solution\n- Shortcut calculation\n- Incorrect reasoning\n- Alternative approach\n\nThese responses form one **group**.\n\n---\n\n### Step 2 — Evaluate Each Response\n\nEach response is evaluated using one or more reward functions.\n\nPossible evaluation criteria include:\n\n- Mathematical correctness\n- Logical consistency\n- Proper formatting\n- Helpfulness\n- Response length\n- XML/JSON structure\n\nInstead of comparing responses independently, GRPO compares them **relative to each other**.\n\nResponses better than the group average receive positive advantages.\n\nPoor responses receive negative advantages.\n\n---", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "paring responses independently, GRPO compares them **relative to each other**.\n\nResponses better than the group average receive positive advantages.\n\nPoor responses receive negative advantages.\n\n---\n\n### Step 3 — Update the Model\n\nThe policy is updated so that:\n\n- Good reasoning becomes more likely\n- Poor reasoning becomes less likely\n\nA KL-divergence penalty prevents the model from changing too aggressively, maintaining training stability.\n\n---\n\n# Reward Functions\n\nThe reward function is one of the most important components of GRPO.\n\n

\n\n

\n\nA reward function can evaluate multiple aspects simultaneously.\n\nExamples include:\n\n### ✔ Correctness\n\nDid the model produce the correct answer?\n\nExample:\n\n```\n24 × 17 = 408\n```\n\nCorrect answer → reward = 1\n\nIncorrect answer → reward = 0\n\n---\n\n### ✔ Formatting\n\nDid the response follow the required format?\n\nExample:\n\n```\n\n...\n\n\n\n...\n\n```\n\n---\n\n### ✔ Helpfulness", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "reward = 1\n\nIncorrect answer → reward = 0\n\n---\n\n### ✔ Formatting\n\nDid the response follow the required format?\n\nExample:\n\n```\n\n...\n\n\n\n...\n\n```\n\n---\n\n### ✔ Helpfulness\n\nDoes the reasoning clearly explain the solution?\n\n---\n\n### ✔ Response Length\n\nIs the answer too short or unnecessarily long?\n\n---\n\nMultiple reward functions can be combined into a single reward score used during reinforcement learning.\n\n---\n\n# Group Relative Advantage\n\nGRPO computes a **relative advantage** for each response:\n\n```\nAdvantage =\n(reward − mean(group rewards))\n/\nstd(group rewards)\n```\n\nThis means the model learns from **relative quality** rather than absolute scores.\n\nResponses better than the average are reinforced.\n\nResponses worse than average become less likely.\n\n---\n\n# DeepSeek R1 Training Pipeline\n\nDeepSeek R1 was trained in several stages:\n\n1. Cold Start Supervised Fine-Tuning\n2. Reinforcement Learning for reasoning\n3. Rejection Sampling\n4.", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "me less likely.\n\n---\n\n# DeepSeek R1 Training Pipeline\n\nDeepSeek R1 was trained in several stages:\n\n1. Cold Start Supervised Fine-Tuning\n2. Reinforcement Learning for reasoning\n3. Rejection Sampling\n4. Diverse Reinforcement Learning\n\nOne of the most interesting discoveries during training was the **\"Aha Moment\"**, where the model spontaneously learned to:\n\n- detect mistakes,\n- reconsider intermediate reasoning,\n- self-correct before producing a final answer.\n\n---\n\n# Notebook Contents\n\nThis notebook covers:\n\n- Introduction to Reinforcement Learning\n- Reinforcement Learning from Human Feedback (RLHF)\n- DeepSeek R1\n- Group Relative Policy Optimization (GRPO)\n- Reward Functions\n- TRL GRPOTrainer\n- LoRA Fine-Tuning\n- Fine-tuning SmolLM\n- Weights & Biases experiment tracking\n- Publishing models to Hugging Face Hub\n- Text generation using the fine-tuned model\n\n---\n\n# Technologies Used\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- TRL\n- PEFT (LoRA)\n- Hugging Face Datasets\n- Accelerate", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "els to Hugging Face Hub\n- Text generation using the fine-tuned model\n\n---\n\n# Technologies Used\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- TRL\n- PEFT (LoRA)\n- Hugging Face Datasets\n- Accelerate\n- BitsAndBytes\n- Flash Attention\n- Weights & Biases\n\n---\n\n# Learning Outcomes\n\nAfter completing this notebook you will understand how to:\n\n- Explain Reinforcement Learning for LLMs.\n- Understand the DeepSeek R1 training strategy.\n- Implement GRPO using Hugging Face TRL.\n- Design custom reward functions.\n- Fine-tune LLMs using LoRA.\n- Monitor reinforcement learning experiments.\n- Publish trained models to the Hugging Face Hub.\n- Generate text using a GRPO fine-tuned model.\n\n---\n\n# References\n\n- Hugging Face Open R1 Course\n- DeepSeek R1 Paper\n- Hugging Face TRL Documentation\n- Hugging Face Transformers Documentation", + "source": "README.md", + "file_type": "md", + "chunk_id": "422-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/README.md", + "section": null + }, + { + "text": "Explanation:\n# Fine-Tuning Reasoning LLMs with Group Relative Policy Optimization (GRPO)\n\n## Overview\n\nThis notebook demonstrates how to fine-tune a reasoning-capable Large Language Model using **Group Relative Policy Optimization (GRPO)**, a reinforcement learning algorithm designed to improve multi-step reasoning without requiring a separate value model.\n\nThe implementation uses the Hugging Face ecosystem, including **TRL**, **Transformers**, **Datasets**, and **PEFT**, to train a compact language model with custom reward functions.\n\nThe workflow follows the same principles used in modern reasoning models such as **DeepSeek R1**.\n\nPython implementation:\nfrom huggingface_hub import notebook_login\n\nnotebook_login()", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "423-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Fine-Tuning Reasoning LLMs with Group Relative Policy Optimization (GRPO)" + }, + { + "text": "Explanation:\n# Training Workflow\n\nThe GRPO training pipeline consists of the following stages:\n\n1. Load a reasoning dataset.\n2. Initialize a pretrained language model.\n3. Define reward functions that evaluate generated responses.\n4. Configure the GRPO training algorithm.\n5. Fine-tune the model using reinforcement learning.\n6. Evaluate the resulting reasoning model.\n\n## Dataset\n\nThe notebook uses a reasoning dataset consisting of prompts and expected responses. During training, the model generates multiple candidate solutions for each prompt, which are evaluated using custom reward functions rather than supervised labels.\n\nPython implementation:\nimport torch\nimport wandb\nfrom datasets import load_dataset\nfrom peft import LoraConfig, get_peft_model\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\nfrom trl import GRPOConfig, GRPOTrainer\n\n# Log to Weights & Biases\nwandb.login()\n\n# Load dataset\ndataset = load_dataset(\"mlabonne/smoltldr\")\nprint(dataset)", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "424-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Training Workflow" + }, + { + "text": "Explanation:\n## Loading the Base Model\n\nA pretrained language model serves as the initialization for reinforcement learning. Instead of training from scratch, GRPO improves the model's reasoning capabilities through reward-driven optimization.\n\nPython implementation:\n# Load model\nmodel_id = \"HuggingFaceTB/SmolLM-135M-Instruct\"\nmodel = AutoModelForCausalLM.from_pretrained(\n model_id,\n torch_dtype=\"auto\",\n device_map=\"auto\",\n)\ntokenizer = AutoTokenizer.from_pretrained(model_id)\n\n# Load LoRA\nlora_config = LoraConfig(\n task_type=\"CAUSAL_LM\",\n r=16,\n lora_alpha=32,\n target_modules=\"all-linear\",\n)\nmodel = get_peft_model(model, lora_config)\nprint(model.print_trainable_parameters())", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "425-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Loading the Base Model" + }, + { + "text": "Explanation:\n## Reward Functions\n\n

\n\n

\n\nReward functions evaluate the quality of generated responses and provide the learning signal for reinforcement learning. Multiple reward functions can be combined to encourage correctness, proper formatting, helpful reasoning, and other desired behaviors.\n\nPython implementation:\n# Reward function\ndef reward_len(completions, **kwargs):\n return [-abs(50 - len(completion)) for completion in completions]", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "426-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Reward Functions" + }, + { + "text": "Explanation:\n## Group Relative Policy Optimization (GRPO)\n\n

\n\n

\n\nUnlike traditional supervised learning, GRPO generates multiple responses for the same prompt and evaluates them as a group. Responses receiving higher relative rewards are reinforced during optimization, allowing the model to progressively improve its reasoning strategies.\n\nPython implementation:\n# Training arguments\ntraining_args = GRPOConfig(\n output_dir=\"GRPO\",\n learning_rate=2e-5,\n per_device_train_batch_size=2,\n gradient_accumulation_steps=2,\n max_prompt_length=512,\n max_completion_length=96,\n num_generations=4,\n #optim=\"adamw_8bit\",\n optim=\"adamw_torch_fused\",\n num_train_epochs=1,\n bf16=True,\n report_to=[\"wandb\"],\n remove_unused_columns=False,\n logging_steps=1,\n)\n\n# Trainer\ntrainer = GRPOTrainer(\n model=model,\n reward_funcs=[reward_len],\n args=training_args,\n train_dataset=dataset[\"train\"],\n)\n\n# Train model\nwandb.init(project=\"GRPO\")\ntrainer.train()", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "427-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Group Relative Policy Optimization (GRPO)" + }, + { + "text": "Explanation:\n## Push Model to Hub\n\nPython implementation:\n# Save model\nmerged_model = trainer.model.merge_and_unload()\nmerged_model.push_to_hub(\n \"Miladsaeedi70/smollm-135m-instruct-grpoV2\",\n private=False)", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "428-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Push Model to Hub" + }, + { + "text": "Explanation:\n## Inference\n\nAfter training, the fine-tuned model is evaluated by generating responses for unseen prompts. Comparing these outputs with those of the base model provides insight into how reinforcement learning improves reasoning quality.\n\n## Generate Text\n\nPython implementation:\nprompt = \"\"\"\n# A long document about the Cat\n\nThe cat (Felis catus), also referred to as the domestic cat or house cat, is a small\ndomesticated carnivorous mammal. It is the only domesticated species of the family Felidae.\nAdvances in archaeology and genetics have shown that the domestication of the cat occurred\nin the Near East around 7500 BC. It is commonly kept as a pet and farm cat, but also ranges\nfreely as a feral cat avoiding human contact. It is valued by humans for companionship and\nits ability to kill vermin. Its retractable claws are adapted to killing small prey species\nsuch as mice and rats. It has a strong, flexible body, quick reflexes, and sharp teeth,", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "429-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Inference" + }, + { + "text": "r companionship and\nits ability to kill vermin. Its retractable claws are adapted to killing small prey species\nsuch as mice and rats. It has a strong, flexible body, quick reflexes, and sharp teeth,\nand its night vision and sense of smell are well developed. It is a social species,\nbut a solitary hunter and a crepuscular predator. Cat communication includes\nvocalizations—including meowing, purring, trilling, hissing, growling, and grunting—as\nwell as body language. It can hear sounds too faint or too high in frequency for human ears,\nsuch as those made by small mammals. It secretes and perceives pheromones.\n\"\"\"\n\nmessages = [\n {\"role\": \"user\", \"content\": prompt},\n]\n\nPython implementation:\n# Generate text\nfrom transformers import pipeline\n\ngenerator = pipeline(\"text-generation\", model=\"Miladsaeedi70/smollm-135m-instruct-grpoV2\")\n\n## Or use the model and tokenizer we defined earlier\n# generator = pipeline(\"text-generation\", model=model, tokenizer=tokenizer)\n\ngenerate_kwargs = {", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "429-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Inference" + }, + { + "text": "del=\"Miladsaeedi70/smollm-135m-instruct-grpoV2\")\n\n## Or use the model and tokenizer we defined earlier\n# generator = pipeline(\"text-generation\", model=model, tokenizer=tokenizer)\n\ngenerate_kwargs = {\n \"max_new_tokens\": 256,\n \"do_sample\": True,\n \"temperature\": 0.5,\n \"min_p\": 0.1,\n}\n\ngenerated_text = generator(\n messages,\n **generate_kwargs\n)\n\nprint(generated_text)", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "429-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Inference" + }, + { + "text": "Explanation:\n# Results\n\nThe GRPO fine-tuning pipeline successfully demonstrates how reinforcement learning can improve the reasoning behavior of a pretrained language model.\n\n### Key Components\n\n- Pretrained language model\n- LoRA parameter-efficient fine-tuning\n- Custom reward functions\n- GRPO optimization\n- Reinforcement learning with TRL\n- Hugging Face Hub integration\n\n### Learning Outcomes\n\n- Fine-tuned an LLM using reinforcement learning.\n- Designed custom reward functions for reasoning tasks.\n- Applied parameter-efficient adaptation using LoRA.\n- Generated responses using a GRPO-trained reasoning model.", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "430-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Results" + }, + { + "text": "Explanation:\n# Conclusion\n\nThis notebook demonstrated the complete workflow for fine-tuning a reasoning language model using Group Relative Policy Optimization (GRPO). By combining pretrained language models, custom reward functions, and reinforcement learning, GRPO enables efficient post-training that improves reasoning capabilities without requiring a separate value model. These techniques form an important part of modern reasoning-oriented LLM development.", + "source": "grpo-finetuning.ipynb", + "file_type": "ipynb", + "chunk_id": "431-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "07-reasoning-models/grpo-finetuning.ipynb", + "section": "Conclusion" + }, + { + "text": "# 🤗 Hugging Face LLM Course Portfolio\n\nA collection of practical implementations completed while following the **Hugging Face Large Language Model (LLM) Course**. This repository has been reorganized into a structured portfolio that demonstrates modern Natural Language Processing (NLP), Transformer models, and Large Language Model (LLM) workflows using the Hugging Face ecosystem.\n\n![Transformer Architecture](01-transformers/transformers_architecture.png)\n\n---\n\n## 🚀 Repository Overview\n\nThis repository documents my hands-on implementation of the Hugging Face LLM Course. Each notebook has been reorganized, documented, and expanded with additional explanations, comments, and examples to create a professional learning portfolio.\n\n---\n\n## 📚 Topics Covered\n\n- Transformer architectures\n- Hugging Face pipelines\n- Transformer model training\n- Fast tokenizers\n- WordPiece, Byte-Pair Encoding (BPE), and Unigram tokenization\n- Dataset preprocessing\n- Question Answering (QA)", + "source": "README.md", + "file_type": "md", + "chunk_id": "432-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "README.md", + "section": null + }, + { + "text": "architectures\n- Hugging Face pipelines\n- Transformer model training\n- Fast tokenizers\n- WordPiece, Byte-Pair Encoding (BPE), and Unigram tokenization\n- Dataset preprocessing\n- Question Answering (QA)\n- Named Entity Recognition (NER)\n- Machine Translation\n- Text Summarization\n- Masked Language Modeling\n- Causal Language Modeling\n- Chat Templates\n- Supervised Fine-Tuning (SFT)\n- Parameter-Efficient Fine-Tuning (LoRA)\n- Reasoning Models\n- Group Relative Policy Optimization (GRPO)\n- Reward Functions\n- Reinforcement Learning for LLMs\n---\n\n## 🛠️ Technologies\n\n- Python\n- PyTorch\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Tokenizers\n- Hugging Face Evaluate\n- Hugging Face Hub\n- Accelerate\n- TRL\n- PEFT\n---\n\n## 📂 Repository Structure\n\n```text\nhuggingface-llm-course/\n│\n├── README.md\n├── requirements.txt\n├── .gitignore\n│\n├── 01-transformers/\n├── 02-transformer-workflow/\n├── 03-model-training/\n├── 04-tokenizers/\n├── 05-nlp-applications/\n│ ├── token-classification", + "source": "README.md", + "file_type": "md", + "chunk_id": "432-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "README.md", + "section": null + }, + { + "text": "/\n│\n├── README.md\n├── requirements.txt\n├── .gitignore\n│\n├── 01-transformers/\n├── 02-transformer-workflow/\n├── 03-model-training/\n├── 04-tokenizers/\n├── 05-nlp-applications/\n│ ├── token-classification\n│ ├── masked-language-modeling\n│ ├── machine-translation\n│ ├── text-summarization\n│ └── causal-language-modeling\n├── 06-fine-tuningLLMs/\n│ ├── chat-templates/\n│ ├── supervised-fine-tuning\n│ └── lora-fine-tuning\n│\n└── 07-reasoning-models/\n └── grpo-finetuning\n```\n\n---\n\n## 📖 Modules\n\n| Module | Description | Status |\n|---------|-------------|:------:|\n| **01 – Transformers** | Introduction to Transformer models and the Hugging Face `pipeline` API | ✅ |\n| **02 – Transformer Workflow** | Understanding the internal workflow of tokenizers, models, and inference pipelines | ✅ |\n| **03 – Model Training** | Dataset preprocessing and supervised fine-tuning using the Hugging Face `Trainer` API | ✅ |", + "source": "README.md", + "file_type": "md", + "chunk_id": "432-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "README.md", + "section": null + }, + { + "text": "ding the internal workflow of tokenizers, models, and inference pipelines | ✅ |\n| **03 – Model Training** | Dataset preprocessing and supervised fine-tuning using the Hugging Face `Trainer` API | ✅ |\n| **04 – Tokenization** | Fast tokenizers, normalization, WordPiece, BPE, Unigram, and custom tokenizer construction | ✅ |\n| **05 – NLP Tasks** | Question Answering, Named Entity Recognition, Translation, and Summarization | ✅ |\n| **06 – LLM Fine-Tuning** | Chat templates, Supervised Fine-Tuning (SFT), and LoRA | ✅ |\n| **07 – Reasoning Models** | Fine-tuning reasoning models using GRPO, custom reward functions, and reinforcement learning for LLMs| ✅ |\n---\n\n## 🎯 Learning Objectives\n\nThroughout this repository, I explore how to:\n\n- Understand the architecture of Transformer models.\n- Build NLP applications using pretrained models.\n- Process and tokenize text efficiently.\n- Train and fine-tune Transformer models.\n- Build custom tokenizers from scratch.", + "source": "README.md", + "file_type": "md", + "chunk_id": "432-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "README.md", + "section": null + }, + { + "text": "ure of Transformer models.\n- Build NLP applications using pretrained models.\n- Process and tokenize text efficiently.\n- Train and fine-tune Transformer models.\n- Build custom tokenizers from scratch.\n- Evaluate NLP models using standard benchmarks.\n- Work effectively with the Hugging Face ecosystem.\n- Understand conversational data formatting using chat templates.\n- Apply parameter-efficient fine-tuning techniques such as LoRA.\n- Understand post-training techniques for reasoning language models.\n- Fine-tune LLMs using Group Relative Policy Optimization (GRPO).\n\n---\n\n## 📚 References\n\n- Hugging Face Large Language Model Course\n- Hugging Face Transformers\n- Hugging Face Datasets\n- Hugging Face Tokenizers\n- Hugging Face Hub\n\n---\n\n## ⭐ About This Repository\n\nThis repository is part of my continuous learning journey in **Large Language Models (LLMs)**, **Generative AI**, and **Natural Language Processing (NLP)**.", + "source": "README.md", + "file_type": "md", + "chunk_id": "432-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "README.md", + "section": null + }, + { + "text": "Face Hub\n\n---\n\n## ⭐ About This Repository\n\nThis repository is part of my continuous learning journey in **Large Language Models (LLMs)**, **Generative AI**, and **Natural Language Processing (NLP)**. Rather than simply reproducing the Hugging Face course notebooks, I reorganized each module into a structured technical portfolio with improved documentation, code comments, and practical implementations. The repository now spans the complete LLM workflow—from Transformer fundamentals and NLP applications to supervised fine-tuning, parameter-efficient adaptation, and reinforcement learning techniques for reasoning models.", + "source": "README.md", + "file_type": "md", + "chunk_id": "432-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "huggingface-llm-exercises", + "relative_path": "README.md", + "section": null + }, + { + "text": "# Natural Language Processing with PyTorch\n\nA collection of practical **Natural Language Processing (NLP)** projects implemented using **PyTorch**. This repository explores word embeddings, recurrent neural networks, language modeling, text generation, and sequence classification through hands-on implementations of Word2Vec, GloVe, RNNs, LSTMs, and generative language models. Each notebook has been reorganized, documented, and expanded to demonstrate practical deep learning workflows for NLP.\n\n---\n\n# 🚀 Repository Overview\n\nThis repository contains practical Natural Language Processing projects implemented in **PyTorch**, covering distributed word representations, sentiment analysis, recurrent neural networks, generative language modeling, and transfer learning for text classification.\n\n

\n \n

\n\n---\n\n# ✨ Project Highlights\n\n- Word Embeddings (Word2Vec & GloVe)\n- Semantic Similarity & Word Analogies\n- Bias in Word Embeddings", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "p align=\"center\">\n \n

\n\n---\n\n# ✨ Project Highlights\n\n- Word Embeddings (Word2Vec & GloVe)\n- Semantic Similarity & Word Analogies\n- Bias in Word Embeddings\n- Recurrent Neural Networks (RNNs)\n- Long Short-Term Memory (LSTM) Networks\n- Sentiment Analysis\n- Character-Level Language Modeling\n- Text Generation\n- Sequence Modeling\n- TorchText\n- Well-documented notebooks suitable for learning and experimentation\n- Transfer Learning\n\n---\n\n# 📚 Topics Covered\n\n- Natural Language Processing (NLP)\n- Word Embeddings (Word2Vec & GloVe)\n- Semantic Similarity\n- Word Analogies\n- Bias in Word Embeddings\n- Recurrent Neural Networks (RNNs)\n- Long Short-Term Memory (LSTMs)\n- Sequence Modeling\n- Character-Level Language Models\n- Text Generation\n- Sentiment Analysis\n- Text Classification\n- Teacher Forcing\n- Variable-Length Sequences\n- TorchText\n- Deep Learning with PyTorch\n- Transfer Learning\n\n---\n\n# 🛠️ Technologies\n\n- Python\n- PyTorch\n- TorchText\n- Torchvision", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "- Text Classification\n- Teacher Forcing\n- Variable-Length Sequences\n- TorchText\n- Deep Learning with PyTorch\n- Transfer Learning\n\n---\n\n# 🛠️ Technologies\n\n- Python\n- PyTorch\n- TorchText\n- Torchvision\n- NumPy\n- Matplotlib\n- Scikit-learn\n\n---\n\n# 📂 Repository Structure\n\n```text\nnatural-language-processing-pytorch/\n│\n├── README.md\n├── requirements.txt\n├── .gitignore\n│\n├���─ notebooks/\n│ ├── 01_word-embeddings-and-rnn.ipynb\n│ ├── 02_generative_rnn.ipynb\n│ └── 03_spam_detection_lstm.ipynb\n│\n├── Images/\n│ └── RNN_LSTM.jpg\n│\n└── models/\n```\n\n---\n\n# 📖 Notebooks\n\n## Notebook 1 — Word Embeddings and Recurrent Neural Networks\n\n### Overview\n\nThis notebook introduces the foundations of Natural Language Processing using distributed word representations and recurrent neural networks. It begins by exploring **Word2Vec** and **GloVe** embeddings to understand semantic relationships between words, including similarity, analogies, and embedding bias.", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "ons and recurrent neural networks. It begins by exploring **Word2Vec** and **GloVe** embeddings to understand semantic relationships between words, including similarity, analogies, and embedding bias. It then builds an end-to-end sentiment analysis pipeline by combining pretrained word embeddings with a recurrent neural network for tweet classification.\n\n### Topics Covered\n\n- Word2Vec\n- GloVe embeddings\n- Cosine similarity\n- Semantic similarity\n- Word analogies\n- Bias in word embeddings\n- Sentiment analysis\n- Recurrent Neural Networks (RNNs)\n- Variable-length sequences\n- Tweet sentiment classification\n- PyTorch implementation\n\n### Notebook\n\n`notebooks/01_word-embeddings-and-rnn.ipynb`\n\n---\n\n## Notebook 2 — Generative Recurrent Neural Networks\n\n### Overview\n\nThis notebook extends recurrent neural networks to character-level language modeling and text generation.", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "dings-and-rnn.ipynb`\n\n---\n\n## Notebook 2 — Generative Recurrent Neural Networks\n\n### Overview\n\nThis notebook extends recurrent neural networks to character-level language modeling and text generation. It demonstrates autoregressive sequence generation, teacher forcing during training, probability-based token sampling, and GPU-accelerated training for generative RNNs.\n\n### Topics Covered\n\n- Character-level language modeling\n- Generative RNNs\n- Text generation\n- Teacher forcing\n- Token sampling\n- Hidden state propagation\n- Sequence generation\n- Autoregressive generation\n- Temperature-based sampling\n- Teacher forcing\n- Hidden state propagation\n- Character embeddings\n- Sequence modeling\n- PyTorch implementation\n\n### Notebook\n\n`notebooks/02_generative_rnn.ipynb`\n\n---\n\n## Notebook 3 — SMS Spam Detection with LSTM Networks\n\n### Overview\n\nThis notebook presents two deep learning approaches for SMS spam detection using PyTorch.", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "`notebooks/02_generative_rnn.ipynb`\n\n---\n\n## Notebook 3 — SMS Spam Detection with LSTM Networks\n\n### Overview\n\nThis notebook presents two deep learning approaches for SMS spam detection using PyTorch. The first implements a Long Short-Term Memory (LSTM) network for sequence classification, while the second applies transfer learning with a pretrained ULMFiT language model. Together, these approaches demonstrate both training recurrent neural networks from scratch and adapting pretrained language models for text classification.\n\n

\n \n

\n\n### Topics Covered\n\n- SMS spam detection\n- Long Short-Term Memory (LSTM)\n- ULMFiT transfer learning\n- Text preprocessing\n- Tokenization\n- Word embeddings\n- Sequence classification\n- Language model fine-tuning\n- Model evaluation\n- Deep learning with PyTorch\n\n### Notebook\n\n`notebooks/03_spam_detection_lstm.ipynb`\n\n---\n\n# 🎯 Learning Objectives\n\nThroughout this repository, I explore how to:", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "model fine-tuning\n- Model evaluation\n- Deep learning with PyTorch\n\n### Notebook\n\n`notebooks/03_spam_detection_lstm.ipynb`\n\n---\n\n# 🎯 Learning Objectives\n\nThroughout this repository, I explore how to:\n\n- Learn distributed word representations using Word2Vec and GloVe embeddings.\n- Analyze semantic similarity and solve word analogy tasks.\n- Understand bias in pretrained word embeddings.\n- Build recurrent neural networks for NLP applications.\n- Process and batch variable-length text sequences.\n- Generate text using character-level recurrent neural networks.\n- Train LSTM models for text classification.\n- Develop complete NLP workflows using PyTorch.\n\n---\n\n# 🚀 Getting Started\n\nClone the repository:\n\n```bash\ngit clone https://github.com/Miladsaeedi70/natural-language-processing-pytorch.git\n```\n\nNavigate to the project:\n\n```bash\ncd natural-language-processing-pytorch\n```\n\nInstall the required packages:\n\n```bash\npip install -r requirements.txt\n```\n\nLaunch Jupyter Notebook:\n\n```bash", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "h.git\n```\n\nNavigate to the project:\n\n```bash\ncd natural-language-processing-pytorch\n```\n\nInstall the required packages:\n\n```bash\npip install -r requirements.txt\n```\n\nLaunch Jupyter Notebook:\n\n```bash\njupyter notebook\n```\n\nOpen the notebooks in numerical order:\n\n1. Word Embeddings and RNNs\n2. Generative RNNs\n3. SMS Spam Detection with LSTMs\n\n---\n\n# ⭐ About\n\nThis repository showcases practical Natural Language Processing projects developed with **PyTorch**, progressing from distributed word representations and recurrent neural networks to generative language modeling and transfer learning for text classification. The notebooks have been reorganized, modernized, and documented to improve readability, reproducibility, and compatibility with recent versions of PyTorch and Python.\n\n---\n\n# 📄 License\n\nThis repository is intended for educational and portfolio purposes.", + "source": "README.md", + "file_type": "md", + "chunk_id": "433-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "README.md", + "section": null + }, + { + "text": "Explanation:\n# Word Embeddings and Recurrent Neural Networks with PyTorch\n\n## Overview\n\nThis notebook introduces modern Natural Language Processing (NLP) techniques using PyTorch. It begins by exploring distributed word representations through **Word2Vec** and **GloVe** embeddings, demonstrating how semantic relationships between words can be captured in vector space. It then applies pretrained embeddings to sentiment analysis and concludes by building a recurrent neural network (RNN) for tweet sentiment classification.\n\nThe notebook progresses from understanding word representations to developing an end-to-end sequence classification pipeline using recurrent neural networks.\n\nThis tutorial is split into three parts. Parts A and B focus on Word2Vec and GloVe word embeddings. Sentiment analysis using GloVe embeddings with a simple ANN classifier is the focus of Part B.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "434-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Word Embeddings and Recurrent Neural Networks with PyTorch" + }, + { + "text": "rks.\n\nThis tutorial is split into three parts. Parts A and B focus on Word2Vec and GloVe word embeddings. Sentiment analysis using GloVe embeddings with a simple ANN classifier is the focus of Part B. Part C introduces recurrent neural networks (RNNs) for sentiment analysis and provides sample code for batching of data.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "434-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Word Embeddings and Recurrent Neural Networks with PyTorch" + }, + { + "text": "Explanation:\n#Part A: Word2Vec and GloVe Vectors\nWe saw how autoencoders are used to learn a latent\n**embedding space**: an alternative, low-dimensional representation\nof a set of data with some appealing properties:\nfor example, we saw that interpolating in the latent space\nis a way of generating new examples. In particular,\ninterpolation in the latent space generates more compelling\nexamples than, say, interpolating in the raw pixel space.\n\nThe idea of learning an alternative representation/features/*embeddings* of data\nis a prevalent one in machine learning. Good representations will\nmake downstream tasks (like generating new data, clustering, computing distances) perform much better.\n\nWith autoencoders, we were able to learn a representation of MNIST digits by using a model that looks like this:\n\n- **Encoder**: data -> embedding\n- **Decoder**: embedding -> data\n\nThis type of architecture works well for certain types of data (e.g. images)", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "435-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "word2vec models" + }, + { + "text": "of MNIST digits by using a model that looks like this:\n\n- **Encoder**: data -> embedding\n- **Decoder**: embedding -> data\n\nThis type of architecture works well for certain types of data (e.g. images)\nthat are easy to generate, and whose meaning is encoded in the input data\nrepresentation (e.g. the pixels).\n\nBut what if we want to train an embedding on words? Words are different\nfrom images, in that the meaning of a word is not represented\nby the letters that make up the word (the same way that the meaning\nof an image is represented by the pixels that make up the pixel). Instead,\nthe meaning of words comes from how they are used in conjunction with other\nwords.\n\n## word2vec models\n\nA word2vec model learns embedding of words using the following architecture:\n\n- **Encoder**: word -> embedding\n- **Decoder**: embedding -> nearby words (context)\n\nSpecific word2vec models differ in the which \"nearby words\" is predicted\nusing the decoder: is it the 3 context words that appeared *before*", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "435-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "word2vec models" + }, + { + "text": "edding\n- **Decoder**: embedding -> nearby words (context)\n\nSpecific word2vec models differ in the which \"nearby words\" is predicted\nusing the decoder: is it the 3 context words that appeared *before*\nthe input word? Is it the 3 words that appeared *after*? Or is it a combination\nof the two words that appeared before and two words that appeared after\nthe input word?\n\nThese models are trained using a large corpus of text: for example the whole\nof Wikipedia or a large collection of news articles. We won't train our\nown word2vec models in this course, so we won't talk about the many considerations involved in training a word2vec model.\n\nInstead, we will use a set of pre-trained word embeddings. These are embeddings\nthat someone else took the time and computational power to train.\nOne of the most commonly-used pre-trained word embeddings are the **GloVe embeddings**.\n\nGloVe is a variation of a word2vec model. Again, the specifics of the algorithm", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "435-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "word2vec models" + }, + { + "text": "computational power to train.\nOne of the most commonly-used pre-trained word embeddings are the **GloVe embeddings**.\n\nGloVe is a variation of a word2vec model. Again, the specifics of the algorithm\nand its training will be beyond the scope of this course.\nYou should think of **GloVe embeddings** similarly to pre-trained AlexNet weights in that they \"may\" provide improvements to the representation of data.\n\nMore information about GloVe is available here: https://nlp.stanford.edu/projects/glove/\n\nUnlike AlexNet, there are several variations of GloVe embeddings. They\ndiffer in the corpus used to train the embedding, and the *size* of the embeddings.\n\n## GloVe Embeddings\n\nTo load pre-trained GloVe embeddings, we'll use a package called `torchtext`.\nThe package `torchtext` contains other useful tools for working with text\nthat we will see later in the course. The documentation for torchtext\nGloVe vectors are available at: https://torchtext.readthedocs.io/en/latest/vocab.html#glove", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "435-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "word2vec models" + }, + { + "text": "r useful tools for working with text\nthat we will see later in the course. The documentation for torchtext\nGloVe vectors are available at: https://torchtext.readthedocs.io/en/latest/vocab.html#glove\n\nWe'll begin by loading a set of GloVe embeddings. The first time you run the code below, Python will download a large file (862MB) containing the pre-trained embeddings.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "435-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "word2vec models" + }, + { + "text": "Explanation:\n## Learning Objectives\n\nBy completing this notebook, you will learn how to:\n\n- Understand distributed word representations\n- Explore Word2Vec and GloVe embeddings\n- Measure semantic similarity using cosine similarity\n- Solve word analogy tasks with vector arithmetic\n- Analyze bias in pretrained word embeddings\n- Build sentiment analysis models using pretrained embeddings\n- Implement recurrent neural networks (RNNs) in PyTorch\n- Process variable-length text sequences\n- Train sequence classification models for tweet sentiment analysis\n\nPython implementation:\nimport torch\nimport torchtext\n\n# The first time you run this will download a ~823MB file\nglove = torchtext.vocab.GloVe(name=\"6B\", # trained on Wikipedia 2014 corpus\n dim=50) # embedding size = 50\n\nExplanation:\nLet's look at what the embedding of the word \"car\" looks like:\n\nPython implementation:\n# Now this will work perfectly!\nglove['car']", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "436-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Learning Objectives" + }, + { + "text": "Explanation:\nIt is a torch tensor with dimension `(50,)`. It is difficult to determine what each\nnumber in this embedding means, if anything. However, we know that there is structure\nin this embedding space. That is, distances in this embedding space is meaningful.\n\n## Measuring Distance\n\nTo explore the structure of the embedding space, it is necessary to introduce\na notion of *distance*. You are probably already familiar with the notion\nof the **Euclidean distance**. The Euclidean distance of two vectors $x = [x_1, x_2, ... x_n]$ and\n$y = [y_1, y_2, ... y_n]$ is just the 2-norm of their difference $x - y$. We can compute\nthe Euclidean distance between $x$ and $y$:\n$\\sqrt{\\sum_i (x_i - y_i)^2}$\n\nThe PyTorch function `torch.norm` computes the 2-norm of a vector for us, so we\ncan compute the Euclidean distance between two vectors like this:\n\nPython implementation:\nx = glove['cat']\ny = glove['dog']\ntorch.norm(y - x)\n\nExplanation:", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "437-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Measuring Distance" + }, + { + "text": "mputes the 2-norm of a vector for us, so we\ncan compute the Euclidean distance between two vectors like this:\n\nPython implementation:\nx = glove['cat']\ny = glove['dog']\ntorch.norm(y - x)\n\nExplanation:\nAn alternative measure of distance is the **Cosine Similarity**.\nThe cosine similarity measures the *angle* between two vectors,\nand has the property that it only considers the *direction* of the\nvectors, not their the magnitudes.\n\nPython implementation:\nx = torch.tensor([1., 1., 1.]).unsqueeze(0)\ny = torch.tensor([2., 2., -2.]).unsqueeze(0)\ntorch.cosine_similarity(x, y)\n\nExplanation:\nThe cosine similarity is a *similarity* measure rather than a *distance* measure:\nThe larger the similarity,\nthe \"closer\" the word embeddings are to each other.\n\nPython implementation:\nx = glove['cat']\ny = glove['dog']\ntorch.cosine_similarity(x.unsqueeze(0), y.unsqueeze(0))\n\nPython implementation:\ntorch.cosine_similarity(glove['good'].unsqueeze(0),\n glove['bad'].unsqueeze(0))\n\nPython implementation:", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "437-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Measuring Distance" + }, + { + "text": "= glove['dog']\ntorch.cosine_similarity(x.unsqueeze(0), y.unsqueeze(0))\n\nPython implementation:\ntorch.cosine_similarity(glove['good'].unsqueeze(0),\n glove['bad'].unsqueeze(0))\n\nPython implementation:\ntorch.cosine_similarity(glove['good'].unsqueeze(0),\n glove['well'].unsqueeze(0))\n\nPython implementation:\ntorch.cosine_similarity(glove['good'].unsqueeze(0),\n glove['perfect'].unsqueeze(0))\n\nPython implementation:\ntorch.cosine_similarity(glove['good'].unsqueeze(0),\n glove['bravo'].unsqueeze(0))", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "437-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Measuring Distance" + }, + { + "text": "Explanation:\n## Word Similarity\n\nNow that we have a notion of distance in our embedding space, we can talk\nabout words that are \"close\" to each other in the embedding space.\nFor now, let's use Euclidean distances to look at how close various words\nare to the word \"cat\".\n\nPython implementation:\nword = 'cat'\nother = ['pet', 'dog', 'bike', 'kitten', 'puppy', 'kite', 'computer', 'neuron']\nfor w in other:\n dist = torch.norm(glove[word] - glove[w]) # euclidean distance\n print(w, float(dist))\n\nExplanation:\nIn fact, we can look through our entire vocabulary for words that are closest\nto a point in the embedding space -- for example, we can look for words\nthat are closest to another word like \"cat\".\n\nPython implementation:\ndef print_closest_words(vec, n=5):\n # 1. Compute Euclidean distances natively in PyTorch\n # (glove.vectors and vec must be on the same device)\n dists = torch.norm(glove.vectors - vec, dim=1)\n\n # 2.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "438-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Word Similarity" + }, + { + "text": "n:\ndef print_closest_words(vec, n=5):\n # 1. Compute Euclidean distances natively in PyTorch\n # (glove.vectors and vec must be on the same device)\n dists = torch.norm(glove.vectors - vec, dim=1)\n\n # 2. Use PyTorch's native sort (ascending=True because smaller distance = closer meaning)\n sorted_dists, sorted_indices = torch.sort(dists)\n\n # 3. Print the top n closest words\n # We skip index 0 because the closest word to \"cat\" is always \"cat\" itself (distance 0)\n for i in range(1, n + 1):\n idx = sorted_indices[i].item() # .item() extracts the raw Python integer\n difference = sorted_dists[i].item() # Extracts the raw distance float\n print(f\"{glove.itos[idx]:<15} {difference:.4f}\")\nprint_closest_words(glove[\"cat\"], n=10)\n\nExplanation:\nWe could also look at which words are closest to the midpoints of two words:", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "438-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Word Similarity" + }, + { + "text": "Explanation:\n## Analogies\n\nOne surprising aspect of GloVe vectors is that the *directions* in the\nembedding space can be meaningful. The structure of the GloVe vectors\ncertain analogy-like relationship like this tend to hold:\n\n$$ king - man + woman \\approx queen $$\n\nExplanation:\nWe get reasonable answers like \"queen\", \"throne\" and the name of\nour current queen.\n\nWe can likewise flip the analogy around:\n\nExplanation:\nOr, try a different but related analogies along the gender axis:\n\nExplanation:\nWe can move an embedding towards the direction of \"goodness\" or \"badness\":", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "439-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Analogies" + }, + { + "text": "Explanation:\n## Biased in Word Vectors\n\nMachine learning models have an air of \"fairness\" about them, since models\nmake decisions without human intervention. However, models can and do learn\nwhatever bias is present in the training data!\n\nGloVe vectors seems innocuous enough: they are just representations of\nwords in some embedding space. Even so, we'll show that the structure\nof the GloVe vectors encodes the everyday biases present in the texts\nthat they are trained on.\n\nWe'll start with an example analogy:\n\n$$doctor - man + woman \\approx ??$$\n\nLet's use GloVe vectors to find the answer to the above analogy:\n\nExplanation:\nThe $$doctor - man + woman \\approx nurse$$ analogy is very concerning.\nJust to verify, the same result does not appear if we flip the gender terms:\n\nExplanation:\nWe see similar types of gender bias with other professions.\n\nExplanation:\nBeyond the first result, none of the other words are even related to", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "440-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Biased in Word Vectors" + }, + { + "text": "es not appear if we flip the gender terms:\n\nExplanation:\nWe see similar types of gender bias with other professions.\n\nExplanation:\nBeyond the first result, none of the other words are even related to\nprogramming! In contrast, if we flip the gender terms, we get very\ndifferent results:\n\nExplanation:\nHere are the results for \"engineer\":", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "440-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Biased in Word Vectors" + }, + { + "text": "Explanation:\n# Part 2 — Sentiment Analysis with GloVe Embeddings\n\n## Sentiment Analysis\n\nSentiment analysis is a Natural Language Processing (NLP) task that aims to determine the emotional tone or opinion expressed in a piece of text. It has numerous real-world applications, including social media monitoring, product review analysis, customer feedback classification, and opinion mining.\n\nTraditional sentiment analysis methods often rely on manually assigning sentiment scores to individual words and aggregating them to estimate the overall sentiment of a document. While simple, these approaches struggle with complex linguistic phenomena such as negation, sarcasm, and contextual meaning.\n\nIn this notebook, sentiment analysis is performed using **pretrained GloVe word embeddings**, which provide dense vector representations that capture semantic relationships between words.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "al meaning.\n\nIn this notebook, sentiment analysis is performed using **pretrained GloVe word embeddings**, which provide dense vector representations that capture semantic relationships between words. These embeddings enable neural networks to learn richer contextual representations than traditional bag-of-words approaches.\n\nThe model is trained on the **Sentiment140** dataset, a widely used benchmark containing approximately **1.6 million** Twitter posts labeled as positive or negative based on the emoticons originally included in each tweet. During preprocessing, the emoticons are removed, and the objective becomes predicting the underlying sentiment from the remaining tweet text.\n\nThe dataset is loaded directly from a publicly shared Google Drive file, eliminating the need for manual downloads or Google Drive mounting.\n\n**Dataset**\n\nhttps://drive.google.com/file/d/1sQuD_H1mMfW7tr42DKUuBLFPMJh5NiS8/view?usp=sharing", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "ublicly shared Google Drive file, eliminating the need for manual downloads or Google Drive mounting.\n\n**Dataset**\n\nhttps://drive.google.com/file/d/1sQuD_H1mMfW7tr42DKUuBLFPMJh5NiS8/view?usp=sharing\n\nThe following section loads the dataset and explores its structure before constructing the sentiment classification pipeline.\n\nPython implementation:\nimport csv\nimport gdown\nimport torch\nimport torchtext\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n# Shared Google Drive link\nurl = \"https://drive.google.com/uc?id=1sQuD_H1mMfW7tr42DKUuBLFPMJh5NiS8\"\n\n# Download only if needed\ngdown.download(\n url,\n \"training.1600000.processed.noemoticon.csv\",\n quiet=False\n)\n\ndef get_data():\n return csv.reader(\n open(\n \"training.1600000.processed.noemoticon.csv\",\n \"rt\",\n encoding=\"latin-1\"\n )\n )\n\ndef split_tweet(tweet):\n tweet = (\n tweet.replace(\".\", \" . \")\n .replace(\",\", \" , \")\n .replace(\";\", \" ; \")\n .replace(\"?\", \" ? \")\n )\n return tweet.lower().split()\n\nglove = torchtext.vocab.GloVe(\n name=\"6B\",", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "plit_tweet(tweet):\n tweet = (\n tweet.replace(\".\", \" . \")\n .replace(\",\", \" , \")\n .replace(\";\", \" ; \")\n .replace(\"?\", \" ? \")\n )\n return tweet.lower().split()\n\nglove = torchtext.vocab.GloVe(\n name=\"6B\",\n dim=50,\n max_vectors=10000\n)\n\nPython implementation:\n# print only the first tweet\nfor i, line in enumerate(get_data()):\n if line[0] != '0':\n print(line[0], line[-1])\n break\n\nExplanation:\nThe columns we care about is the first one and the last one. The first column is the\nlabel (the label `0` means \"sad\" tweet, `4` means \"happy\" tweet), and the last column\ncontains the tweet. Our task is to predict the sentiment of the tweet given the text.\n\nThe approach today is as follows, for each tweet:\n\n1. We will split the text into words. We will do so by splitting at all whitespace\n characters. There are better ways to perform the split, but let's keep our\n dependencies light.\n2. We will look up the GloVe embedding of each word.\n Words that do not have a GloVe vector will be ignored.\n3.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "ers. There are better ways to perform the split, but let's keep our\n dependencies light.\n2. We will look up the GloVe embedding of each word.\n Words that do not have a GloVe vector will be ignored.\n3. We will sum up all the embeddings, to get an embedding for an entire tweet.\n4. Finally, we will use a fully-connected neural network\n to predict whether the tweet has positive or negative sentiment.\n\nFirst, let's sanity check that there are enough words for us to work with.\n\nPython implementation:\nimport torchtext\nglove = torchtext.vocab.GloVe(name=\"6B\", dim=50)\n\ndef split_tweet(tweet):\n # separate punctuations\n tweet = tweet.replace(\".\", \" . \") \\\n .replace(\",\", \" , \") \\\n .replace(\";\", \" ; \") \\\n .replace(\"?\", \" ? \")\n return tweet.lower().split()\n\nsplit_tweet(\"hello; don't you know?\")\n\nPython implementation:\n# verify that each tweet has a reasonable number of words\n# that have GloVe embeddings\nfor i, line in enumerate(get_data()):\n if i > 30:\n break", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "tweet(\"hello; don't you know?\")\n\nPython implementation:\n# verify that each tweet has a reasonable number of words\n# that have GloVe embeddings\nfor i, line in enumerate(get_data()):\n if i > 30:\n break\n print(sum(int(w in glove.stoi) for w in split_tweet(line[-1])))\n\nExplanation:\nLooks like each tweet has at least one word that has an embedding.\n\nNow, steps 1-3 from above can be done ahead of time, just like the transfer learning\nportion of Lab 2. So, we will write a function that will take the tweets data\nfile, computes the tweet embeddings, and splits the data into train/validation/test.\n\nWe will only use $\\frac{1}{59}$ of the data in the file, so that this demo runs\nrelatively quickly.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\n\ndef get_tweet_vectors(glove_vector):\n train, valid, test = [], [], []\n for i, line in enumerate(get_data()):\n tweet = line[-1]\n if i % 59 == 0:\n # obtain an embedding for the entire tweet", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "h.nn as nn\n\ndef get_tweet_vectors(glove_vector):\n train, valid, test = [], [], []\n for i, line in enumerate(get_data()):\n tweet = line[-1]\n if i % 59 == 0:\n # obtain an embedding for the entire tweet\n tweet_emb = sum(glove_vector[w] for w in split_tweet(tweet))\n # generate a label: 1 = happy, 0 = sad\n label = torch.tensor(int(line[0] == \"4\")).long()\n # place the data set in either the training, validation, or test set\n if i % 5 < 3:\n train.append((tweet_emb, label)) # 60% training\n elif i % 5 == 4:\n valid.append((tweet_emb, label)) # 20% validation\n else:\n test.append((tweet_emb, label)) # 20% test\n return train, valid, test\n\nExplanation:\nI'm making the `glove_vector` a parameter so that we can test the effect\nof using a higher dimensional GloVe\nembedding later. Now, let's get our training, validation, and test set.\nThe format is what `torch.utils.data.DataLoader` expects.\n\nPython implementation:\nimport torchtext\n\nglove = torchtext.vocab.GloVe(name=\"6B\", dim=50)", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "let's get our training, validation, and test set.\nThe format is what `torch.utils.data.DataLoader` expects.\n\nPython implementation:\nimport torchtext\n\nglove = torchtext.vocab.GloVe(name=\"6B\", dim=50)\n\ntrain, valid, test = get_tweet_vectors(glove)\n\ntrain_loader = torch.utils.data.DataLoader(train, batch_size=128, shuffle=True)\nvalid_loader = torch.utils.data.DataLoader(valid, batch_size=128, shuffle=True)\ntest_loader = torch.utils.data.DataLoader(test, batch_size=128, shuffle=True)\n\nExplanation:\nNow, our actual training script! Note that will we use `CrossEntropyLoss`,\nhave two neurons in the final layer of our output layer, and use softmax instead of\na sigmoid activation. This is different from our choice in the earlier weeks!\nTypically, machine learning practitioners will choose to use two output\nneurons instead of one, even in a binary classification task. The reason is that\nan extra neuron adds some more parameters to the network, and makes the network", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "itioners will choose to use two output\nneurons instead of one, even in a binary classification task. The reason is that\nan extra neuron adds some more parameters to the network, and makes the network\na little easier to train (performs better).\n\nPython implementation:\nimport matplotlib.pyplot as plt\n\ndef train_network(model, train_loader, valid_loader, num_epochs=5, learning_rate=1e-5):\n criterion = nn.CrossEntropyLoss()\n optimizer = torch.optim.Adam(model.parameters(), lr=learning_rate)\n losses, train_acc, valid_acc = [], [], []\n epochs = []\n for epoch in range(num_epochs):\n for tweets, labels in train_loader:\n optimizer.zero_grad()\n pred = model(tweets)\n loss = criterion(pred, labels)\n loss.backward()\n optimizer.step()\n\n losses.append(float(loss))\n if epoch % 5 == 4:\n epochs.append(epoch)\n train_acc.append(get_accuracy(model, train_loader))\n valid_acc.append(get_accuracy(model, valid_loader))\n print(\"Epoch %d; Loss %f; Train Acc %f; Val Acc %f\" % (", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "poch % 5 == 4:\n epochs.append(epoch)\n train_acc.append(get_accuracy(model, train_loader))\n valid_acc.append(get_accuracy(model, valid_loader))\n print(\"Epoch %d; Loss %f; Train Acc %f; Val Acc %f\" % (\n epoch+1, loss, train_acc[-1], valid_acc[-1]))\n\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(losses, label=\"Train\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(epochs, train_acc, label=\"Train\")\n plt.plot(epochs, valid_acc, label=\"Validation\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\ndef get_accuracy(model, data_loader):\n correct, total = 0, 0\n for tweets, labels in data_loader:\n output = model(tweets)\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += labels.shape[0]\n return correct / total\n\nExplanation:\nAs for the actual mode, we will start with a 3-layer neural network.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-9", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "im=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += labels.shape[0]\n return correct / total\n\nExplanation:\nAs for the actual mode, we will start with a 3-layer neural network.\nWe won't create our own class since this is a fairly straightforward neural\nnetwork, so an `nn.Sequential` object will do.\nLet's build and train our network.\n\nPython implementation:\nmymodel = nn.Sequential(nn.Linear(50, 30),\n nn.ReLU(),\n nn.Linear(30, 10),\n nn.ReLU(),\n nn.Linear(10, 2))\ntrain_network(mymodel, train_loader, valid_loader, num_epochs=100, learning_rate=1e-4)\nprint(\"Final test accuracy:\", get_accuracy(mymodel, test_loader))\n\nPython implementation:\ndef test_model(model, glove_vector, tweet):\n emb = sum(glove_vector[w] for w in split_tweet(tweet))\n out = mymodel(emb.unsqueeze(0))\n pred = out.max(1, keepdim=True)[1]\n return pred\n\ntest_model(mymodel, glove, \"very happy\")\n\nExplanation:\nNote that the model does not perform very well, but it does get it right from time to time.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-10", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "0))\n pred = out.max(1, keepdim=True)[1]\n return pred\n\ntest_model(mymodel, glove, \"very happy\")\n\nExplanation:\nNote that the model does not perform very well, but it does get it right from time to time. There are a number of things we could try to improve the performance, for example changing the number of hidden units, number of layers, or change the dimension of the Glove embeddings.\n\nAnother option is to use a more powerful architecture.\n\nExplanation:\n#Part C: Recurrent Neural Networks\n\nOne of the drawbacks of the previous approach is that the order of\nwords is lost. The tweets \"the cat likes the dog\" and \"the dog likes the cat\"\nwould have the exact same embedding, even though the sentences have different\nmeanings.\n\nFor this part we will use a **recurrent neural network**. We will treat each tweet\nas a **sequence** of words. Like before, we will use GloVe embeddings as inputs\nto the recurrent network. (As a sidenote, not all recurrent neural networks use\nword embeddings as input.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-11", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "eat each tweet\nas a **sequence** of words. Like before, we will use GloVe embeddings as inputs\nto the recurrent network. (As a sidenote, not all recurrent neural networks use\nword embeddings as input. If we had a small enough vocabulary, we could have used\na one-hot embedding of the words.)\n\nPython implementation:\nimport csv\nimport gdown\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport torchtext\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n# Download the Sentiment140 dataset from a shared Google Drive link\nurl = \"https://drive.google.com/uc?id=1sQuD_H1mMfW7tr42DKUuBLFPMJh5NiS8\"\noutput = \"training.1600000.processed.noemoticon.csv\"\n\ngdown.download(url, output, quiet=False)\n\ndef get_data():\n \"\"\"Load the Sentiment140 dataset.\"\"\"\n return csv.reader(open(output, \"rt\", encoding=\"latin-1\"))\n\ndef split_tweet(tweet):\n \"\"\"Tokenize a tweet by separating punctuation and converting to lowercase.\"\"\"\n tweet = (\n tweet.replace(\".\", \" . \")", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-12", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "return csv.reader(open(output, \"rt\", encoding=\"latin-1\"))\n\ndef split_tweet(tweet):\n \"\"\"Tokenize a tweet by separating punctuation and converting to lowercase.\"\"\"\n tweet = (\n tweet.replace(\".\", \" . \")\n .replace(\",\", \" , \")\n .replace(\";\", \" ; \")\n .replace(\"?\", \" ? \")\n )\n return tweet.lower().split()\n\n# Load the 10,000 most common pretrained GloVe word vectors\nglove = torchtext.vocab.GloVe(\n name=\"6B\",\n dim=50,\n max_vectors=10000\n)\n\nExplanation:\nSince we are going to store the individual words in a tweet,\nwe will defer looking up the word embeddings.\nInstead, we will store the **index** of each word in a PyTorch tensor.\nOur choice is the most memory-efficient, since it takes fewer bits to\nstore an integer index than a 50-dimensional vector or a word.\n\nPython implementation:\ndef get_tweet_words(glove_vector):\n train, valid, test = [], [], []\n for i, line in enumerate(get_data()):\n if i % 29 == 0:\n tweet = line[-1]\n idxs = [glove_vector.stoi[w] # lookup the index of word", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-13", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "def get_tweet_words(glove_vector):\n train, valid, test = [], [], []\n for i, line in enumerate(get_data()):\n if i % 29 == 0:\n tweet = line[-1]\n idxs = [glove_vector.stoi[w] # lookup the index of word\n for w in split_tweet(tweet)\n if w in glove_vector.stoi] # keep words that has an embedding\n if not idxs: # ignore tweets without any word with an embedding\n continue\n idxs = torch.tensor(idxs) # convert list to pytorch tensor\n label = torch.tensor(int(line[0] == \"4\")).long()\n if i % 5 < 3:\n train.append((idxs, label))\n elif i % 5 == 4:\n valid.append((idxs, label))\n else:\n test.append((idxs, label))\n return train, valid, test\n\ntrain, valid, test = get_tweet_words(glove)\n\nExplanation:\nHere's what an element of the training set looks like:\n\nExplanation:\nUnlike in the past, each element of the training set will have a\ndifferent shape. The difference will present some difficulties when\nwe discuss batching later on.\n\nPython implementation:\nfor i in range(10):\n tweet, label = train[i]", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-14", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "of the training set will have a\ndifferent shape. The difference will present some difficulties when\nwe discuss batching later on.\n\nPython implementation:\nfor i in range(10):\n tweet, label = train[i]\n print(tweet.shape)", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "441-15", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Part 2 — Sentiment Analysis with GloVe Embeddings" + }, + { + "text": "Explanation:\n## Embedding\n\nWe are also going to use an `nn.Embedding` layer, instead of using the variable\n`glove` directly. The reason is that the `nn.Embedding` layer allows us look up\nthe embeddings of multiple words simultaneously.\n\nPython implementation:\nglove_emb = nn.Embedding.from_pretrained(glove.vectors)\n\n# Example: we use the forward function of glove_emb to lookup the\n# embedding of each word in `tweet`\ntweet_emb = glove_emb(tweet)\ntweet_emb.shape", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "442-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Embedding" + }, + { + "text": "Explanation:\n## Recurrent Neural Network Module\n\nPyTorch has variations of recurrent neural network modules.\nThese modules computes the following:\n\n$$hidden = updatefn(hidden, input)$$\n$$output = outputfn(hidden)$$\n\nThese modules are more complex and less intuitive than the usual\nneural network layers, so let's take a look:\n\nPython implementation:\nrnn_layer = nn.RNN(input_size=50, # dimension of the input repr\n hidden_size=50, # dimension of the hidden units\n batch_first=True) # input format is [batch_size, seq_len, repr_dim]\n\nExplanation:\nNow, let's try and run this untrained `rnn_layer` on `tweet_emb`.\nWe will need to add an extra dimension to `tweet_emb` to account for\nbatching. We will also need to initialize a set of hidden units of size\n`[batch_size, 1, repr_dim]`, to be used for the *first* set of computations.\n\n![](imgs/rnn.png)\n\nPython implementation:\ntweet_input = tweet_emb.unsqueeze(0) # add the batch_size dimension\nh0 = torch.zeros(1, 1, 50) # initial hidden state", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "443-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Recurrent Neural Network Module" + }, + { + "text": "d for the *first* set of computations.\n\n![](imgs/rnn.png)\n\nPython implementation:\ntweet_input = tweet_emb.unsqueeze(0) # add the batch_size dimension\nh0 = torch.zeros(1, 1, 50) # initial hidden state\nout, last_hidden = rnn_layer(tweet_input, h0)\n\nExplanation:\nWe don't technically have to explictly provide the initial hidden state,\nif we want to use an initial state of zeros. Just for today, we will be\nexplicit about the hidden states that we provide.\n\nExplanation:\nNow, let's look at the output and hidden dimensions that we have:\n\nExplanation:\nThe shape of the hidden units is the same as our initial `h0`.\nThe variable `out`, though, has the same shape as our `input`.\nThe variable contains the concatenation of all of the output units\nfor each word (i.e. at each time point).\n\nNormally, we only care about the output at the **final** time point,\nwhich we can extract like this.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "443-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Recurrent Neural Network Module" + }, + { + "text": "Explanation:\nThis tensor summarizes the entire tweet, and can be used as an input\nto a classifier.\n\n## Building a Model\n\nLet's put both the embedding layer, the RNN and the classifier into one model:\n\nPython implementation:\nclass TweetRNN(nn.Module):\n def __init__(self, input_size, hidden_size, num_classes):\n super(TweetRNN, self).__init__()\n self.emb = nn.Embedding.from_pretrained(glove.vectors)\n self.hidden_size = hidden_size\n self.rnn = nn.RNN(input_size, hidden_size, batch_first=True)\n self.fc = nn.Linear(hidden_size, num_classes)\n\n def forward(self, x):\n # Look up the embedding\n x = self.emb(x)\n # Set an initial hidden state\n h0 = torch.zeros(1, x.size(0), self.hidden_size)\n # Forward propagate the RNN\n out, _ = self.rnn(x, h0)\n # Pass the output of the last time step to the classifier\n out = self.fc(out[:, -1, :])\n return out\n\nmodel = TweetRNN(50, 50, 2)", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "444-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Building a Model" + }, + { + "text": "Explanation:\nWe should be able to train this model similar to any other model that we have trained before. However, there is one caveat that we have been avoiding this entire time, **batching**.\n\n## Batching\n\nUnfortunately, we will not be able to use `DataLoader` with a\n`batch_size` of greater than one. This is because each tweet has\na different shaped tensor.\n\nPython implementation:\nfor i in range(10):\n tweet, label = train[i]\n print(tweet.shape)\n\nExplanation:\nPyTorch implementation of `DataLoader` class expects all data samples\nto have the same shape. So, if we create a DataLoader like below,\nit will throw an error when we try to iterate over its elements.\n\nExplanation:\nSo, we will need a different way of batching.\n\nOne strategy is to **pad shorter sequences with zero inputs**, so that\nevery sequence is the same length. The following PyTorch utilities\nare helpful.\n\n- `torch.nn.utils.rnn.pad_sequence`\n- `torch.nn.utils.rnn.pad_packed_sequence`\n- `torch.nn.utils.rnn.pack_sequence`", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "at\nevery sequence is the same length. The following PyTorch utilities\nare helpful.\n\n- `torch.nn.utils.rnn.pad_sequence`\n- `torch.nn.utils.rnn.pad_packed_sequence`\n- `torch.nn.utils.rnn.pack_sequence`\n- `torch.nn.utils.rnn.pack_padded_sequence`\n\n(Actually, there are more powerful helpers in the `torchtext` module\nthat we will use in Lab 4. We'll stick to these in this demo, so that\nyou can see what's actually going on under the hood.)\n\nPython implementation:\nfrom torch.nn.utils.rnn import pad_sequence\n\ntweet_padded = pad_sequence([tweet for tweet, label in train[:10]],\n batch_first=True)\nprint(tweet_padded.shape)\nprint(tweet_padded[0:2])\n\nExplanation:\nNow, we can pass multiple tweets in a batch through the RNN at once!\n\nExplanation:\nOne issue we overlooked was that in our `TweetRNN` model, we always\ntake the **last output unit** as input to the final classifier. Now\nthat we are padding the input sequences, we should really be using\nthe output at a previous time step.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "r `TweetRNN` model, we always\ntake the **last output unit** as input to the final classifier. Now\nthat we are padding the input sequences, we should really be using\nthe output at a previous time step. Recurrent neural networks therefore\nrequire much more record keeping than ANNs or even CNNs.\n\nThere is yet another problem:\nthe longest tweet has many, many more words than the shortest.\nPadding tweets so that every tweet has the same length as the longest\ntweet is impractical. Padding tweets in a mini-batch, however, is much\nmore reasonable.\n\nIn practice, practitioners will batch together tweets with the same\nlength. For simplicity, we will do the same. We will implement a (more or less)\nstraightforward way to batch tweets.\n\nPython implementation:\nimport random\n\nclass TweetBatcher:\n def __init__(self, tweets, batch_size=32, drop_last=False):\n # store tweets by length\n self.tweets_by_length = {}\n for words, label in tweets:\n # compute the length of the tweet\n wlen = words.shape[0]", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "ef __init__(self, tweets, batch_size=32, drop_last=False):\n # store tweets by length\n self.tweets_by_length = {}\n for words, label in tweets:\n # compute the length of the tweet\n wlen = words.shape[0]\n # put the tweet in the correct key inside self.tweet_by_length\n if wlen not in self.tweets_by_length:\n self.tweets_by_length[wlen] = []\n self.tweets_by_length[wlen].append((words, label),)\n\n # create a DataLoader for each set of tweets of the same length\n self.loaders = {wlen : torch.utils.data.DataLoader(\n tweets,\n batch_size=batch_size,\n shuffle=True,\n drop_last=drop_last) # omit last batch if smaller than batch_size\n for wlen, tweets in self.tweets_by_length.items()}\n\n def __iter__(self): # called by Python to create an iterator\n # make an iterator for every tweet length\n iters = [iter(loader) for loader in self.loaders.values()]\n while iters:\n # pick an iterator (a length)\n im = random.choice(iters)\n try:\n yield next(im)\n except StopIteration:", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "for every tweet length\n iters = [iter(loader) for loader in self.loaders.values()]\n while iters:\n # pick an iterator (a length)\n im = random.choice(iters)\n try:\n yield next(im)\n except StopIteration:\n # no more elements in the iterator, remove it\n iters.remove(im)\n\nExplanation:\nLet's take a look at our batcher in action. We will set `drop_last` to be true for training,\nso that all of our batches have exactly the same size.\n\nPython implementation:\nfor i, (tweets, labels) in enumerate(TweetBatcher(train, drop_last=True)):\n if i > 20: break\n print(tweets.shape, labels.shape)\n\nExplanation:\nJust to verify that our batching is reasonable, here is a modification of the\n`get_accuracy` function we wrote last time.\n\nPython implementation:\ndef get_accuracy(model, data_loader):\n correct, total = 0, 0\n for tweets, labels in data_loader:\n output = model(tweets)\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += labels.shape[0]", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "rrect, total = 0, 0\n for tweets, labels in data_loader:\n output = model(tweets)\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += labels.shape[0]\n return correct / total\n\ntest_loader = TweetBatcher(test, batch_size=64, drop_last=False)\nget_accuracy(model, test_loader)\n\nExplanation:\nOur training code will also be very similar to the code we wrote last time:\n\nPython implementation:\ndef train_rnn_network(model, train, valid, num_epochs=5, learning_rate=1e-5):\n criterion = nn.CrossEntropyLoss()\n optimizer = torch.optim.Adam(model.parameters(), lr=learning_rate)\n losses, train_acc, valid_acc = [], [], []\n epochs = []\n for epoch in range(num_epochs):\n for tweets, labels in train:\n optimizer.zero_grad()\n pred = model(tweets)\n loss = criterion(pred, labels)\n loss.backward()\n optimizer.step()\n losses.append(float(loss))\n\n epochs.append(epoch)\n train_acc.append(get_accuracy(model, train_loader))", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "grad()\n pred = model(tweets)\n loss = criterion(pred, labels)\n loss.backward()\n optimizer.step()\n losses.append(float(loss))\n\n epochs.append(epoch)\n train_acc.append(get_accuracy(model, train_loader))\n valid_acc.append(get_accuracy(model, valid_loader))\n print(\"Epoch %d; Loss %f; Train Acc %f; Val Acc %f\" % (\n epoch+1, loss, train_acc[-1], valid_acc[-1]))\n # plotting\n plt.title(\"Training Curve\")\n plt.plot(losses, label=\"Train\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"Loss\")\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(epochs, train_acc, label=\"Train\")\n plt.plot(epochs, valid_acc, label=\"Validation\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\nExplanation:\nLet's train our model. Note that there will be some inaccuracies in computing the training loss.\nWe are dropping some data from the training set by setting `drop_last=True`. Again, the choice is\nnot ideal, but simplifies our code.\n\nPython implementation:\nmodel = TweetRNN(50, 50, 2)", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "raining loss.\nWe are dropping some data from the training set by setting `drop_last=True`. Again, the choice is\nnot ideal, but simplifies our code.\n\nPython implementation:\nmodel = TweetRNN(50, 50, 2)\ntrain_loader = TweetBatcher(train, batch_size=64, drop_last=True)\nvalid_loader = TweetBatcher(valid, batch_size=64, drop_last=False)\ntrain_rnn_network(model, train_loader, valid_loader, num_epochs=20, learning_rate=2e-4)\nget_accuracy(model, test_loader)\n\nExplanation:\nThe hidden size and the input embedding size don't have to be the same.\n\nPython implementation:\nmodel = TweetRNN(50, 100, 2)\ntrain_rnn_network(model, train_loader, valid_loader, num_epochs=20, learning_rate=2e-4)\nget_accuracy(model, test_loader)", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "445-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Batching" + }, + { + "text": "Explanation:\n##### Sentiment Analysis on New Tweets\n\nPython implementation:\ndef get_new_tweet(glove_vector, sample_tweet):\n tweet = sample_tweet\n idxs = [glove_vector.stoi[w] # lookup the index of word\n for w in split_tweet(tweet)\n if w in glove_vector.stoi] # keep words that has an embedding\n idxs = torch.tensor(idxs) # convert list to pytorch tensor\n return idxs\n\nPython implementation:\nnew_tweet = get_new_tweet(glove, \"This is a terrible tragedy\")\nprint(new_tweet.shape)\n\nout = torch.sigmoid(model(new_tweet.unsqueeze(0)))\npred = out.max(1, keepdim=True)[1]\nprint(pred)\n\nPython implementation:\nnew_tweet = get_new_tweet(glove, \"This is the best day of my life\")\nprint(new_tweet.shape)\n\nout = torch.sigmoid(model(new_tweet.unsqueeze(0)))\npred = out.max(1, keepdim=True)[1]\nprint(pred)\n\nExplanation:\nSentiment analysis is not a straight forward problem and it's impressive that without too much tunning we are able to get validation accuracy close to 70%.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "446-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Sentiment Analysis on New Tweets" + }, + { + "text": ", keepdim=True)[1]\nprint(pred)\n\nExplanation:\nSentiment analysis is not a straight forward problem and it's impressive that without too much tunning we are able to get validation accuracy close to 70%. There are many ways that we can improve our performance:\n\n* Tune hyperparameters\n* Use all of the available data\n* Apply more advanced RNN architectures (ex. LSTM and GRU)\n\nPython implementation:\nclass TweetLSTM(nn.Module):\n def __init__(self, input_size, hidden_size, num_classes):\n super(TweetLSTM, self).__init__()\n self.emb = nn.Embedding.from_pretrained(glove.vectors)\n self.hidden_size = hidden_size\n self.rnn = nn.LSTM(input_size, hidden_size, batch_first=True)\n self.fc = nn.Linear(hidden_size, num_classes)\n\n def forward(self, x):\n # Look up the embedding\n x = self.emb(x)\n # Set an initial hidden state and cell state\n h0 = torch.zeros(1, x.size(0), self.hidden_size)\n c0 = torch.zeros(1, x.size(0), self.hidden_size)\n # Forward propagate the LSTM\n out, _ = self.rnn(x, (h0, c0))", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "446-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Sentiment Analysis on New Tweets" + }, + { + "text": "n initial hidden state and cell state\n h0 = torch.zeros(1, x.size(0), self.hidden_size)\n c0 = torch.zeros(1, x.size(0), self.hidden_size)\n # Forward propagate the LSTM\n out, _ = self.rnn(x, (h0, c0))\n # Pass the output of the last time step to the classifier\n out = self.fc(out[:, -1, :])\n return out\n\nmodel = TweetLSTM(50, 50, 2)\ntrain_rnn_network(model, train_loader, valid_loader, num_epochs=20, learning_rate=2e-5)\nget_accuracy(model, test_loader)", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "446-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Sentiment Analysis on New Tweets" + }, + { + "text": "Explanation:\n# Results\n\n## Summary\n\nThis notebook explored modern NLP techniques for text representation and sequence modeling using PyTorch.\n\n### Implemented Components\n\n- Word2Vec embeddings\n- GloVe embeddings\n- Semantic similarity\n- Word analogy reasoning\n- Bias analysis\n- Sentiment analysis\n- Recurrent Neural Networks (RNNs)\n- Variable-length sequence batching\n\n### Key Outcomes\n\n- Demonstrated how word embeddings capture semantic relationships.\n- Measured semantic similarity using cosine similarity.\n- Solved analogy tasks using vector arithmetic.\n- Investigated bias present in pretrained embeddings.\n- Built an end-to-end tweet sentiment classifier using recurrent neural networks.\n- Explored batching strategies for variable-length text sequences.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "447-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Results" + }, + { + "text": "Explanation:\n# Conclusion\n\nThis notebook introduced the core components of modern Natural Language Processing using PyTorch. Beginning with distributed word representations, it demonstrated how pretrained embeddings encode semantic information that can be leveraged for downstream tasks. The notebook concluded by implementing a recurrent neural network for tweet sentiment classification, illustrating how sequential neural networks can effectively model natural language data.", + "source": "01_word_embeddings_and_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "448-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/01_word_embeddings_and_rnn.ipynb", + "section": "Conclusion" + }, + { + "text": "Explanation:\n# Generative Recurrent Neural Networks with PyTorch\n\n## Overview\n\nThis notebook explores **generative recurrent neural networks (RNNs)** for character-level text generation using PyTorch. Unlike predictive sequence models that classify or label existing text, generative RNNs learn the probability distribution of sequential data and generate new text one character at a time.\n\nThe notebook introduces the complete generative modeling pipeline, including text preprocessing, character-level tokenization, recurrent neural network architecture, teacher forcing, autoregressive sequence generation, and sampling from learned probability distributions.\n\nMuch of content is an adaptation of the \"Practical PyTorch\" github\nrepository [1].\n\n[1] https://github.com/spro/practical-pytorch/blob/master/char-rnn-generation/char-rnn-generation.ipynb", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "449-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generative Recurrent Neural Networks with PyTorch" + }, + { + "text": "Explanation:\n## Learning Objectives\n\nBy completing this notebook, you will learn how to:\n\n- Build character-level recurrent neural networks\n- Represent text as sequences of tokens\n- Construct character vocabularies\n- Encode and decode text sequences\n- Train generative RNNs using teacher forcing\n- Predict the next character in a sequence\n- Generate text autoregressively\n- Sample from learned probability distributions\n- Implement generative language models in PyTorch", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "450-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Learning Objectives" + }, + { + "text": "Explanation:\n## Character-Level Sequence Modeling\n\nIn recurrent neural networks the input sequence is broken down into tokens. We could choose whether to tokenize based on words, or based on characters. The representation of each token (GloVe or one-hot) is processed by the RNN one step at a time to update the hidden (or context) state.\n\nIn a predictive RNN, the value of the hidden states is a representation of **all the text that was processed thus far**. Similarly, in a generative RNN, The value of the hidden state will be a representation of **all the text that still needs to be generated**. We will use this hidden state to produce the sequence, one token at a time.\n\nSimilar to the last tutorial we will break up the problem of generating text\nto generating one token at a time.\n\nWe will do so with the help of two functions:\n\n1. We need to be able to generate the *next* token, given the current\n hidden state. In practice, we get a probability distribution over", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "451-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Sequence Modeling" + }, + { + "text": "ken at a time.\n\nWe will do so with the help of two functions:\n\n1. We need to be able to generate the *next* token, given the current\n hidden state. In practice, we get a probability distribution over\n the next token, and sample from that probability distribution.\n2. We need to be able to update the hidden state somehow. To do so,\n we need two piece of information: the old hidden state, and the actual\n token that was generated in the previous step. The actual token generated\n will inform the subsequent tokens.\n\nWe will repeat both functions until a special \"END OF SEQUENCE\" token is\ngenerated.\n\nNote that there are several tricky things that we will have to figure out.\nFor example, how do we actually sample the actual token from the probability\ndistribution over tokens? What would we do during training, and how might\nthat be different from during testing/evaluation? We will answer those\nquestions as we implement the RNN.\n\nFor now, let's start with our training data.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "451-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Sequence Modeling" + }, + { + "text": "hat would we do during training, and how might\nthat be different from during testing/evaluation? We will answer those\nquestions as we implement the RNN.\n\nFor now, let's start with our training data.\n\n## Data: Donald Trump's Tweets from 2018\n\nThe training set we use is a collection of Donald Trump's tweets from 2018.\nWe will only use tweets that are 140 characters or shorter, and tweets\nthat contains more than just a URL.\nSince tweets often contain creative spelling and numbers, and upper vs lower\ncase characters are read very differently, we will use a character-level RNN.\n\nTo start, let us load the trump.csv file to Google Colab and provide access to the drive. The file can be obtained from Quercus.\n\nPython implementation:\n# Install gdown (run once)\n!pip install -q gdown\n\nimport csv\nimport gdown\n\n# Google Drive file ID\nfile_id = \"1LinYW9nmPRKLbXSCCGSKJ9MvgcsijDly\"\n\n# Download the dataset\ngdown.download(\n f\"https://drive.google.com/uc?id={file_id}\",\n output=\"trump.csv\",\n quiet=False\n)", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "451-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Sequence Modeling" + }, + { + "text": "t gdown\n\n# Google Drive file ID\nfile_id = \"1LinYW9nmPRKLbXSCCGSKJ9MvgcsijDly\"\n\n# Download the dataset\ngdown.download(\n f\"https://drive.google.com/uc?id={file_id}\",\n output=\"trump.csv\",\n quiet=False\n)\n\n# Load the tweets\nwith open(\"trump.csv\", \"r\", encoding=\"utf-8\") as f:\n tweets = [row[0] for row in csv.reader(f)]\n\nprint(f\"Number of tweets: {len(tweets)}\")\n\nExplanation:\nThere are over 20000 tweets in this collection.\nLet's look at a few of them, just to get a sense of the kind of text\nwe're dealing with:", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "451-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Sequence Modeling" + }, + { + "text": "Explanation:\n## Character-Level Text Generation\nTo introduce the training procedure, we first demonstrate how a recurrent neural network can learn to generate a single tweet. This simplified example illustrates the complete sequence generation pipeline before extending the approach to larger datasets.\n\nNormally, when we build a new machine learn model, we want to make sure\nthat our model can overfit. To that end, we will first build a neural network\nthat can generate _one_ tweet really well. We can choose any tweet (or any other text)\nwe want. Let's choose to build an RNN that generates `tweet[100]`.\n\nExplanation:\nFirst, we will need to encode this tweet using a one-hot encoding.\nWe'll build dictionary mappings\nfrom the character to the index of that character (a unique integer identifier),\nand from the index to the character. We'll use the same naming scheme that `torchtext`\nuses (`stoi` and `itos`).\n\nFor simplicity, we'll work with a limited vocabulary containing", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "452-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Text Generation" + }, + { + "text": "integer identifier),\nand from the index to the character. We'll use the same naming scheme that `torchtext`\nuses (`stoi` and `itos`).\n\nFor simplicity, we'll work with a limited vocabulary containing\njust the characters in `tweet[100]`, plus two special tokens:\n\n- `` represents \"End of String\", which we'll append to the end of our tweet.\n Since tweets are variable-length, this is a way for the RNN to signal\n that the entire sequence has been generated.\n- `` represents \"Beginning of String\", which we'll prepend to the beginning of\n our tweet. This is the first token that we will feed into the RNN.\n\nThe way we use these special tokens will become more clear as we build the model.\n\nPython implementation:\nvocab = list(set(tweet)) + [\"\", \"\"]\nvocab_stoi = {s: i for i, s in enumerate(vocab)}\nvocab_itos = {i: s for i, s in enumerate(vocab)}\nvocab_size = len(vocab)\n\nExplanation:\nNow that we have our vocabulary, we can build the PyTorch model\nfor this problem.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "452-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Text Generation" + }, + { + "text": "for i, s in enumerate(vocab)}\nvocab_itos = {i: s for i, s in enumerate(vocab)}\nvocab_size = len(vocab)\n\nExplanation:\nNow that we have our vocabulary, we can build the PyTorch model\nfor this problem.\nThe actual model is not as complex as you might think. We actually\nalready learned about all the components that we need. (Using and training\nthe model is the hard part)\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\n\nPython implementation:\nclass TextGenerator(nn.Module):\n def __init__(self, vocab_size, hidden_size, n_layers=1):\n super(TextGenerator, self).__init__()\n\n # identiy matrix for generating one-hot vectors\n self.ident = torch.eye(vocab_size)\n\n # recurrent neural network\n self.rnn = nn.GRU(vocab_size, hidden_size, n_layers, batch_first=True)\n\n # a fully-connect layer that outputs a distribution over\n # the next token, given the RNN output\n self.decoder = nn.Linear(hidden_size, vocab_size)", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "452-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Text Generation" + }, + { + "text": "b_size, hidden_size, n_layers, batch_first=True)\n\n # a fully-connect layer that outputs a distribution over\n # the next token, given the RNN output\n self.decoder = nn.Linear(hidden_size, vocab_size)\n\n def forward(self, inp, hidden=None):\n inp = self.ident[inp] # generate one-hot vectors of input\n output, hidden = self.rnn(inp, hidden) # get the next output and hidden state\n output = self.decoder(output) # predict distribution over next tokens\n return output, hidden\n\nmodel = TextGenerator(vocab_size, 64)", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "452-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Character-Level Text Generation" + }, + { + "text": "Explanation:\n## Training with Teacher Forcing\n\nAt a very high level, we want our RNN model to have a high probability\nof generating the tweet. An RNN model generates text\none character at a time based on the hidden state value.\nAt each time step, we will check whether the mdoel generated the\ncorrect character. That is, at each time step,\nwe are trying to select the correct next character out of all the\ncharacters in our vocabulary. Recall that this problem is a multi-class\nclassification problem, and we can use Cross-Entropy loss to train our\nnetwork to become better at this type of problem.\n\nExplanation:\nHowever, we don't just have a single multi-class classification problem.\nInstead, we have **one classification problem per time-step** (per token)!\nSo, how do we predict the first token in the sequence?\nHow do we predict the second token in the sequence?\n\nTo help you understand what happens durign RNN training, we'll start with a", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": "** (per token)!\nSo, how do we predict the first token in the sequence?\nHow do we predict the second token in the sequence?\n\nTo help you understand what happens durign RNN training, we'll start with a\ninefficient training code that shows you what happens step-by-step. We'll\nstart with computing the loss for the first token generated, then the second token,\nand so on.\nLater on, we'll switch to a simpler and more performant version of the code.\n\nSo, let's start with the first classification problem: the problem of generating\nthe **first** token (`tweet[0]`).\n\nTo generate the first token, we'll feed the RNN network (with an initial, empty\nhidden state) the \"\" token. Then, the output\n\nPython implementation:\nbos_input = torch.Tensor([vocab_stoi[\"\"]])\nprint(bos_input.shape, type(bos_input))\nbos_input = bos_input.long()\nprint(bos_input.shape, type(bos_input))\nbos_input = bos_input.unsqueeze(0)\nprint(bos_input.shape, type(bos_input))\noutput, hidden = model(bos_input, hidden=None)", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": "_input))\nbos_input = bos_input.long()\nprint(bos_input.shape, type(bos_input))\nbos_input = bos_input.unsqueeze(0)\nprint(bos_input.shape, type(bos_input))\noutput, hidden = model(bos_input, hidden=None)\noutput # distribution over the first token\n\nExplanation:\nWe can compute the loss using `criterion`. Since the model is untrained,\nthe loss is expected to be high. (For now, we won't do anything\nwith this loss, and omit the backward pass.)\n\nPython implementation:\ntarget = torch.Tensor([vocab_stoi[tweet[0]]]).long().unsqueeze(0)\ncriterion(output.reshape(-1, vocab_size), # reshape to 2D tensor\n target.reshape(-1)) # reshape to 1D tensor\n\nExplanation:\nNow, we need to update the hidden state and generate a prediction\nfor the next token. To do so, we need to provide the current token to\nthe RNN. We already said that during test time, we'll need to sample\nfrom the predicted probabilty over tokens that the neural network\njust generated.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": "do so, we need to provide the current token to\nthe RNN. We already said that during test time, we'll need to sample\nfrom the predicted probabilty over tokens that the neural network\njust generated.\n\nRight now, we can do something better: we can **use the ground-truth,\nactual target token**. This technique is called **teacher-forcing**,\nand generally speeds up training. The reason is that right now,\nsince our model does not perform well, the predicted probability\ndistribution is pretty far from the ground truth. So, it is very,\nvery difficult for the neural network to get back on track given bad\ninput data.\n\nPython implementation:\n# Use teacher-forcing: we pass in the ground truth `target`,\n# rather than using the NN predicted distribution\noutput, hidden = model(target, hidden)\noutput # distribution over the second token\n\nExplanation:\nSimilar to the first step, we can compute the loss, quantifying the\ndifference between the predicted distribution and the actual next\ntoken.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": ")\noutput # distribution over the second token\n\nExplanation:\nSimilar to the first step, we can compute the loss, quantifying the\ndifference between the predicted distribution and the actual next\ntoken. This loss can be used to adjust the weights of the neural\nnetwork (which we are not doing yet).\n\nPython implementation:\ntarget = torch.Tensor([vocab_stoi[tweet[1]]]).long().unsqueeze(0)\ncriterion(output.reshape(-1, vocab_size), # reshape to 2D tensor\n target.reshape(-1)) # reshape to 1D tensor\n\nExplanation:\nWe can continue this process of:\n\n- feeding the previous ground-truth token to the RNN,\n- obtaining the prediction distribution over the next token, and\n- computing the loss,\n\nfor as many steps as there are tokens in the ground-truth tweet.\n\nPython implementation:\nfor i in range(2, len(tweet)):\n output, hidden = model(target, hidden)\n target = torch.Tensor([vocab_stoi[tweet[1]]]).long().unsqueeze(0)\n loss = criterion(output.reshape(-1, vocab_size), # reshape to 2D tensor", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": "nge(2, len(tweet)):\n output, hidden = model(target, hidden)\n target = torch.Tensor([vocab_stoi[tweet[1]]]).long().unsqueeze(0)\n loss = criterion(output.reshape(-1, vocab_size), # reshape to 2D tensor\n target.reshape(-1)) # reshape to 1D tensor\n print(i, output, loss)\n\nExplanation:\nFinally, with our final token, we should expect to output the \"\"\ntoken, so that our RNN learns when to stop generating characters.\n\nPython implementation:\noutput, hidden = model(target, hidden)\ntarget = torch.Tensor([vocab_stoi[\"\"]]).long().unsqueeze(0)\nloss = criterion(output.reshape(-1, vocab_size), # reshape to 2D tensor\n target.reshape(-1)) # reshape to 1D tensor\nprint(i, output, loss)\n\nExplanation:\nIn practice, we don't really need a loop. Recall that in a predictive RNN,\nthe `nn.RNN` module can take an entire sequence as input. We can do the\nsame thing here:\n\nPython implementation:\ntweet_ch = [\"\"] + list(tweet) + [\"\"]\ntweet_indices = [vocab_stoi[ch] for ch in tweet_ch]", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": "module can take an entire sequence as input. We can do the\nsame thing here:\n\nPython implementation:\ntweet_ch = [\"\"] + list(tweet) + [\"\"]\ntweet_indices = [vocab_stoi[ch] for ch in tweet_ch]\ntweet_tensor = torch.Tensor(tweet_indices).long().unsqueeze(0)\n\nprint(tweet_tensor.shape)\n\noutput, hidden = model(tweet_tensor[:,:-1]) # is never an input token\ntarget = tweet_tensor[:,1:] # is never a target token\nloss = criterion(output.reshape(-1, vocab_size), # reshape to 2D tensor\n target.reshape(-1)) # reshape to 1D tensor\n\nExplanation:\nHere, the input to our neural network model is the *entire*\nsequence of input tokens (everything from \"\" to the\nlast character of the tweet). The neural network generates a prediction distribution\nof the next token at each step. We can compare each of these with the ground-truth\n`target`.\n\nOur training loop (for learning to generate the single `tweet`) will therefore\nlook something like this:\n\nPython implementation:", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": "ch step. We can compare each of these with the ground-truth\n`target`.\n\nOur training loop (for learning to generate the single `tweet`) will therefore\nlook something like this:\n\nPython implementation:\noptimizer = torch.optim.Adam(model.parameters(), lr=0.001)\ncriterion = nn.CrossEntropyLoss()\nfor it in range(500):\n optimizer.zero_grad()\n output, _ = model(tweet_tensor[:,:-1])\n loss = criterion(output.reshape(-1, vocab_size),\n target.reshape(-1))\n loss.backward()\n optimizer.step()\n\n if (it+1) % 100 == 0:\n print(\"[Iter %d] Loss %f\" % (it+1, float(loss)))", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "453-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training with Teacher Forcing" + }, + { + "text": "Explanation:\nThe training loss is decreasing with training, which is what we expect.\n\n## Generating a Token\n\nAt this point, we want to see whether our model is actually learning\nsomething. So, we need to talk about how to\nactually use the RNN model to generate text. If we can\ngenerate text, we can make a qualitative asssessment of how well\nour RNN is performing.\n\nThe main difference between training and test-time (generation time)\nis that we don't have the ground-truth tokens to feed as inputs\nto the RNN. Instead, we need to actually **sample** a token based\non the neural network's prediction distribution.\n\nBut how can we sample a token from a distribution?\n\nOn one extreme, we can always take\nthe token with the largest probability (argmax). This has been our\ngo-to technique in other classification tasks. However, this idea\nwill fail here. The reason is that in practice,\n**we want to be able to generate a variety of different sequences from\nthe same model**.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "454-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generating a Token" + }, + { + "text": "-to technique in other classification tasks. However, this idea\nwill fail here. The reason is that in practice,\n**we want to be able to generate a variety of different sequences from\nthe same model**. An RNN that can only generate a single new Trump Tweet\nis fairly useless.\n\nIn short, we want some randomness. We can do so by using the logit\noutputs from our model to construct a multinomial distribution over\nthe tokens, then and sample a random token from that multinomial distribution.\n\nOne natural multinomial distribution we can choose is the\ndistribution we get after applying the softmax on the outputs.\nHowever, we will do one more thing: we will add a **temperature**\nparameter to manipulate the softmax outputs. We can set a\n**higher temperature** to make the probability of each token\n**more even** (more random), or a **lower temperature** to assign\nmore probability to the tokens with a higher logit (output).\nA **higher temperature** means that we will get a more diverse sample,", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "454-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generating a Token" + }, + { + "text": "n\n**more even** (more random), or a **lower temperature** to assign\nmore probability to the tokens with a higher logit (output).\nA **higher temperature** means that we will get a more diverse sample,\nwith potentially more mistakes. A **lower temperature** means that we\nmay see repetitions of the same high probability sequence.\n\nPython implementation:\ndef sample_sequence(model, max_len=100, temperature=0.8):\n generated_sequence = \"\"\n\n inp = torch.Tensor([vocab_stoi[\"\"]]).long()\n hidden = None\n for p in range(max_len):\n output, hidden = model(inp.unsqueeze(0), hidden)\n # Sample from the network as a multinomial distribution\n output_dist = output.data.view(-1).div(temperature).exp()\n top_i = int(torch.multinomial(output_dist, 1)[0])\n # Add predicted character to string and use as next input\n predicted_char = vocab_itos[top_i]\n\n if predicted_char == \"\":\n break\n generated_sequence += predicted_char\n inp = torch.Tensor([top_i]).long()\n return generated_sequence", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "454-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generating a Token" + }, + { + "text": "and use as next input\n predicted_char = vocab_itos[top_i]\n\n if predicted_char == \"\":\n break\n generated_sequence += predicted_char\n inp = torch.Tensor([top_i]).long()\n return generated_sequence\n\nprint(sample_sequence(model, temperature=0.8))\nprint(sample_sequence(model, temperature=1.0))\nprint(sample_sequence(model, temperature=1.5))\nprint(sample_sequence(model, temperature=2.0))\nprint(sample_sequence(model, temperature=5.0))", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "454-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generating a Token" + }, + { + "text": "Explanation:\nSince we only trained the model on a single sequence, we won't see\nthe effect of the temperature parameter yet.\n\nFor now, the output of the calls to the `sample_sequence` function\nassures us that our training code looks reasonable, and we can\nproceed to training on our full dataset!\n\n## Training the Trump Tweet Generator\n\nFor the actual training, let's use `torchtext` so that we can use\nthe `BucketIterator` to make batches. Like in Lab 5, we'll create a\n`torchtext.data.Field` to use `torchtext` to read the CSV file, and convert\ncharacters into indices. The object has convient parameters to specify\nthe BOS and EOS tokens.\n\nPython implementation:\nimport torchtext\n\ntext_field = torchtext.data.Field(sequential=True, # text sequence\n tokenize=lambda x: x, # because are building a character-RNN\n include_lengths=True, # to track the length of sequences, for batching\n batch_first=True,\n use_vocab=True, # to turn each character into an integer index", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "455-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training the Trump Tweet Generator" + }, + { + "text": "x: x, # because are building a character-RNN\n include_lengths=True, # to track the length of sequences, for batching\n batch_first=True,\n use_vocab=True, # to turn each character into an integer index\n init_token=\"\", # BOS token\n eos_token=\"\") # EOS token\n\nPython implementation:\nimport gdown\nimport torchtext\n\n# Google Drive file ID\nfile_id = \"1LinYW9nmPRKLbXSCCGSKJ9MvgcsijDly\"\n\n# Download the dataset\ngdown.download(\n f\"https://drive.google.com/uc?id={file_id}\",\n output=\"trump.csv\",\n quiet=False\n)\n\n# Define the fields\nfields = [\n ('text', text_field),\n ('created_at', None),\n ('id_str', None)\n]\n\n# Load the dataset\ntrump_tweets = torchtext.data.TabularDataset(\n path=\"trump.csv\",\n format=\"csv\",\n fields=fields\n)\n\nprint(f\"Number of tweets: {len(trump_tweets)}\")\n\nPython implementation:\ntext_field.build_vocab(trump_tweets)\nvocab_stoi = text_field.vocab.stoi # so we don't have to rewrite sample_sequence\nvocab_itos = text_field.vocab.itos # so we don't have to rewrite sample_sequence", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "455-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training the Trump Tweet Generator" + }, + { + "text": "text_field.build_vocab(trump_tweets)\nvocab_stoi = text_field.vocab.stoi # so we don't have to rewrite sample_sequence\nvocab_itos = text_field.vocab.itos # so we don't have to rewrite sample_sequence\nvocab_size = len(text_field.vocab.itos)\nvocab_size\n\nExplanation:\nLet's just verify that the `BucketIterator` works as expected, but start with batch_size of 1.\n\nPython implementation:\ndata_iter = torchtext.data.BucketIterator(\n trump_tweets,\n batch_size=1,\n sort_key=lambda x: len(x.text)\n)\n\nfor batch in data_iter:\n print(batch)\n break\n\nPython implementation:\nfor batch in data_iter:\n print(batch.text)\n break\n\nExplanation:\nTo account for batching, our actual training code will change, but just a little bit.\nIn fact, our training code from before will work with a batch size larger than one!\n\nPython implementation:\ndef train(model, data, batch_size=1, num_epochs=1, lr=0.001, print_every=100):\n optimizer = torch.optim.Adam(model.parameters(), lr=lr)\n criterion = nn.CrossEntropyLoss()", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "455-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training the Trump Tweet Generator" + }, + { + "text": "ne!\n\nPython implementation:\ndef train(model, data, batch_size=1, num_epochs=1, lr=0.001, print_every=100):\n optimizer = torch.optim.Adam(model.parameters(), lr=lr)\n criterion = nn.CrossEntropyLoss()\n\n model.train()\n it = 0\n\n data_iter = torchtext.data.BucketIterator(\n data,\n batch_size=batch_size,\n sort_key=lambda x: len(x.text),\n device=-1,\n shuffle=True\n )\n\n for e in range(num_epochs):\n avg_loss = 0\n\n for batch in data_iter:\n tweet, lengths = batch.text\n\n target = tweet[:, 1:]\n inp = tweet[:, :-1]\n\n optimizer.zero_grad()\n\n output, _ = model(inp)\n\n loss = criterion(\n output.reshape(-1, vocab_size),\n target.reshape(-1)\n )\n\n loss.backward()\n optimizer.step()\n\n avg_loss += loss.item()\n it += 1\n\n if it % print_every == 0:\n print(f\"[Iter {it}] Loss {avg_loss / print_every:.4f}\")\n print(\" \" + sample_sequence(model, 140, 0.8))\n avg_loss = 0", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "455-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Training the Trump Tweet Generator" + }, + { + "text": "Explanation:\n## Generative RNN using GPU\nTraining a generative RNN can be a slow process. Here's a sample GPU implementation to speed up the training. The changes required to enable GPU are provided in the comments below.\n\nPython implementation:\n# Generative Recurrent Neural Network Implementation with GPU\n\ndef sample_sequence_cuda(model, max_len=100, temperature=0.8):\n generated_sequence = \"\"\n\n inp = torch.Tensor([vocab_stoi[\"\"]]).long().cuda() # <----- GPU\n hidden = None\n for p in range(max_len):\n output, hidden = model(inp.unsqueeze(0), hidden)\n # Sample from the network as a multinomial distribution\n output_dist = output.data.view(-1).div(temperature).exp().cpu()\n top_i = int(torch.multinomial(output_dist, 1)[0])\n # Add predicted character to string and use as next input\n predicted_char = vocab_itos[top_i]\n\n if predicted_char == \"\":\n break\n generated_sequence += predicted_char\n inp = torch.Tensor([top_i]).long().cuda() # <----- GPU\n return generated_sequence", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "456-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generative RNN using GPU" + }, + { + "text": "ut\n predicted_char = vocab_itos[top_i]\n\n if predicted_char == \"\":\n break\n generated_sequence += predicted_char\n inp = torch.Tensor([top_i]).long().cuda() # <----- GPU\n return generated_sequence\n\ndef train_cuda(model, data, batch_size=1, num_epochs=1, lr=0.001, print_every=100):\n optimizer = torch.optim.Adam(model.parameters(), lr=lr)\n criterion = nn.CrossEntropyLoss()\n it = 0\n data_iter = torchtext.data.BucketIterator(data,\n batch_size=batch_size,\n sort_key=lambda x: len(x.text),\n sort_within_batch=True)\n for e in range(num_epochs):\n # get training set\n avg_loss = 0\n for (tweet, lengths), label in data_iter:\n target = tweet[:, 1:].cuda() # <------- GPU\n inp = tweet[:, :-1].cuda() # <------- GPU\n # cleanup\n optimizer.zero_grad()\n # forward pass\n output, _ = model(inp)\n loss = criterion(output.reshape(-1, vocab_size), target.reshape(-1))\n # backward pass\n loss.backward()\n optimizer.step()\n\n avg_loss += loss\n it += 1 # increment iteration count\n if it % print_every == 0:", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "456-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generative RNN using GPU" + }, + { + "text": "= criterion(output.reshape(-1, vocab_size), target.reshape(-1))\n # backward pass\n loss.backward()\n optimizer.step()\n\n avg_loss += loss\n it += 1 # increment iteration count\n if it % print_every == 0:\n print(\"[Iter %d] Loss %f\" % (it+1, float(avg_loss/print_every)))\n print(\" \" + sample_sequence_cuda(model, 140, 0.8))\n avg_loss = 0\n\nmodel = TextGenerator(vocab_size, 64)\nmodel = model.cuda()\nmodel.ident = model.ident.cuda()\ntrain_cuda(model, trump_tweets, batch_size=32, num_epochs=1, lr=0.004, print_every=100)\n\nExplanation:\nLet's generate some results using different levels of temperature.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "456-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Generative RNN using GPU" + }, + { + "text": "Explanation:\n# PART 2 - Transfer Learning for Text\n\nExplanation:\nhttps://github.com/karthikayan4u/Simplified-ULMFIT-for-Twitter-Airlines-Sentiment-Classification/blob/master/SimplifiedULMFIT.ipynb\n\nPython implementation:\nimport pandas as pd\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom fastai.text import *\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.metrics import confusion_matrix\n\nPython implementation:\nimport gdown\nimport pandas as pd\n\n# Google Drive file ID\nfile_id = \"1xN4IO6-NieQ5hBYnZiCEDSmLkbePIXKE\"\n\n# Download the file\ngdown.download(\n f\"https://drive.google.com/uc?id={file_id}\",\n output=\"Tweets.csv\",\n quiet=False\n)\n\n# Load the CSV\ndf = pd.read_csv(\"Tweets.csv\")\n\n# Check the data\nprint(df.head())\nprint(df.shape)\n\nPython implementation:\n# Distribution of the respective 3 sentiments.\nsns.set(style=\"darkgrid\")\nax = sns.countplot(x=\"airline_sentiment\", data=df)\n\nPython implementation:\n# Distribution of the negative sentiment.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "entation:\n# Distribution of the respective 3 sentiments.\nsns.set(style=\"darkgrid\")\nax = sns.countplot(x=\"airline_sentiment\", data=df)\n\nPython implementation:\n# Distribution of the negative sentiment.\nax = sns.countplot(x='negativereason',data=df)\nax.set_xticklabels(ax.get_xticklabels(), rotation=60)\n\nPython implementation:\nimport re\n\ndef removeUnicode(text):\n \"\"\" Removes unicode strings like \"\\u002c\" and \"x96\" \"\"\"\n text = re.sub(r'(\\\\u[0-9A-Fa-f]+)',r'', text)\n text = re.sub(r'[^\\x00-\\x7f]',r'',text)\n return text\n\ndef replaceURL(text):\n \"\"\"Replaces url address with \"url\" \"\"\"\n text = re.sub('((www\\.[^\\s]+)|(https?://[^\\s]+))','url',text)\n text = re.sub(r'#([^\\s]+)', r'\\1', text)\n return text\n\ndef replaceAtUser(text):\n \"\"\" Replaces \"@user\" with \"atUser\" \"\"\"\n # text = re.sub('@[^\\s]+','atUser',text)\n text = re.sub('@[^\\s]+','',text)\n return text\n\ndef removeHashtagInFrontOfWord(text):\n \"\"\" Removes hastag in front of a word \"\"\"\n text = re.sub(r'#([^\\s]+)', r'\\1', text)\n return text", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "'atUser',text)\n text = re.sub('@[^\\s]+','',text)\n return text\n\ndef removeHashtagInFrontOfWord(text):\n \"\"\" Removes hastag in front of a word \"\"\"\n text = re.sub(r'#([^\\s]+)', r'\\1', text)\n return text\n\ndef removeNumbers(text):\n \"\"\" Removes integers \"\"\"\n text = ''.join([i for i in text if not i.isdigit()])\n return text\n\ndef removeEmoticons(text):\n \"\"\" Removes emoticons from text \"\"\"\n text = re.sub(':\\)|;\\)|:-\\)|\\(-:|:-D|=D|:P|xD|X-p|\\^\\^|:-*|\\^\\.\\^|\\^\\-\\^|\\^\\_\\^|\\,-\\)|\\)-:|:\\'\\(|:\\(|:-\\(|:\\S|T\\.T|\\.\\_\\.|:<|:-\\S|:-<|\\*\\-\\*|:O|=O|=\\-O|O\\.o|XO|O\\_O|:-\\@|=/|:/|X\\-\\(|>\\.<|>=\\(|D:', '', text)\n return text\n\n\"\"\" Replaces contractions from a string to their equivalents \"\"\"\ncontraction_patterns = [ (r'won\\'t', 'will not'), (r'can\\'t', 'cannot'), (r'i\\'m', 'i am'), (r'ain\\'t', 'is not'), (r'(\\w+)\\'ll', '\\g<1> will'), (r'(\\w+)n\\'t', '\\g<1> not'),", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "to their equivalents \"\"\"\ncontraction_patterns = [ (r'won\\'t', 'will not'), (r'can\\'t', 'cannot'), (r'i\\'m', 'i am'), (r'ain\\'t', 'is not'), (r'(\\w+)\\'ll', '\\g<1> will'), (r'(\\w+)n\\'t', '\\g<1> not'),\n (r'(\\w+)\\'ve', '\\g<1> have'), (r'(\\w+)\\'s', '\\g<1> is'), (r'(\\w+)\\'re', '\\g<1> are'), (r'(\\w+)\\'d', '\\g<1> would'), (r'&', 'and'), (r'dammit', 'damn it'), (r'dont', 'do not'), (r'wont', 'will not') ]\ndef replaceContraction(text):\n patterns = [(re.compile(regex), repl) for (regex, repl) in contraction_patterns]\n for (pattern, repl) in patterns:\n (text, count) = re.subn(pattern, repl, text)\n return text\n\nPython implementation:\ndef preprocessTwitterData(df):\n \"\"\"Function to apply text preprocessing functions to a dataframe\"\"\"\n\n # remove unicode\n df['text'] = df['text'].apply(removeUnicode)\n\n # replace url\n df['text'] = df['text'].apply(replaceURL)\n\n # replace '@' signs\n df['text'] = df['text'].apply(replaceAtUser)\n\n # replace hastags", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "nicode\n df['text'] = df['text'].apply(removeUnicode)\n\n # replace url\n df['text'] = df['text'].apply(replaceURL)\n\n # replace '@' signs\n df['text'] = df['text'].apply(replaceAtUser)\n\n # replace hastags\n df['text'] = df['text'].apply(removeHashtagInFrontOfWord)\n\n # remove numbers in the tweets\n df['text'] = df['text'].apply(removeNumbers)\n\n # remove the emoticons\n df['text'] = df['text'].apply(removeEmoticons)\n\n # replace contractions\n df['text'] = df['text'].apply(replaceContraction)\n\n# Call the function and preprocess the data\npreprocessTwitterData(df)\n\n# Since we don't need the rest of the columns in the data, subindex the relevant columns and make this the new dataframe\ndf = df[['text','airline_sentiment']]\n\n# Split the dataset into a train and test set.\n# Using a validation set is built into the fastai API, so we don't need to do this split ourselves\n\n# use an 80-20 split for the train and test sets\ndf_train, df_test = train_test_split(df,test_size=0.1,random_state=20)", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-4", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "set is built into the fastai API, so we don't need to do this split ourselves\n\n# use an 80-20 split for the train and test sets\ndf_train, df_test = train_test_split(df,test_size=0.1,random_state=20)\n\n# Convert the cleaned training and testing data into their own CSV files which we can import later to perform modeling on them\n\ndf_train.to_csv('twitter_data_cleaned_train.csv')\ndf_test.to_csv('twitter_data_cleaned_test.csv')\n\nPython implementation:\n# Create a 'TextLMDataBunch' from a csv file.\n# We specify 'valid=0.1' to signify that when we want to actually put this into our language model, we'll be setting off 10% of it for a validation set\ndata_batch = TextLMDataBunch.from_csv(path='',csv_name='twitter_data_cleaned_train.csv',valid_pct=0.1)\n\n# run this to see how the batch looks like\ndata_batch.show_batch()\n\nPython implementation:\n# pass in our 'data_lm' objet to specify our Twitter data\n# pass in AWD_LSTM to specify that we're using this particular language model", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-5", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "atch looks like\ndata_batch.show_batch()\n\nPython implementation:\n# pass in our 'data_lm' objet to specify our Twitter data\n# pass in AWD_LSTM to specify that we're using this particular language model\ntweet_model = language_model_learner(data_batch, AWD_LSTM, drop_mult=0.3)\n\nPython implementation:\n# fastai learning-rate finding\n# implemented using fastai callbacks\n#learning_rate_finder is used to find the optimal learning rate .\ntweet_model.lr_find()\n\nPython implementation:\n# plot the graph we were talking about earlier\ntweet_model.recorder.plot()\n\nPython implementation:\n# We set cycle_len to 1 because we only train with one epoch 'moms' refers to a tuple with the form (max_momentum,min_momentum)\ntweet_model.fit_one_cycle(cyc_len=1,max_lr=1e-1,moms=(0.85,0.75))\n\nPython implementation:\n# unfreeze the LSTM layers of the model\ntweet_model.unfreeze()\n\nPython implementation:\n#Now let's train the model\ntweet_model.fit_one_cycle(cyc_len=5, max_lr=slice(1e-1/(2.6**4),1e-1), moms=(0.85, 0.75))", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-6", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "reeze the LSTM layers of the model\ntweet_model.unfreeze()\n\nPython implementation:\n#Now let's train the model\ntweet_model.fit_one_cycle(cyc_len=5, max_lr=slice(1e-1/(2.6**4),1e-1), moms=(0.85, 0.75))\n\nPython implementation:\n# save the encoder\ntweet_model.save_encoder('encoder')\n\nPython implementation:\n# create 'TextClasDataBunch'\n# pass in vocab to ensure the vocab is the\n# same one that was modified in the fine-tuned LM\ndata_class = TextClasDataBunch.from_csv(path='',csv_name='twitter_data_cleaned_train.csv',\n vocab=data_batch.train_ds.vocab,bs=32,text_cols='text',label_cols='airline_sentiment')\n\n# show what our batch looks like\ndata_class.show_batch()\n\nPython implementation:\n# create new learner object with the 'text_classifier_learner' object.\n# The concept behind this learner is the same as the 'language_model_learner'.\n# It can similarly take in callbacks that allow us to train with special optimization methods. We use a slightly bigger dropout this time", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-7", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "ind this learner is the same as the 'language_model_learner'.\n# It can similarly take in callbacks that allow us to train with special optimization methods. We use a slightly bigger dropout this time\ntweet_model = text_classifier_learner(data_class, AWD_LSTM, drop_mult=0.5)\n\n# load the fine-tuned encoder onto the learner\ntweet_model.load_encoder('encoder')\n\n# look at the model\ntweet_model.model\n\nPython implementation:\n# find the optimal learning rate, just like we did before\ntweet_model.lr_find()\n\n# plot it\ntweet_model.recorder.plot()\n\nPython implementation:\n# like we did before, we choose a learning rate before\n# the minimum of the graph and use the 1cycle policy\ntweet_model.fit_one_cycle(5,1e-1,moms=(0.8,0.7))\n\nPython implementation:\n# unfreeze next layer\ntweet_model.freeze_to(-2)\n\n# train with next layer unfrozen, apply discriminative fine-tuning\ntweet_model.fit_one_cycle(5,slice(1e-2/(2.6**4),1e-2))\n\nPython implementation:\n# repeat the process\ntweet_model.freeze_to(-3)", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-8", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "o(-2)\n\n# train with next layer unfrozen, apply discriminative fine-tuning\ntweet_model.fit_one_cycle(5,slice(1e-2/(2.6**4),1e-2))\n\nPython implementation:\n# repeat the process\ntweet_model.freeze_to(-3)\ntweet_model.fit_one_cycle(5,slice(1e-2/(2.6**4),1e-2))\n\nPython implementation:\n# now unfreeze everything\ntweet_model.unfreeze()\ntweet_model.fit_one_cycle(5,slice(1e-2/(2.6**4),1e-2))", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "457-9", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "PART 2 - Transfer Learning for Text" + }, + { + "text": "Explanation:\n# Prediction Phase\n\nPython implementation:\n# put test data in test df\ndf_test = pd.read_csv('twitter_data_cleaned_test.csv')\nprint(df_test[['text','airline_sentiment']])\ndf_test.head()\n\nPython implementation:\n# add a column with the predictions on the test set\n\ndf_test['sentiment_pred'] = df_test['text'].apply(lambda row:str(tweet_model.predict(row)[0]))\n\nPython implementation:\n# print the accuracy against the test set\nprint(\"Accuracy: {}\".format(accuracy_score(df_test['airline_sentiment'],df_test[\n 'sentiment_pred'])))\n\nPython implementation:\nimport matplotlib.pyplot as plt\n# Taken from the scikit-learn documentation\ndef plot_confusion_matrix(y_true, y_pred, classes,\n normalize=False,\n title=None,\n cmap=plt.cm.Blues):\n \"\"\"\n This function prints and plots the confusion matrix.\n Normalization can be applied by setting `normalize=True`.\n \"\"\"\n if not title:\n if normalize:\n title = 'Normalized confusion matrix'\n else:\n title = 'Confusion matrix, without normalization'", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "458-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Prediction Phase" + }, + { + "text": "matrix.\n Normalization can be applied by setting `normalize=True`.\n \"\"\"\n if not title:\n if normalize:\n title = 'Normalized confusion matrix'\n else:\n title = 'Confusion matrix, without normalization'\n\n # Compute confusion matrix\n cm = confusion_matrix(y_true, y_pred)\n # Only use the labels that appear in the data\n #classes = classes[unique_labels(y_true, y_pred)]\n\n fig, ax = plt.subplots()\n im = ax.imshow(cm, interpolation='nearest', cmap=cmap)\n ax.figure.colorbar(im, ax=ax)\n # We want to show all ticks...\n ax.set(xticks=np.arange(cm.shape[1]),\n yticks=np.arange(cm.shape[0]),\n # ... and label them with the respective list entries\n xticklabels=classes, yticklabels=classes,\n title=title,\n ylabel='True label',\n xlabel='Predicted label')\n\n # Rotate the tick labels and set their alignment.\n plt.setp(ax.get_xticklabels(), rotation=45, ha=\"right\",\n rotation_mode=\"anchor\")\n\n # Loop over data dimensions and create text annotations.\n fmt = '.2f' if normalize else 'd'\n thresh = cm.max() / 2.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "458-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Prediction Phase" + }, + { + "text": "plt.setp(ax.get_xticklabels(), rotation=45, ha=\"right\",\n rotation_mode=\"anchor\")\n\n # Loop over data dimensions and create text annotations.\n fmt = '.2f' if normalize else 'd'\n thresh = cm.max() / 2.\n for i in range(cm.shape[0]):\n for j in range(cm.shape[1]):\n ax.text(j, i, format(cm[i, j], fmt),\n ha=\"center\", va=\"center\",\n color=\"white\" if cm[i, j] > thresh else \"black\")\n fig.tight_layout()\n return ax\n\nPython implementation:\n# plot the confusion matrix for the test set\nplot_confusion_matrix(df_test['airline_sentiment'],df_test['sentiment_pred'],\n classes=['negative','neutral','positive'])\nplt.show()", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "458-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Prediction Phase" + }, + { + "text": "Explanation:\n# Results\n\n## Summary\n\nThis notebook implemented a complete character-level language model using recurrent neural networks.\n\n### Implemented Components\n\n- Character tokenization\n- Vocabulary construction\n- Character embeddings\n- Recurrent Neural Networks\n- Teacher forcing\n- Sequence generation\n- Probabilistic sampling\n\n### Key Outcomes\n\n- Successfully trained a recurrent neural network for character-level language modeling.\n- Learned sequential dependencies between characters.\n- Generated realistic text by sampling from the learned probability distribution.\n- Demonstrated autoregressive sequence generation using recurrent hidden states.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "459-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Results" + }, + { + "text": "Explanation:\n# Conclusion\n\nThis notebook demonstrated how recurrent neural networks can be adapted from sequence prediction to sequence generation. By learning the probability distribution over characters, the model generates coherent text one character at a time using autoregressive decoding. These concepts provide the foundation for more advanced generative language models and modern sequence generation techniques.", + "source": "02_generative_rnn.ipynb", + "file_type": "ipynb", + "chunk_id": "460-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/02_generative_rnn.ipynb", + "section": "Conclusion" + }, + { + "text": "Explanation:\n# Spam Detection with LSTM Networks and Transfer Learning\n\nThis lab is based on an assignment developed by Prof. Lisa Zhang.\n\nIn this assignment, we will build a recurrent neural network to classify a SMS text message\nas \"spam\" or \"not spam\". In the process, you will\n\n1. Clean and process text data for machine learning.\n2. Understand and implement a character-level recurrent neural network.\n3. Use torchtext to build recurrent neural network models.\n4. Understand batching for a recurrent neural network, and use torchtext to implement RNN batching.\n5. Understand how transfer learning can be applied to NLP projects.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "461-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Spam Detection with LSTM Networks and Transfer Learning" + }, + { + "text": "Explanation:\n## Overview\n\nThis notebook explores deep learning approaches for SMS spam detection using recurrent neural networks and transfer learning. It presents two complementary text classification pipelines.\n\nThe first approach builds a Long Short-Term Memory (LSTM) network from scratch using character-level representations to classify SMS messages as spam or legitimate.\n\nThe second approach applies transfer learning using a pretrained language model (ULMFiT), demonstrating how pretrained representations can improve classification accuracy while reducing training requirements.\n\nThroughout the notebook, the complete NLP workflow is covered, including data preprocessing, vocabulary construction, sequence modeling, model training, evaluation, and comparison of both approaches.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "462-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Overview" + }, + { + "text": "Explanation:\n## Learning Objectives\n\nBy completing this notebook you will learn how to:\n\n- Build character-level LSTM classifiers\n- Preprocess SMS text datasets\n- Encode variable-length text sequences\n- Train recurrent neural networks for text classification\n- Evaluate binary classification performance\n- Interpret confusion matrices\n- Apply transfer learning using ULMFiT\n- Compare recurrent networks with pretrained language models", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "463-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Learning Objectives" + }, + { + "text": "Explanation:\n# PART A - Spam Detection\n\nIn this part we will construct a LSTM model for identifying spam from non spam messages.\n\nPython implementation:\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport numpy as np", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "464-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "PART A - Spam Detection" + }, + { + "text": "Explanation:\n## Part 1. Data Cleaning\n\nWe will be using the \"SMS Spam Collection Data Set\" available at http://archive.ics.uci.edu/ml/datasets/SMS+Spam+Collection\n\nThere is a link to download the \"Data Folder\" at the very top of the webpage. Download the zip file, unzip it, and upload the file `SMSSpamCollection` to Colab.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "465-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 1. Data Cleaning" + }, + { + "text": "Explanation:\n### Part (a)\n\nOpen up the file in Python, and print out one example of a spam SMS, and one example of a non-spam SMS.\n\n$\\color{blue}{\\text{ }}$\n\n- the label value for a spam message is spam \n\n- the label value for non-spam values is ham\n\nPython implementation:\nfor line in open('SMSSpamCollection'):\n if 'ham' in line:\n print(line)\n break\n\nfor line in open('SMSSpamCollection'):\n if 'spam' in line:\n print(line)\n break", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "466-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (a)" + }, + { + "text": "Explanation:\n### Part (b)\n\n$\\color{blue}{\\text{ }}$\n\n- There are 747 spam messages. \n\n- There are 4827 spam non-spam messages. \n\nPython implementation:\ni=0\nh=0\n\nfor line in open('SMSSpamCollection'):\n i+=1\n if 'spam' in line:\n h+=1\n\nprint(i-h, h)", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "467-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c) [2 pt]\n\nWe will be using the package `torchtext` to load, process, and batch the data.\nA tutorial to torchtext is available below. This tutorial uses the same\nSentiment140 data set that we explored during lecture.\n\nhttps://medium.com/@sonicboom8/sentiment-analysis-torchtext-55fb57b1fab8\n\nWe will be building a **character level RNN**.\nThat is, we will treat each **character** as a token in our sequence,\nrather than each **word**.\n\n$\\color{blue}{\\text{}}$\n\nAdvantages: \n- The discrete space we're working with is much smaller -- there are about 97 English-language characters in common usage if we include all punctuation marks. By contrast, a vocabulary is many thousands of words.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "468-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c) [2 pt]" + }, + { + "text": "te space we're working with is much smaller -- there are about 97 English-language characters in common usage if we include all punctuation marks. By contrast, a vocabulary is many thousands of words. This implies that just storing the word embeddings will require a lot of memory, and including word embeddings in a model adds many, many parameters to the model so the computational cost is much higher on this account than the character-level model. \n- Modeling as a series of characters can be used to simulate the correct sequence in a variety of languages.\n\nDisadvantages: \n- It's harder to train. The cross entropy loss takes the sum over all elements in each sentence. For a word-level RNN, the number of elements equals the number of words in the sentence, whereas for a character-level RNN, the number of elements equals the number of characters in the sentence.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "468-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c) [2 pt]" + }, + { + "text": "nce. For a word-level RNN, the number of elements equals the number of words in the sentence, whereas for a character-level RNN, the number of elements equals the number of characters in the sentence. Thus, in training, it takes a longer path to propagate the error from the softmax at the last time step back to the beginning.. \n- Character level RNN should make more prediction. The more predictions the RNN has to make, the more error-prone the result is. \n\n- Character level models give up the semantic information that words have, as well as the plug and play ecosystem of pre-trained word vectors", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "468-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c) [2 pt]" + }, + { + "text": "Explanation:\n### Part (d)\nWe will be loading our data set using `torchtext.data.TabularDataset`. The\nconstructor will read directly from the `SMSSpamCollection` file.\n\nFor the data file to be read successfuly, we\nneed to specify the **fields** (columns) in the file.\nIn our case, the dataset has two fields:\n\n- a text field containing the sms messages,\n- a label field which will be converted into a binary label.\n\nSplit the dataset into `train`, `valid`, and `test`. Use a 60-20-20 split.\nYou may find this torchtext API page helpful:\nhttps://torchtext.readthedocs.io/en/latest/data.html#dataset\n\nHint: There is a `Dataset` method that can perform the random split for you.\n\nPython implementation:\nimport torchtext\n\ntext_field = torchtext.legacy.data.Field(sequential=True, # text sequence\n tokenize=lambda x: x, # because are building a character-RNN\n include_lengths=True, # to track the length of sequences, for batching\n batch_first=True,", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "469-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (d)" + }, + { + "text": "cy.data.Field(sequential=True, # text sequence\n tokenize=lambda x: x, # because are building a character-RNN\n include_lengths=True, # to track the length of sequences, for batching\n batch_first=True,\n use_vocab=True) # to turn each character into an integer index\nlabel_field = torchtext.legacy.data.Field(sequential=False, # not a sequence\n use_vocab=False, # don't need to track vocabulary\n is_target=True,\n batch_first=True,\n preprocessing=lambda x: int(x == 'spam')) # convert text to 0 and 1\n\nfields = [('label', label_field), ('sms', text_field)]\ndataset = torchtext.legacy.data.TabularDataset(\"SMSSpamCollection\", # name of the file\n \"tsv\", # fields are separated by a tab\n fields)\n\nprint(dataset[0].sms)\nprint(dataset[0].label)\ntrain, valid, test = dataset.split(split_ratio=[0.6,0.2,0.2])", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "469-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n### Part (e)\n\nYou saw in part (b) that there are many more non-spam messages than spam messages.\nThis **imbalance** in our training data will be problematic for training.\nWe can fix this disparity by duplicating spam messages in the training set,\nso that the training set is roughly **balanced**.\n\nExplain why having a balanced training set is helpful for training our neural network.\n\nNote: if you are not sure, try removing the below code and train your mode.\n\n$\\color{blue}{\\text{ }}$\n\n- Imbalanced data is one of the potential problems in the field of data mining and machine learning. Balancing training data is an important part of data preprocessing. Data imbalance refers to when the classes in a dataset are not equally distributed, which can then lead to potential risks in training a model. By balancing the data, our model would be exposed and trained on almost as many spam messages as non-spam messages\n\nPython implementation:", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "470-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (e)" + }, + { + "text": "ad to potential risks in training a model. By balancing the data, our model would be exposed and trained on almost as many spam messages as non-spam messages\n\nPython implementation:\n# save the original training examples\nold_train_examples = train.examples\n# get all the spam messages in `train`\ntrain_spam = []\nfor item in train.examples:\n if item.label == 1:\n train_spam.append(item)\n# duplicate each spam message 6 more times\ntrain.examples = old_train_examples + train_spam * 6", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "470-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (e)" + }, + { + "text": "Explanation:\n### Part (f)\n\nWe need to build the vocabulary on the training data by running the below code.\nThis finds all the possible character tokens in the training set.\n\nExplain what the variables `text_field.vocab.stoi` and `text_field.vocab.itos` represent.\n\n$\\color{blue}{\\text{}}$\n\n- **stoi** is mapping token strings to numerical identifiers.\n\n- **itos** is a list of token strings indexed by their numerical identifiers.\n\nPython implementation:\ntext_field.build_vocab(train)\n#text_field.vocab.stoi\n#text_field.vocab.itos", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "471-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (f)" + }, + { + "text": "Explanation:\n### Part (g)\n\nThe tokens `` and `` were not in our SMS text messages.\nWhat do these two values represent?\n\n$\\color{blue}{\\text{ }}$\n\n- **UNK** means unknown word, a word that doesn't exist the the vocabulary set. It represents unknown or unseen characters\n\n- **pad** means adding padding token before or after a character to allow sentences' lengths to be equal.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "472-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (g)" + }, + { + "text": "Explanation:\n### Part (h)\n\nSince text sequences are of variable length, `torchtext` provides a `BucketIterator` data loader,\nwhich batches similar length sequences together. The iterator also provides functionalities to\npad sequences automatically.\n\nTake a look at 10 batches in `train_iter`. What is the maximum length of the\ninput sequence in each batch? How many `` tokens are used in each of the 10\nbatches?\n\nPython implementation:\ntrain_iter = torchtext.legacy.data.BucketIterator(train,\n batch_size=32,\n sort_key=lambda x: len(x.sms), # to minimize padding\n sort_within_batch=True, # sort within each batch\n repeat=False) # repeat the iterator for many epochs\n\nPython implementation:\ni=0\nfor batch in train_iter:\n print('The maximum length of the input sequence in batch', i+1, 'is equal to', torch.max(batch.sms[1]))\n N_pad=0\n for j in range(len(batch.sms[0])):\n N_pad=N_pad+np.count_nonzero(np.array(batch.sms[0][j])==1) #pad token has index of 1", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "473-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (h)" + }, + { + "text": "ut sequence in batch', i+1, 'is equal to', torch.max(batch.sms[1]))\n N_pad=0\n for j in range(len(batch.sms[0])):\n N_pad=N_pad+np.count_nonzero(np.array(batch.sms[0][j])==1) #pad token has index of 1\n\n print(\"Number of tokens are used=\", N_pad)\n i+=1\n if i==10:\n break", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "473-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (h)" + }, + { + "text": "Explanation:\n## Part 2. Model Building\n\nBuild a recurrent neural network model, using an architecture of your choosing.\nUse the one-hot embedding of each character as input to your recurrent network.\nUse one or more fully-connected layers to make the prediction based on your\nrecurrent network output.\n\nInstead of using the RNN output value for the final token, another often used\nstrategy is to max-pool over the entire output array. That is, instead of calling\nsomething like:\n\n```\nout, _ = self.rnn(x)\nself.fc(out[:, -1, :])\n```\n\nwhere `self.rnn` is an `nn.RNN`, `nn.GRU`, or `nn.LSTM` module, and `self.fc` is a\nfully-connected\nlayer, we use:\n\n```\nout, _ = self.rnn(x)\nself.fc(torch.max(out, dim=1)[0])\n```\n\nThis works reasonably in practice. An even better alternative is to concatenate the\nmax-pooling and average-pooling of the RNN outputs:\n\n```\nout, _ = self.rnn(x)\nout = torch.cat([torch.max(out, dim=1)[0],\n torch.mean(out, dim=1)], dim=1)\nself.fc(out)\n```", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "474-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 2. Model Building" + }, + { + "text": "ative is to concatenate the\nmax-pooling and average-pooling of the RNN outputs:\n\n```\nout, _ = self.rnn(x)\nout = torch.cat([torch.max(out, dim=1)[0],\n torch.mean(out, dim=1)], dim=1)\nself.fc(out)\n```\n\nWe encourage you to try out all these options. The way you pool the RNN outputs\nis one of the \"hyperparameters\" that you can choose to tune later on.\n\nPython implementation:\ndef one_hot_encoding(x):\n text_field.build_vocab(train)\n ident = torch.eye(len(text_field.vocab.stoi))\n for i in range(0,len(x)):\n a = ident[x[i]].unsqueeze(0)\n if i==0:\n a_new=a\n else:\n a_new=torch.cat((a_new,a),dim=0)\n return a_new\n\nPython implementation:\nclass RNN_char(nn.Module):\n def __init__(self, input_size, hidden_size, num_classes, method=1):\n super(RNN_char, self).__init__()\n self.ohe = one_hot_encoding\n self.hidden_size = hidden_size\n self.rnn = nn.RNN(input_size, hidden_size, batch_first=True)\n if method==1:\n self.fc = nn.Linear(hidden_size, num_classes)\n else:", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "474-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 2. Model Building" + }, + { + "text": "nit__()\n self.ohe = one_hot_encoding\n self.hidden_size = hidden_size\n self.rnn = nn.RNN(input_size, hidden_size, batch_first=True)\n if method==1:\n self.fc = nn.Linear(hidden_size, num_classes)\n else:\n self.fc = nn.Linear(hidden_size*2, num_classes)\n self.method=method\n\n def forward(self, x):\n method=self.method\n x = self.ohe(x)\n h0 = torch.zeros(1, x.size(0), self.hidden_size)\n out, _ = self.rnn(x, h0)\n if method==1:\n out = self.fc(torch.max(out, dim=1)[0])\n else:\n out = torch.cat([torch.max(out, dim=1)[0], torch.mean(out, dim=1)], dim=1)\n self.fc(out)\n return out", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "474-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 2. Model Building" + }, + { + "text": "Explanation:\n## Part 3. Training\n\n### Part (a)\n\nComplete the `get_accuracy` function, which will compute the\naccuracy (rate) of your model across a dataset (e.g. validation set).\nYou may modify `torchtext.data.BucketIterator` to make your computation\nfaster.\n\nPython implementation:\nvalid_iter = torchtext.legacy.data.BucketIterator(valid,\n batch_size=32,\n sort_key=lambda x: len(x.sms), # to minimize padding\n sort_within_batch=True, # sort within each batch\n repeat=False)\n\nPython implementation:\ndef get_accuracy(model, data, criterion):\n \"\"\" Compute the accuracy of the `model` across a dataset `data`\n\n Example usage:\n\n >>> model = MyRNN() # to be defined\n >>> get_accuracy(model, valid) # the variable `valid` is from above\n \"\"\"\n correct, total = 0, 0\n total_loss = 0.0\n i=0\n for batch in data:\n length=len(batch)\n text=batch.sms[0]\n labels=batch.label\n output = model(text)\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += length", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "475-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 3. Training" + }, + { + "text": "in data:\n length=len(batch)\n text=batch.sms[0]\n labels=batch.label\n output = model(text)\n pred = output.max(1, keepdim=True)[1]\n correct += pred.eq(labels.view_as(pred)).sum().item()\n total += length\n loss = criterion(output, labels)\n total_loss += loss.item()\n i+=1\n loss = float(total_loss) / (i + 1)\n\n return correct/total, loss", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "475-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 3. Training" + }, + { + "text": "Explanation:\n### Part (b)\n\nTrain your model. Plot the training curve of your final model.\nYour training curve should have the training/validation loss and\naccuracy plotted periodically.\n\nNote: Not all of your batches will have the same batch size.\nIn particular, if your training set does not divide evenly by\nyour batch size, there will be a batch that is smaller than\nthe rest.\n\nPython implementation:\nimport matplotlib.pyplot as plt\n\ndef train_rnn_network(model, train, valid, num_epochs=5, learning_rate=1e-5):\n criterion = nn.CrossEntropyLoss()\n optimizer = torch.optim.Adam(model.parameters(), lr=learning_rate)\n losses, train_acc, valid_acc = [], [], []\n train_loss, valid_loss=[], []\n epochs = []\n for epoch in range(num_epochs):\n for batch in train:\n optimizer.zero_grad()\n pred = model(batch.sms[0])\n loss = criterion(pred, batch.label)\n loss.backward()\n optimizer.step()\n losses.append(float(loss))\n\n epochs.append(epoch)\n train_acc.append(get_accuracy(model, train, criterion)[0])", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "476-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "el(batch.sms[0])\n loss = criterion(pred, batch.label)\n loss.backward()\n optimizer.step()\n losses.append(float(loss))\n\n epochs.append(epoch)\n train_acc.append(get_accuracy(model, train, criterion)[0])\n valid_acc.append(get_accuracy(model, valid, criterion)[0])\n train_loss.append(get_accuracy(model, train, criterion)[1])\n valid_loss.append(get_accuracy(model, valid, criterion)[1])\n print(\"Epoch %d; Loss %f; Train Acc %f; Val Acc %f; average Train loss %f; average Val loss %f\" % (\n epoch+1, loss, train_acc[-1], valid_acc[-1], train_loss[-1], valid_loss[-1]))\n # plotting\n\n plt.title(\"Training Curve\")\n plt.plot(epochs, train_acc, label=\"Train\")\n plt.plot(epochs, valid_acc, label=\"Validation\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"Accuracy\")\n plt.legend(loc='best')\n plt.show()\n\n plt.title(\"Training Curve\")\n plt.plot(epochs, train_loss, label=\"Train\")\n plt.plot(epochs, valid_loss, label=\"Validation\")\n plt.xlabel(\"Epoch\")\n plt.ylabel(\"average loss\")\n plt.legend(loc='best')\n plt.show()", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "476-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\n\nChoose at least 4 hyperparameters to tune. Explain how you tuned the hyperparameters.\nYou don't need to include your training curve for every model you trained.\nInstead, explain what hyperparemters you tuned, what the best validation accuracy was,\nand the reasoning behind the hyperparameter decisions you made.\n\n$\\color{blue}{\\text{ }}$\n\n- Hyperparameter-1: learning rate\n\n\nDecreasing learning rate decrease accuracy and increasing learning rate will not change the accuracy significantly.\n\n\nPython implementation:\nmodel3=RNN_char(len(text_field.vocab.stoi),50,2, 1)\ntrain_rnn_network(model3, train_iter, valid_iter, num_epochs=20, learning_rate=1e-6)\n\nExplanation:\n$\\color{blue}{\\text{Hyperparameter-2: }}$\n\n- Number of epoch ", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "477-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": "1)\ntrain_rnn_network(model3, train_iter, valid_iter, num_epochs=20, learning_rate=1e-6)\n\nExplanation:\n$\\color{blue}{\\text{Hyperparameter-2: }}$\n\n- Number of epoch \n\n- Decreasing Number of epoch will decrease accuracy and increasing number of epoch more than 20 will not change accuracy of the model \n\nPython implementation:\nmodel4=RNN_char(len(text_field.vocab.stoi),50,2, 1)\ntrain_rnn_network(model4, train_iter, valid_iter, num_epochs=5, learning_rate=1e-4)\n\nExplanation:\n$\\color{blue}{\\text{Hyperparameter-3: }}$\n\n- Two ways that we pool the RNN outputs is one of the \"hyperparameters\" that you can choose to tune later on.\nmethod 1 \n```\nout, _ = self.rnn(x)\nself.fc(torch.max(out, dim=1)[0])\n```\n\nmethod 2 \n\n```\nout, _ = self.rnn(x)\nout = torch.cat([torch.max(out, dim=1)[0],\n torch.mean(out, dim=1)], dim=1)\nself.fc(out)\n```\n", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "477-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": "dim=1)[0])\n```\n\nmethod 2 \n\n```\nout, _ = self.rnn(x)\nout = torch.cat([torch.max(out, dim=1)[0],\n torch.mean(out, dim=1)], dim=1)\nself.fc(out)\n```\n\n\n- In this part the second method has been used to see the imapct of that method on the output. This method has not improved the model.\n\nExplanation:\n$\\color{blue}{\\text{Hyperparameter-4: }}$\n\n- I use a multi-layer long short-term memory (LSTM) RNN in this section. The accuracy of model over train set decreased and it increased over validation set. So, we would choose base model as the best model for next part.\n \n\nPython implementation:\nclass RNN_LSTM(nn.Module):\n def __init__(self, input_size, hidden_size, num_classes, method=1):\n super(RNN_LSTM, self).__init__()\n self.ohe = one_hot_encoding\n self.hidden_size = hidden_size\n self.lstm = nn.LSTM(input_size, hidden_size, batch_first=True)\n if method==1:", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "477-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": ", num_classes, method=1):\n super(RNN_LSTM, self).__init__()\n self.ohe = one_hot_encoding\n self.hidden_size = hidden_size\n self.lstm = nn.LSTM(input_size, hidden_size, batch_first=True)\n if method==1:\n self.fc = nn.Linear(hidden_size, num_classes)\n else:\n self.fc = nn.Linear(hidden_size*2, num_classes)\n self.method=method\n\n def forward(self, x):\n method=self.method\n x = self.ohe(x)\n h0 = torch.zeros(1, x.size(0), self.hidden_size)\n c0 = torch.zeros(1, x.size(0), self.hidden_size)\n out, _ = self.lstm(x, (h0,c0))\n if method==1:\n out = self.fc(torch.max(out, dim=1)[0])\n else:\n out = torch.cat([torch.max(out, dim=1)[0], torch.mean(out, dim=1)], dim=1)\n self.fc(out)\n return out\n\nExplanation:\n$\\color{blue}{\\text{Hyperparameter-5: Increasing Hidden Size }}$\n\n- I want to see the impact of increasing hidden size on the result. This is the best model based on accuracy.\n ", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "477-3", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (d)\n\nBefore we deploy a machine learning model, we usually want to have a better understanding\nof how our model performs beyond its validation accuracy. An important metric to track is\n*how well our model performs in certain subsets of the data*.\n\nIn particular, what is the model's error rate amongst data with negative labels?\nThis is called the **false positive rate**.\n\nWhat about the model's error rate amongst data with positive labels?\nThis is called the **false negative rate**.\n\nReport your final model's false positive and false negative rate across the\nvalidation set.\n\nPython implementation:\n# Create a Dataset of only spam validation examples\nvalid_spam = torchtext.legacy.data.Dataset(\n [e for e in valid.examples if e.label == 1],\n valid.fields)\n# Create a Dataset of only non-spam validation examples\nvalid_nospam = torchtext.legacy.data.Dataset(\n [e for e in valid.examples if e.label == 0],\n valid.fields) # TODO\n\nPython implementation:", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "478-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (d)" + }, + { + "text": "elds)\n# Create a Dataset of only non-spam validation examples\nvalid_nospam = torchtext.legacy.data.Dataset(\n [e for e in valid.examples if e.label == 0],\n valid.fields) # TODO\n\nPython implementation:\nvalid_spam_iter=torchtext.legacy.data.BucketIterator(valid_spam,\n batch_size=32,\n sort_key=lambda x: len(x.sms), # to minimize padding\n sort_within_batch=True, # sort within each batch\n repeat=False)\n\nPython implementation:\ncriterion = nn.CrossEntropyLoss()\nfalse_positive_rate=get_accuracy(model7, valid_spam_iter, criterion)[0]\nprint(\"The false positive rate of the best model is:\", 1-false_positive_rate)\n\nPython implementation:\nvalid_nospam_iter=torchtext.legacy.data.BucketIterator(valid_nospam,\n batch_size=32,\n sort_key=lambda x: len(x.sms), # to minimize padding\n sort_within_batch=True, # sort within each batch\n repeat=False)", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "478-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n### Part (e)\n\nThe impact of a false positive vs a false negative can be drastically different.\nIf our spam detection algorithm was deployed on your phone, what is the impact\nof a false positive on the phone's user? What is the impact of a false negative?\n\n$\\color{blue}{\\text{ }}$\n\n- The effect of a **false positive** is more dangerous because it would consider an sms that is not spam to be spam. It could be inefficient because if we had an important sms that was sent to spam when it was actually not spam, we may never know about that important message or we may retrieve and read that message so late.\n\n- On the other hand, the effect of a **false negative** is not alarming, and this is where it would allow a spam sms to pass through. So having a false negative rate of 0.0317 could be considered good, beacuse people can read the spams that has been labeled as non spam one and delete them without any consequence.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "479-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (e)" + }, + { + "text": "ms to pass through. So having a false negative rate of 0.0317 could be considered good, beacuse people can read the spams that has been labeled as non spam one and delete them without any consequence. ", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "479-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (e)" + }, + { + "text": "Explanation:\n## Part 4. Evaluation\n\n### Part (a)\n\nReport the final test accuracy of your model.\n\nPython implementation:\ntest_iter = torchtext.legacy.data.BucketIterator(test,\n batch_size=32,\n sort_key=lambda x: len(x.sms),\n sort_within_batch=True,\n repeat=False)\n\nPython implementation:\ncriterion = nn.CrossEntropyLoss()\ntest_acc=get_accuracy(model7, test_iter, criterion)[0]\nprint(\"The accuracy of the best model for test set is:\", test_acc)", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "480-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 4. Evaluation" + }, + { + "text": "Explanation:\n### Part (b)\n\nReport the false positive rate and false negative rate of your model across the test set.\n\nPython implementation:\n# Create a Dataset of only spam validation examples\ntest_spam = torchtext.legacy.data.Dataset(\n [e for e in test.examples if e.label == 1],\n valid.fields)\n# Create a Dataset of only non-spam validation examples\ntest_nospam = torchtext.legacy.data.Dataset(\n [e for e in test.examples if e.label == 0],\n valid.fields) # TODO\n\nPython implementation:\ntest_spam_iter=torchtext.legacy.data.BucketIterator(test_spam,\n batch_size=32,\n sort_key=lambda x: len(x.sms), # to minimize padding\n sort_within_batch=True, # sort within each batch\n repeat=False)\n\nPython implementation:\ncriterion = nn.CrossEntropyLoss()\nfalse_positive_rate=get_accuracy(model7, test_spam_iter, criterion)[0]\nprint(\"The false positive rate of the best model is:\", 1-false_positive_rate)\n\nPython implementation:\ntest_nospam_iter=torchtext.legacy.data.BucketIterator(test_nospam,\n batch_size=32,", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "481-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "criterion)[0]\nprint(\"The false positive rate of the best model is:\", 1-false_positive_rate)\n\nPython implementation:\ntest_nospam_iter=torchtext.legacy.data.BucketIterator(test_nospam,\n batch_size=32,\n sort_key=lambda x: len(x.sms), # to minimize padding\n sort_within_batch=True, # sort within each batch\n repeat=False)", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "481-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\n\nWhat is your model's prediction of the **probability** that\nthe SMS message \"machine learning is sooo cool!\" is spam?\n\nHint: To begin, use `text_field.vocab.stoi` to look up the index\nof each character in the vocabulary.\n\nPython implementation:\nmsg = \"machine learning is sooo cool!\"\nmsg_new=list(tuple(msg))\nmsg_new2=[]\nfor i in enumerate(msg_new):\n msg_new2.append(text_field.vocab.stoi[i[1]])\n\nmsg_new2=torch.tensor(msg_new2).unsqueeze(0)\nprob=model7(msg_new2)\n\nprob\n\nPython implementation:\nimport math\n\ndef sigmoid(x):\n return 1 / (1 + math.exp(-x))", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "482-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n### Part (d)\n\nDo you think detecting spam is an easy or difficult task?\n\nSince machine learning models are expensive to train and deploy, it is very\nimportant to compare our models against baseline models: a simple\nmodel that is easy to build and inexpensive to run that we can compare our\nrecurrent neural network model against.\n\nExplain how you might build a simple baseline model. This baseline model\ncan be a simple neural network (with very few weights), a hand-written algorithm,\nor any other strategy that is easy to build and test.\n\n**Do not actually build a baseline model. Instead, provide instructions on\nhow to build it.**\n\n$\\color{blue}{\\text{ }}$\n\n- Detecting Spam is a difficult task because companies and scammers always try to write their messages in a way to not be classified as spam and hackers try to find new ways to send messages that can be labeled as no spam.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "483-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (d)" + }, + { + "text": "se companies and scammers always try to write their messages in a way to not be classified as spam and hackers try to find new ways to send messages that can be labeled as no spam.\n\n- There are different ways for detecting spam emails. Some of these ways like support vector machines (SVM), Content Based Filtering Technique, and etc has been mentioned in articles.\n\n- One of the methods that has been used widely is solving the problem using a classifier method like SVM.\n\n- I think using SVM would solve this problem properly.\n\n\nThe way we would build the simple baseline model using SVM is the following:\n\n1. Load the dataset containing the spam/ham emails (make sure the labels and emails are in separate dataframes).\n\n2. Split data into training, validation and testing.\n\n3.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "483-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (d)" + }, + { + "text": "using SVM is the following:\n\n1. Load the dataset containing the spam/ham emails (make sure the labels and emails are in separate dataframes).\n\n2. Split data into training, validation and testing.\n\n3. Convert the text to integers using CountVectorizer() from sklearn (similar to count number of differet words in every email).\n\n4. Perform support vector classification using SVC from sklearn, which is classification method for classifing spam/non-spam.\n\nThis method has been used in the following link:\nhttps://www.youtube.com/watch?v=exHwwy9kVcg\n\nIn the link they achieved 98 accuracy.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "483-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (d)" + }, + { + "text": "Explanation:\n# PART B - Transfer Learning\n\nIn this part we will compare our earlier model with one that takes advantage of a generative RNN model to improve the prediction. There are several ways to implement transfer learning with RNNs, here we will use an approach known as ULMFiT developed by fastai. Rather than rebuilding the model from scratch, we will take advantage of the fastai library.\n\nProvided below is some helper code to get you started.\n\n#### Helper Code\n\nPython implementation:\n# install relevant libraries\n!pip install fastai\n\nPython implementation:\n# load relevant libraries\nfrom fastai import *\nimport pandas as pd\nimport numpy as np\nfrom functools import partial\nimport io\nimport os\nfrom fastai.text import *\n\nPython implementation:\n# download SPAM data\n!wget https://archive.ics.uci.edu/ml/machine-learning-databases/00228/smsspamcollection.zip\n!unzip smsspamcollection.zip\n\nExplanation:\nThis time we will load the data using pandas.\n\nPython implementation:", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "484-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "PART B - Transfer Learning" + }, + { + "text": "https://archive.ics.uci.edu/ml/machine-learning-databases/00228/smsspamcollection.zip\n!unzip smsspamcollection.zip\n\nExplanation:\nThis time we will load the data using pandas.\n\nPython implementation:\n# set up data and verify\ndf1 = pd.read_csv('SMSSpamCollection', sep='\\t', header=None, names=['target', 'text'])\ndf1.head()\n\nPython implementation:\n# check distribution\ndf1['target'].value_counts()\n\nExplanation:\nSplit the data into training and validation datasets.\n\nPython implementation:\n# split the data and check dimensions\n\nfrom sklearn.model_selection import train_test_split\n\n# split data into training and validation set\ndf_trn, df_val = train_test_split(df1, stratify = df1['target'], test_size = 0.3, random_state = 999)", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "484-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "PART B - Transfer Learning" + }, + { + "text": "Explanation:\n### Create the language model\nEsentially, the language model contains the structure of the language (English in this case), allowing us to quickly use in a classification model, skipping the part of learning the semantics of the language from scratch.\n\nCreating a language model from scratch can be intensive due to the sheer size of data. Instead we will download the pre-trained model, which is a neural network (NN) with an AWD_LSTM architecture. By setting pretrained = True we say to fastai to download the weights from the trained model (a corpus of 103 MM of wikipedia articles).\n\nPython implementation:\n# create pretrained language model data\ndata_lm = TextLMDataBunch.from_df(train_df = df_trn, valid_df = df_val, path = \"\")\nlang_mod = language_model_learner(data_lm, arch = AWD_LSTM, pretrained = True, drop_mult=1.)", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "485-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Create the language model" + }, + { + "text": "Explanation:\n### Testing the language model\n\nExplanation:\nEach time we excecute the `predict`, we get a different random sentence, completed with the number of choosen words (`n_words`).\n\nTry your own sentences!", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "486-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Testing the language model" + }, + { + "text": "Explanation:\n### Fine-tuning the language model\nThe language model that we have \"loaded\" is great for generating wikipedia-like sentances, but here we're more interested in generating data like our email dataset.\n\nMake sure to enable GPU for this step or it make takes several hours to train.\n\nPython implementation:\n# fine-tune language model data\nlang_mod.fit_one_cycle(4, max_lr= 5e-02)\nlang_mod.freeze_to(-1)\nlang_mod.fit_one_cycle(3, slice(1e-2/(2.6**4), 1e-2))\nlang_mod.freeze_to(-2)\nlang_mod.fit_one_cycle(3, slice(3e-3/(2.6**4), 1e-3))\nlang_mod.unfreeze()\nlang_mod.fit_one_cycle(3, slice(3e-3/(2.6**4), 1e-3))\n\n# save language model\nlang_mod.save_encoder('my_awsome_encoder')", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "487-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Fine-tuning the language model" + }, + { + "text": "Explanation:\n### Classification model\nNow we can train a classification model that will identify spam and non-spam messages. Since we used a fastai language model, it will be easier to just continue working with the fastai library.\n\nPython implementation:\n# Classifier model data\ndata_clas = TextClasDataBunch.from_df(path = \"\", train_df = df_trn, valid_df = df_val, vocab=data_lm.train_ds.vocab, bs=32)\n\nPython implementation:\n# create the classifier\nlearn_classifier = text_classifier_learner(data_clas, drop_mult=0.7, arch = AWD_LSTM)\n\nPython implementation:\n# load language model\nlearn_classifier.load_encoder('my_awsome_encoder')\n\nPython implementation:\n# train classifier\nlearn_classifier.lr_find()\nlearn_classifier.recorder.plot(suggestion=True)\n\nPython implementation:\nlang_mod.freeze_to(-1)\n\nlearn_classifier.lr_find()\nlearn_classifier.recorder.plot(suggestion=True)\n\nExplanation:\nTest out the classification model on spam and non-spam examples.\n\nPython implementation:\n# predict", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "488-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Classification model" + }, + { + "text": "eze_to(-1)\n\nlearn_classifier.lr_find()\nlearn_classifier.recorder.plot(suggestion=True)\n\nExplanation:\nTest out the classification model on spam and non-spam examples.\n\nPython implementation:\n# predict\nlearn_classifier.predict('did you buy the groceries for dinner? :)')\n\nPython implementation:\n# predict\nlearn_classifier.predict('Free entry call back now')\n\nExplanation:\nNext we will evaluate on all of our validation data.\n\nPython implementation:\n# get predictions from validation\nvalid_preds, valid_label=learn_classifier.get_preds(ds_type=DatasetType.Valid, ordered=True)\nvalid_preds.shape", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "488-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Classification model" + }, + { + "text": "Explanation:\n## Part 1. Evaluate Performance\n\n### Part (a)\n\nImplement the above helper code for spam detection.\n\nWhat is the accuracy obtained with ULMFiT? How does ULMFiT compare to the approach in the first part using only LSTM?\n\n$\\color{blue}{\\text{ }}$\n\n- The accuracy obtained with ULMFiT is 0.9802 that is higher than the accuracy from previous part .", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "489-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 1. Evaluate Performance" + }, + { + "text": "Explanation:\n### Part (b)\nProvide a confusion matrix of the performance for the two models. How do they compare? Are there any qualitative differences between the performances (i.e. examine the samples for which the models differred)?\n\n$\\color{blue}{\\text{ }}$\n\n- The Model from part-B (using transfer learning) is doing better than the model from part-A with RNN , we can see qualitatively that the accuracy is higher. The qualitative difference is that the model from Part_B is performing better for detecting the no-spam messages while the model from part_A is performing better in detecting the spam messages. The false positive rate of part-B is performing way better. The false negative rate, only part_A model is performing better.\n\n- AS we discussed **False positive** is more dangerous. So, The model from Part_B with transfer learning is far better.\n\nPython implementation:", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "490-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "\n\n- AS we discussed **False positive** is more dangerous. So, The model from Part_B with transfer learning is far better.\n\nPython implementation:\nfrom sklearn.metrics import confusion_matrix\npred_label=torch.max(valid_preds,1)[1]\n\ncm1=confusion_matrix(list(valid_label),list(pred_label))\ncm1\n\nPython implementation:\nimport torchtext\n\nvalid_iter3= torchtext.legacy.data.BucketIterator(df_val,\n batch_size=32,\n sort_key=lambda x: len(x.sms),\n sort_within_batch=True,\n repeat=False)\n\nPython implementation:\ni=0\nfor batch in valid_iter:\n y=batch.sms[0]\n label=batch.label\n output = model7(y)\n pred = output.max(1, keepdim=True)[1]\n\n if i==0:\n valid_label_LSTM=batch.label\n valid_pred_LSTM=pred.reshape(1,-1)[0]\n else:\n valid_label_LSTM=torch.cat((valid_label_LSTM,batch.label),0)\n valid_pred_LSTM=torch.cat((valid_pred_LSTM,pred.reshape(1,-1)[0]),0)\n i+=1\n\nPython implementation:\ncm2=confusion_matrix(list(valid_label_LSTM),list(valid_pred_LSTM))\ncm2", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "490-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n## Part 2. Evaluate on New Data\n\n### Part (a)\nWhat is your model's prediction of the probability that the SMS message \"machine learning is sooo cool!\" is spam?\n\nPython implementation:\npred_msgml=learn_classifier.predict('machine learning is sooo cool!')\npred_msgml", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "491-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part 2. Evaluate on New Data" + }, + { + "text": "Explanation:\n### Part (b)\nLoad 5 sample sentences from your spam mail and test it out out the two models you created. How well do they perform?\n\nPython implementation:\nmsgs_spam_email=['You have won a £2,000 price! To claim, call 09050000301.',\n 'You have won a guaranteed 32000 award or maybe even £1000 cash to claim ur award call',\n 'You are awarded a SiPix Digital Camera! call 09061221061 from landline. Delivery within 28days.!',\n 'ITS YOUR CHANCE to win 100$ in bitcoin',\n 'Show ur colours! Euro 2004 2-4-1 Offer! Get an England Flag & 3Lions tone on ur phone! ']\n\nExplanation:\n **Part B model: using transfer learning:** \n\nPython implementation:\nspam_num=0\nfor i in msgs_spam_email:\n pred=learn_classifier.predict(i)\n if pred[1]==1:\n spam_num+=1\nspam_num=spam_num/5\nprint(\"The accuracy of detecting spam emails is:\", spam_num)\n\nExplanation:\n **Part A model: using RNN:** \n\nPython implementation:\nspam_num_model=0", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "492-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "am_num=spam_num/5\nprint(\"The accuracy of detecting spam emails is:\", spam_num)\n\nExplanation:\n **Part A model: using RNN:** \n\nPython implementation:\nspam_num_model=0\nfor j in msgs_spam_email:\n j_new=list(tuple(i))\n j_new2=[]\n for i in enumerate(j_new):\n j_new2.append(text_field.vocab.stoi[i[1]])\n\n j_new2=torch.tensor(j_new2).unsqueeze(0)\n prob=model7(j_new2)\n pred=prob.max(1,keepdim=True)[1]\n if pred.item()==1:\n spam_num_model+=1\nspam_num_model=spam_num_model/5\nprint(\"The accuracy of detecting spam emails is:\", spam_num_model)\n\nExplanation:\n As we can see the accuracy of the model in part B (using transfer learning) is better than the RNN model from part A in detecting spam emails. ", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "492-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (b)" + }, + { + "text": "Explanation:\n### Part (c)\nLoad 5 sample sentences from your regular mail and test it out out the two models you created. How well do they perform?\n\nPython implementation:\nmsgs_nospam_partB=['Job Opportunities in First Year Office-Summer 2022. The First Year Office has at least 2 job opportunities that are available to all returning graduate students. Details for each position can be found below.',\n 'Dear Students, If you have not yet taken JDE1000H as part of your research program, please sign up for the spring offering by registering via Microsoft form (see below).',\n 'I will run an introduction session on the Final Project tomorrow, Saturday, March 19 from 7-8 PM.',\n 'Thank you so much for the reply and detailed explanation.',\n 'Hi all, Our undergraduate students are organizing their formal dinner dance, details below. This will be the first dinner dance in two years']\n\nExplanation:\n **Part B model: using transfer learning:** ", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "493-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": "organizing their formal dinner dance, details below. This will be the first dinner dance in two years']\n\nExplanation:\n **Part B model: using transfer learning:** \n\nPython implementation:\nnospam_num=0\nfor i in msgs_nospam_partB:\n pred=learn_classifier.predict(i)\n if pred[1]==0:\n nospam_num+=1\nnospam_num=nospam_num/5\nprint(\"The accuracy of predicting non-spam is:\",nospam_num)\n\nExplanation:\n **Part A model: using RNN:** \n\nPython implementation:\nnospam_num2=0\nfor j in msgs_nospam_partB:\n j_new=list(tuple(i))\n j_new2=[]\n for i in enumerate(j_new):\n j_new2.append(text_field.vocab.stoi[i[1]])\n\n j_new2=torch.tensor(j_new2).unsqueeze(0)\n prob=model7(j_new2)\n pred=prob.max(1,keepdim=True)[1]\n if pred.item()==0:\n nospam_num2+=1\nnospam_num2=nospam_num2/5\nprint(\"The accuracy of detecting non-spam emails is:\", nospam_num2)\n\nExplanation:", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "493-1", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": "prob=model7(j_new2)\n pred=prob.max(1,keepdim=True)[1]\n if pred.item()==0:\n nospam_num2+=1\nnospam_num2=nospam_num2/5\nprint(\"The accuracy of detecting non-spam emails is:\", nospam_num2)\n\nExplanation:\n As we can see the accuracy of the model in part B (using transfer learning) and the RNN model from part A works good in detecting no-spam emails. ", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "493-2", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Part (c)" + }, + { + "text": "Explanation:\n# Results\n\n## Summary\n\nTwo spam detection approaches were implemented and compared.\n\n### Model 1\n\n- Character-level LSTM\n- Built entirely from scratch\n- Learns sequential text representations directly from SMS messages\n\n### Model 2\n\n- Transfer Learning (ULMFiT)\n- Fine-tuned pretrained language model\n- Uses learned language representations for spam classification\n\n### Key Findings\n\n- Both models successfully distinguish spam from legitimate SMS messages.\n- Transfer learning improved classification performance by leveraging pretrained language representations.\n- Confusion matrices and probability predictions provide additional insight beyond overall accuracy.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "494-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Results" + }, + { + "text": "Explanation:\n# Conclusion\n\nThis notebook demonstrated two modern approaches to spam detection using deep learning. A character-level LSTM established a strong baseline for sequential text classification, while transfer learning with ULMFiT further improved predictive performance by leveraging pretrained language representations. Together, these experiments illustrate the effectiveness of recurrent neural networks and transfer learning for practical NLP classification tasks.", + "source": "3_spam_detection_lstm.ipynb", + "file_type": "ipynb", + "chunk_id": "495-0", + "page": null, + "document_type": "github_project", + "page_header": null, + "repository": "sequential-deep-learning-pytorch", + "relative_path": "notebooks/3_spam_detection_lstm.ipynb", + "section": "Conclusion" + } +] \ No newline at end of file