{ "cells": [ { "cell_type": "markdown", "id": "2f04eee0-5928-4e74-a754-6dc2e528810c", "metadata": {}, "source": [ "# SmokingMcCartney" ] }, { "cell_type": "markdown", "id": "a3f514a3-772c-4a14-afdf-5a8376851ff4", "metadata": {}, "source": [ "## Index\n", "1. [Instantiate model class](#Instantiate-model-class)\n", "2. [Define clock metadata](#Define-clock-metadata)\n", "3. [Download clock dependencies](#Download-clock-dependencies)\n", "5. [Load features](#Load-features)\n", "6. [Load weights into base model](#Load-weights-into-base-model)\n", "7. [Load reference values](#Load-reference-values)\n", "8. [Load preprocess and postprocess objects](#Load-preprocess-and-postprocess-objects)\n", "10. [Check all clock parameters](#Check-all-clock-parameters)\n", "10. [Basic test](#Basic-test)\n", "11. [Save torch model](#Save-torch-model)\n", "12. [Clear directory](#Clear-directory)\n" ] }, { "cell_type": "markdown", "id": "d95fafdc-643a-40ea-a689-200bd132e90c", "metadata": {}, "source": [ "Let's first import some packages:" ] }, { "cell_type": "code", "execution_count": 1, "id": "4adfb4de-cd79-4913-a1af-9e23e9e236c9", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:34.133350Z", "iopub.status.busy": "2025-04-07T17:47:34.132902Z", "iopub.status.idle": "2025-04-07T17:47:35.476506Z", "shell.execute_reply": "2025-04-07T17:47:35.476150Z" } }, "outputs": [], "source": [ "import os\n", "import inspect\n", "import shutil\n", "import json\n", "import torch\n", "import pandas as pd\n", "import pyaging as pya" ] }, { "cell_type": "markdown", "id": "145082e5-ced4-47ae-88c0-cb69773e3c5a", "metadata": {}, "source": [ "## Instantiate model class" ] }, { "cell_type": "code", "execution_count": 2, "id": "8aa77372-7ed3-4da7-abc9-d30372106139", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:35.478214Z", "iopub.status.busy": "2025-04-07T17:47:35.477997Z", "iopub.status.idle": "2025-04-07T17:47:35.484695Z", "shell.execute_reply": "2025-04-07T17:47:35.484415Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "class McCartneySmoking(pyagingModel):\n", " def __init__(self):\n", " super().__init__()\n", "\n", " def preprocess(self, x):\n", " return x\n", "\n", " def postprocess(self, x):\n", " return x\n", "\n" ] } ], "source": [ "def print_entire_class(cls):\n", " source = inspect.getsource(cls)\n", " print(source)\n", "\n", "print_entire_class(pya.models.McCartneySmoking)\n" ] }, { "cell_type": "code", "execution_count": 3, "id": "78536494-f1d9-44de-8583-c89a310d2307", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:35.485963Z", "iopub.status.busy": "2025-04-07T17:47:35.485866Z", "iopub.status.idle": "2025-04-07T17:47:35.487577Z", "shell.execute_reply": "2025-04-07T17:47:35.487295Z" } }, "outputs": [], "source": [ "model = pya.models.McCartneySmoking()" ] }, { "cell_type": "markdown", "id": "51f8615e-01fa-4aa5-b196-3ee2b35d261c", "metadata": {}, "source": [ "## Define clock metadata" ] }, { "cell_type": "code", "execution_count": 4, "id": "6601da9e-8adc-44ee-9308-75e3cd31b816", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:35.488812Z", "iopub.status.busy": "2025-04-07T17:47:35.488727Z", "iopub.status.idle": "2025-04-07T17:47:35.490718Z", "shell.execute_reply": "2025-04-07T17:47:35.490459Z" } }, "outputs": [], "source": [ "model.metadata[\"clock_name\"] = \"mccartneysmoking\"\n", "model.metadata[\"data_type\"] = \"DNA methylation\" # Paper: The predictors use DNA methylation states at CpG sites.\n", "model.metadata[\"species\"] = \"Homo sapiens\" # Paper: Generation Scotland is a population-based cohort of human participants.\n", "model.metadata[\"year\"] = 2018\n", "model.metadata[\"approved_by_author\"] = \"⌛\"\n", "model.metadata[\"citation\"] = \"McCartney, D. L., et al. “Epigenetic prediction of complex traits and death.” Genome Biology 19, 136 (2018).\"\n", "model.metadata[\"doi\"] = \"https://doi.org/10.1186/s13059-018-1514-1\"\n", "model.metadata[\"notes\"] = \"Whole-blood DNAm LASSO score for smoking exposure (pack-years), trained in Generation Scotland on an age-, sex-, and ancestry-adjusted phenotype residual and evaluated out of sample in LBC1936.\"\n", "model.metadata[\"research_only\"] = None\n", "model.metadata[\"tissue\"] = [\"whole blood\"] # Paper: Stored DNA from baseline blood samples was used to build the predictors.\n", "model.metadata[\"predicts\"] = [\"smoking exposure\"] # Paper: The study developed a DNAm predictor for smoking exposure (pack-years).\n", "model.metadata[\"training_target\"] = [\"smoking exposure\"] # Paper: The smoking exposure (pack-years) phenotype was regressed on age, sex and ten genetic principal components; its residual was the LASSO outcome.\n", "model.metadata[\"unit\"] = [\"pack-years\"] # Paper: No output transformation is applied; the score retains the pack-years residual scale.\n", "model.metadata[\"model_type\"] = \"LASSO regression\" # Paper: The glmnet mixing parameter alpha was set to 1, applying a LASSO penalty with tenfold cross-validation.\n", "model.metadata[\"platform\"] = [\"Illumina EPIC\"] # Paper: Predictor training used quality-controlled HumanMethylationEPIC blood data; probes absent from 450K were filtered only to enable LBC1936 prediction.\n", "model.metadata[\"population\"] = \"adults\" # Paper: The predictors were built on a subset of 5,087 Generation Scotland participants; the parent cohort spans ages 18–99.\n", "model.metadata[\"journal\"] = \"Genome Biology\"\n", "model.metadata[\"last_author\"] = \"Riccardo E. Marioni\"\n", "model.metadata[\"n_features\"] = 233\n", "model.metadata[\"citations\"] = 301\n", "model.metadata[\"citations_date\"] = \"2026-07-05\"\n" ] }, { "cell_type": "markdown", "id": "74492239-5aae-4026-9d90-6bc9c574c110", "metadata": {}, "source": [ "## Download clock dependencies" ] }, { "cell_type": "markdown", "id": "7bec474f-80ce-4884-9472-30c193327117", "metadata": {}, "source": [ "#### Download coefficient file" ] }, { "cell_type": "code", "execution_count": 5, "id": "aa4a1b59-dda3-4ea8-8f34-b3c53ecbc310", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:35.492080Z", "iopub.status.busy": "2025-04-07T17:47:35.491992Z", "iopub.status.idle": "2025-04-07T17:47:36.204837Z", "shell.execute_reply": "2025-04-07T17:47:36.204380Z" } }, "outputs": [ { "data": { "text/plain": [ "0" ] }, "execution_count": 5, "metadata": {}, "output_type": "execute_result" } ], "source": [ "coeff_url = \"https://raw.githubusercontent.com/bio-learn/biolearn/master/biolearn/data/Smoking.csv\"\n", "os.system(f\"curl -L {coeff_url} -o Smoking.csv\")\n" ] }, { "cell_type": "markdown", "id": "5035b180-3d1b-4432-8ebe-b9c92bd93a7f", "metadata": {}, "source": [ "## Load features" ] }, { "cell_type": "markdown", "id": "15f4af76-b93c-438c-b57f-f129d6e9ec99", "metadata": {}, "source": [ "#### From CSV file" ] }, { "cell_type": "code", "execution_count": 6, "id": "f26a49e3-7389-416c-9080-539f50e9abd0", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.206985Z", "iopub.status.busy": "2025-04-07T17:47:36.206832Z", "iopub.status.idle": "2025-04-07T17:47:36.211174Z", "shell.execute_reply": "2025-04-07T17:47:36.210783Z" } }, "outputs": [], "source": [ "coeffs = pd.read_csv('Smoking.csv')\n", "coeffs['feature'] = coeffs['CpGmarker']\n", "coeffs['coefficient'] = coeffs['CoefficientTraining']\n", "\n", "model.features = coeffs['feature'].tolist()\n" ] }, { "cell_type": "markdown", "id": "ee6d8fa0-4767-4c45-9717-eb1c95e2ddc0", "metadata": {}, "source": [ "## Load weights into base model" ] }, { "cell_type": "markdown", "id": "d79e5690-e284-4de6-8460-d3545a8192af", "metadata": {}, "source": [ "#### From CSV file" ] }, { "cell_type": "code", "execution_count": 7, "id": "7f6187ed-fcff-4ff2-bcb1-b5bcef8190e8", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.213121Z", "iopub.status.busy": "2025-04-07T17:47:36.212960Z", "iopub.status.idle": "2025-04-07T17:47:36.216800Z", "shell.execute_reply": "2025-04-07T17:47:36.216396Z" } }, "outputs": [], "source": [ "weights = torch.tensor(coeffs['coefficient'].tolist()).unsqueeze(0)\n", "intercept = torch.tensor([0.0])\n" ] }, { "cell_type": "markdown", "id": "ad261636-5b00-4979-bb1d-67a851f7aa19", "metadata": {}, "source": [ "#### Linear model" ] }, { "cell_type": "code", "execution_count": 8, "id": "d7f43b99-26f2-4622-9a76-316712058877", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.218646Z", "iopub.status.busy": "2025-04-07T17:47:36.218507Z", "iopub.status.idle": "2025-04-07T17:47:36.221310Z", "shell.execute_reply": "2025-04-07T17:47:36.220958Z" } }, "outputs": [], "source": [ "base_model = pya.models.LinearModel(input_dim=len(model.features))\n", "\n", "base_model.linear.weight.data = weights.float()\n", "base_model.linear.bias.data = intercept.float()\n", "\n", "model.base_model = base_model" ] }, { "cell_type": "markdown", "id": "ad8b4c1d-9d57-48b7-9a30-bcfea7b747b1", "metadata": {}, "source": [ "## Load reference values" ] }, { "cell_type": "code", "execution_count": 9, "id": "90d45266-962d-41b6-927c-6a147ed41305", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.222862Z", "iopub.status.busy": "2025-04-07T17:47:36.222749Z", "iopub.status.idle": "2025-04-07T17:47:36.224649Z", "shell.execute_reply": "2025-04-07T17:47:36.224309Z" } }, "outputs": [], "source": [ "model.reference_values = None" ] }, { "cell_type": "markdown", "id": "af3bcf7b-74a8-4d21-9ccb-4de0c2b0516b", "metadata": {}, "source": [ "## Load preprocess and postprocess objects" ] }, { "cell_type": "code", "execution_count": 10, "id": "f7d32b69-e20e-42ff-aba9-d07b9b44dbd1", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.226270Z", "iopub.status.busy": "2025-04-07T17:47:36.226156Z", "iopub.status.idle": "2025-04-07T17:47:36.228064Z", "shell.execute_reply": "2025-04-07T17:47:36.227750Z" } }, "outputs": [], "source": [ "model.preprocess_name = None\n", "model.preprocess_dependencies = None\n" ] }, { "cell_type": "code", "execution_count": 11, "id": "7a22fb20-c605-424d-8efb-7620c2c0755c", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.229604Z", "iopub.status.busy": "2025-04-07T17:47:36.229468Z", "iopub.status.idle": "2025-04-07T17:47:36.231233Z", "shell.execute_reply": "2025-04-07T17:47:36.230916Z" } }, "outputs": [], "source": [ "model.postprocess_name = None\n", "model.postprocess_dependencies = None\n" ] }, { "cell_type": "markdown", "id": "86e3d6b1-e67e-4f3d-bd39-0ebec5726c3c", "metadata": {}, "source": [ "## Check all clock parameters" ] }, { "cell_type": "code", "execution_count": 12, "id": "2168355c-47d9-475d-b816-49f65e74887c", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.232808Z", "iopub.status.busy": "2025-04-07T17:47:36.232688Z", "iopub.status.idle": "2025-04-07T17:47:36.236913Z", "shell.execute_reply": "2025-04-07T17:47:36.236596Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "\n", "%==================================== Model Details ====================================%\n", "Model Attributes:\n", "\n", "training: True\n", "metadata: {'approved_by_author': '⌛',\n", " 'citation': 'McCartney, Daniel L., et al. \"Epigenetic prediction of complex '\n", " 'traits and death.\" Genome biology 19.1 (2018): 136.',\n", " 'clock_name': 'mccartneysmoking',\n", " 'data_type': 'methylation',\n", " 'doi': 'https://doi.org/10.1186/s13059-018-1514-1',\n", " 'notes': 'Linear CpG score; higher values indicate smoking exposure.',\n", " 'research_only': None,\n", " 'species': 'Homo sapiens',\n", " 'version': None,\n", " 'year': 2018}\n", "reference_values: None\n", "preprocess_name: None\n", "preprocess_dependencies: None\n", "postprocess_name: None\n", "postprocess_dependencies: None\n", "features: ['cg10573386', 'cg13560072', 'cg11084015', 'cg10321266', 'cg07597069', 'cg05218653', 'cg00077898', 'cg11207515', 'cg18392085', 'cg06088918', 'cg20077343', 'cg24169820', 'cg17833862', 'cg23079012', 'cg10696199', 'cg05442408', 'cg08038054', 'cg26775087', 'cg22371743', 'cg04136748', 'cg23288337', 'cg04561727', 'cg09648091', 'cg14142965', 'cg14708990', 'cg01558110', 'cg11704876', 'cg03616722', 'cg24714011', 'cg00231810']... [Total elements: 233]\n", "base_model_features: None\n", "\n", "%==================================== Model Details ====================================%\n", "Model Structure:\n", "\n", "base_model: LinearModel(\n", " (linear): Linear(in_features=233, out_features=1, bias=True)\n", ")\n", "\n", "%==================================== Model Details ====================================%\n", "Model Parameters and Weights:\n", "\n", "base_model.linear.weight: [1.5160139799118042, 1.3958498239517212, 1.2834880352020264, 1.2621612548828125, 0.9191465377807617, 0.8301234841346741, 0.8246065378189087, 0.8208099007606506, 0.7282519936561584, 0.7219155430793762, 0.6905287504196167, 0.6882893443107605, 0.6189409494400024, 0.5594330430030823, 0.49723178148269653, 0.48923414945602417, 0.47508591413497925, 0.4401831328868866, 0.36066848039627075, 0.3536936938762665, 0.3378012776374817, 0.3341846168041229, 0.33264631032943726, 0.318849116563797, 0.31472083926200867, 0.308366596698761, 0.30539241433143616, 0.3030904531478882, 0.302953839302063, 0.28863614797592163]... [Tensor of shape torch.Size([1, 233])]\n", "base_model.linear.bias: tensor([0.])\n", "\n", "%==================================== Model Details ====================================%\n", "\n" ] } ], "source": [ "pya.utils.print_model_details(model)" ] }, { "cell_type": "markdown", "id": "01aa118d", "metadata": {}, "source": [ "## Basic Test" ] }, { "cell_type": "code", "execution_count": 13, "id": "936b9877-d076-4ced-99aa-e8d4c58c5caf", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.238465Z", "iopub.status.busy": "2025-04-07T17:47:36.238354Z", "iopub.status.idle": "2025-04-07T17:47:36.242868Z", "shell.execute_reply": "2025-04-07T17:47:36.242565Z" } }, "outputs": [ { "data": { "text/plain": [ "tensor([[ -1.4350],\n", " [ -4.1281],\n", " [-10.1821],\n", " [ -5.1272],\n", " [-11.3244],\n", " [ -3.2461],\n", " [ 8.5989],\n", " [ 3.6743],\n", " [ 6.9403],\n", " [-17.4587]], dtype=torch.float64, grad_fn=)" ] }, "execution_count": 13, "metadata": {}, "output_type": "execute_result" } ], "source": [ "torch.manual_seed(42)\n", "input = torch.randn(10, len(model.features), dtype=float)\n", "model.eval()\n", "model.to(float)\n", "pred = model(input)\n", "pred" ] }, { "cell_type": "markdown", "id": "fe8299d7-9285-4e22-82fd-b664434b4369", "metadata": {}, "source": [ "## Save torch model\n" ] }, { "cell_type": "code", "execution_count": 14, "id": "5ef2fa8d-c80b-4fdd-8555-79c0d541788e", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.244263Z", "iopub.status.busy": "2025-04-07T17:47:36.244159Z", "iopub.status.idle": "2025-04-07T17:47:36.246682Z", "shell.execute_reply": "2025-04-07T17:47:36.246395Z" } }, "outputs": [], "source": [ "torch.save(model, f\"../weights/{model.metadata['clock_name']}.pt\")" ] }, { "cell_type": "markdown", "id": "bac6257b-8d08-4a90-8d0b-7f745dc11ac1", "metadata": {}, "source": [ "\n", "## Clear directory\n", "\n" ] }, { "cell_type": "code", "execution_count": 15, "id": "11aeaa70-44c0-42f9-86d7-740e3849a7a6", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:47:36.248184Z", "iopub.status.busy": "2025-04-07T17:47:36.248077Z", "iopub.status.idle": "2025-04-07T17:47:36.255709Z", "shell.execute_reply": "2025-04-07T17:47:36.255449Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "Deleted file: Smoking.csv\n" ] } ], "source": [ "# Function to remove a folder and all its contents\n", "def remove_folder(path):\n", " try:\n", " shutil.rmtree(path)\n", " print(f\"Deleted folder: {path}\")\n", " except Exception as e:\n", " print(f\"Error deleting folder {path}: {e}\")\n", "\n", "# Get a list of all files and folders in the current directory\n", "all_items = os.listdir('.')\n", "\n", "# Loop through the items\n", "for item in all_items:\n", " # Check if it's a file and does not end with .ipynb\n", " if os.path.isfile(item) and not item.endswith('.ipynb'):\n", " os.remove(item)\n", " print(f\"Deleted file: {item}\")\n", " # Check if it's a folder\n", " elif os.path.isdir(item):\n", " remove_folder(item)" ] } ], "metadata": { "kernelspec": { "display_name": "research", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.9.7" } }, "nbformat": 4, "nbformat_minor": 5 }