{ "cells": [ { "cell_type": "markdown", "id": "2f04eee0-5928-4e74-a754-6dc2e528810c", "metadata": {}, "source": [ "# CpGPTGrimAge3" ] }, { "cell_type": "markdown", "id": "a3f514a3-772c-4a14-afdf-5a8376851ff4", "metadata": {}, "source": [ "## Index\n", "1. [Instantiate model class](#Instantiate-model-class)\n", "2. [Define clock metadata](#Define-clock-metadata)\n", "3. [Download clock dependencies](#Download-clock-dependencies)\n", "5. [Load features](#Load-features)\n", "6. [Load weights into base model](#Load-weights-into-base-model)\n", "7. [Load reference values](#Load-reference-values)\n", "8. [Load preprocess and postprocess objects](#Load-preprocess-and-postprocess-objects)\n", "10. [Check all clock parameters](#Check-all-clock-parameters)\n", "10. [Basic test](#Basic-test)\n", "11. [Save torch model](#Save-torch-model)\n", "12. [Clear directory](#Clear-directory)" ] }, { "cell_type": "markdown", "id": "d95fafdc-643a-40ea-a689-200bd132e90c", "metadata": {}, "source": [ "Let's first import some packages:" ] }, { "cell_type": "code", "execution_count": 1, "id": "4adfb4de-cd79-4913-a1af-9e23e9e236c9", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:41.316681Z", "iopub.status.busy": "2025-04-07T17:51:41.316440Z", "iopub.status.idle": "2025-04-07T17:51:42.738147Z", "shell.execute_reply": "2025-04-07T17:51:42.737780Z" } }, "outputs": [], "source": [ "import os\n", "import inspect\n", "import shutil\n", "import json\n", "import torch\n", "import pandas as pd\n", "import pyaging as pya\n", "import numpy as np" ] }, { "cell_type": "markdown", "id": "145082e5-ced4-47ae-88c0-cb69773e3c5a", "metadata": {}, "source": [ "## Instantiate model class" ] }, { "cell_type": "code", "execution_count": 2, "id": "8aa77372-7ed3-4da7-abc9-d30372106139", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:42.740018Z", "iopub.status.busy": "2025-04-07T17:51:42.739761Z", "iopub.status.idle": "2025-04-07T17:51:42.750935Z", "shell.execute_reply": "2025-04-07T17:51:42.750574Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "class CpGPTGrimAge3(pyagingModel):\n", " def __init__(self):\n", " super().__init__()\n", "\n", " def preprocess(self, x):\n", " \"\"\"\n", " Scales an array based on the median and standard deviation.\n", " \"\"\"\n", " median = torch.tensor(self.preprocess_dependencies[0], device=x.device, dtype=x.dtype)\n", " std = torch.tensor(self.preprocess_dependencies[1], device=x.device, dtype=x.dtype)\n", " x = (x - median) / std\n", " return x\n", "\n", " def postprocess(self, x):\n", " \"\"\"\n", " Converts from a Cox parameter to age in units of years.\n", " \"\"\"\n", " cox_mean = self.postprocess_dependencies[0]\n", " cox_std = self.postprocess_dependencies[1]\n", " age_mean = self.postprocess_dependencies[2]\n", " age_std = self.postprocess_dependencies[3]\n", "\n", " # Normalize\n", " x = (x - cox_mean) / cox_std\n", "\n", " # Scale\n", " x = (x * age_std) + age_mean\n", "\n", " return x\n", "\n" ] } ], "source": [ "def print_entire_class(cls):\n", " source = inspect.getsource(cls)\n", " print(source)\n", "\n", "print_entire_class(pya.models.CpGPTGrimAge3)" ] }, { "cell_type": "code", "execution_count": 3, "id": "78536494-f1d9-44de-8583-c89a310d2307", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:42.753015Z", "iopub.status.busy": "2025-04-07T17:51:42.752877Z", "iopub.status.idle": "2025-04-07T17:51:42.754599Z", "shell.execute_reply": "2025-04-07T17:51:42.754314Z" } }, "outputs": [], "source": [ "model = pya.models.CpGPTGrimAge3()" ] }, { "cell_type": "markdown", "id": "51f8615e-01fa-4aa5-b196-3ee2b35d261c", "metadata": {}, "source": [ "## Define clock metadata" ] }, { "cell_type": "code", "execution_count": 4, "id": "135ce001-03f7-4025-bceb-01a3e2e2b0ef", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:42.755964Z", "iopub.status.busy": "2025-04-07T17:51:42.755873Z", "iopub.status.idle": "2025-04-07T17:51:42.758035Z", "shell.execute_reply": "2025-04-07T17:51:42.757775Z" } }, "outputs": [], "source": [ "model.metadata[\"clock_name\"] = \"cpgptgrimage3\"\n", "model.metadata[\"data_type\"] = \"DNA methylation\" # Paper: CpGPT models DNA-methylation profiles and derives methylation-based aging and mortality predictors.\n", "model.metadata[\"species\"] = \"Homo sapiens\" # Paper: The assigned mortality task uses human DNA-methylation cohorts.\n", "model.metadata[\"year\"] = 2024\n", "model.metadata[\"approved_by_author\"] = \"✅\"\n", "model.metadata[\"citation\"] = \"de Lima Camillo, L. P., Sehgal, R., Armstrong, J., Higgins-Chen, A. T., Horvath, S., & Wang, B. CpGPT: a foundation model for DNA methylation. bioRxiv 2024.10.24.619766 (2024).\"\n", "model.metadata[\"doi\"] = \"https://doi.org/10.1101/2024.10.24.619766\"\n", "model.metadata[\"notes\"] = \"CpGPTGrimAge3 implementation combining chronological age, GrimAge2 DNAm proxies, and CpGPT-predicted plasma-protein proxies in a Cox linear predictor that is calibrated to years.\"\n", "model.metadata[\"research_only\"] = True\n", "model.metadata[\"tissue\"] = [\"whole blood\"] # Paper: cpgptgrimage3 and cpgptpcgrimage3 were trained in blood with the 450k array to predict mortality. it was trained in the FHS cohort with methylation data.\n", "model.metadata[\"predicts\"] = [\"biological age\", \"mortality risk\"] # Paper: The implementation converts the Cox parameter to an age-scaled output.\n", "model.metadata[\"training_target\"] = [\"mortality\"] # Paper: The mortality model was trained with a modified Cox proportional-hazards loss against time-to-mortality data.\n", "model.metadata[\"unit\"] = [\"years\"] # Paper: The implementation normalizes a Cox score and rescales it with an age mean and standard deviation.\n", "model.metadata[\"model_type\"] = \"Cox proportional hazards regression\" # Paper: The implementation applies a linear Cox score to standardized age and biomarker proxies and calibrates the score to years.\n", "model.metadata[\"platform\"] = [\"Illumina 450K\"] # Paper: cpgptgrimage3 and cpgptpcgrimage3 were trained in blood with the 450k array to predict mortality. it was trained in the FHS cohort with methylation data.\n", "model.metadata[\"population\"] = \"adults\" # Paper: cpgptgrimage3 and cpgptpcgrimage3 were trained in blood with the 450k array to predict mortality. it was trained in the FHS cohort with methylation data.\n", "model.metadata[\"journal\"] = \"bioRxiv\"\n", "model.metadata[\"last_author\"] = \"Bo Wang\"\n", "model.metadata[\"n_features\"] = 24\n", "model.metadata[\"citations\"] = 30\n", "model.metadata[\"citations_date\"] = \"2026-07-05\"\n" ] }, { "cell_type": "markdown", "id": "74492239-5aae-4026-9d90-6bc9c574c110", "metadata": {}, "source": [ "## Download clock dependencies" ] }, { "cell_type": "code", "execution_count": 5, "id": "95f6ba57", "metadata": {}, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "|-----------> Data found in ./cpgpt_grimage3_weights_all_datasets_reliable.csv\n", "|-----------> Data found in ./input_scaler_mean_all_datasets_reliable.npy\n", "|-----------> Data found in ./input_scaler_scale_all_datasets_reliable.npy\n" ] } ], "source": [ "logger = pya.logger.Logger()\n", "urls = [\n", " \"https://huggingface.co/lucascamillomd/pyaging-data/resolve/main/supporting_files/cpgpt_grimage3_dependencies/reliable/cpgpt_grimage3_weights_all_datasets_reliable.csv\",\n", " \"https://huggingface.co/lucascamillomd/pyaging-data/resolve/main/supporting_files/cpgpt_grimage3_dependencies/reliable/input_scaler_mean_all_datasets_reliable.npy\",\n", " \"https://huggingface.co/lucascamillomd/pyaging-data/resolve/main/supporting_files/cpgpt_grimage3_dependencies/reliable/input_scaler_scale_all_datasets_reliable.npy\"\n", "]\n", "dir = \".\"\n", "for url in urls:\n", " pya.utils.download(url, dir, logger, indent_level=1)" ] }, { "cell_type": "markdown", "id": "a14c7fc1-abe5-42a3-8bc9-0987521ddf33", "metadata": {}, "source": [ "## Load features" ] }, { "cell_type": "markdown", "id": "3e737582-3a28-4f55-8da9-3e34125362cc", "metadata": {}, "source": [ "#### From CSV" ] }, { "cell_type": "code", "execution_count": 6, "id": "f1486db4", "metadata": {}, "outputs": [], "source": [ "df = pd.read_csv('cpgpt_grimage3_weights_all_datasets_reliable.csv')\n", "model.features = df['feature'].tolist()" ] }, { "cell_type": "code", "execution_count": 7, "id": "8c43149c", "metadata": {}, "outputs": [ { "data": { "text/html": [ "
\n", "\n", "\n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", " \n", "
featurecoefficient
0age0.845167
1grimage2timp10.318954
2grimage2packyrs0.385882
3grimage2logcrp0.404675
4grimage2adm0.180551
\n", "
" ], "text/plain": [ " feature coefficient\n", "0 age 0.845167\n", "1 grimage2timp1 0.318954\n", "2 grimage2packyrs 0.385882\n", "3 grimage2logcrp 0.404675\n", "4 grimage2adm 0.180551" ] }, "execution_count": 7, "metadata": {}, "output_type": "execute_result" } ], "source": [ "df.head()" ] }, { "cell_type": "markdown", "id": "ee6d8fa0-4767-4c45-9717-eb1c95e2ddc0", "metadata": {}, "source": [ "## Load weights into base model" ] }, { "cell_type": "markdown", "id": "3958ba73-42e8-40a5-94a1-4f4b8ae05dca", "metadata": {}, "source": [ "#### Linear model" ] }, { "cell_type": "code", "execution_count": 8, "id": "321a437c-8888-4e10-96e9-5ed2826a8f74", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.647294Z", "iopub.status.busy": "2025-04-07T17:51:44.647116Z", "iopub.status.idle": "2025-04-07T17:51:44.688112Z", "shell.execute_reply": "2025-04-07T17:51:44.687757Z" } }, "outputs": [], "source": [ "weights = torch.tensor(df['coefficient'].tolist()).unsqueeze(0)\n", "intercept = torch.tensor([0.0])" ] }, { "cell_type": "markdown", "id": "5742dc16-e063-414f-a38e-9721beb11351", "metadata": {}, "source": [ "#### Linear model" ] }, { "cell_type": "code", "execution_count": 9, "id": "c2e54115-c17b-48ce-88f1-de546c90d2b3", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.689783Z", "iopub.status.busy": "2025-04-07T17:51:44.689683Z", "iopub.status.idle": "2025-04-07T17:51:44.692010Z", "shell.execute_reply": "2025-04-07T17:51:44.691738Z" } }, "outputs": [], "source": [ "base_model = pya.models.LinearModel(input_dim=len(model.features))\n", "\n", "base_model.linear.weight.data = weights.float()\n", "base_model.linear.bias.data = intercept.float()\n", "\n", "model.base_model = base_model" ] }, { "cell_type": "markdown", "id": "ad8b4c1d-9d57-48b7-9a30-bcfea7b747b1", "metadata": {}, "source": [ "## Load reference values" ] }, { "cell_type": "code", "execution_count": 11, "id": "2089b66f-9cc4-4528-9bdc-5e45efc6d06b", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.693409Z", "iopub.status.busy": "2025-04-07T17:51:44.693309Z", "iopub.status.idle": "2025-04-07T17:51:44.708180Z", "shell.execute_reply": "2025-04-07T17:51:44.707878Z" } }, "outputs": [], "source": [ "scale_mean = np.load('input_scaler_mean_all_datasets_reliable.npy')\n", "scale_std = np.load('input_scaler_scale_all_datasets_reliable.npy')\n", "\n", "model.reference_values = None" ] }, { "cell_type": "markdown", "id": "af3bcf7b-74a8-4d21-9ccb-4de0c2b0516b", "metadata": {}, "source": [ "## Load preprocess and postprocess objects" ] }, { "cell_type": "code", "execution_count": 12, "id": "7a22fb20-c605-424d-8efb-7620c2c0755c", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.709686Z", "iopub.status.busy": "2025-04-07T17:51:44.709594Z", "iopub.status.idle": "2025-04-07T17:51:44.711225Z", "shell.execute_reply": "2025-04-07T17:51:44.710973Z" } }, "outputs": [], "source": [ "model.preprocess_name = 'scale'\n", "model.preprocess_dependencies = [scale_mean, scale_std]" ] }, { "cell_type": "code", "execution_count": 13, "id": "ff4a21cb-cf41-44dc-9ed1-95cf8aa15772", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.712430Z", "iopub.status.busy": "2025-04-07T17:51:44.712349Z", "iopub.status.idle": "2025-04-07T17:51:44.713842Z", "shell.execute_reply": "2025-04-07T17:51:44.713585Z" } }, "outputs": [], "source": [ "model.postprocess_name = 'cox_to_years'\n", "model.postprocess_dependencies = [\n", " 0.54372919,\n", " 1.52036698,\n", " 64.94560376271838,\n", " 11.920838151170104\n", "]" ] }, { "cell_type": "markdown", "id": "86e3d6b1-e67e-4f3d-bd39-0ebec5726c3c", "metadata": {}, "source": [ "## Check all clock parameters" ] }, { "cell_type": "code", "execution_count": 14, "id": "2168355c-47d9-475d-b816-49f65e74887c", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.715112Z", "iopub.status.busy": "2025-04-07T17:51:44.715032Z", "iopub.status.idle": "2025-04-07T17:51:44.726874Z", "shell.execute_reply": "2025-04-07T17:51:44.726577Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "\n", "%==================================== Model Details ====================================%\n", "Model Attributes:\n", "\n", "training: True\n", "metadata: {'approved_by_author': '✅',\n", " 'citation': 'de Lima Camillo, Lucas Paulo, et al. \"CpGPT: a foundation model '\n", " 'for DNA methylation.\" bioRxiv (2024): 2024-10.',\n", " 'clock_name': 'cpgptgrimage3',\n", " 'data_type': 'methylation',\n", " 'doi': 'https://doi.org/10.1101/2024.10.24.619766',\n", " 'notes': None,\n", " 'research_only': True,\n", " 'species': 'Homo sapiens',\n", " 'version': None,\n", " 'year': 2025}\n", "reference_values: None\n", "preprocess_name: 'scale'\n", "preprocess_dependencies: [array([ 6.50000000e+01, 3.49212152e+04, 1.21734902e+01, 2.73993813e-01,\n", " 3.51222301e+02, 8.51761217e+03, 8.85501049e+02, -5.21484375e-01,\n", " -2.49755859e-01, -2.58056641e-01, -7.65991211e-02, -1.37939453e-01,\n", " 2.53173828e-01, 1.33399963e-02, -4.01245117e-01, 1.90368652e-01,\n", " -3.27301025e-02, 1.29127502e-02, -2.78564453e-01, 1.92277772e+04,\n", " 2.87399292e-02, -3.61083984e-01, -1.25961304e-02, -2.21801758e-01]),\n", " array([1.52000000e+01, 2.39220372e+03, 1.43614564e+01, 7.58010775e-01,\n", " 3.24518463e+01, 5.54305012e+03, 2.81677880e+02, 1.54296875e-01,\n", " 1.58935547e-01, 3.54949951e-01, 1.87866211e-01, 4.43069458e-01,\n", " 2.97714233e-01, 3.31024170e-01, 3.81805420e-01, 1.43981934e-01,\n", " 1.65519714e-01, 2.10388184e-01, 1.30737305e-01, 3.97488350e+03,\n", " 3.68286133e-01, 7.37304688e-02, 3.28125000e-01, 4.43408966e-01])]\n", "postprocess_name: 'cox_to_years'\n", "postprocess_dependencies: [0.54372919, 1.52036698, 64.94560376271838, 11.920838151170104]\n", "features: ['age',\n", " 'grimage2timp1',\n", " 'grimage2packyrs',\n", " 'grimage2logcrp',\n", " 'grimage2adm',\n", " 'grimage2leptin',\n", " 'grimage2gdf15',\n", " 'cpgpt_s100a9',\n", " 'cpgpt_tnfrsf13c',\n", " 'cpgpt_tgfb1',\n", " 'cpgpt_tek',\n", " 'cpgpt_ccl14',\n", " 'cpgpt_tnfsf15',\n", " 'cpgpt_lilrb2',\n", " 'cpgpt_tnf',\n", " 'cpgpt_chit1',\n", " 'cpgpt_postn',\n", " 'cpgpt_il34',\n", " 'cpgpt_pdcd1',\n", " 'grimage2pai1',\n", " 'cpgpt_cst3',\n", " 'cpgpt_cxcl2',\n", " 'cpgpt_gzma',\n", " 'cpgpt_il5']\n", "base_model_features: None\n", "\n", "%==================================== Model Details ====================================%\n", "Model Structure:\n", "\n", "base_model: LinearModel(\n", " (linear): Linear(in_features=24, out_features=1, bias=True)\n", ")\n", "\n", "%==================================== Model Details ====================================%\n", "Model Parameters and Weights:\n", "\n", "base_model.linear.weight: tensor([[ 0.8452, 0.3190, 0.3859, 0.4047, 0.1806, -0.2435, 0.0367, -0.0855,\n", " 2.0569, -3.9567, 1.8897, -2.3948, -3.8697, 4.5260, 0.0498, 1.4570,\n", " -1.7014, -1.5117, -1.5438, 0.1255, 5.3232, -0.4491, 0.6656, 0.7276]])\n", "base_model.linear.bias: tensor([0.])\n", "\n", "%==================================== Model Details ====================================%\n", "\n" ] } ], "source": [ "pya.utils.print_model_details(model)" ] }, { "cell_type": "markdown", "id": "986d0262-e0c7-4036-b687-dee53ba392fb", "metadata": {}, "source": [ "## Basic test" ] }, { "cell_type": "code", "execution_count": 15, "id": "936b9877-d076-4ced-99aa-e8d4c58c5caf", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.728239Z", "iopub.status.busy": "2025-04-07T17:51:44.728153Z", "iopub.status.idle": "2025-04-07T17:51:44.733560Z", "shell.execute_reply": "2025-04-07T17:51:44.733262Z" } }, "outputs": [ { "data": { "text/plain": [ "tensor([[-425.9024],\n", " [-312.5518],\n", " [-563.8393],\n", " [-260.0897],\n", " [ -69.1534],\n", " [ 30.5343],\n", " [-252.0608],\n", " [-445.2048],\n", " [ -64.2164],\n", " [ 103.0451]], dtype=torch.float64, grad_fn=)" ] }, "execution_count": 15, "metadata": {}, "output_type": "execute_result" } ], "source": [ "torch.manual_seed(42)\n", "input = torch.randn(10, len(model.features), dtype=float).double()\n", "model.eval()\n", "model.to(float)\n", "pred = model(input)\n", "pred" ] }, { "cell_type": "markdown", "id": "fe8299d7-9285-4e22-82fd-b664434b4369", "metadata": {}, "source": [ "## Save torch model" ] }, { "cell_type": "code", "execution_count": 16, "id": "5ef2fa8d-c80b-4fdd-8555-79c0d541788e", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.734916Z", "iopub.status.busy": "2025-04-07T17:51:44.734814Z", "iopub.status.idle": "2025-04-07T17:51:44.739278Z", "shell.execute_reply": "2025-04-07T17:51:44.738936Z" } }, "outputs": [], "source": [ "torch.save(model, f\"../weights/{model.metadata['clock_name']}.pt\")" ] }, { "cell_type": "markdown", "id": "bac6257b-8d08-4a90-8d0b-7f745dc11ac1", "metadata": {}, "source": [ "## Clear directory\n", "" ] }, { "cell_type": "code", "execution_count": 17, "id": "11aeaa70-44c0-42f9-86d7-740e3849a7a6", "metadata": { "execution": { "iopub.execute_input": "2025-04-07T17:51:44.740582Z", "iopub.status.busy": "2025-04-07T17:51:44.740494Z", "iopub.status.idle": "2025-04-07T17:51:44.743819Z", "shell.execute_reply": "2025-04-07T17:51:44.743572Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "Deleted file: cpgpt_grimage3_weights_all_datasets_reliable.csv\n", "Deleted file: input_scaler_mean_all_datasets_reliable.npy\n", "Deleted file: input_scaler_scale_all_datasets_reliable.npy\n" ] } ], "source": [ "# Function to remove a folder and all its contents\n", "def remove_folder(path):\n", " try:\n", " shutil.rmtree(path)\n", " print(f\"Deleted folder: {path}\")\n", " except Exception as e:\n", " print(f\"Error deleting folder {path}: {e}\")\n", "\n", "# Get a list of all files and folders in the current directory\n", "all_items = os.listdir('.')\n", "\n", "# Loop through the items\n", "for item in all_items:\n", " # Check if it's a file and does not end with .ipynb\n", " if os.path.isfile(item) and not item.endswith('.ipynb'):\n", " os.remove(item)\n", " print(f\"Deleted file: {item}\")\n", " # Check if it's a folder\n", " elif os.path.isdir(item):\n", " remove_folder(item)" ] } ], "metadata": { "kernelspec": { "display_name": ".venv", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.13.7" } }, "nbformat": 4, "nbformat_minor": 5 }