Rubinov Lab · Vanderbilt University

AGENT-P

The Association of Gene Expression to Neuroimaging Traits Pipeline

A Python toolkit for running transcriptome-wide association studies (TWAS) on brain phenotypes in a standardized way, and for comparing the results against GWAS. It integrates four of the most widely used TWAS methods for eight cortical and subcortical brain regions, and supports both individual and summary TWAS.

Paper: Associations of estimated gene expression with neuroimaging phenotypes across people and brain regions with AGENT-P

Overview

For each phenotype, AGENT-P can run 64 association studies. Each one combines a TWAS framework, a TWAS method and a regional model of gene expression. The 32 sets of models (four methods × eight GTEx brain regions) cover 20,305 unique genes. Outputs are standardized into one directory layout, so you can compare phenotypes and TWAS approaches directly.

2 frameworks

  • Individual TWAS
  • Summary TWAS
×

4 TWAS methods

  • PrediXcan / S-PrediXcan
  • FUSION
  • UTMOST
  • JTI
×

8 GTEx brain regions

  • DLPFC
  • Anterior cingulate
  • Amygdala
  • Hippocampus
  • Caudate
  • Putamen
  • Nucleus accumbens
  • Cerebellar hemisphere
=
64studies per
phenotype

Analyses

Individual TWAS

Input: genotypes, phenotypes and covariates

Estimates genetically regulated gene expression for every person in a cohort, then associates it directly with their phenotypes.

Summary TWAS

Input: GWAS summary statistics

Enhances existing GWAS by converting their results into gene-level associations, with no need for restricted individual-level genotype data.

GWAS

Input: genotypes, phenotypes and covariates

Runs REGENIE on the same cohort, so TWAS results can be evaluated against variant-level associations.

How it works

Processing and quality control

Filters for common SNPs supported by TWAS, with filters on missingness rate, minor allele frequency and Hardy–Weinberg equilibrium. For subject-level data, it also regresses out covariates such as age, sex and genotype principal components.

Estimation of gene expression

Estimates genetically regulated gene expression for each gene in a selected brain region of each person, using any of the 32 sets of models.

Association analysis

Associates estimated gene expression directly with phenotypes, or converts GWAS results directly into TWAS.

Installation

AGENT-P relies on external tools (REGENIE, PLINK2, bgenix and R). The container image already includes them, so it is the quickest way to get started.

Pull the image and start a shell inside it.

docker pull --platform linux/amd64 zoeytang/agentp
docker run --platform linux/amd64 -it zoeytang/agentp

Then run the bundled example from inside the container:

cd /usr/home/example
python run_example.py

Keep --platform linux/amd64 on Apple Silicon Macs, because PLINK2 and REGENIE are x86_64 only. To build the image yourself, use the Dockerfile in the repository.

Quick start

Every analysis starts with a Project, a directory that holds all GWAS and TWAS results for one cohort. This example mirrors example/run_example.py.

from agentp import Project, SummaryTWAS, IndividualTWAS, GWAS

project = Project('example_project')

# Summary TWAS: JTI hippocampus models applied to an amygdala-volume GWAS
stwas = SummaryTWAS(project, 'JTI', 'hippocampus')
stwas.add_gwas('amygdala_volume', 'example_data/vol_mean_amygdala.regenie')
stwas.run_twas('amygdala_volume')

# Individual TWAS: FUSION caudate models on subject-level data
project.set_subjects('example_data/subjects.txt')
project.add_covariates('example_data/covariates.csv')
project.add_phenotypes('example_data/volumes.csv')
project.add_genotypes('example_data/genotypes', 'test_c*.bgen')

itwas = IndividualTWAS(project, 'FUS', 'caudate')
itwas.predict_grex()
itwas.run_twas('putamen_volume')

# GWAS with REGENIE on the same cohort
gwas = GWAS(project)
gwas.run('hippocampus_volume')

Output layout

example_project/
├── sTWAS/JTI_hippocampus/amygdala_volume.csv
├── iTWAS/FUS_caudate/putamen_volume.csv
└── GWAS/hippocampus_volume.regenie

Supported models

Pass these abbreviations when you create a SummaryTWAS or IndividualTWAS, for example SummaryTWAS(project, 'PDX', 'dlpfc').

TWAS methodTissue contextAbbreviation
PrediXcan / S-PrediXcanSingle-tissuePDX
FUSIONSingle-tissueFUS
UTMOSTMulti-tissueUTM
Joint-Tissue ImputationJoint-tissueJTI
GTEx brain regionAbbreviation
Dorsolateral prefrontal cortexdlpfc
Anterior cingulateant-cingulate
Amygdalaamygdala
Hippocampushippocampus
Caudatecaudate
Putamenputamen
Nucleus accumbensnuc-accumbens
Cerebellar hemispherecerebellar-hemi

Main classes

Parameters for every method are documented in the usage notes.

Project

Creates the cohort directory and loads subjects, covariates, phenotypes and genotypes.

BGEN

Applies quality control to input genotypes in BGEN format.

GWAS

Runs REGENIE on the project's genotypes and traits, in parallel across chromosomes.

GREX

Predicts genetically regulated expression for each subject from pre-trained models.

SummaryTWAS

Tests gene–trait associations from GWAS summary statistics.

IndividualTWAS

Tests gene–trait associations from subject-level genotypes and phenotypes.

Citation

If you use AGENT-P in your work, please cite:

Hoang N.*, Tang K.*, Sardaripour N., Capra T., Rubinov M. Associations of
estimated gene expression with neuroimaging phenotypes across people and brain
regions with AGENT-P. Manuscript in preparation.

* These authors contributed equally.

Contact

Mika Rubinovmika.rubinov at vanderbilt.edu
Nhung Hoangnhung.t.hoang at outlook.com
Keyi Tangkeyi.tang at vanderbilt.edu

Bug reports and questions: open an issue on GitHub.