Structural variants (Cancer GIAB)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.13. Desktop beta builds are coming soon.
The Cancer Genome in a Bottle project publishes HG008, a matched tumor/normal PDAC cell line, as PacBio HiFi reads, a draft SV and CNV benchmark, and a telomere-to-telomere assembly of the tumor. We load them as JBrowse tracks and read each benchmark call against the alignments and the copy number under it. We:
- load the benchmark SV and CNV calls beside four further published callsets
- read one chr3-chr13 translocation three ways: the caller's breakend, the reads that span it, and the tumor assembly that resolves it onto one contig
- read a small heterozygous deletion in CUZD1 at base level, against ClinVar's copy-number variants
- check copy number at four driver genes, CDKN2A, TP53, SMAD4 and KRAS, against depth and B-allele frequency
- align the tumor assembly to GRCh38 and view the same rearrangement as synteny and a dotplot
Prerequisites
The walkthroughs at the end also run on the hosted C-GIAB demo. Building your own instance needs:
- A machine with HTTP access, either a public URL or
http://localhost - ~1 TB of free disk for the tracks, or ~1.5 TB for the full pipeline
- At least 32 GB of RAM for the minimap2 step; only data preparation needs it.
- The command-line tools below, with the versions tested in parentheses:
- JBrowse CLI (
@jbrowse/cliv3.6.5 or later) - Node.js (v18 minimum, v24.1.0 used for this tutorial)
- tabix (v1.21 or later)
- samtools (v1.21 or later)
- minimap2
- megadepth (v1.2.0 or later), for the coverage tracks
- HiFiCNV (v1.0 or later), for the binned depth track
- bcftools (v1.21 or later), for the B-allele frequency track
bedGraphToBigWigfrom the UCSC utilities
- JBrowse CLI (
Where the data comes from
HG008, the C-GIAB matched tumor/normal pair (NCBI BioProject PRJNA200694), is on the C-GIAB FTP. The assemblies are on NIST's S3 bucket.
The build script takes these files from their URLs, so there is nothing to download by hand.
The reference and the reads:
- the C-GIAB reference build (GRCh38 with decoys and masked regions): ftp-trace.ncbi.nlm.nih.gov/…/GRCh38_GIABv3_no_alt_analysis_set_maskedGRC_decoys_MAP2K3_KMT2C_KCNJ18.fasta.gzhttps://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh38/GRCh38_GIABv3_no_alt_analysis_set_maskedGRC_decoys_MAP2K3_KMT2C_KCNJ18.fasta.gz
- the tumor/normal PacBio HiFi reads (Revio run, 116x tumor, 35x normal): ftp-trace.ncbi.nlm.nih.gov/…/PacBio_Revio_20240125https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/PacBio_Revio_20240125/
The somatic call sets, one per group that published one:
- the V0.5 draft benchmark SV and CNV calls: ftp-trace.ncbi.nlm.nih.gov/…/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318/
- Severus somatic SVs (HiFi): ftp-trace.ncbi.nlm.nih.gov/…/severus_somatic.vcf.gzhttps://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi_Severus-SV_20240308/somatic_SVs/severus_somatic.vcf.gz
- the minda ensemble SVs (HiFi, ONT and Illumina callers): ftp-trace.ncbi.nlm.nih.gov/…/HG008_minda_ensemble.vcfhttps://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH-NCI_minda-ensemble_20240710/HG008_minda_ensemble.vcf
- DRAGEN's somatic SV and CNV calls (Illumina): ftp-trace.ncbi.nlm.nih.gov/…/standardhttps://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/DRAGEN-v4.2.4_ILMN-WGS_20240312/standard/
- NYGC's somatic SVs, annotated CNV segments and BIC-seq2 log2 ratio (Illumina): ftp-trace.ncbi.nlm.nih.gov/…/GRCh38-GIABv3https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NYGC-somatic-pipeline_20240412/GRCh38-GIABv3/
- Wakhan's haplotype-specific copy number and LOH segments (HiFi phased with Hi-C): ftp-trace.ncbi.nlm.nih.gov/…/bed_outputhttps://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi-HiC_Wakhan-CNA_20240424/bed_output/
- the earlier Wakhan run's Clair3 tumor small-variant calls: ftp-trace.ncbi.nlm.nih.gov/…/merge_output_tumor.vcf.gzhttps://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi_Wakhan-CNA_20240308/vcf_inputs/merge_output_tumor.vcf.gz
- the normal's germline calls: ftp-trace.ncbi.nlm.nih.gov/…/HG008-N-P.GRCh38.deepvariant.phased.vcf.gzhttps://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/PacBio_Revio_20240125/pacbio-wgs-wdl_germline_20240206/HG008-N-P.GRCh38.deepvariant.phased.vcf.gz
The assemblies:
- the T2T tumor assembly, v3.2: nist-giab.s3.us-east-1.amazonaws.com/…/HG008T_v3.2.fasta.gzhttps://nist-giab.s3.us-east-1.amazonaws.com/giab_tumor-normal/analysis/HG008/NIST_asm_dev/HG008T_v3.2/HG008T_v3.2.fasta.gz
- the matched normal assembly, v6.3, which the build script leaves out: nist-giab.s3.us-east-1.amazonaws.com/…/HG008N_v6.3.fasta.gzhttps://nist-giab.s3.us-east-1.amazonaws.com/giab_tumor-normal/analysis/HG008/NIST_asm_dev/HG008N_v6.3/HG008N_v6.3.fasta.gz
HG008, a pancreatic tumor and its matched normal
C-GIAB publishes HG008 from one pancreatic ductal adenocarcinoma (PDAC) donor, as the tumor (HG008-T) and normal pancreatic tissue (HG008-N-P). HG008-T is hypodiploid, with 35 tumor chromosomes down from 46, widespread arm-level loss, and truncal (present in every tumor cell) interchromosomal rearrangements (Wagner et al. 2026).
Installing megadepth and HiFiCNV
Set up JBrowse with the web quickstart. Two of the prerequisites install from release binaries:
wget https://github.com/ChristopherWilks/megadepth/releases/download/1.2.0/megadepth
chmod +x megadepth && sudo mv megadepth /usr/local/bin/
# --strip-components=1: drop the top-level folder of the release tarball
# --wildcards '*/hificnv': extract the binary alone
curl -L https://github.com/PacificBiosciences/HiFiCNV/releases/download/v1.0.1/hificnv-v1.0.1-x86_64-unknown-linux-gnu.tar.gz \
| tar xz --strip-components=1 -C /usr/local/bin --wildcards '*/hificnv'Assemblies: the C-GIAB GRCh38 reference and the tumor assembly
Every track below names one of two assemblies. The calls and reads sit on the
C-GIAB GRCh38 build, GRCh38_GIABv3. GIAB hosts it BGZF-compressed with its
.fai and .gzi, so the assembly reads it in place:
Goes in the assemblies array of config.json. See Assemblies.
{
"name": "GRCh38_GIABv3",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh38/GRCh38_GIABv3_no_alt_analysis_set_maskedGRC_decoys_MAP2K3_KMT2C_KCNJ18.fasta.gz"
}jbrowse add-assembly https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh38/GRCh38_GIABv3_no_alt_analysis_set_maskedGRC_decoys_MAP2K3_KMT2C_KCNJ18.fasta.gz \
--name GRCh38_GIABv3In JBrowse Desktop, Open new genome on the start screen (or File → Open genome... in a session), then Open from a URL and paste, one per line:
https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh38/GRCh38_GIABv3_no_alt_analysis_set_maskedGRC_decoys_MAP2K3_KMT2C_KCNJ18.fasta.gz
https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh38/GRCh38_GIABv3_no_alt_analysis_set_maskedGRC_decoys_MAP2K3_KMT2C_KCNJ18.fasta.gz.fai
https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh38/GRCh38_GIABv3_no_alt_analysis_set_maskedGRC_decoys_MAP2K3_KMT2C_KCNJ18.fasta.gz.gziJBrowse reads the format off the file name. Then fill in:
- Genome name:
GRCh38_GIABv3
The T2T tumor assembly HG008T_v3.2 is the other, used by the synteny and
dotplot sections. NIST ships it BGZF-compressed, so samtools faidx writes the
.fai and .gzi beside it without decompressing:
curl -LO https://nist-giab.s3.us-east-1.amazonaws.com/giab_tumor-normal/analysis/HG008/NIST_asm_dev/HG008T_v3.2/HG008T_v3.2.fasta.gz
samtools faidx HG008T_v3.2.fasta.gzGoes in the assemblies array of config.json. See Assemblies.
{
"name": "HG008T_v3.2",
"uri": "HG008T_v3.2.fasta.gz"
}jbrowse add-assembly HG008T_v3.2.fasta.gz \
--name HG008T_v3.2 \
--load copyIn JBrowse Desktop, Open new genome on the start screen (or File → Open genome... in a session), then Open from a URL and paste, one per line:
HG008T_v3.2.fasta.gz
HG008T_v3.2.fasta.gz.fai
HG008T_v3.2.fasta.gz.gziJBrowse reads the format off the file name. Then fill in:
- Genome name:
HG008T_v3.2
HG008T_v3.2.fasta.gz is relative to a config.json. Replace it with its URL or its path on this computer.
The benchmark SV and CNV calls
The V0.5 HG008-T draft benchmark is two files, the SV calls as an indexed VCF and the CNV calls as a BED, both loaded straight from their FTP URL.
Goes in the tracks array of config.json. See Tracks.
{
"trackId": "hg008t_benchmark_sv",
"name": "HG008-T V0.5 draft benchmark somatic SVs",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318/GRCh38_HG008-T-V0.5_somatic-stvar_PASS.draftbenchmark.vcf.gz",
"assemblyNames": ["GRCh38_GIABv3"]
}jbrowse add-track https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318/GRCh38_HG008-T-V0.5_somatic-stvar_PASS.draftbenchmark.vcf.gz \
--trackId hg008t_benchmark_sv \
--name "HG008-T V0.5 draft benchmark somatic SVs" \
--assemblyNames GRCh38_GIABv3In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track... and, in Add a track from file or URL, enter:
- Main file:
https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318/GRCh38_HG008-T-V0.5_somatic-stvar_PASS.draftbenchmark.vcf.gz
Click Next. JBrowse reads the adapter and track type off the file name. Then fill in:
- Track name:
HG008-T V0.5 draft benchmark somatic SVs - Assembly:
GRCh38_GIABv3
Click Add.
The CNV BED has no header, so name its columns with
columnNames:
Goes in the tracks array of config.json. See Tracks.
{
"type": "FeatureTrack",
"trackId": "hg008t_somatic_cnv",
"name": "HG008-T somatic CNV",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318/GRCh38_HG008-T-V0.5_somatic-CNV_PASS.draftbenchmark.calls.bed",
"columnNames": [
"chrom",
"start",
"end",
"total_copy_number",
"hap1_copy_number",
"hap2_copy_number",
"name"
]
},
"displayDefaults": {
"labels": {
"name": "jexl:'CN '+feature.total_copy_number+' ('+feature.hap1_copy_number+'|'+feature.hap2_copy_number+')'"
},
"color": {
"field": "total_copy_number",
"scale": "threshold",
"domain": ["1", "2", "3", "4"],
"range": ["#2166ac", "#92c5de", "#e0e0e0", "#f4a582", "#b2182b"],
"labels": ["CN 0", "CN 1", "CN 2", "CN 3", "CN 4+"],
"title": "Copy number"
}
}
}jbrowse add-track-json '{
"type": "FeatureTrack",
"trackId": "hg008t_somatic_cnv",
"name": "HG008-T somatic CNV",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318/GRCh38_HG008-T-V0.5_somatic-CNV_PASS.draftbenchmark.calls.bed",
"columnNames": [
"chrom",
"start",
"end",
"total_copy_number",
"hap1_copy_number",
"hap2_copy_number",
"name"
]
},
"displayDefaults": {
"labels": {
"name": "jexl:'\''CN '\''+feature.total_copy_number+'\'' ('\''+feature.hap1_copy_number+'\''|'\''+feature.hap2_copy_number+'\'')'\''"
},
"color": {
"field": "total_copy_number",
"scale": "threshold",
"domain": ["1", "2", "3", "4"],
"range": ["#2166ac", "#92c5de", "#e0e0e0", "#f4a582", "#b2182b"],
"labels": ["CN 0", "CN 1", "CN 2", "CN 3", "CN 4+"],
"title": "Copy number"
}
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "FeatureTrack",
"trackId": "hg008t_somatic_cnv",
"name": "HG008-T somatic CNV",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIST_HG008-T_somatic-stvar-CNV_DraftBenchmark_V0.5-20260318/GRCh38_HG008-T-V0.5_somatic-CNV_PASS.draftbenchmark.calls.bed",
"columnNames": [
"chrom",
"start",
"end",
"total_copy_number",
"hap1_copy_number",
"hap2_copy_number",
"name"
]
},
"displayDefaults": {
"labels": {
"name": "jexl:'CN '+feature.total_copy_number+' ('+feature.hap1_copy_number+'|'+feature.hap2_copy_number+')'"
},
"color": {
"field": "total_copy_number",
"scale": "threshold",
"domain": ["1", "2", "3", "4"],
"range": ["#2166ac", "#92c5de", "#e0e0e0", "#f4a582", "#b2182b"],
"labels": ["CN 0", "CN 1", "CN 2", "CN 3", "CN 4+"],
"title": "Copy number"
}
}
}Converting the HiFi reads to CRAM with coverage
The tumor and normal BAMs have no MD tags. Convert each to a local CRAM
against the reference above and write a coverage bigWig beside it:
# -T: the reference for the CRAM, the same build the assembly was loaded from
# --write-index: write the .crai in the same pass
samtools view HG008-T.bam --write-index -o HG008-T.cram -T GRCh38.fa
megadepth HG008-T.cram --bigwigEach CRAM is a read track that decodes against GRCh38_GIABv3; the normal is
the same with HG008-N.cram:
Goes in the tracks array of config.json. See Tracks.
{
"trackId": "HG008-T_hifi",
"name": "HG008-T PacBio HiFi",
"uri": "HG008-T.cram",
"assemblyNames": ["GRCh38_GIABv3"]
}jbrowse add-track HG008-T.cram \
--trackId HG008-T_hifi \
--name "HG008-T PacBio HiFi" \
--assemblyNames GRCh38_GIABv3 \
--load copyIn JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track... and, in Add a track from file or URL, enter:
- Main file:
HG008-T.cram
Click Next. JBrowse reads the adapter and track type off the file name. Then fill in:
- Track name:
HG008-T PacBio HiFi - Assembly:
GRCh38_GIABv3
Click Add.
HG008-T.cram is relative to a config.json. Replace it with its URL or its path on this computer.
Structural variants from the published callsets
C-GIAB publishes five somatic SV callsets for this pair, the benchmark among them:
| Callset | Called from |
|---|---|
| NIST V0.5 draft benchmark | assembly comparison plus read support |
| Severus | PacBio HiFi |
| minda ensemble | eleven caller runs over HiFi, ONT and Illumina |
| DRAGEN | Illumina WGS |
| NYGC somatic pipeline (Manta and GRIDSS) | Illumina WGS |
Loading Severus and DRAGEN SV calls
Severus and DRAGEN are both indexed VCFs with a record at each breakend, so they load with no display settings:
Goes in the tracks array of config.json. See Tracks.
{
"type": "VariantTrack",
"trackId": "hg008t_severus_sv",
"name": "HG008-T Severus somatic SVs (HiFi)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi_Severus-SV_20240308/somatic_SVs/severus_somatic.vcf.gz"
}
}jbrowse add-track https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi_Severus-SV_20240308/somatic_SVs/severus_somatic.vcf.gz \
--trackId hg008t_severus_sv \
--name "HG008-T Severus somatic SVs (HiFi)" \
--assemblyNames GRCh38_GIABv3In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track... and, in Add a track from file or URL, enter:
- Main file:
https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi_Severus-SV_20240308/somatic_SVs/severus_somatic.vcf.gz
Click Next. JBrowse reads the adapter and track type off the file name. Then fill in:
- Track name:
HG008-T Severus somatic SVs (HiFi) - Assembly:
GRCh38_GIABv3
Click Add.
Open this track in JBrowse Web ↗
Open this track in JBrowse Desktop ↗ (JBrowse Desktop 5.0+)
Goes in the tracks array of config.json. See Tracks.
{
"type": "VariantTrack",
"trackId": "hg008t_dragen_sv",
"name": "HG008-T DRAGEN somatic SVs (Illumina)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/DRAGEN-v4.2.4_ILMN-WGS_20240312/standard/dragen_4.2.4_HG008-mosaic_tumor.sv.vcf.gz"
}
}jbrowse add-track https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/DRAGEN-v4.2.4_ILMN-WGS_20240312/standard/dragen_4.2.4_HG008-mosaic_tumor.sv.vcf.gz \
--trackId hg008t_dragen_sv \
--name "HG008-T DRAGEN somatic SVs (Illumina)" \
--assemblyNames GRCh38_GIABv3In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track... and, in Add a track from file or URL, enter:
- Main file:
https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/DRAGEN-v4.2.4_ILMN-WGS_20240312/standard/dragen_4.2.4_HG008-mosaic_tumor.sv.vcf.gz
Click Next. JBrowse reads the adapter and track type off the file name. Then fill in:
- Track name:
HG008-T DRAGEN somatic SVs (Illumina) - Assembly:
GRCh38_GIABv3
Click Add.
Open this track in JBrowse Web ↗
Open this track in JBrowse Desktop ↗ (JBrowse Desktop 5.0+)
Loading the minda ensemble SV calls
The minda ensemble callset is a plain VCF, so
VcfAdapter loads it whole, with no index:
Goes in the tracks array of config.json. See Tracks.
{
"type": "VariantTrack",
"trackId": "hg008t_minda_sv",
"name": "HG008-T minda ensemble SVs (HiFi, ONT, Illumina)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "VcfAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH-NCI_minda-ensemble_20240710/HG008_minda_ensemble.vcf"
}
}jbrowse add-track https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH-NCI_minda-ensemble_20240710/HG008_minda_ensemble.vcf \
--trackId hg008t_minda_sv \
--name "HG008-T minda ensemble SVs (HiFi, ONT, Illumina)" \
--assemblyNames GRCh38_GIABv3In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track... and, in Add a track from file or URL, enter:
- Main file:
https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH-NCI_minda-ensemble_20240710/HG008_minda_ensemble.vcf
Click Next. JBrowse reads the adapter and track type off the file name. Then fill in:
- Track name:
HG008-T minda ensemble SVs (HiFi, ONT, Illumina) - Assembly:
GRCh38_GIABv3
Click Add.
Open this track in JBrowse Web ↗
Open this track in JBrowse Desktop ↗ (JBrowse Desktop 5.0+)
Drawing the NYGC SV calls (BEDPE) as arcs
BedpeAdapter reads a paired-end BED whole, with
no index, and serves it to a variant track:
Goes in the tracks array of config.json. See Tracks.
{
"type": "VariantTrack",
"trackId": "hg008t_nygc_sv",
"name": "HG008-T NYGC somatic SVs (Manta, GRIDSS)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedpeAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NYGC-somatic-pipeline_20240412/GRCh38-GIABv3/HG008-T--HG008-N.sv.annotated.v7.somatic.high_confidence.final.bedpe"
},
"displays": [
{
"type": "LinearMarkDisplay",
"marks": [
{
"mark": "link",
"encoding": { "size": 2 },
"transform": [{ "type": "mate" }]
}
]
}
]
}jbrowse add-track-json '{
"type": "VariantTrack",
"trackId": "hg008t_nygc_sv",
"name": "HG008-T NYGC somatic SVs (Manta, GRIDSS)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedpeAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NYGC-somatic-pipeline_20240412/GRCh38-GIABv3/HG008-T--HG008-N.sv.annotated.v7.somatic.high_confidence.final.bedpe"
},
"displays": [
{
"type": "LinearMarkDisplay",
"marks": [
{
"mark": "link",
"encoding": { "size": 2 },
"transform": [{ "type": "mate" }]
}
]
}
]
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "VariantTrack",
"trackId": "hg008t_nygc_sv",
"name": "HG008-T NYGC somatic SVs (Manta, GRIDSS)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedpeAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NYGC-somatic-pipeline_20240412/GRCh38-GIABv3/HG008-T--HG008-N.sv.annotated.v7.somatic.high_confidence.final.bedpe"
},
"displays": [
{
"type": "LinearMarkDisplay",
"marks": [
{
"mark": "link",
"encoding": { "size": 2 },
"transform": [{ "type": "mate" }]
}
]
}
]
}Open chr3:139,970,000-140,005,000 with the five callsets loaded. A link whose
other breakend lies outside the view draws as a short stem at the end in view,
as both NYGC records do in the figure below.
Copy number from the published callsets
C-GIAB publishes copy-number calls on this pair from four groups:
| Callset | Called from | Each segment has |
|---|---|---|
| NIST V0.5 draft benchmark | assembly comparison plus read support | absolute total and per-haplotype copy number |
| Wakhan | PacBio HiFi, phased with Hi-C | copy number per parental haplotype, with loss-of-heterozygosity (LOH) intervals in a second file |
| NYGC somatic pipeline, BIC-seq2 | Illumina WGS | log2 tumor-versus-normal copy ratio, and the genes the segment covers |
| DRAGEN | Illumina WGS | integer copy number, minor-haplotype copy number and minor allele frequency |
Read the other callsets against the benchmark CNV BED added above, whose copy numbers are absolute, so CN 2 marks a diploid region.
Loading DRAGEN copy-number segments
DRAGEN's CNV calls load as an indexed VCF, one record per segment. Record IDs
read DRAGEN:CNLOH:chr9:22631070-22939213 (copy-neutral loss of
heterozygosity), so the label takes the class from the second field:
Goes in the tracks array of config.json. See Tracks.
{
"type": "VariantTrack",
"trackId": "hg008t_dragen_cnv",
"name": "HG008-T DRAGEN somatic CNV (Illumina)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/DRAGEN-v4.2.4_ILMN-WGS_20240312/standard/dragen_4.2.4_HG008-mosaic_tumor.cnv.vcf.gz"
},
"displayDefaults": {
"labels": { "name": "jexl:split(feature.name,':')[1]" }
}
}jbrowse add-track https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/DRAGEN-v4.2.4_ILMN-WGS_20240312/standard/dragen_4.2.4_HG008-mosaic_tumor.cnv.vcf.gz \
--trackId hg008t_dragen_cnv \
--name "HG008-T DRAGEN somatic CNV (Illumina)" \
--assemblyNames GRCh38_GIABv3 \
--displayDefaults "{\"labels\":{\"name\":\"jexl:split(feature.name,':')[1]\"}}"In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "VariantTrack",
"trackId": "hg008t_dragen_cnv",
"name": "HG008-T DRAGEN somatic CNV (Illumina)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/DRAGEN-v4.2.4_ILMN-WGS_20240312/standard/dragen_4.2.4_HG008-mosaic_tumor.cnv.vcf.gz"
},
"displayDefaults": {
"labels": { "name": "jexl:split(feature.name,':')[1]" }
}
}NYGC copy-number segments and log2 copy ratio
C-GIAB publishes the NYGC CNV output in two forms.
HG008-T--HG008-N.cnv.annotated.v7.final.bed loads as is, because the adapter
reads column names from its # header line.
Goes in the tracks array of config.json. See Tracks.
{
"type": "FeatureTrack",
"trackId": "hg008t_nygc_cnv",
"name": "HG008-T NYGC CNV calls, annotated (BIC-seq2)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NYGC-somatic-pipeline_20240412/GRCh38-GIABv3/HG008-T--HG008-N.cnv.annotated.v7.final.bed"
},
"displayDefaults": {
"displayMode": "compact",
"color": {
"field": "type",
"domain": ["DEL", "DUP"],
"range": ["#2166ac", "#b2182b"],
"labels": ["Loss (DEL)", "Gain (DUP)"],
"title": "Call"
},
"labels": { "name": "jexl:feature.type+' '+feature.cytoband" }
}
}jbrowse add-track https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NYGC-somatic-pipeline_20240412/GRCh38-GIABv3/HG008-T--HG008-N.cnv.annotated.v7.final.bed \
--trackId hg008t_nygc_cnv \
--name "HG008-T NYGC CNV calls, annotated (BIC-seq2)" \
--assemblyNames GRCh38_GIABv3 \
--displayDefaults "{\"displayMode\":\"compact\",\"color\":{\"field\":\"type\",\"domain\":[\"DEL\",\"DUP\"],\"range\":[\"#2166ac\",\"#b2182b\"],\"labels\":[\"Loss (DEL)\",\"Gain (DUP)\"],\"title\":\"Call\"},\"labels\":{\"name\":\"jexl:feature.type+' '+feature.cytoband\"}}"In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "FeatureTrack",
"trackId": "hg008t_nygc_cnv",
"name": "HG008-T NYGC CNV calls, annotated (BIC-seq2)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NYGC-somatic-pipeline_20240412/GRCh38-GIABv3/HG008-T--HG008-N.cnv.annotated.v7.final.bed"
},
"displayDefaults": {
"displayMode": "compact",
"color": {
"field": "type",
"domain": ["DEL", "DUP"],
"range": ["#2166ac", "#b2182b"],
"labels": ["Loss (DEL)", "Gain (DUP)"],
"title": "Call"
},
"labels": { "name": "jexl:feature.type+' '+feature.cytoband" }
}
}HG008-T--HG008-N.bicseq2.txt is the same segmentation with one log2 ratio per
segment. One awk command converts it to a bedGraph:
# column 9 is log2.copyRatio; $2-1 converts the 1-based start to 0-based
awk 'NR>1 {printf "%s\t%d\t%d\t%.4f\n", $1, $2-1, $3, $9}' \
HG008-T--HG008-N.bicseq2.txt > HG008-T_bicseq2_log2ratio.bedgraphWe'll plot it as bars from zero on a fixed axis, so a step means the same thing from one window to the next:
- A homozygous deletion has no reads and so no finite ratio.
- BIC-seq2 normalizes on total read counts and HG008-T is hypodiploid, so a balanced region sits above zero.
Goes in the tracks array of config.json. See Tracks.
{
"type": "QuantitativeTrack",
"trackId": "hg008t_bicseq2",
"name": "HG008-T copy ratio, segmented (NYGC BIC-seq2, log2 T/N)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedGraphAdapter",
"uri": "HG008-T_bicseq2_log2ratio.bedgraph"
},
"displayDefaults": {
"mark": "bar",
"scales": {
"y": {
"domainMin": -2,
"domainMax": 2,
"grid": true,
"title": "log2 tumor/normal"
}
}
}
}jbrowse add-track HG008-T_bicseq2_log2ratio.bedgraph \
--trackId hg008t_bicseq2 \
--name "HG008-T copy ratio, segmented (NYGC BIC-seq2, log2 T/N)" \
--assemblyNames GRCh38_GIABv3 \
--displayDefaults '{"mark":"bar","scales":{"y":{"domainMin":-2,"domainMax":2,"grid":true,"title":"log2 tumor/normal"}}}' \
--load copyIn JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "QuantitativeTrack",
"trackId": "hg008t_bicseq2",
"name": "HG008-T copy ratio, segmented (NYGC BIC-seq2, log2 T/N)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedGraphAdapter",
"uri": "HG008-T_bicseq2_log2ratio.bedgraph"
},
"displayDefaults": {
"mark": "bar",
"scales": {
"y": {
"domainMin": -2,
"domainMax": 2,
"grid": true,
"title": "log2 tumor/normal"
}
}
}
}HG008-T_bicseq2_log2ratio.bedgraph is relative to a config.json. Replace it with its URL or its path on this computer.
Loading Wakhan copy number per haplotype
Wakhan's HG008_HiFi_HiC_copynumbers_segments.bed is long format, one row per
haplotype, with no # on its column-name line:
chr start end copynumber_state coverage haplotype
chr1 0 23750000 2 106.025 1
chr1 23750001 119650000 0.72 58.025 1Because the haplotype column already assigns each segment to a row, we use
LinearMultiRowFeatureDisplay and
set rows to
haplotype. The adapter names its columns with
columnNames:
Goes in the tracks array of config.json. See Tracks.
{
"type": "FeatureTrack",
"trackId": "hg008t_wakhan_hifi_hic",
"name": "HG008-T Wakhan copy number per haplotype (HiFi + Hi-C)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi-HiC_Wakhan-CNA_20240424/bed_output/HG008_HiFi_HiC_copynumbers_segments.bed",
"columnNames": [
"chrom",
"start",
"end",
"copynumber_state",
"coverage",
"haplotype"
]
},
"displays": [
{
"type": "LinearMultiRowFeatureDisplay",
"rows": "haplotype",
"color": {
"field": "copynumber_state",
"scale": "threshold",
"domain": ["0.5", "1.5"],
"range": ["#2166ac", "#bdbdbd", "#f4a582"],
"labels": ["Haplotype lost (0)", "One copy", "Two or more copies"],
"title": "Copy number per haplotype"
}
}
]
}jbrowse add-track-json '{
"type": "FeatureTrack",
"trackId": "hg008t_wakhan_hifi_hic",
"name": "HG008-T Wakhan copy number per haplotype (HiFi + Hi-C)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi-HiC_Wakhan-CNA_20240424/bed_output/HG008_HiFi_HiC_copynumbers_segments.bed",
"columnNames": [
"chrom",
"start",
"end",
"copynumber_state",
"coverage",
"haplotype"
]
},
"displays": [
{
"type": "LinearMultiRowFeatureDisplay",
"rows": "haplotype",
"color": {
"field": "copynumber_state",
"scale": "threshold",
"domain": ["0.5", "1.5"],
"range": ["#2166ac", "#bdbdbd", "#f4a582"],
"labels": ["Haplotype lost (0)", "One copy", "Two or more copies"],
"title": "Copy number per haplotype"
}
}
]
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "FeatureTrack",
"trackId": "hg008t_wakhan_hifi_hic",
"name": "HG008-T Wakhan copy number per haplotype (HiFi + Hi-C)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BedAdapter",
"uri": "https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data_somatic/HG008/Liss_lab/analysis/NIH_HiFi-HiC_Wakhan-CNA_20240424/bed_output/HG008_HiFi_HiC_copynumbers_segments.bed",
"columnNames": [
"chrom",
"start",
"end",
"copynumber_state",
"coverage",
"haplotype"
]
},
"displays": [
{
"type": "LinearMultiRowFeatureDisplay",
"rows": "haplotype",
"color": {
"field": "copynumber_state",
"scale": "threshold",
"domain": ["0.5", "1.5"],
"range": ["#2166ac", "#bdbdbd", "#f4a582"],
"labels": ["Haplotype lost (0)", "One copy", "Two or more copies"],
"title": "Copy number per haplotype"
}
}
]
}A copynumber_state of 0 marks a haplotype lost across an arm (LOH), and 1
is the normal single copy.
Building depth-per-bin and B-allele frequency tracks
HiFiCNV writes a binned depth track from the tumor reads:
# the depth output is <prefix>.<sample>.depth.bw, named for the --bam sample,
# so this run writes hificnv.<sample>.depth.bw
# the hosted demo renames it HG008-T.hificnv.depth.bw, the uri below
# --maf: the Clair3 tumor small-variant VCF from the C-GIAB Wakhan run;
# HiFiCNV reads AD from it for the allele-frequency output
# --bam: the reads, for depth
hificnv --bam HG008-T.cram --ref GRCh38.fa --maf tumor_smallvariants.vcf.gz \
--output-prefix hificnvWe'll load the depth as a quantitative track drawn as points, the Scatter plot type:
Goes in the tracks array of config.json. See Tracks.
{
"type": "QuantitativeTrack",
"trackId": "hg008_depth",
"name": "HG008-T HiFiCNV depth",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BigWigAdapter",
"uri": "HG008-T.hificnv.depth.bw"
},
"displayDefaults": { "mark": "point", "size": 1 }
}jbrowse add-track HG008-T.hificnv.depth.bw \
--trackId hg008_depth \
--name "HG008-T HiFiCNV depth" \
--assemblyNames GRCh38_GIABv3 \
--displayDefaults '{"mark":"point","size":1}' \
--load copyIn JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "QuantitativeTrack",
"trackId": "hg008_depth",
"name": "HG008-T HiFiCNV depth",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BigWigAdapter",
"uri": "HG008-T.hificnv.depth.bw"
},
"displayDefaults": { "mark": "point", "size": 1 }
}HG008-T.hificnv.depth.bw is relative to a config.json. Replace it with its URL or its path on this computer.
To compare the depth with the callsets, draw it with Plot type → XY plot and
a fixed axis from Y axis... → Range, so a halving reads as a level. Open
chr9:21,850,000-23,050,000, around CDKN2A, with the depth above the four
published CNV callsets:
B-allele frequency (BAF) is the fraction of tumor reads with the alternate allele, drawn unfolded:
- A balanced region is one band at 0.5.
- A region with LOH (one parental copy gone) splits into two bands at 0 and 1.
Build it by piling up the tumor reads at the sites the normal calls heterozygous and taking the alt fraction:
# het sites come from the normal, because LOH sites are homozygous in the tumor
bcftools view -g het -Oz -o hets.vcf.gz normal.deepvariant.vcf.gz
tabix -p vcf hets.vcf.gz
cut -f1,2 GRCh38.fa.fai > GRCh38.chrom.sizes
# -q 1 drops multi-mapped reads, -Q 0 leaves HiFi base qualities alone
bcftools mpileup -f GRCh38.fa -T hets.vcf.gz -a AD -q 1 -Q 0 tumor.bam |
bcftools query -f '%CHROM\t%POS\t[%AD]\n' |
# unfolded alt fraction at sites with depth 10 or more
awk -F'[\t,]' '{d=$3+$4; if (d>=10) printf "%s\t%d\t%d\t%.4f\n",$1,$2-1,$2,$4/d}' |
LC_COLLATE=C sort -k1,1 -k2,2n > baf.bedgraph
bedGraphToBigWig baf.bedgraph GRCh38.chrom.sizes HG008-T_baf.bcftools.bwPlot it with Scatter over a fixed 0 to 1 range.
Drawing individual BAF points when zoomed out
Zoomed out, the default bigWig summary paints an LOH arm as a solid block of
full height. A small
resolutionMultiplier
makes the adapter fetch raw per-site values at the zoom levels of these figures:
Goes in the tracks array of config.json. See Tracks.
{
"type": "QuantitativeTrack",
"trackId": "HG008-T_baf",
"name": "HG008-T B-allele frequency (BAF)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BigWigAdapter",
"uri": "HG008-T_baf.bcftools.bw",
"resolutionMultiplier": 0.001
},
"displayDefaults": {
"mark": "point",
"size": 1,
"scales": { "y": { "domainMin": 0, "domainMax": 1 } }
}
}jbrowse add-track-json '{
"type": "QuantitativeTrack",
"trackId": "HG008-T_baf",
"name": "HG008-T B-allele frequency (BAF)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BigWigAdapter",
"uri": "HG008-T_baf.bcftools.bw",
"resolutionMultiplier": 0.001
},
"displayDefaults": {
"mark": "point",
"size": 1,
"scales": { "y": { "domainMin": 0, "domainMax": 1 } }
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "QuantitativeTrack",
"trackId": "HG008-T_baf",
"name": "HG008-T B-allele frequency (BAF)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BigWigAdapter",
"uri": "HG008-T_baf.bcftools.bw",
"resolutionMultiplier": 0.001
},
"displayDefaults": {
"mark": "point",
"size": 1,
"scales": { "y": { "domainMin": 0, "domainMax": 1 } }
}
}HG008-T_baf.bcftools.bw is relative to a config.json. Replace it with its URL or its path on this computer.
Aligning the tumor assembly to GRCh38
The tumor assembly is haplotype-resolved into T2T scaffolds. The synteny and dotplot views draw from its alignment to GRCh38, a PAF file:
# asm5: minimap2's preset for an assembly against a reference of the same species
# -c: write base-level CIGARs
minimap2 -cx asm5 GRCh38.fa HG008T_v3.2.fasta > HG008T_v3.2.pafThe track lists the query assembly (the tumor assembly) first and the target (GRCh38) second, the reverse of the minimap2 argument order; reversed, the view opens empty and reports no error.
Goes in the tracks array of config.json. See Tracks.
{
"type": "SyntenyTrack",
"trackId": "HG008T_v3.2_paf",
"name": "HG008T_v3.2 vs GRCh38_GIABv3",
"assemblyNames": ["HG008T_v3.2", "GRCh38_GIABv3"],
"adapter": {
"type": "PAFAdapter",
"uri": "HG008T_v3.2.paf"
}
}jbrowse add-track HG008T_v3.2.paf \
--trackId HG008T_v3.2_paf \
--name "HG008T_v3.2 vs GRCh38_GIABv3" \
--assemblyNames HG008T_v3.2,GRCh38_GIABv3 \
--load copyIn JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "SyntenyTrack",
"trackId": "HG008T_v3.2_paf",
"name": "HG008T_v3.2 vs GRCh38_GIABv3",
"assemblyNames": ["HG008T_v3.2", "GRCh38_GIABv3"],
"adapter": {
"type": "PAFAdapter",
"uri": "HG008T_v3.2.paf"
}
}HG008T_v3.2.paf is relative to a config.json. Replace it with its URL or its path on this computer.
That shows it in the linear view. For the synteny view, open Add → Linear synteny view, pick the track under Quick start, and click Launch.
The matched normal assembly (HG008N_v6.3.fasta.gz, same S3 path) loads the
same way. See the
synteny track config guide.
Checking SV and CNV calls against reads, depth and assembly
A chr3-chr13 translocation: breakpoint reads and the tumor contig
Add → SV inspector, then Open from track to pick the C-GIAB benchmark VCF loaded earlier.
SV_20 and SV_190 are the two records of one breakpoint, where
chr3:139,976,414 joins chr13:114,353,244. The benchmark files them under
EVENT=cluster_3 with two further breakends and tags the cluster
EVENTTYPE=CHROMOPLEXY (rearrangements chained across chromosomes). To read the
junction in the reads:
- Choose
cluster_3under Filter by event, then click the chord joining chr3 and chr13 to launch the breakpoint split view. - Set Read height → Compact on the tumor PacBio HiFi reads opened on each panel. The cancer SV tutorial covers opening the view from a record and following the further breakends.
The same junction reads three ways:
- Tumor reads: each split read has one half on chr13 and one on chr3, joined by a spline. The read runs forward on chr13 into the junction, then continues on the reverse strand of chr3.
- Matched normal: open
HG008-N-P_PacBio-HiFi-Revio_20240125_35x_GRCh38-GIABv3from the hosted config on the same two panels. The reads run through the locus with no split. - Tumor assembly: the synteny track loaded earlier shows the same junction with no reads. The C-GIAB assembly resolves both loci onto one tumor contig, named for the two chromosomes it fuses.
CUZD1: a small heterozygous deletion
Use the search (magnifying glass) button in the SV inspector to find
SV_85, a heterozygous deletion of two exons of CUZD1
(NCBI Gene 50624). At about 1.8 kb,
the whole deletion fits in a pileup at base level.
The ClinVar CNVs track holds submitted copy-number variants and their clinical significance, served by UCSC as a bigBed:
Goes in the tracks array of config.json. See Tracks.
{
"type": "FeatureTrack",
"trackId": "hg38_clinvar_cnv_ucsc",
"name": "ClinVar CNVs (UCSC)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BigBedAdapter",
"uri": "https://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/clinvar/clinvarCnv.bb"
},
"displayDefaults": {
"displayMode": "compact",
"filter": ["jexl:feature._varLen < 50000"]
}
}jbrowse add-track https://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/clinvar/clinvarCnv.bb \
--trackId hg38_clinvar_cnv_ucsc \
--name "ClinVar CNVs (UCSC)" \
--assemblyNames GRCh38_GIABv3 \
--displayDefaults '{"displayMode":"compact","filter":["jexl:feature._varLen < 50000"]}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "FeatureTrack",
"trackId": "hg38_clinvar_cnv_ucsc",
"name": "ClinVar CNVs (UCSC)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "BigBedAdapter",
"uri": "https://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/clinvar/clinvarCnv.bb"
},
"displayDefaults": {
"displayMode": "compact",
"filter": ["jexl:feature._varLen < 50000"]
}
}The filter keeps ClinVar CNVs under 50 kb, and none of those covers this deletion.
Open NCBI RefSeq genes (hg38) from the hosted config and the tumor PacBio HiFi reads, set Read height → Compact and Sort by... → Base pair from the track menu, and center the deletion. Show... → Show center line in the view menu helps line up the breakpoint.
Checking copy number at four pancreatic cancer driver genes
In pancreatic ductal adenocarcinoma the recurrently altered genes are KRAS, CDKN2A, TP53 and SMAD4 (Waddell et al. 2015, Bailey et al. 2016). Each copy-number figure below draws the gene's MANE Select (standard reference) transcript under the tracks.
For a first check, load the tumor and normal coverage from
goleft indexcov, one
bigWig per sample, as one multi-wiggle track. indexcov scales each sample by
that sample's median depth, so the two samples compare directly. Its output is a
per-sample table, so the track below points at the demo's bigWig copies; swap
each uri for your sample's bigWig:
Goes in the tracks array of config.json. See Tracks.
{
"type": "MultiQuantitativeTrack",
"trackId": "hg008_cnv_indexcov",
"name": "HG008 normal vs tumor coverage (indexcov)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"name": "HG008-N (normal)",
"type": "BigWigAdapter",
"bigWigLocation": {
"uri": "https://jbrowse.org/demos/cgiab/HG008-N_indexcov.bw"
}
},
{
"name": "HG008-T (tumor)",
"type": "BigWigAdapter",
"bigWigLocation": {
"uri": "https://jbrowse.org/demos/cgiab/HG008-T_indexcov.bw"
}
}
]
}
}jbrowse add-track-json '{
"type": "MultiQuantitativeTrack",
"trackId": "hg008_cnv_indexcov",
"name": "HG008 normal vs tumor coverage (indexcov)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"name": "HG008-N (normal)",
"type": "BigWigAdapter",
"bigWigLocation": {
"uri": "https://jbrowse.org/demos/cgiab/HG008-N_indexcov.bw"
}
},
{
"name": "HG008-T (tumor)",
"type": "BigWigAdapter",
"bigWigLocation": {
"uri": "https://jbrowse.org/demos/cgiab/HG008-T_indexcov.bw"
}
}
]
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "MultiQuantitativeTrack",
"trackId": "hg008_cnv_indexcov",
"name": "HG008 normal vs tumor coverage (indexcov)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"name": "HG008-N (normal)",
"type": "BigWigAdapter",
"bigWigLocation": {
"uri": "https://jbrowse.org/demos/cgiab/HG008-N_indexcov.bw"
}
},
{
"name": "HG008-T (tumor)",
"type": "BigWigAdapter",
"bigWigLocation": {
"uri": "https://jbrowse.org/demos/cgiab/HG008-T_indexcov.bw"
}
}
]
}
}Plot type → Overlapping → Scatter in its track menu draws the two samples as points in one band, with a key naming each sample's color. Y axis... → Range with 0 and 3 stops indexcov's centromere spikes from flattening every plateau.
Zoom to a region and open the benchmark CNV BED. Coverage shows where the copy number steps, and the BAF track shows the allelic balance across each step.
CDKN2A: homozygous deletion
Navigate to CDKN2A on chr9: the benchmark calls a focal ~20 kb homozygous
deletion over the gene (SV_75, CN 0, no copies left), inside a larger
single-copy-loss arm (CNA_14, 0+1 copies on the two haplotypes) where depth is
already halved
(Wagner et al. 2026).
Load the tumor and matched normal per-base coverage as one
multi-quantitative track, one row per
sample, with an explicit score range. The two bigWigs are the megadepth
outputs for each sample's CRAM; swap in yours:
Goes in the tracks array of config.json. See Tracks.
{
"type": "MultiQuantitativeTrack",
"trackId": "hg008_tn_perbase",
"name": "HG008 tumor vs matched normal coverage (per-base)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"name": "HG008-T (tumor)",
"type": "BigWigAdapter",
"color": "#e41a1c",
"bigWigLocation": { "uri": "HG008-T.cram.all.bw" }
},
{
"name": "HG008-N (normal)",
"type": "BigWigAdapter",
"color": "#377eb8",
"bigWigLocation": { "uri": "HG008-N.cram.all.bw" }
}
]
},
"displayDefaults": {
"mark": "bar",
"summaryScoreMode": "mean",
"scales": { "y": { "domainMin": 0, "domainMax": 80, "grid": false } },
"height": 280
}
}jbrowse add-track-json '{
"type": "MultiQuantitativeTrack",
"trackId": "hg008_tn_perbase",
"name": "HG008 tumor vs matched normal coverage (per-base)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"name": "HG008-T (tumor)",
"type": "BigWigAdapter",
"color": "#e41a1c",
"bigWigLocation": { "uri": "HG008-T.cram.all.bw" }
},
{
"name": "HG008-N (normal)",
"type": "BigWigAdapter",
"color": "#377eb8",
"bigWigLocation": { "uri": "HG008-N.cram.all.bw" }
}
]
},
"displayDefaults": {
"mark": "bar",
"summaryScoreMode": "mean",
"scales": { "y": { "domainMin": 0, "domainMax": 80, "grid": false } },
"height": 280
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "MultiQuantitativeTrack",
"trackId": "hg008_tn_perbase",
"name": "HG008 tumor vs matched normal coverage (per-base)",
"assemblyNames": ["GRCh38_GIABv3"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"name": "HG008-T (tumor)",
"type": "BigWigAdapter",
"color": "#e41a1c",
"bigWigLocation": { "uri": "HG008-T.cram.all.bw" }
},
{
"name": "HG008-N (normal)",
"type": "BigWigAdapter",
"color": "#377eb8",
"bigWigLocation": { "uri": "HG008-N.cram.all.bw" }
}
]
},
"displayDefaults": {
"mark": "bar",
"summaryScoreMode": "mean",
"scales": { "y": { "domainMin": 0, "domainMax": 80, "grid": false } },
"height": 280
}
}HG008-T.cram.all.bw, HG008-N.cram.all.bw are relative to a config.json. Replace each with its URL or its path on this computer.
Thin lines across the gap in the read pileup are reads with the deletion.
The benchmark's total_copy_number is absolute: CN 2 is diploid, and 9p has
already lost a copy, so CN 1 is the local background. Widen the view several
hundred kilobases right to read CN 2 against it.
Chromosome 17: two kinds of loss of heterozygosity
Chromosome 17 has two LOH states, with the boundary partway along the q-arm. Open the whole chromosome with the depth track above the BAF:
- The p-arm (covering TP53) and the start of the q-arm are a single-copy
loss with LOH (
CNA_20, CN 1, 1+0). - The rest of the q-arm is copy-neutral LOH (
CNA_21, CN 2, 2+0): one haplotype lost, the other duplicated.
Depth and BAF together separate four states:
| depth | BAF | Interpretation |
|---|---|---|
| flat (CN 2) | one band at 0.5 | balanced diploid |
| flat (CN 2) | split to 0, 1 | copy-neutral LOH |
| halved | split to 0, 1 | single-copy loss with LOH |
| raised | 1/3 and 2/3 | allelic gain |
The hap1_copy_number and hap2_copy_number columns of the benchmark BED give
the 1+0 and 2+0 splits. Click a CNV feature to see both.
KRAS gain and SMAD4 loss
KRAS on chr12 sits in a gain (SV_101, CN 3, 2+1): a 2 Mb tandem duplication
that includes the G12V-mutated copy
(Wagner et al. 2026). At
whole-chromosome scale it is a handful of pixels wide.
SMAD4 on 18q is a single-copy loss with LOH (CNA_48, CN 1, 0+1), like
TP53. The balanced p-arm and the matched normal are the two controls. Leave
the copy-ratio track's color unset (one color above the origin of 0, another
below) with a symmetric axis, so a step down and a step up fill equally.
Dotplot and synteny views of the tumor assembly against GRCh38
Add → Dotplot view, set the tumor assembly as one axis and GRCh38 as the other, and pick the matching synteny track.
HG008-T v3.2's scaffold names end in _hap1 or _hap2, so one plot stacks both
haplotypes and doubles every diagonal. In the import form, tick Plot only
certain chromosomes and type *_hap1 into the chromosomes box of the
HG008T_v3.2 axis, then *_hap2 in a second view, for a plain
assembly-vs-reference diagonal per haplotype.
Drag over a region and take Launch → Linear synteny view, keeping
HG008T_v3.2 vs GRCh38_GIABv3 (HG008T v3.2 in the hosted config) as the
synteny track, then enter chr13 chr3 in the GRCh38 search box.
Add row in the import form gives each haplotype a separate row: hap2, GRCh38, hap1, with ribbons between each adjacent pair.
For more on these views, see the dotplot view guide and the linear synteny view guide.
Reproduce it end to end
build_sv_visualization_cgiab.sh
runs the data preparation above in one shot. The script:
- Fetches the C-GIAB GRCh38 build and the V0.5 benchmark calls.
- Runs the CRAM, coverage, HiFiCNV, BAF and minimap2 steps above.
- Downloads JBrowse and writes a
config.jsonwith all of it loaded beside the published Wakhan segments.
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_sv_visualization_cgiab.sh
bash build_sv_visualization_cgiab.sh # builds ./cgiab_build/jbrowse2
npx --yes serve cgiab_build/jbrowse2It downloads more than 200 GB and takes hours.
See also
- Synteny visualization (pairwise minimap2)
- Reviewing a whole SV callset
- Complex rearrangements and derivative alleles
- Gene fusion calls and the DNA behind them
- SV visualization
- SV inspector view
- User guide: Quantitative track
Citations
- Bailey et al. (2016). Genomic analyses identify molecular subtypes of pancreatic cancer
- Diesh et al. (2023). JBrowse 2: A Modular Genome Browser with Views of Synteny and Structural Variation
- McDaniel et al. (2025). Development and Extensive Sequencing of a Broadly-Consented Genome in a Bottle Matched Tumor-Normal Pair
- Rautiainen et al. (2023). Verkko: telomere-to-telomere assembly of diploid chromosomes
- Waddell et al. (2015). Whole genomes redefine the mutational landscape of pancreatic cancer
- Wagner et al. (2026). A complete human pancreatic cancer genome
Feedback on this tutorial is welcome: contact us.