Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Altered Cancer Pathways

Maintained by the JuliaQUBO organization
SECQUOIA  ·  PSR Energy

Open In Colab

An original Julia case study balancing mutation coverage and pairwise co-mutation.

Educational use only. This notebook reproduces an optimization formulation on a fixed public-data aggregate. It does not identify a clinically validated pathway and is not medical guidance.

Setup

Local installation

From the repository root, instantiate the shared Julia environment before opening this notebook:

julia --project=notebooks_jl -e 'using Pkg; Pkg.instantiate()'

Restart the kernel and run every cell with the targeted command:

make verify-cancer-genomics-julia

Ordinary execution reads only the committed files under notebooks_data. It requires no account, credential, live data request, or quantum service. The optional make refresh-tcga-aml maintenance command is documented in notebooks_data/README.md and is not run here.

Google Colab

Open the badge above, select a Julia runtime, and run the setup cells. The bootstrap clones this repository only when Colab does not already have it, then activates the same checked-in notebook project and reads the same committed aggregate.

Notebook Cell
Notebook Cell

Learning objectives

By the end of this notebook you will be able to:

  1. construct coverage and symmetric co-mutation matrices from a patient-by-gene incidence matrix;

  2. explain the double-counting convention in the matrix QUBO;

  3. exhaustively verify a tiny fixture against an independent analytic energy;

  4. decode coverage, distinct-patient coverage, co-mutation, pathway size, and raw energy safely; and

  5. run and validate a seeded local sampler on a provenance-checked TCGA AML aggregate.

Prerequisites

Prior notebooks: Notebook 2 introduces JuMP/QUBO models, and Notebook 7 develops exhaustive verification. This case study defines the helpers it needs so it remains independently runnable.

Mathematical background: Binary incidence matrices, matrix multiplication, quadratic objectives, and finite enumeration.

Data background: The committed TCGA AML files are gene-level reproducibility aggregates, not raw clinical records. Patient and sample identifiers are intentionally absent.

Software and accounts: Julia 1.10+ with the shared notebook project instantiated. No external account is required.

From incidence data to coverage and co-mutation

Let rows of the binary matrix BB represent patients or samples and columns represent genes. An entry Bp,g=1B_{p,g}=1 means gene gg is observed as mutated for patient pp. The cited formulation uses the transposed orientation (genes by patients); either convention produces the same gene-by-gene Gram matrix when used consistently.

We start with seven synthetic patients and four synthetic genes. The fixture is deliberately small enough to inspect without biological interpretation.

Tiny fixture: 7 patients × 4 genes

For each gene, the coverage value is the number of patients containing that gene. The diagonal matrix DD stores those counts. For each pair i<ji<j, AijA_{ij} is the number of patients containing both genes. We assign both AijA_{ij} and AjiA_{ji}, and leave the diagonal at zero. Thus

BTB=D+A.B^{\mathsf{T}}B = D + A.

The explicit pair loop below prevents the stale-index defect found in the companion notebooks.

coverage_and_comutation
D diagonal = [3, 2, 2, 2]
A =
4×4 Matrix{Int64}: 0 1 0 0 1 0 0 0 0 0 0 1 0 0 1 0

Build and interpret the QUBO

For binary xgx_g, where one selects gene gg, the tutorial model is

Q(x)=xT(A−αD)x=2∑i<jAijxixj−α∑iDiixi.Q(x)=x^{\mathsf{T}}(A-\alpha D)x = 2\sum_{i<j}A_{ij}x_ix_j-\alpha\sum_iD_{ii}x_i.

The negative diagonal term rewards gene coverage. The positive off-diagonal term penalizes pairwise co-mutation. Because AA is symmetric, xTAxx^{\mathsf{T}}Ax intentionally counts each undirected gene pair twice; code, analytic energy, and reported metrics all use that convention.

The named parameter α\alpha controls the modeling trade-off. It is not a confidence score and does not make the sampled genes clinically meaningful.

validate_decoded_metrics (generic function with 1 method)
same_states (generic function with 1 method)

Exhaustive validation on the tiny fixture

Four genes give 24=162^4=16 assignments. For every state, we compare the JuMP objective against the independent matrix calculation and against the explicit pairwise expansion. This proves the factor-of-two convention before any sampler is trusted.

All 16 assignments match the analytic matrix and explicit pairwise energies.
Enumerated optimum energy: -3.75
ExactSampler returns the same 2 global minima: [[1, 0, 0, 1], [1, 0, 1, 0]]

Decode pathway metrics safely

The coverage reward is the sum of per-gene patient counts and can count a patient more than once. Distinct-patient coverage is the union of affected patients. The tiny incidence matrix lets us compute that union exactly. Pairwise co-mutation is reported once per undirected gene pair, while the QUBO uses twice that value.

A zero- or one-gene pathway has no gene pairs, so the pair fraction is reported as missing instead of dividing by zero.

Selected genes: G-A, G-C
Coverage sum: 5; distinct-patient coverage: 5
Pairwise co-mutation: 0; raw energy: -3.75
Empty pathway decoded safely: size=0, coverage=0, co-mutation=0, energy=0.

Provenance-checked TCGA AML aggregate

The full educational example loads three committed CSV files and their JSON provenance record. The data refresh was performed separately by the reviewed workflow in issue #91; this notebook never contacts cBioPortal. The files retain only gene labels, coverage counts, pairwise co-mutation counts, and aggregate provenance. They contain no patient or sample identifiers.

The loader cross-checks gene ordering across all files, integer bounds, symmetry, zero diagonal, and the selected-gene and patient counts recorded in provenance.

load_tcga_aml_aggregate
Loaded 33 genes for 196 aggregate patients.
Source retrieval recorded at 2026-07-20T13:05:32Z; default execution made no network request.

What the aggregate can and cannot decode

For a selected set SS, the files determine the coverage reward ∑g∈SDg\sum_{g\in S}D_g, every pairwise co-mutation AijA_{ij}, and the QUBO energy exactly. They do not determine the exact union of affected patients for arbitrary ∣S∣>2|S|>2: triple and higher intersections were intentionally not committed.

We therefore report a rigorous interval. The largest selected-gene coverage and the second-order Bonferroni expression ∑Dg−∑i<jAij\sum D_g-\sum_{i<j}A_{ij} give lower bounds; the patient count and ∑Dg\sum D_g give upper bounds. This preserves the privacy/minimal-data design of issue #91 and avoids fabricating a distinct-patient count.

sample_aggregate (generic function with 1 method)
Seeded local Neal sampler: seed=94094, reads=500, sweeps=2000
alpha  states checked  size  coverage sum  distinct-patient bounds  pair co-mutation  raw energy
0.25      13      13     4            91        91–91                  0      -22.75
      selected genes: NPM1, RUNX1, TP53, FCGBP
0.45       8       8     6           109       106–109                 3      -43.05
      selected genes: FLT3, RUNX1, TP53, KIT, KRAS, FCGBP
1.00      23      23    11           152       133–152                19     -114.00
      selected genes: FLT3, RUNX1, TP53, CEBPA, IDH1, KIT, KRAS, RAD21, CROCC, CSMD1, FCGBP

The sampled pathway size and composition change across the three α\alpha values. These are reproducible outputs from a locked, seeded local sampler, not proofs of global optimality. The validity claim is narrower and checked for every returned state: its bit vector, decoded metrics, and independently recomputed raw energy agree.

Increasing α\alpha makes coverage more valuable relative to the pairwise co-mutation penalty, so larger selected sets become more attractive. No clinical interpretation is made from any selected name.

Practice checkpoints

  1. For the tiny state selecting G-A and G-B, compute the undirected co-mutation count and predict the contribution of xTAxx^{\mathsf{T}}Ax. Confirm the factor of two.

  2. Enumerate the tiny fixture at α=0.5\alpha=0.5 and α=1.5\alpha=1.5. Compare the optimal pathway sizes without relying on stochastic output.

  3. Select G-A, G-B, and G-C. Derive the aggregate-style distinct-patient bounds, then use the incidence matrix to locate the exact union inside them.

Notebook Cell
One undirected co-mutation contributes 2 to the symmetric matrix energy.
Notebook Cell
Tiny optimum sizes: alpha=0.5 → [2], alpha=1.5 → [4]
Notebook Cell
Aggregate-style bounds: 6–7 patients; incidence-derived exact union: 6 patients.

Summary

Learning objectives met:

  • A patient-by-gene incidence matrix yields coverage DD and a zero-diagonal symmetric co-mutation matrix AA by visiting every i<ji<j pair.

  • The QUBO xT(A−αD)xx^{\mathsf{T}}(A-\alpha D)x counts each undirected co-mutation twice, consistently in code and mathematics.

  • All 16 tiny states agree across the analytic matrix formula, explicit pair expansion, JuMP objective, and ExactSampler optimum set.

  • Empty-pathway metrics are safe, and exact distinct-patient coverage is computed only when incidence rows are available.

  • The committed 33-gene aggregate loads without a service call; every state returned by the seeded local sampler passes independent energy and metric checks.

  • Aggregate distinct-patient coverage is reported as a rigorous interval because higher-order patient intersections are not present.

Next steps: Notebook 11 will compare solver interfaces and optional D-Wave hardware. This modeling notebook intentionally stops at a local, credential-free sampler and makes no performance claim.

Further reading:

  • The Five Starter Problems tutorial gives the displayed QUBO formulation and its coverage/exclusivity motivation.

  • Alghassi et al. develop quantum and quantum-inspired formulations for altered-pathway discovery.

  • Ley et al. describe the underlying TCGA AML study; the committed provenance explains the exact cBioPortal aggregate used here.

References

  1. A. R. Mazumder and S. Tayur, Five Starter Problems: Solving Quadratic Unconstrained Binary Optimization Models on Quantum Computers, TutORials in Operations Research (2025), pp. 145–183, Mazumder & Tayur (2025).

  2. H. Alghassi, R. Dridi, A. G. Robertson, and S. Tayur, Quantum and Quantum-inspired Methods for de novo Discovery of Altered Cancer Pathways, bioRxiv 845719 (2019), Alghassi et al. (2019).

  3. T. J. Ley et al., Genomic and Epigenomic Landscapes of Adult De Novo Acute Myeloid Leukemia, New England Journal of Medicine 368 (2013), pp. 2059–2074, Massachusetts Medical Society (2013).

  4. E. Cerami et al., The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data, Cancer Discovery 2 (2012), pp. 401–404, Cerami et al. (2012); current portal: https://www.cbioportal.org/.

  5. Companion materials: https://github.com/arulrhikm/Solving-QUBOs-on-Quantum-Computers.

  6. D-Wave Neal simulated annealing documentation: https://docs.dwavequantum.com/en/latest/ocean/api_ref_samplers/index.html.

The committed artifact’s exact query, retrieval time, aggregation rules, license, and checksums are recorded in notebooks_data/9-CancerGenomics_provenance.json. This notebook contains original Julia code and prose. The unlicensed companion repository is cited as context; no cells, prose, outputs, figures, or assets were copied.

This educational reproduction is not a clinically validated pathway analysis, biomarker study, diagnostic tool, or medical guidance.

References
  1. Mazumder, A. R., & Tayur, S. (2025). Five Starter Problems: Solving Quadratic Unconstrained Binary Optimization Models on Quantum Computers. In Tutorials in Operations Research: Advances in Analytics and Operations Research: Improving Decisions to Secure the Future (pp. 145–183). INFORMS. 10.1287/educ.2025.0288
  2. Alghassi, H., Dridi, R., Robertson, A. G., & Tayur, S. (2019). Quantum and Quantum-inspired Methods for de novo Discovery of Altered Cancer Pathways. 10.1101/845719
  3. (2013). New England Journal of Medicine, 368(22), 2059–2074. 10.1056/nejmoa1301689
  4. Cerami, E., Gao, J., Dogrusoz, U., Gross, B. E., Sumer, S. O., Aksoy, B. A., Jacobsen, A., Byrne, C. J., Heuer, M. L., Larsson, E., Antipin, Y., Reva, B., Goldberg, A. P., Sander, C., & Schultz, N. (2012). The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data. Cancer Discovery, 2(5), 401–404. 10.1158/2159-8290.cd-12-0095