Data Science – Medicine Application Project 3, Part 2

Technical Replicates in qPCR: Why Correlation Is Not Enough

Two measurements can move together without being close enough to substitute for one another.

That distinction is easy to miss in qPCR data. If replicate 1 and replicate 2 produce a correlation near 1, the assay may look almost perfect. But correlation answers a limited question: **do high measurements in one replicate tend to remain high in the other?**

Quality control asks something more practical:

**How far apart are the two measurements for the same sample and assay?**

In [Part 1](/synthetic-qpcr-dataset-technical-replicates-qc/), we generated 96 biological samples, five qPCR assays, and two technical replicates per sample–assay combination. We also declared the QC rule before normalization: a complete pair is flagged when its absolute Cq difference exceeds 0.75 cycles.

Now we can examine what the replicates actually show.

## Start with the paired measurements

The synthetic dataset contains 480 planned replicate pairs:

– 471 pairs have both measurements available;

– 9 pairs contain only one detected measurement;

– 12 of the complete pairs exceed the 0.75-cycle disagreement threshold.

The first panel below plots replicate 1 against replicate 2. The dashed diagonal represents perfect equality.

Scatterplot and difference plot assessing agreement between duplicate qPCR technical replicates

*Figure 1. Agreement between qPCR technical replicates. The left panel compares replicate 1 with replicate 2 and includes the line of perfect equality. The right panel plots the pairwise mean against the replicate difference, with the mean difference and 95% limits of agreement.*

At first glance, the left panel is reassuring. Across all complete pairs, the Pearson correlation is **r = 0.995**.

That number is real, but it is not the whole story.

## Why the correlation looks so impressive

The plot pools five assays with different typical Cq ranges. Reference RNAs tend to have lower Cq values, while the candidate miRNAs tend to have higher values. This wide between-assay range makes it easy to preserve the overall ranking.

For example, a reference measurement around Cq 24 will usually remain lower than a candidate measurement around Cq 31 even if one pair contains a sizeable technical error. The pooled correlation therefore benefits from variation between assays that is unrelated to repeatability within a sample–assay pair.

The assay-specific correlations are lower, ranging from approximately **0.90 to 0.95**. They are still strong, but they illustrate how a pooled correlation can look better than the agreement within each assay.

More importantly, correlation does not express disagreement in cycle units. A consistent shift, or a small number of large replicate differences, can coexist with a very high correlation.

## Look at the differences directly

The second panel uses a Bland–Altman-style display:

– the horizontal position is the mean Cq of each pair;

– the vertical position is replicate 1 minus replicate 2;

– the solid line is the average difference;

– the dashed lines are the mean difference ± 1.96 standard deviations.

The original [Bland and Altman paper](https://pubmed.ncbi.nlm.nih.gov/2868172/) proposed examining paired differences because association alone does not establish agreement.

In this synthetic dataset:

– the mean replicate difference is **−0.01 cycles**;

– the approximate 95% limits of agreement are **−0.56 to 0.54 cycles**.

The average bias is therefore negligible. Replicate 1 is not systematically higher or lower than replicate 2.

However, a near-zero average can hide individual disagreements. Positive and negative errors cancel when averaged. That is why the individual points and the absolute differences still matter.

## Agreement limits are not the QC threshold

The dashed agreement limits summarize the observed spread of the differences. They do not automatically define what is analytically acceptable.

Our 0.75-cycle QC threshold was specified separately. It represents a rule for this teaching dataset, not a universal qPCR standard. In a real experiment, the acceptable difference should be justified using assay validation, expected precision, laboratory procedures, and the consequences of measurement error.

This separation prevents circular reasoning. If we defined the threshold only after seeing the data, we could move it until the number of failures looked convenient.

## Which assays produced the flags?

The assay-level summary makes the QC pattern easier to inspect.

Table showing complete, single-detected, and discordant qPCR replicate pairs for three candidate miRNAs and two reference RNAs
Table illustrating the quality of technical replicates in qPCR assays, including planned and complete pairs, and correlation data.

Most pairs are close. Across the five assays, the median absolute replicate difference ranges from **0.14 to 0.18 cycles**.

The deliberately introduced problems are spread across the assay panel:

– candidate-miR-A has 3 discordant pairs;

– candidate-miR-B has 4 discordant pairs;

– candidate-miR-C has 1 discordant pair and all 9 single-detected pairs;

– ref-RNA-1 has 2 discordant pairs;

– ref-RNA-2 has 2 discordant pairs.

Candidate-miR-C has fewer complete pairs because the simulation assigned its nine non-detect measurements to different sample pairs. This is a detection issue, not evidence that its complete pairs are necessarily less precise.

## Do not average first and inspect later

Suppose a pair contains Cq values of 27.1 and 28.3. Their mean is 27.7, which looks like an ordinary value in a sample-level spreadsheet. Once the original measurements are hidden, there may be no indication that the two replicates differed by 1.2 cycles.

For that reason, the workflow calculates the replicate difference before creating the analysis-level mean.

The primary rule is:

1. **Passing pair:** average the two measurements.

2. **Discordant pair:** retain the original values and QC flag, but do not create a primary-analysis mean.

3. **Single-detected pair:** retain the observed measurement and missing replicate, but do not silently treat the single value as a clean pair.

This rule is intentionally conservative. Later in the series, a sensitivity analysis will compare the primary QC-filtered result with an analysis that uses all available replicate means. If both approaches lead to similar effect estimates, that is reassuring. If they disagree, the QC decision becomes part of the scientific interpretation rather than a hidden preprocessing detail.

## Should the correlation p value be reported?

Not here.

The question is not whether the replicate correlation differs from zero. With hundreds of paired measurements spread across a wide Cq range, that null hypothesis is uninformative. A tiny p value would not tell us whether the replicates are close enough for their intended use.

The useful quantities are instead:

– the number of complete and incomplete pairs;

– the distribution of pairwise differences;

– the mean difference;

– the limits of agreement;

– the number exceeding the predefined QC threshold;

– the assay in which each problem occurs.

This is consistent with the broader emphasis on documenting intra-assay repeatability in [MIQE 2.0](https://academic.oup.com/clinchem/article/71/6/634/8119148).

## What this check does—and does not—prove

Technical repeatability tells us how consistently the same prepared material was measured under the same experimental setting. It does not capture every source of uncertainty.

Duplicate wells cannot, by themselves, evaluate:

– biological variability between people;

– RNA extraction differences between separate extractions;

– reverse-transcription variability when the same cDNA preparation is reused;

– reproducibility across laboratories, operators, or instruments;

– whether a reference RNA is stable across the groups being compared.

The last point is where the next post begins. Two references may each have tidy technical replicates and still behave very differently across study groups or batches.

## The practical takeaway

A high replicate correlation is encouraging, but it should not be used as a shortcut for agreement.

Keep the replicate-level data. Plot the paired differences. Define the QC rule before testing group effects. Record every exclusion, and later show whether the result changes when the rule changes.

That creates a traceable path from raw measurement to normalized result.

**Next:** *Reference RNA Stability: The Assumption Behind Every ΔCq Result*

Leave a Reply

Create a website or blog at WordPress.com

Up ↑

Discover more from Writing my way through ideas.

Subscribe now to keep reading and get access to the full archive.

Continue reading